HARP: The Human--AI Research Platform

📅 2026-07-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of fine-grained investigation into dynamic user interactions with large language models (LLMs) in realistic yet controlled settings, particularly the lack of joint control over prompt construction behaviors and model responses. To bridge this gap, the authors present a configurable experimental platform that enables precise manipulation of LLM behavioral parameters while simultaneously capturing multimodal data—including keystroke-level prompt editing actions (e.g., typing, deletions, pauses), interaction timing, and subjective user feedback. The platform uniquely supports coordinated control over model behavior, user input, and contextual conditions during LLM interactions, integrating behavioral logs with self-report measures to facilitate causal inference in AI interface design. Empirical validation demonstrates that technical detail level and response length significantly affect user memory retention, confirming the platform’s efficacy for systematically evaluating design variables in AI-human interaction.
📝 Abstract
Large language models (LLMs) have shifted human--computer interaction from `traditional'' interface journeys toward more conversational exchanges. Researchers studying HCI and UI use moderated usability sessions, interviews, surveys, transcript analysis, and static prototypes. However, static prototypes provide limited opportunities to study interaction with live AI systems or systematically control how an LLM behaves across participants and scenarios. Conversation transcripts reveal little about how users formulate, revise, and hesitate over prompts before submission. We designed the Human--AI Research Platform (HARP) for researchers, designers, and anyone who has ever wondered, `What if AI did this?' HARP places participants in controlled mock scenarios with live, configurable AI agents. Researchers can control agent prompts, model parameters, response characteristics, and experimental conditions; trigger surveys at predefined moments; and record prompt composition time, response latency, deletions, and keystroke pauses. Planned capabilities include voice, facial expression, gesture, and, where legally and ethically appropriate, emotion analysis. We illustrate HARP through a study examining how technical specificity and response length affect retention of LLM output. By pairing controllable live agents with behavioral and self-report measures, HARP enables systematic testing of how AI design choices affect users.
Problem

Research questions and friction points this paper is trying to address.

human-AI interaction
large language models
usability research
conversational AI
experimental control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Human-AI Interaction
Controllable LLM Agents
Behavioral Logging
Experimental Platform
Prompt Engineering
🔎 Similar Papers
2024-07-31arXiv.orgCitations: 2
Z
Zeshu Zhu
BTPX Innovation Lab
N
Natalie Friedman
BTPX Innovation Lab
K
Kevin Weatherwax
BTPX Innovation Lab
E
Emily Eiben
BTPX User Assistance