AgentHabit: Characterizing Distinct Behaviors of Agents on Everyday Tasks

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue that large language model (LLM) agents frequently exhibit behaviors deviating from user preferences without systematic characterization. To this end, it proposes HABIT, a taxonomy comprising 5 categories and 23 axes, along with the AgentHABIT benchmark. Through bottom-up induction, trajectory analysis, and prompt-tuning experiments, we conduct quantitative evaluations across 18 models. Results demonstrate that the proposed taxonomy achieves superior model discriminability and annotation consistency compared to existing approaches. Furthermore, the identified behavioral discrepancies exhibit cross-task stability, reflecting general model tendencies rather than task-specific biases. These findings offer critical guidance for optimizing agent alignment.
📝 Abstract
Large language model (LLM) agents assist users with everyday tasks that can be completed in many reasonable ways. Even when their answers are useful, how agents carry out these tasks may not match users'preferences and needs. For example, agents differ in whether they ask clarifying questions or search the web. We introduce HABIT, a taxonomy of 23 behavioral axes in five categories, which three authors and three LLMs derive bottom-up from 408 agent trajectories across 17 domains. On held-out tasks, HABIT distinguishes models more clearly than existing taxonomies of human values and agent actions while supporting comparably consistent annotation. Building on HABIT, we construct AgentHABIT, a benchmark that profiles each agent's behavioral tendencies from its trajectories on 86 everyday tasks. Profiling 18 models with AgentHABIT reveals a range of distinctive tendencies. For example, most GPT and Claude models state their assumptions and offer alternatives when requirements conflict, whereas Qwen and Google's models more often leave assumptions or changes to requirements unstated. These profiles remain recognizable even when built from entirely different sets of tasks, indicating that they reflect general tendencies rather than task-specific behavior. Prompting agents to adopt specific behaviors shifts some axes readily but barely changes others, while fine-tuning on another model's trajectories changes only part of a model's profile and leaves much of it intact. Overall, HABIT and AgentHABIT provide a systematic framework for characterizing how agents carry out everyday tasks beyond task success, offering insights to guide the development of agents whose behavior better fits users'needs.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
behavioral characterization
everyday tasks
user preferences
agent profiling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Behavioral Taxonomy
Agent Profiling
LLM Agents
Benchmark
Behavioral Tendencies
🔎 Similar Papers
W
Woojung Song
Seoul National University, Seoul, Republic of Korea
H
Hoyeol Yang
Seoul National University, Seoul, Republic of Korea
Jeonghoon Shim
Jeonghoon Shim
SNU Ph.D Student
Dialogue SystemLLM
S
Sungjib Lim
Seoul National University, Seoul, Republic of Korea
J
Jonggeun Lee
Seoul National University, Seoul, Republic of Korea
Yunho Choi
Yunho Choi
Gwangju Institute of Science and Technology
AI
Yohan Jo
Yohan Jo
Seoul National University
Natural Language ProcessingAgentsComputational PsychologyReasoning