Towards Understanding the Cognitive Habits of Large Reasoning Models

📅 2026-04-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper investigates whether large reasoning models (LRMs) exhibit human-like cognitive habits and how such habits influence their behavior and safety. Method: Grounded in cognitive science’s “habits of mind” theory, we introduce CogTest—the first systematic benchmark for evaluating 16 distinct cognitive habits across 400+ tasks—and propose evidence-prioritized extraction to analyze habit deployment patterns within chain-of-thought (CoT) reasoning. Our approach uniquely integrates human cognitive frameworks into LLM interpretability analysis via multi-model comparison, pattern mining, and correlation assessment with safety-sensitive tasks. Contribution/Results: Experiments across 13 LRMs reveal widespread task-adaptive habit activation and cross-architectural consistency; notably, habits such as “taking responsible risks” correlate significantly with harmful outputs. The CogTest benchmark and code are publicly released to advance research in LLM cognitive modeling, interpretability, and safety.

Technology Category

Natural Language Processing: Safety and RobustnessCognitive Modeling & Cognitive Systems: Simulating Human BehaviorMachine Learning: Large Multimodal Models (LMMs)

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
Large Reasoning Models (LRMs), which autonomously produce a reasoning Chain of Thought (CoT) before producing final responses, offer a promising approach to interpreting and monitoring model behaviors. Inspired by the observation that certain CoT patterns -- e.g., ``Wait, did I miss anything?'' -- consistently emerge across tasks, we explore whether LRMs exhibit human-like cognitive habits. Building on Habits of Mind, a well-established framework of cognitive habits associated with successful human problem-solving, we introduce CogTest, a principled benchmark designed to evaluate LRMs' cognitive habits. CogTest includes 16 cognitive habits, each instantiated with 25 diverse tasks, and employs an evidence-first extraction method to ensure reliable habit identification. With CogTest, we conduct a comprehensive evaluation of 16 widely used LLMs (13 LRMs and 3 non-reasoning ones). Our findings reveal that LRMs, unlike conventional LLMs, not only exhibit human-like habits but also adaptively deploy them according to different tasks. Finer-grained analyses further uncover patterns of similarity and difference in LRMs' cognitive habit profiles, particularly certain inter-family similarity (e.g., Qwen-3 models and DeepSeek-R1). Extending the study to safety-related tasks, we observe that certain habits, such as Taking Responsible Risks, are strongly associated with the generation of harmful responses. These findings suggest that studying persistent behavioral patterns in LRMs' CoTs is a valuable step toward deeper understanding of LLM misbehavior. The code is available at: https://github.com/jianshuod/CogTest.
Problem

Research questions and friction points this paper is trying to address.

Exploring whether Large Reasoning Models exhibit human-like cognitive habits
Evaluating LRMs' cognitive habits using CogTest benchmark with 16 habits
Identifying associations between cognitive habits and harmful response generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Introduces CogTest benchmark for LRM habits
Employs evidence-first extraction for habit identification
Evaluates 16 LLMs on 16 cognitive habits
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.