Institution profile

DeepAuto.ai

Industry researchnorthamerica · us
Official website
Research library9linked papers
Opportunities0open roles
Selected work

Representative Papers

RunningTab: Direct Workspace Interaction with Environment-Side Tabs

Oct 07, 2026

This study addresses the loss of task requirements and file states caused by limited context windows when large language model (LLM) agents interact directly with workspaces. To mitigate this, we propose RunningTab, a framework that introduces an environment-side tab-based persistent recording mechanism to decouple agent memory from environmental states. This approach dynamically tracks task progress, file read/write statuses, and unprocessed candidates in real time, enabling continuous alignment between task requirements and workspace content. Extensive experiments conducted across three benchmarks and three LLMs demonstrate that RunningTab consistently outperforms both direct interaction paradigms and existing baselines, significantly enhancing the completeness of delivered artifacts.

0 citationsRead paper

Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge

Sep 28, 2026

This study addresses the inefficiency of small reasoning models constrained by parametric knowledge, where blindly increasing inference compute yields diminishing returns. We propose FlyBy, a framework that trains models to “reason first, diagnose later” by distinguishing execution from knowledge bottlenecks, enabling on-demand queries to stronger models when encountering knowledge deficits. The method integrates supervised fine-tuning, multi-depth query actions, and cost-aware reinforcement learning, alongside intermediate-state intervention analysis for selective querying. Experimental results demonstrate that FlyBy-4B surpasses Qwen3-14B in performance at lower serving costs, while the 8B variant achieves 51.81% Pass@8. This work establishes a new paradigm for efficient reasoning in small language models.

0 citationsRead paper

Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning

Sep 27, 2026

This study addresses the limitations of fine-grained credit assignment in large language model (LLM) reasoning, which typically relies on auxiliary models, and the tendency of conventional entropy-guided methods to suppress exploration. To overcome these challenges, this work proposes Entropy-Advantage Policy Optimization (EAPO), introducing a novel asymmetric token-level reward redistribution mechanism based on policy entropy and advantage signs. Without requiring additional supervision, EAPO reinforces unexpected successes and rectifies repetitive failures, thereby effectively balancing exploration and exploitation. Experimental results demonstrate that EAPO achieves state-of-the-art overall performance across diverse reasoning tasks and base models, significantly improving both problem coverage and answer diversity.

0 citationsRead paper
Recent publications

Latest Papers

RunningTab: Direct Workspace Interaction with Environment-Side Tabs

Oct 07, 2026

This study addresses the loss of task requirements and file states caused by limited context windows when large language model (LLM) agents interact directly with workspaces. To mitigate this, we propose RunningTab, a framework that introduces an environment-side tab-based persistent recording mechanism to decouple agent memory from environmental states. This approach dynamically tracks task progress, file read/write statuses, and unprocessed candidates in real time, enabling continuous alignment between task requirements and workspace content. Extensive experiments conducted across three benchmarks and three LLMs demonstrate that RunningTab consistently outperforms both direct interaction paradigms and existing baselines, significantly enhancing the completeness of delivered artifacts.

0 citationsRead paper

Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge

Sep 28, 2026

This study addresses the inefficiency of small reasoning models constrained by parametric knowledge, where blindly increasing inference compute yields diminishing returns. We propose FlyBy, a framework that trains models to “reason first, diagnose later” by distinguishing execution from knowledge bottlenecks, enabling on-demand queries to stronger models when encountering knowledge deficits. The method integrates supervised fine-tuning, multi-depth query actions, and cost-aware reinforcement learning, alongside intermediate-state intervention analysis for selective querying. Experimental results demonstrate that FlyBy-4B surpasses Qwen3-14B in performance at lower serving costs, while the 8B variant achieves 51.81% Pass@8. This work establishes a new paradigm for efficient reasoning in small language models.

0 citationsRead paper

Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning

Sep 27, 2026

This study addresses the limitations of fine-grained credit assignment in large language model (LLM) reasoning, which typically relies on auxiliary models, and the tendency of conventional entropy-guided methods to suppress exploration. To overcome these challenges, this work proposes Entropy-Advantage Policy Optimization (EAPO), introducing a novel asymmetric token-level reward redistribution mechanism based on policy entropy and advantage signs. Without requiring additional supervision, EAPO reinforces unexpected successes and rectifies repetitive failures, thereby effectively balancing exploration and exploitation. Experimental results demonstrate that EAPO achieves state-of-the-art overall performance across diverse reasoning tasks and base models, significantly improving both problem coverage and answer diversity.

0 citationsRead paper