Institution profile

DeepWisdom

Industry researchasia · cn
Official website
Research library25linked papers
Opportunities0open roles
Selected work

Representative Papers

MindFlow: Mind Supernet Powered Thinking Flows for Research Idea Innovation

Oct 08, 2026

This study addresses the limitations of static workflows, the trade-off between novelty and feasibility, and uncontrollable evaluation in scientific idea generation by proposing an explicitly controllable graph-structured flow-of-thought framework. This method models ideation as a directed graph, incorporating modular cognitive operators and a probabilistic supernetwork. A controller dynamically samples high-quality reasoning paths via tournament-based relative ranking optimization, while a comprehensive evaluation protocol is established to balance problem discovery with resolution. Multi-topic experiments demonstrate the superiority of this framework, achieving explicit generation, controllable optimization, and high-quality innovation of scientific research ideas.

0 citationsRead paper

ScopeSAE: Model-Scope Feature Discovery with Interpretable Layer Selection

Oct 04, 2026

This study addresses the limitations of sparse autoencoders (SAEs) arising from fixed hierarchical subspaces, including low feature utilization, high redundancy, and a lack of behavioral correlation. To overcome these issues, we propose a dynamic subspace selection mechanism based on predictive relevance. Methodologically, we employ normalized gradient attribution to achieve token-level state selection, integrating SAE architectures with intervention analysis for modeling. Notably, this work identifies and validates, for the first time, the phenomenon wherein reconstructions outperform original activations. Experimental results demonstrate that the proposed approach significantly increases the number of effective features while enhancing interpretability and dictionary efficiency. Furthermore, it reduces next-token cross-entropy loss, comprehensively outperforming hierarchical baseline methods across all evaluated metrics.

0 citationsRead paper

MESH-Harness: Self-Improving Agent Harnesses via Bandit-Guided Compositional Evolution

Oct 04, 2026

This study addresses the challenge of optimizing agent frameworks under limited evaluation budgets by proposing an efficient self-improvement method that freezes model weights and optimizes only the framework. Through modular design and combinatorial evolution, it overcomes enumeration bottlenecks while synergizing module-level reuse with feedback-driven optimization via full-covariance LinUCB-guided search, mixed-init coordinate ascent, and validation-trajectory-driven code iteration. Experimental results demonstrate that the proposed approach significantly outperforms the Meta-Harness baseline across multiple tasks, reducing testing costs by 44.2% and total costs by 14.6%.

0 citationsRead paper

CEO Arena: Evaluating Long-Horizon Multi-Agent Decision-Making in Competitive Markets

Sep 28, 2026

This study addresses the challenge of evaluating agents’ long-term strategic decision-making and competitive adaptability in uncertain markets. To this end, it constructs an eight-firm simulated market benchmark and introduces a “match-replacement evaluation” mechanism, wherein large language model-driven CEO agents engage in long-horizon multi-agent games involving pricing and R&D decisions, effectively disentangling individual returns from market externalities. The findings reveal that most agents yield negative average returns, indicating that private gains frequently entail aggregate market losses. Furthermore, the analysis identifies critical interaction patterns, such as demand capture and competitor pricing responses. Collectively, this work establishes a rigorous, controlled evaluation framework for investigating the economic behavior of large language models.

0 citationsRead paper

StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

May 26, 2026

This work addresses the credit assignment mismatch between sparse trajectory rewards and critical local actions in multi-turn agent reinforcement learning. It proposes StepOPSD, a novel framework that refines credit assignment to the level of individual agent steps for the first time. StepOPSD guides GRPO policy updates through step-level trajectory decomposition, hindsight-augmented teacher context re-scoring, sign-preserving advantage shaping, and normalized credit budgeting. Additionally, it introduces a dual-parameter control mechanism—α_clip and λ_mix—to modulate learning dynamics. The method achieves state-of-the-art performance on ALFWorld (e.g., 79.1% on Heat, 95.0% on PickTwo) and Search-QA (61.6% on TriviaQA). Empirical analysis further reveals that α_clip stabilizes local trust regions, while λ_mix exhibits task-dependent tuning characteristics.

0 citationsRead paper
Recent publications

Latest Papers

MindFlow: Mind Supernet Powered Thinking Flows for Research Idea Innovation

Oct 08, 2026

This study addresses the limitations of static workflows, the trade-off between novelty and feasibility, and uncontrollable evaluation in scientific idea generation by proposing an explicitly controllable graph-structured flow-of-thought framework. This method models ideation as a directed graph, incorporating modular cognitive operators and a probabilistic supernetwork. A controller dynamically samples high-quality reasoning paths via tournament-based relative ranking optimization, while a comprehensive evaluation protocol is established to balance problem discovery with resolution. Multi-topic experiments demonstrate the superiority of this framework, achieving explicit generation, controllable optimization, and high-quality innovation of scientific research ideas.

0 citationsRead paper

ScopeSAE: Model-Scope Feature Discovery with Interpretable Layer Selection

Oct 04, 2026

This study addresses the limitations of sparse autoencoders (SAEs) arising from fixed hierarchical subspaces, including low feature utilization, high redundancy, and a lack of behavioral correlation. To overcome these issues, we propose a dynamic subspace selection mechanism based on predictive relevance. Methodologically, we employ normalized gradient attribution to achieve token-level state selection, integrating SAE architectures with intervention analysis for modeling. Notably, this work identifies and validates, for the first time, the phenomenon wherein reconstructions outperform original activations. Experimental results demonstrate that the proposed approach significantly increases the number of effective features while enhancing interpretability and dictionary efficiency. Furthermore, it reduces next-token cross-entropy loss, comprehensively outperforming hierarchical baseline methods across all evaluated metrics.

0 citationsRead paper

MESH-Harness: Self-Improving Agent Harnesses via Bandit-Guided Compositional Evolution

Oct 04, 2026

This study addresses the challenge of optimizing agent frameworks under limited evaluation budgets by proposing an efficient self-improvement method that freezes model weights and optimizes only the framework. Through modular design and combinatorial evolution, it overcomes enumeration bottlenecks while synergizing module-level reuse with feedback-driven optimization via full-covariance LinUCB-guided search, mixed-init coordinate ascent, and validation-trajectory-driven code iteration. Experimental results demonstrate that the proposed approach significantly outperforms the Meta-Harness baseline across multiple tasks, reducing testing costs by 44.2% and total costs by 14.6%.

0 citationsRead paper

CEO Arena: Evaluating Long-Horizon Multi-Agent Decision-Making in Competitive Markets

Sep 28, 2026

This study addresses the challenge of evaluating agents’ long-term strategic decision-making and competitive adaptability in uncertain markets. To this end, it constructs an eight-firm simulated market benchmark and introduces a “match-replacement evaluation” mechanism, wherein large language model-driven CEO agents engage in long-horizon multi-agent games involving pricing and R&D decisions, effectively disentangling individual returns from market externalities. The findings reveal that most agents yield negative average returns, indicating that private gains frequently entail aggregate market losses. Furthermore, the analysis identifies critical interaction patterns, such as demand capture and competitor pricing responses. Collectively, this work establishes a rigorous, controlled evaluation framework for investigating the economic behavior of large language models.

0 citationsRead paper

StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

May 26, 2026

This work addresses the credit assignment mismatch between sparse trajectory rewards and critical local actions in multi-turn agent reinforcement learning. It proposes StepOPSD, a novel framework that refines credit assignment to the level of individual agent steps for the first time. StepOPSD guides GRPO policy updates through step-level trajectory decomposition, hindsight-augmented teacher context re-scoring, sign-preserving advantage shaping, and normalized credit budgeting. Additionally, it introduces a dual-parameter control mechanism—α_clip and λ_mix—to modulate learning dynamics. The method achieves state-of-the-art performance on ALFWorld (e.g., 79.1% on Heat, 95.0% on PickTwo) and Search-QA (61.6% on TriviaQA). Empirical analysis further reveals that α_clip stabilizes local trust regions, while λ_mix exhibits task-dependent tuning characteristics.

0 citationsRead paper