Increasing Skill Level Recruits Deeper Attention Layers in a Frozen Chess Transformer
研究通过调整Maia-3棋类转换器的技能水平,发现提高技能会促使更深层次注意力层的参与,特别是在特定战术中。
研究通过调整Maia-3棋类转换器的技能水平,发现提高技能会促使更深层次注意力层的参与,特别是在特定战术中。
This study investigates task-dependent functional redundancy and approximate invariance in recurrent neural networks. By applying the real Schur decomposition, the recurrent weight matrix is decoupled into spectral blocks and non-normal coupling structures. Structured ablation is then performed while preserving the input–output mapping to identify weight perturbations that do not significantly affect task performance. The work introduces the notion of “task-dependent approximate functional invariance,” revealing non-universal yet task-specific symmetries within recurrent architectures. Experiments across dynamic tasks—including copying, flip-flop triggering, sine wave generation, and context-dependent integration—demonstrate that certain non-normal Schur couplings can be safely removed, while others are essential for autonomous replay. These findings validate the existence of task-constrained symmetries and provide an interpretable diagnostic framework for analyzing recurrent network dynamics.
This work investigates how reinforcement learning (RL) during cross-domain transfer suppresses exploratory reasoning primitives, thereby impairing performance on complex mathematical problem solving. By exclusively employing constraint-satisfaction puzzles to conduct supervised fine-tuning and RL post-training on a 7B language model, the study introduces a reasoning-primitive-based analytical framework that, for the first time, reveals RL’s detrimental effect on exploratory reasoning. To mitigate this issue, the authors propose a novelty reward mechanism grounded in reference model perplexity, combined with a reasoning primitive segmentation approach leveraging nine span classifiers and motive extraction, as well as the GSPO algorithm. Remarkably, without using any mathematical training data, this method boosts pass@32 on OlymMATH-Hard from 16.0% to 36.0%, achieving a 20-percentage-point improvement over the baseline.
This study addresses the poor performance of large language models in spatial reasoning tasks involving global topological properties—such as connectivity, loop closure, and regional symmetry—in grid-based puzzles. To systematically evaluate and analyze this limitation, the authors introduce TopoBench, the first standardized benchmark dedicated to hard topological reasoning, encompassing six puzzle types across three difficulty levels, along with a fine-grained error taxonomy. Through chain-of-thought annotations, intervention simulations, and tool-augmented constraint verification, the work diagnoses the root causes of model failures. Experimental results reveal that state-of-the-art models solve fewer than 25% of challenging instances, with the primary bottleneck lying in extracting valid constraints from spatial representations rather than in subsequent deductive reasoning.
This study addresses the longstanding lack of deep integration between artificial intelligence and neuroscience, which has hindered both algorithmic efficiency and our understanding of biological intelligence. The authors propose and advocate for a “NeuroAI” paradigm grounded in bidirectional synergy: on one hand, leveraging neuroscientific principles—such as embodied cognition and neuromorphic computing—to advance next-generation AI systems; on the other, employing AI models to enhance the interpretation of biological neural computation. Through a systematic review of the current landscape, synthesis of expert perspectives, and a SWOT analysis, the work identifies concrete collaborative pathways in domains including language communication, robotics, and human-AI co-learning, thereby offering a clear roadmap for future interdisciplinary research at the intersection of neuroscience and artificial intelligence.
研究通过调整Maia-3棋类转换器的技能水平,发现提高技能会促使更深层次注意力层的参与,特别是在特定战术中。
This study investigates task-dependent functional redundancy and approximate invariance in recurrent neural networks. By applying the real Schur decomposition, the recurrent weight matrix is decoupled into spectral blocks and non-normal coupling structures. Structured ablation is then performed while preserving the input–output mapping to identify weight perturbations that do not significantly affect task performance. The work introduces the notion of “task-dependent approximate functional invariance,” revealing non-universal yet task-specific symmetries within recurrent architectures. Experiments across dynamic tasks—including copying, flip-flop triggering, sine wave generation, and context-dependent integration—demonstrate that certain non-normal Schur couplings can be safely removed, while others are essential for autonomous replay. These findings validate the existence of task-constrained symmetries and provide an interpretable diagnostic framework for analyzing recurrent network dynamics.
This work investigates how reinforcement learning (RL) during cross-domain transfer suppresses exploratory reasoning primitives, thereby impairing performance on complex mathematical problem solving. By exclusively employing constraint-satisfaction puzzles to conduct supervised fine-tuning and RL post-training on a 7B language model, the study introduces a reasoning-primitive-based analytical framework that, for the first time, reveals RL’s detrimental effect on exploratory reasoning. To mitigate this issue, the authors propose a novelty reward mechanism grounded in reference model perplexity, combined with a reasoning primitive segmentation approach leveraging nine span classifiers and motive extraction, as well as the GSPO algorithm. Remarkably, without using any mathematical training data, this method boosts pass@32 on OlymMATH-Hard from 16.0% to 36.0%, achieving a 20-percentage-point improvement over the baseline.
This study addresses the poor performance of large language models in spatial reasoning tasks involving global topological properties—such as connectivity, loop closure, and regional symmetry—in grid-based puzzles. To systematically evaluate and analyze this limitation, the authors introduce TopoBench, the first standardized benchmark dedicated to hard topological reasoning, encompassing six puzzle types across three difficulty levels, along with a fine-grained error taxonomy. Through chain-of-thought annotations, intervention simulations, and tool-augmented constraint verification, the work diagnoses the root causes of model failures. Experimental results reveal that state-of-the-art models solve fewer than 25% of challenging instances, with the primary bottleneck lying in extracting valid constraints from spatial representations rather than in subsequent deductive reasoning.
This study addresses the longstanding lack of deep integration between artificial intelligence and neuroscience, which has hindered both algorithmic efficiency and our understanding of biological intelligence. The authors propose and advocate for a “NeuroAI” paradigm grounded in bidirectional synergy: on one hand, leveraging neuroscientific principles—such as embodied cognition and neuromorphic computing—to advance next-generation AI systems; on the other, employing AI models to enhance the interpretation of biological neural computation. Through a systematic review of the current landscape, synthesis of expert perspectives, and a SWOT analysis, the work identifies concrete collaborative pathways in domains including language communication, robotics, and human-AI co-learning, thereby offering a clear roadmap for future interdisciplinary research at the intersection of neuroscience and artificial intelligence.