Score
Designs and implements systems, prompts, or algorithms that produce explicit step‑by‑step tactile reasoning traces (chain‑of‑thought) for models that process touch‑related inputs or tactile‑language tasks. Builds and evaluates the representations and control procedures that structure intermediate tactile inference steps to guide model decisions and reduce hallucination in tactile‑language outputs.
This work addresses the challenge that multimodal large language models struggle to correct vision-based priors using physical tactile evidence during reasoning. To this end, the authors introduce TouchReason-Bench, a large-scale tactile–language dataset and evaluation benchmark, and propose the first reinforcement learning–based tactile reasoning framework built upon Qwen2.5-VL-7B. The approach incorporates a tactile-grounded GRPO training objective and a tactile-utilization reward mechanism to ensure effective integration of real tactile inputs. Experimental results demonstrate that the resulting model, Touch-R1-7B, substantially outperforms Octopi-13B by 18.4% and GPT-4o by 24.7% on average across TouchReason-Bench, marking a significant advance in tactile reasoning capabilities.
Humanoid robots struggle with long-horizon, multimodal coordination of locomotion-manipulation tasks in complex, unstructured environments. Method: This paper proposes the Embodied Chain-of-Action (ECOA) framework—a novel, humanoid-specific chain-of-thought paradigm grounded in multimodal foundation models. ECOA jointly models object affordances, whole-body kinematics, and spatial reasoning to enable cross-modal action generation under occlusion or for unseen objects; it further employs upper-lower limb decoupled control and joint vision-language-action representation learning to achieve end-to-end mapping from natural language instructions to coordinated whole-body actions. Contribution/Results: Evaluated on real-world object rearrangement and loco-manipulation tasks, ECOA significantly improves instruction understanding accuracy and task success rate, effectively bridging the semantic gap between high-level planning and low-level execution.
Existing structured prompting paradigms—such as Chain-of-Thought (CoT), Tree-of-Thought (ToT), and Graph-of-Thought (GoT)—lack a unified theoretical foundation, suffering from conceptual conflation and an absence of systematic taxonomy. Method: We propose the first comprehensive taxonomy for structured prompting, formally defining the notion of “reasoning topology,” constructing its spatial representation, and unifying CoT, ToT, and GoT through pipeline-based execution analysis, structural modeling, behavioral interpretation, and cross-paradigm empirical comparison. Contribution/Results: (1) We establish the first principled taxonomy for structured-prompt reasoning; (2) we uncover intrinsic relationships between topological structure and both reasoning performance and computational cost; and (3) we provide a theoretically grounded framework and design principles for scalable, interpretable prompt engineering.
This work addresses the limitations of existing tactile foundation models, which suffer from hallucinations due to the absence of explicit reasoning mechanisms and struggle to model dynamic tactile signals. To overcome these challenges, the authors propose a dynamic tactile-language reasoning framework that integrates a dynamics-aware tactile encoder with a structured chain-of-thought reasoning mechanism. They further introduce TouchCoT-10k, the first tactile chain-of-thought dataset, and DynTac-Bench, a benchmark for evaluating dynamic tactile reasoning. Leveraging a 7B-parameter large language model, the proposed approach significantly outperforms current methods—including the 14B-parameter VTV-LLM—across multiple tactile commonsense reasoning tasks, achieving more accurate and efficient multimodal interactive reasoning.
This work addresses the challenge faced by blind and low-vision students in accessing statistical graphics, a barrier exacerbated by the inefficiency and specialized CAD expertise required by conventional 3D printing approaches, which hinder classroom-scale deployment. To overcome this, the authors propose a reusable, three-tier software pipeline that, for the first time, automatically integrates haptic perceptual parameters into the generation workflow. Leveraging a multimodal large language model to parse chart structures directly from images, the system supports automated tactile rendering of scatter plots, bar charts, histograms, line graphs, and box plots. Implemented in JavaScript with a modular architecture, the pipeline produces print-ready STL files in under 250 milliseconds, substantially lowering production barriers and enabling educators to rapidly review and deploy accessible data representations, thereby enhancing data accessibility in inclusive education.
It remains challenging to determine whether chain-of-thought (CoT) reasoning in large language models reflects genuine internal reasoning or merely superficial, post-hoc rationalization. Method: This paper introduces Concept Walk, the first framework to model CoT reasoning explicitly in semantic concept space. It employs contrastive learning to extract interpretable concept directions, then projects hidden-layer activations onto these directions to dynamically track the evolution of internal representations throughout reasoning. Contribution/Results: Concept Walk enables fine-grained diagnosis of reasoning faithfulness: on simple tasks, perturbations decay rapidly—indicating decorative CoT; on difficult tasks, perturbations induce sustained, directional shifts in concept activation—confirming substantive reasoning. Experiments on Qwen3-4B validate its effectiveness, establishing a novel methodology for model interpretability grounded in dynamic concept-level analysis.
This work addresses the significant yet poorly understood performance variations of Chain-of-Thought (CoT) reasoning across different tasks by providing the first theoretical framework for its step-by-step inference process. The authors model CoT as a Markov chain and propose that its effectiveness hinges on the consistency of the transition kernels between reasoning steps, while also quantifying how noise in intermediate steps degrades performance. Through rigorous theoretical analysis, they prove that consistent transition kernels substantially reduce sample complexity. To validate these predictions, they construct synthetic benchmark experiments that align with their theoretical findings and offer a principled explanation for the observed disparities in CoT’s empirical success across real-world tasks. This study thus establishes a novel theoretical lens and analytical framework for understanding and improving CoT reasoning.
Existing tactile commonsense reasoning systems struggle to generalize to open-world scenarios due to limited data scale and insufficient modeling of the action-dependence and redundancy inherent in tactile signals. To address this, this work introduces TouchThinker-1M, a million-scale, multi-source tactile reasoning dataset, along with TouchThinker-Bench, a comprehensive open-world evaluation benchmark. Furthermore, we propose an action-aware tactile representation method that leverages a tactile-language fusion framework to explicitly model action context, significantly enhancing both representational efficiency and semantic expressiveness. The proposed approach achieves state-of-the-art performance across multiple benchmarks, establishing a scalable data and modeling foundation for tactile-driven embodied intelligence.
This work addresses the challenge of simultaneously maintaining contact and accurately tracking object contours in robotic contour-following tasks by proposing a vision-based tactile model predictive control framework (VBT-MPC). For the first time, model predictive control is directly applied in the contour feature space extracted from an eye-in-hand visuotactile sensor, eliminating the need for separate pose estimation or complex force-control modules. By integrating visuotactile perception with feature-based visual servoing, VBT-MPC achieves high-precision and stable contour tracking across objects with diverse geometries and material properties in both simulation and real-world experiments. This approach substantially simplifies the system architecture while significantly enhancing tracking performance.
Current large reasoning models (LRMs) exhibit emergent hierarchical reasoning capabilities via chain-of-thought (CoT) prompting, yet their internal reasoning dynamics remain poorly understood and lack interpretable, formal modeling. Method: We propose a memoryless finite-state machine (FSM)-based framework to model CoT reasoning dynamics, abstracting CoT trajectories into high-level semantic state sequences—automatically identifying critical states including initialization, derivation, strategy enhancement, uncertainty estimation, and backtracking—and integrating human annotation for structured parsing and visualization. Contribution/Results: Unlike black-box analyses, our framework systematically characterizes cross-model differences in reasoning pathways, pinpoints typical failure patterns (e.g., strategy rigidity, insufficient backtracking), and exposes underlying mechanistic deficiencies. It establishes a novel, interpretable, and computationally tractable paradigm for evaluating reasoning capability, guiding training optimization, and enhancing model robustness.