Score
Designs and implements a control policy that selects cutting motions, contact forces, and tool poses to execute and adapt cutting tasks in real time. The policy is learned (commonly via reinforcement learning) to optimize objectives such as speed–energy trade-offs, perform online force and pose adjustments, and generalize across variations in workpiece properties.
This work addresses the challenge of dynamically selecting cutting tools and strategies for robotic cooking systems when handling food items with diverse physical properties. The authors propose an integrated perception-action framework for adaptive food cutting, which leverages force feedback from preliminary incision trials to autonomously select appropriate cutting tools and employs reinforcement learning to optimize cutting policies in real time. By uniquely combining force-guided tool selection with reinforcement learning–based control, the system achieves 100% tool selection accuracy and high-precision fully automated cutting on previously unseen ingredients. The overall performance matches that of human operators while effectively balancing cutting speed and energy consumption.
本文提出了一种名为Agent as Policy (AGP)的方法,通过让通用智能体直接控制物理机器人完成任务,无需特定任务或环境的额外训练,解决了机器人操作中的泛化问题。
Dexterous manipulation of articulated tools (e.g., tweezers, scissors) by anthropomorphic robotic hands is challenging due to dynamic tool configuration changes, which complicate perception and control. Method: We propose a hierarchical goal-conditioned reinforcement learning framework. A low-level policy executes fine-grained motion control, while a high-level policy, guided by a tool-efficiency state encoder, perceives and adapts to real-time tool configuration. A privileged-information-augmented heuristic replay buffer accelerates training. The method integrates point-cloud encoding, synthetic-data pretraining, and goal-conditioned policies. Contribution/Results: This design significantly improves generalization to unseen object shapes and sizes. On a physical robot platform, the system achieves a 70.8% success rate in tweezer-like tool manipulation, demonstrating effectiveness and practicality for dexterous handling of complex articulated tools.
This work addresses the challenge of enabling robots to autonomously perform contact-rich, force-sensitive fine manipulation tasks—such as peeling—whose success criteria are inherently subjective. To this end, the authors propose a two-stage learning framework: first, a robust initial policy is acquired through force-aware imitation learning; then, a reward model is constructed using quantitative metrics and human preference feedback to fine-tune the policy via preference optimization, aligning its behavior with human judgments of task quality. This study presents the first application of human preference learning to fine manipulation. Using only 50–200 demonstration trajectories, the approach achieves over 90% average success rates across diverse produce—including cucumbers, apples, and potatoes—and demonstrates approximately 40% performance improvement from preference fine-tuning, along with strong zero-shot generalization across object types.
This work addresses robotic assembly tasks requiring precise force application under pose uncertainty and frequent physical contact. We propose a force-aware sim-to-real transfer framework that introduces a novel force-guided exploration paradigm. Our method integrates a force-threshold-conditioned policy network, dynamics randomization during training, and a task-success probability prediction model to enable online, adaptive tuning of force constraints during deployment. Unlike conventional controllers relying on fixed gains and accurate pose estimates, our approach eliminates strict pose dependency and supports safe, multi-stage adaptive manipulation. We validate the framework on fully autonomous planetary gear assembly—including snap-fit insertion, nut threading, and gear meshing—under significant pose errors. Results demonstrate high task success rates and robust contact safety, even with substantial pose uncertainty. This work establishes a new paradigm for dexterous force-controlled assembly in unstructured, uncertain environments.
该研究通过结合模型预测控制(MPC)与强化学习(RL),解决了真实世界中灵巧操作任务早期探索效率低下的问题。
This study addresses the performance limitations of contact-rich manipulation caused by mismatches between preset stiffness or directions in traditional control and environmental constraints. We propose a proprioception-based reflex strategy trained in simulation with a frozen execution layer. By leveraging interaction primitives such as springs and planes alongside state-history mapping, the method translates task commands into joint targets, decoupling contact responses from command generation without requiring direct force or geometric measurements. Experimental results demonstrate that this approach maintains low contact forces during box lifting and outperforms baselines in surface following. Furthermore, it improves peg-in-hole insertion success rates by 36–58% while reducing contact forces by approximately 50%, thereby providing robust low-level execution capabilities for high-level planning.
This study addresses the inefficiency and poor generalization of traditional trial-and-error approaches in laser cutting of optical thin films. To overcome these limitations, the authors propose RL²C, a reinforcement learning framework that integrates Q-learning with an ε-greedy strategy and incorporates a dynamic state-space expansion mechanism to adaptively tune critical parameters such as focal length and laser power. This work represents the first application of reinforcement learning with dynamic environmental adaptability to industrial laser cutting, significantly enhancing optimization efficiency and cross-material generalization. Experimental results demonstrate that RL²C reduces the number of optimization steps by 12.5% and processing time by 81.8% compared to existing methods, while effectively minimizing cut taper and film loss.
This study addresses the difficulty of traditional pose-based action representations in precisely regulating interaction forces during contact-rich manipulation. We propose a torque-control imitation learning policy built upon the Action Chunking with Transformers (ACT) architecture. This method pioneers the use of predicted torques as the sole action output to drive a pure torque controller, and introduces bidirectional force-feedback teleoperation to ensure high-quality demonstration data. Experimental results demonstrate that the proposed policy matches or surpasses position-control baselines across five contact-intensive tasks. Furthermore, we release an open-source dataset comprising over 1,000 high-quality torque demonstrations, establishing a significant foundation for future research in fine-grained robotic force control.
研究通过结合扩散策略与自适应控制器解决建筑中机器人接触丰富操作的挑战,实现零样本模拟到现实迁移,提高组装成功率和稳定性。