MintAct: A Unified Visual Agent for Digital Environments

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
MintAct通过统一视觉代理解决跨移动、桌面和网页的多步骤导航问题,使用强化学习方法在不同环境中训练模型。
📝 Abstract
We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance of per-domain specialists across all of these capabilities. To enable this, we develop a scalable environment and reinforcement learning (RL) infrastructure. On the environment side, we host hundreds of concurrent instances across heterogeneous per-domain backends, serving both trajectory data collection and online RL. To enable efficient and scalable RL training, an asynchronous framework keeps explicit control over the cross-domain training distribution and remains stable under noisy environment feedback and off-policy drift. Experimental results show that MintAct achieves state-of-the-art performance (48.9 on OSWorld-Verified) across a wide range of benchmarks at comparable model sizes.
Problem

Research questions and friction points this paper is trying to address.

UI grounding
multi-step navigation
visual tool use
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-language models
reinforcement learning
scalable infrastructure
asynchronous framework
cross-domain performance
🔎 Similar Papers
No similar papers found.