CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-offs among coverage, interaction fidelity, and trajectory supervision in existing datasets, which constrain the ability of computer agents to solve interactive CAPTCHAs. To this end, we construct the first large-scale, fine-grained CAPTCHA dataset comprising 50,000 puzzles with pixel-level annotations, providing complete execution trajectories and step-by-step reasoning labels. Furthermore, we propose a method combining supervised fine-tuning with reinforcement learning, leveraging an environment verifier to directly supply reward signals for processing pixel masks and screenshot-based action trajectories. Experimental results demonstrate that a single policy model can generalize across 20 CAPTCHA categories, achieving a Pass@1 accuracy of 71.7% and significantly improving performance on external benchmarks. These findings validate the effectiveness of large-scale, fine-grained supervision in training agents capable of solving interactive CAPTCHAs.
📝 Abstract
Interactive CAPTCHAs remain challenging for computer-use agents, while existing datasets face trade-offs among type coverage, interaction fidelity, and trajectory supervision. To address these gaps, we present CaptchaArena, the first large-scale, fine-grained training dataset for interactive CAPTCHA solving. It contains 50K puzzles across 20 CAPTCHA types and 5 interaction modes, with every solution verified through execution. CaptchaArena provides 50K screenshot-action trajectories, including 46K with step-by-step reasoning annotations. It also includes fine-grained pixel-mask annotations for irregular targets. Using CaptchaArena, we train CaptchaAgent, a single 9B policy for all 20 CAPTCHA types, with supervised fine-tuning followed by reinforcement learning. The environment verifier directly provides the RL reward. Supervised fine-tuning reaches 70.5 Pass@1, and reinforcement learning further improves it to 71.7, while also improving performance on two external benchmarks. These results demonstrate the value of large-scale, fine-grained computer-use supervision for training interactive CAPTCHA agents. We release CaptchaArena and CaptchaAgent at https://github.com/X0X0X00/CaptchaArena.
Problem

Research questions and friction points this paper is trying to address.

Interactive CAPTCHAs
Computer-Use Agents
Dataset
Trajectory Supervision
Interaction Fidelity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interactive CAPTCHA
Computer-Use Agents
Trajectory Supervision
Reinforcement Learning
Fine-Grained Dataset
🔎 Similar Papers
No similar papers found.
Z
Zhenhao Zhang
Columbia University
Z
Zhaoyu Fan
Zhejiang University
H
Haohan Ying
University of Rochester
Jingwen Hu
Jingwen Hu
University of Rochester
H
Hancen Fan
Columbia University
J
Junhao Zhou
University of Illinois at Urbana-Champaign
Zitian Chen
Zitian Chen
University of Massachusetts Amherst
Computer vision
L
Linchao Zhu
Zhejiang University