CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs
Summary
CaptchaArena is a large-scale, fine-grained dataset for training computer-use agents to solve interactive CAPTCHAs, containing 50K puzzles and annotations. The associated CaptchaAgent model achieves 71.7% accuracy using supervised and reinforcement learning.
View Cached Full Text
Cached at: 09/30/26, 04:15 AM
Paper page - CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs
Source: https://huggingface.co/papers/2609.31957
Abstract
InteractiveCAPTCHAsremainchallengingforcomputer-useagents,whileexistingdatasetsfacetrade-offsamongtypecoverage,interactionfidelity,andtrajectorysupervision.Toaddressthesegaps,wepresentCaptchaArena,thefirstlarge-scale,fine-grainedtrainingdatasetforinteractiveCAPTCHAsolving.Itcontains50Kpuzzlesacross20CAPTCHAtypesand5interactionmodes,witheverysolutionverifiedthroughexecution.CaptchaArenaprovides50Kscreenshot-actiontrajectories,including46Kwithstep-by-stepreasoningannotations.Italsoincludesfine-grainedpixel-maskannotationsforirregulartargets.UsingCaptchaArena,wetrainCaptchaAgent,asingle9Bpolicyforall20CAPTCHAtypes,withsupervisedfine-tuningfollowedbyreinforcementlearning.TheenvironmentverifierdirectlyprovidestheRLreward.Supervisedfine-tuningreaches70.5Pass@1,andreinforcementlearningfurtherimprovesitto71.7,whilealsoimprovingperformanceontwoexternalbenchmarks.Theseresultsdemonstratethevalueoflarge-scale,fine-grainedcomputer-usesupervisionfortraininginteractiveCAPTCHAagents.WereleaseCaptchaArenaandCaptchaAgentathttps://github.com/X0X0X00/CaptchaArena.
View arXiv pageView PDFGitHub7Add to collection
Models citing this paper1
#### ZHEN-04/CaptchaAgent Image-Text-to-Text• Updated23 minutes ago • 1
Datasets citing this paper2
#### ZHEN-04/CaptchaArena Viewer• Updatedabout 8 hours ago • 89.6k • 24 • 3 #### ZHEN-04/CaptchaArena-Trajectories Updatedabout 8 hours ago • 14 • 1
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.31957 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
CAPTCHAs have failed for 20 years
A deep dive into the 20-year arms race between CAPTCHAs and automated solvers, culminating in Browserbase's new approach of agent identity to bypass CAPTCHAs entirely by verifying browser identity.
@rohanpaul_ai: Today’s AI agents still struggle to pass real human-verification checks (CAPTCHAs) on websites. The paper proposes HLL,…
A new paper introduces HLL, a benchmark for testing AI agents on real CAPTCHA tasks, showing that even strong agents fail on cluttered pages and recovery from mistakes.
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
CUA-Gym introduces a scalable pipeline for generating verifiable training environments and tasks for computer-use agents, addressing data scarcity. The resulting dataset and models achieve strong performance on benchmarks like OSWorld-Verified and WebArena.
CAPTCHAs can still detect AI agents
A research paper shows that while AI can solve CAPTCHAs as well as humans, behavioral differences in interaction patterns can still reliably distinguish bots from people, leading to the proposal of a 'Process Turing Test'.
Prove you are a robot: CAPTCHAs for agents
Browser Use launched agent-native signup using reverse-CAPTCHAs that are designed to keep humans out and let AI agents in. Agents solve obfuscated math problems to gain API key access and free tier benefits.