BIABench: Evaluating AI agents on real-world bioimage analysis tasks
Summary
The paper introduces BIABench, an open benchmark of 16 real-world bioimage analysis tasks reconstructed from published studies, evaluating AI agents end-to-end with outcome and process scores. Agents solved routine 2D tasks well but failed on 3D and time-lapse tasks, with neither specialization, stronger models, nor expert instructions closing the reliability gap.
View Cached Full Text
Cached at: 10/01/26, 08:24 PM
Paper page - BIABench: Evaluating AI agents on real-world bioimage analysis tasks
Source: https://huggingface.co/papers/2609.34274
Abstract
Artificial-intelligence(AI)agentsholdpromiseforautomatingbioimageanalysis,yetnobenchmarkevaluateswhethertheycancarryoutreal-worldanalysesendtoend.Suchanalysesarehardforagentsbecause2Dimages,3Dvolumesandtime-lapsesequencesareoftentoolargetoreadascontext,soanagentmustchooseandrunananalysisthroughcode,specializedsoftwareandrenderedviews.Publishedstudiesmakethiscapabilitytestable,becauseeachpairsrawimageswithapeer-reviewedresult.WeintroduceBIABench,abenchmarkof16tasksreconstructedfrompublishedbiologicalstudiesthatretaintheirscientificquestions,imagingdataandgroundtruth.ThetasksspanelevenanalysissubtasksandmodalitiesfromH&Ehistologytosingle-moleculelocalizationmicroscopy.Eachsubmissionreceivesanoutcomescore,whichcomparestheoutputfileswiththegroundtruthusingfield-standardmetrics,andaprocessscore,inwhichavision-languagemodeljudgesmethodchoiceandqualitycontrolagainstanexpert-writtenrubric.Weevaluatedgeneral-purposeandbiology-specificagentsacrossseverallanguagemodels,withrepeatedrunsofeverytask.Routinetwo-dimensionaltasksweresolvedwell,butonsometasksthataddedathirddimensionoratimeaxisnoagentscoredabove0.19.Neitherbiologicalspecialization,strongermodelsnordetailedexpertinstructionsclosedthisgap.Theagentswerealsounreliable,withscoresvaryingmorebetweenrepeatedrunsofoneagentthanbetweendifferentagents,andwithoutgroundtruthacorrectruncouldnotbetoldfromawrongonebyitsprocessscoreorbythetimespent.Releasedopenlywithitsdataandcode,BIABenchprovidesaverifiableframeworkforevaluating,andeventuallytraining,agentsforreliablelong-horizonbioimageanalysis.
View arXiv pageView PDFProject pageGitHub6Add to collection
Get this paper in your agent:
hf papers read 2609\.34274
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.34274 in a model README.md to link it from this page.
Datasets citing this paper1
#### BIABench/BIABench Updated3 days ago • 871 • 2
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.34274 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
BixBench3: Frontier AI Agents Can Now Reproduce ~48% of Real Computational Biology Research Workflows
This paper introduces BixBench3, a benchmark for evaluating AI agents on computational biology tasks, revealing that frontier LLMs can reproduce approximately 48% of real research workflows but struggle with large datasets and sequential steps.
BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence
This paper introduces BI-Bench, the first benchmark for evaluating LLMs on end-to-end business intelligence tasks, and BI-Agent, a tool-augmented agent that decomposes workflows and uses post-training to improve accuracy significantly.
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
This paper introduces HealthAgentBench, a suite of 54 realistic healthcare tasks for evaluating frontier AI agents. It finds that even the best agent (Codex GPT-5.5) achieves only ~42% success, highlighting substantial room for improvement.
AutoMedBench: Towards Medical AutoResearch with Agentic AI Models
AutoMedBench is a workflow-aware benchmark for autonomous medical-AI research, evaluating agents across five stages on diverse medical imaging tasks. Stage-level scoring reveals validation as the weakest stage, highlighting the need for reliable verification in agentic workflows.
K-Bench: measuring model performance on real scientific agent requests
The paper introduces K-Bench 01, a benchmark for evaluating AI agents on real scientific requests, revealing that no model consistently meets the threshold for acceptable performance, with overclaiming as a common failure.