Tag
The paper introduces BIABench, an open benchmark of 16 real-world bioimage analysis tasks reconstructed from published studies, evaluating AI agents end-to-end with outcome and process scores. Agents solved routine 2D tasks well but failed on 3D and time-lapse tasks, with neither specialization, stronger models, nor expert instructions closing the reliability gap.