self-driving-labs

Tag

Cards List
#self-driving-labs

Agentic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasks

arXiv cs.AI · 2d ago Cached

This paper presents a benchmark and trace-logging framework for evaluating LLM-based agents that control microscopes, comparing 105 agent configurations and finding that benchmarks support qualification but do not reliably predict performance on unseen tasks.

0 favorites 0 likes
← Back to home

Submit Feedback