multi-step-evaluation

Tag

Cards List
#multi-step-evaluation

Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

arXiv cs.AI · yesterday Cached

Sci-MMR is a benchmark for evaluating multi-step evidence-grounded scientific reasoning in multimodal agents, revealing gaps where answer accuracy exceeds evidence recovery by over 20%.

0 favorites 0 likes
← Back to home

Submit Feedback