multimodal-audit

Tag

Cards List
#multimodal-audit

Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

arXiv cs.CL · 2d ago Cached

This study evaluates two multimodal LLMs as peer reviewers for ICLR 2026 submissions, finding high scoring calibration but low error detection, with author identity having no effect and figures reducing error detection.

0 favorites 0 likes
← Back to home

Submit Feedback