Tag
This paper presents SynthAVE, a large-scale human-validated benchmark for attribute value extraction in e-commerce, using a multi-LLM arena framework with 21 judge configurations to validate synthetic labels efficiently and cost-effectively while maintaining quality parity with human review.
TeachObs introduces a human-validated benchmark for multimodal teaching observation, consisting of 30 classroom videos annotated with segment-level binary codes and lesson-level expert ratings, and evaluates five frontier LLMs across three tracks, finding no single model consistently outperforms and that model evaluations overrate procedurally clear lessons.