Tag
The Last Translation Benchmark introduces a live dataset of peer-reviewed, multimodal examples designed to evaluate and break leading machine translation models, with handcrafted verification rules for reliable assessment. It addresses the saturation of current benchmarks and the unreliability of automatic metrics.
A new peer-reviewed study highlights a critical lack of legal and ethical guidance for using AI in citizen science, particularly regarding transparency about training data, and offers recommendations for addressing this gap.