crosslingual-evaluation

Tag

Cards List
#crosslingual-evaluation

Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation

arXiv cs.CL · 2d ago Cached

This paper investigates the fairness of crosslingual evaluation methods for language models, showing that common normalized metrics can be biased due to tokenization and orthographic differences, and proposes using sentence-level negative log likelihood on semantically equivalent sequences for more consistent crosslingual comparisons.

0 favorites 0 likes
← Back to home

Submit Feedback