Tag
This paper introduces a comprehensive benchmark for evaluating LLMs in key-value extraction from documents under OCR noise, revealing substantial performance degradation and emphasizing the need for joint optimization of OCR quality and LLM reasoning.
This paper presents the results of HIPE-2026, the third edition of the HIPE evaluation series, which focuses on temporally grounded person-place relation extraction from multilingual historical documents in French, German, and English. Seventeen participating teams were evaluated on predictive accuracy, computational efficiency, and cross-domain generalization.