Tag
Introduces DocAnnot, a framework that uses a large vision-language model, OCR, and a spatially informed contextual matching algorithm to automatically generate training datasets for key information extraction from documents, reducing manual annotation effort. Evaluated on CORD and SROIE benchmarks, it achieves reasonable F1-scores and enables efficient human verification.
Introduces BaFCo, a benchmark dataset for Bangla form comprehension focusing on Document Layout Analysis (DLA) and Key Information Extraction (KIE). It includes 200 multi-page complex Bangladeshi government forms with fine-grained annotations across 26 entity types and evaluates multiple MLLMs, revealing limitations in understanding complex Bangla forms.