Tag
The paper introduces KhatianDoc, a human-verified benchmark for diagnosing multimodal LLM failures on Bengali legal land records, revealing that current models fail on tasks like symbol recognition and document QA.
This paper presents an LLM-driven pipeline using GPT-5, GPT-4o, and Claude Sonnet 4 to automatically design neural network architectures for cross-lingual handwritten OCR, achieving over 93% accuracy across Arabic, English, and Persian scripts without human intervention.