Tag
This paper proposes an end-to-end sequence-to-sequence approach for contextual Tamil spelling and grammar correction, using progressively fine-tuned mT5 and mBART models on synthetic data, achieving 69.3% exact-match accuracy on a diagnostic set.
The paper presents an auditable reliability layer for biomedical text classification that uses deterministic spell-correction to address OCR artifacts, improving classifier performance while ensuring safety through abstention under uncertainty.