Tag
DharmaOCR, a model specialized for Brazilian Portuguese OCR, outperforms newer models like Mistral OCR4 through domain-specific fine-tuning and Direct Preference Optimization. The article explains the training pipeline and presents benchmark results.
This paper presents SAMPA, a Whisper-based segmenter fine-tuned on Brazilian Portuguese speech data to automatically mark terminal prosodic boundaries, achieving competitive F1 scores on held-out and out-of-distribution datasets.
TOTEN is a knowledge-based ontological tokenization framework that replaces statistical tokenization with declarative classification grounded in a formal ontology of engineering entities, achieving high ontological atomicity and numerical reconstruction for physical quantities and technical notation in Brazilian Portuguese.