Tag
This paper analyzes how image tokenizer design affects joint text-image modeling in multimodal models using a controlled autoregressive testbed, showing distinct scaling behaviors and correlations with downstream performance.
An 8B reasoning OCR model achieves near-perfect 99% accuracy, instantly extracting text from images and converting complex layouts into clean Markdown, outperforming larger models like GPT-4o on document tasks.
Vik Paruchuri shares that a new model can now read an unreadable Dr. Bronner's soap label, showcasing improved OCR capabilities.