1/100 → 44/100: fine-tuning a 450M VLM on 50K browser screenshots

Reddit r/LocalLLaMA Papers

Summary

The article describes the process of fine-tuning a 450 million parameter vision-language model using 50,000 browser screenshots, showing progress from initial to current stages.

No content available
Original Article

Similar Articles

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ

arXiv cs.CL

This paper investigates test-time scaling techniques for small open vision-language models (≤7B parameters) on the multilingual visual MCQ benchmark EXAMS-V, finding that inference budget and parseability matter more than complex search or verification methods. The best configuration achieves 84.1% on the ImageCLEF 2026 test split, ranking first on the leaderboard.