1/100 → 44/100: fine-tuning a 450M VLM on 50K browser screenshots
Summary
The article describes the process of fine-tuning a 450 million parameter vision-language model using 50,000 browser screenshots, showing progress from initial to current stages.
Similar Articles
@h100envy: Google engineer explained how to fine-tune a tiny LLM from 46% to 90% accuracy on your phone in 21 minutes - better tha…
A Google engineer shares a method to fine-tune a Gemma 270M model from 46% to 90% accuracy in 21 minutes on a phone, using synthetic data, LoRA, int4 quantization, achieving 2000 tokens per second offline.
@maximelabonne: Neat app to understand and explore VLM evals
A shared app for exploring and understanding vision language model evaluations, referencing a thread analyzing popular vision benchmarks.
A 460M VLM gets first-token latency down to 0.3s on an iPhone by using only 64 visual tokens
VisionPsy-Nano-460M-Flash is a new 460M vision-language model that uses only 64 visual tokens per image, cutting first-token latency to 0.3s on iPhone while retaining roughly 99% of the full model's benchmark score. The optimization trades off some OCR and fine-detail quality.
I pretrained and post trained a 500M parameter LLM and 330M parameter Image generator from scratch
The author details the process of pretraining and post-training a 500M parameter language model and a 330M parameter image generator entirely from scratch.
Test-Time Scaling for Small VLMs on Multilingual Visual MCQ
This paper investigates test-time scaling techniques for small open vision-language models (≤7B parameters) on the multilingual visual MCQ benchmark EXAMS-V, finding that inference budget and parseability matter more than complex search or verification methods. The best configuration achieves 84.1% on the ImageCLEF 2026 test split, ranking first on the leaderboard.