@0xCodez: Andrej Karpathy predicted the future of AI once again: “Everyone’s renting frontier models for jobs a 3B model could do…
Summary
The post highlights Andrej Karpathy's argument that small 3B models can replace frontier-model calls, and introduces TwIL-LM3-Pro, a 3.6B-parameter open-source reasoning model from webAI that matches Qwen3-8B on formal logic tasks while running locally on CPU or 4GB VRAM.
View Cached Full Text
Cached at: 10/03/26, 04:51 AM
Andrej Karpathy predicted the future of AI once again:
“Everyone’s renting frontier models for jobs a 3B model could do. Small models are the future.”
this 18-page PDF breaks down Karpathy’s case for working with small LLMs.
the real question isn’t “Is the small model as good?” It’s “Which of my 1,000 calls ever needed a frontier model?”
And @thewebai just answered it for formal logic.
TwIL-LM3-Pro: → 3.6B params, on par with Qwen3-8B on formal logic → leads VibeThinker-3B on all 6 formal-logic tasks tested → 95.4% on BBH logic, 95% on SVAMP → 2.09 GiB in Q4, runs on CPU or 4GB VRAM → no API bill, no data leaving your machine
The secret isn’t size. It’s post-training.
PDF below. Model 👇 http://huggingface.co/webAI-Official/TwIL-LM3-Pro…
David Stout (@Davidstout): Half a million downloads in a month. Today, our open source family takes another step forward.
Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro.
At just 3.6 billion parameters, it brings powerful reasoning to
Similar Articles
@rohanpaul_ai: China's best small reasoning model (VibeThinker-3B) just got beaten by a 3.6B model from Austin. > webAI's TwIL-LM3-Pro…
webAI released TwIL-LM3-Pro, a 3.66B open-source reasoning model that surpasses VibeThinker-3B, Qwen3.5-4B, and LFM2.5-8B-A1B on formal logic benchmarks and runs locally via llama.cpp in a ~2 GiB Q4 build.
@omarsar0: Small models are getting really good at reasoning. It's exciting because SLMs can unlock so much at the harness layer. …
TheWebAI released TwIL-LM3-Pro, a 3.6B-parameter open-source SLM that runs locally and achieves 95.4 on BIG-Bench Hard, outperforming Qwen3-8B's 63.7, using a post-training recipe of formal-logic fine-tuning, weight merging, and RL with a programmatic verifier.
@andrewchen: “don’t bet against AI models improving dramatically over time” used to be something you’d say in support of the frontie…
Andrew Chen argues that open weight AI models are improving rapidly and will cover most consumer/prosumer use cases, leaving frontier models to compete for the remaining high-value 10% of applications like coding, science, and math.
@Davidstout: Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incred…
webAI introduces TwIL-LM3-Pro, a 3.6B-parameter open-source reasoning model that runs locally on everyday computers with quantized builds, claiming top small-model scores on formal logic, SVAMP, MuSR, and BIG-Bench Hard. They also tease Meridian, an upcoming family of frontier-class on-device models.
@Zephyr_hg: UC Berkeley just open-sourced frontier-class AI that runs on a gaming PC. Models keep getting cheaper, faster, freer. A…
UC Berkeley has open-sourced a frontier-class AI model that runs on gaming PCs, making advanced AI more accessible and affordable, with the focus shifting to user setup for real work.