@rohanpaul_ai: China's best small reasoning model (VibeThinker-3B) just got beaten by a 3.6B model from Austin. > webAI's TwIL-LM3-Pro…
Summary
webAI released TwIL-LM3-Pro, a 3.66B open-source reasoning model that surpasses VibeThinker-3B, Qwen3.5-4B, and LFM2.5-8B-A1B on formal logic benchmarks and runs locally via llama.cpp in a ~2 GiB Q4 build.
View Cached Full Text
Cached at: 10/01/26, 12:22 PM
China’s best small reasoning model (VibeThinker-3B) just got beaten by a 3.6B model from Austin.
webAI’s TwIL-LM3-Pro, a 3.66B local model, lifts IBM Granite’s formal-logic score by 28%
the recommended Q4 build is a 2.09GiB file that runs through llama.cpp on CPU or local GPU, so private data can stay on the device.
TwIL-LM3-Pro now leads every small model webAI compared on formal logic, scoring roughly 35% above Weibo’s VibeThinker-3B, 24% above Qwen3.5-4B and 47% above Liquid AI’s LFM2.5-8B-A1B.
David Stout (@Davidstout): Half a million downloads in a month. Today, our open source family takes another step forward.
Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro.
At just 3.6 billion parameters, it brings powerful reasoning to
Similar Articles
@aijoey: WeiboAI dropped VibeThinker-3B, so I had to try it locally. this is a 3B model, not a giant frontier system. in the vid…
WeiboAI released VibeThinker-3B, a small 3B reasoning model tested locally on coding tasks, achieving 3/3 on algorithm problems.
@0xCodez: Andrej Karpathy predicted the future of AI once again: “Everyone’s renting frontier models for jobs a 3B model could do…
The post highlights Andrej Karpathy's argument that small 3B models can replace frontier-model calls, and introduces TwIL-LM3-Pro, a 3.6B-parameter open-source reasoning model from webAI that matches Qwen3-8B on formal logic tasks while running locally on CPU or 4GB VRAM.
@omarsar0: Small models are getting really good at reasoning. It's exciting because SLMs can unlock so much at the harness layer. …
TheWebAI released TwIL-LM3-Pro, a 3.6B-parameter open-source SLM that runs locally and achieves 95.4 on BIG-Bench Hard, outperforming Qwen3-8B's 63.7, using a post-training recipe of formal-logic fine-tuning, weight merging, and RL with a programmatic verifier.
Why Weibo's tiny VibeThinker-3B has the AI world arguing over benchmarks again (15 minute read)
Weibo's VibeThinker-3B, a 3B parameter model, claims to match or exceed the reasoning performance of much larger models like DeepSeek V3.2 and Gemini 3 Pro on math and coding benchmarks, sparking debate over benchmark reliability and the necessity of scaling.
@Davidstout: Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incred…
webAI introduces TwIL-LM3-Pro, a 3.6B-parameter open-source reasoning model that runs locally on everyday computers with quantized builds, claiming top small-model scores on formal logic, SVAMP, MuSR, and BIG-Bench Hard. They also tease Meridian, an upcoming family of frontier-class on-device models.