@rohanpaul_ai: China's best small reasoning model (VibeThinker-3B) just got beaten by a 3.6B model from Austin. > webAI's TwIL-LM3-Pro…

X AI KOLs Timeline Models

Summary

webAI released TwIL-LM3-Pro, a 3.66B open-source reasoning model that surpasses VibeThinker-3B, Qwen3.5-4B, and LFM2.5-8B-A1B on formal logic benchmarks and runs locally via llama.cpp in a ~2 GiB Q4 build.

China's best small reasoning model (VibeThinker-3B) just got beaten by a 3.6B model from Austin. > webAI's TwIL-LM3-Pro, a 3.66B local model, lifts IBM Granite's formal-logic score by 28% > the recommended Q4 build is a 2.09GiB file that runs through llama.cpp on CPU or local GPU, so private data can stay on the device. > TwIL-LM3-Pro now leads every small model webAI compared on formal logic, scoring roughly 35% above Weibo's VibeThinker-3B, 24% above Qwen3.5-4B and 47% above Liquid AI's LFM2.5-8B-A1B.
Original Article
View Cached Full Text

Cached at: 10/01/26, 12:22 PM

China’s best small reasoning model (VibeThinker-3B) just got beaten by a 3.6B model from Austin.

webAI’s TwIL-LM3-Pro, a 3.66B local model, lifts IBM Granite’s formal-logic score by 28%

the recommended Q4 build is a 2.09GiB file that runs through llama.cpp on CPU or local GPU, so private data can stay on the device.

TwIL-LM3-Pro now leads every small model webAI compared on formal logic, scoring roughly 35% above Weibo’s VibeThinker-3B, 24% above Qwen3.5-4B and 47% above Liquid AI’s LFM2.5-8B-A1B.

David Stout (@Davidstout): Half a million downloads in a month. Today, our open source family takes another step forward.

Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro.

At just 3.6 billion parameters, it brings powerful reasoning to

Similar Articles