@Davidstout: Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incred…
Summary
webAI introduces TwIL-LM3-Pro, a 3.6B-parameter open-source reasoning model that runs locally on everyday computers with quantized builds, claiming top small-model scores on formal logic, SVAMP, MuSR, and BIG-Bench Hard. They also tease Meridian, an upcoming family of frontier-class on-device models.
View Cached Full Text
Cached at: 10/01/26, 08:20 AM
Half a million downloads in a month. Today, our open source family takes another step forward.
Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro.
At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required.
In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%.
Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared.
BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%.
We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own.
That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab.
Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application.
Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸
Similar Articles
@rohanpaul_ai: China's best small reasoning model (VibeThinker-3B) just got beaten by a 3.6B model from Austin. > webAI's TwIL-LM3-Pro…
webAI released TwIL-LM3-Pro, a 3.66B open-source reasoning model that surpasses VibeThinker-3B, Qwen3.5-4B, and LFM2.5-8B-A1B on formal logic benchmarks and runs locally via llama.cpp in a ~2 GiB Q4 build.
@omarsar0: Small models are getting really good at reasoning. It's exciting because SLMs can unlock so much at the harness layer. …
TheWebAI released TwIL-LM3-Pro, a 3.6B-parameter open-source SLM that runs locally and achieves 95.4 on BIG-Bench Hard, outperforming Qwen3-8B's 63.7, using a post-training recipe of formal-logic fine-tuning, weight merging, and RL with a programmatic verifier.
@0xCodez: Andrej Karpathy predicted the future of AI once again: “Everyone’s renting frontier models for jobs a 3B model could do…
The post highlights Andrej Karpathy's argument that small 3B models can replace frontier-model calls, and introduces TwIL-LM3-Pro, a 3.6B-parameter open-source reasoning model from webAI that matches Qwen3-8B on formal logic tasks while running locally on CPU or 4GB VRAM.
webAI released a formal reasoning model family, TwIL, that's worth a look if you're doing verification pipelines
webAI has released the TwIL family of formal reasoning models, with TwIL-LM3 outperforming a 120B model on specific benchmarks while being 40x smaller and more efficient.
@svpino: Tiny, specialized, and open models are the future! The TwiL-LM family of models is now available on HuggingFace for Tra…
TwiL-LM, a family of tiny specialized open models, is released on HuggingFace. The 3B version outperforms OpenAI's 120B gpt-oss on formal reasoning benchmarks and runs efficiently on consumer hardware.