@omarsar0: Small models are getting really good at reasoning. It's exciting because SLMs can unlock so much at the harness layer. …
Summary
TheWebAI released TwIL-LM3-Pro, a 3.6B-parameter open-source SLM that runs locally and achieves 95.4 on BIG-Bench Hard, outperforming Qwen3-8B's 63.7, using a post-training recipe of formal-logic fine-tuning, weight merging, and RL with a programmatic verifier.
View Cached Full Text
Cached at: 10/02/26, 06:36 AM
Small models are getting really good at reasoning.
It’s exciting because SLMs can unlock so much at the harness layer.
TwIL-LM3-Pro from @thewebAI has 3.6B parameters and runs locally on everyday computers. It scores 95.4 on BIG-Bench Hard, well ahead of Qwen3-8B at 63.7.
I like their post-training recipe. They fine-tune on formal logic, merge the weights back toward the base model, and then run RL against a programmatic verifier. Logic scores go up, and general reasoning holds steady.
Great to see more of this work released as open source.
David Stout (@Davidstout): Half a million downloads in a month. Today, our open source family takes another step forward.
Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro.
At just 3.6 billion parameters, it brings powerful reasoning to
Similar Articles
@svpino: This open-source model scores 95.4% on BIG-Bench Hard’s logic subset with just 3.66B parameters. TwIL-LM3-Pro is a mode…
TwIL-LM3-Pro, an open-source 3.66B-parameter model, scores 95.4% on BIG-Bench Hard's logic subset and is small enough to run locally on a laptop, showcasing the power of specialized post-training for task-focused small models.
@rohanpaul_ai: China's best small reasoning model (VibeThinker-3B) just got beaten by a 3.6B model from Austin. > webAI's TwIL-LM3-Pro…
webAI released TwIL-LM3-Pro, a 3.66B open-source reasoning model that surpasses VibeThinker-3B, Qwen3.5-4B, and LFM2.5-8B-A1B on formal logic benchmarks and runs locally via llama.cpp in a ~2 GiB Q4 build.
@svpino: Tiny, specialized, and open models are the future! The TwiL-LM family of models is now available on HuggingFace for Tra…
TwiL-LM, a family of tiny specialized open models, is released on HuggingFace. The 3B version outperforms OpenAI's 120B gpt-oss on formal reasoning benchmarks and runs efficiently on consumer hardware.
@Davidstout: Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incred…
webAI introduces TwIL-LM3-Pro, a 3.6B-parameter open-source reasoning model that runs locally on everyday computers with quantized builds, claiming top small-model scores on formal logic, SVAMP, MuSR, and BIG-Bench Hard. They also tease Meridian, an upcoming family of frontier-class on-device models.
@0xCodez: Andrej Karpathy predicted the future of AI once again: “Everyone’s renting frontier models for jobs a 3B model could do…
The post highlights Andrej Karpathy's argument that small 3B models can replace frontier-model calls, and introduces TwIL-LM3-Pro, a 3.6B-parameter open-source reasoning model from webAI that matches Qwen3-8B on formal logic tasks while running locally on CPU or 4GB VRAM.