@omarsar0: Small models are getting really good at reasoning. It's exciting because SLMs can unlock so much at the harness layer. …

X AI KOLs Timeline Models

Summary

TheWebAI released TwIL-LM3-Pro, a 3.6B-parameter open-source SLM that runs locally and achieves 95.4 on BIG-Bench Hard, outperforming Qwen3-8B's 63.7, using a post-training recipe of formal-logic fine-tuning, weight merging, and RL with a programmatic verifier.

Small models are getting really good at reasoning. It's exciting because SLMs can unlock so much at the harness layer. TwIL-LM3-Pro from @thewebAI has 3.6B parameters and runs locally on everyday computers. It scores 95.4 on BIG-Bench Hard, well ahead of Qwen3-8B at 63.7. I like their post-training recipe. They fine-tune on formal logic, merge the weights back toward the base model, and then run RL against a programmatic verifier. Logic scores go up, and general reasoning holds steady. Great to see more of this work released as open source.
Original Article
View Cached Full Text

Cached at: 10/02/26, 06:36 AM

Small models are getting really good at reasoning.

It’s exciting because SLMs can unlock so much at the harness layer.

TwIL-LM3-Pro from @thewebAI has 3.6B parameters and runs locally on everyday computers. It scores 95.4 on BIG-Bench Hard, well ahead of Qwen3-8B at 63.7.

I like their post-training recipe. They fine-tune on formal logic, merge the weights back toward the base model, and then run RL against a programmatic verifier. Logic scores go up, and general reasoning holds steady.

Great to see more of this work released as open source.

David Stout (@Davidstout): Half a million downloads in a month. Today, our open source family takes another step forward.

Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro.

At just 3.6 billion parameters, it brings powerful reasoning to

Similar Articles