@LiTianleli: Incredibly proud of the team. After countless late nights, Inkling is out, and I especially want to highlight the post-…

X AI KOLs Timeline Models

Summary

Thinking Machines releases Inkling, an open-source multi-modal reasoning model with innovations in post-training RL, achieving stable scaling to 30M+ rollouts and controllable thinking effort. The model exhibits compressed reasoning and will be available soon for fine-tuning.

Incredibly proud of the team. After countless late nights, Inkling is out, and I especially want to highlight the post-training stack and RL recipes behind it. A few of my favorite details: We scaled our largest RL run to 30M+ rollouts and thousands of continuous training steps—with no collapse, no restarts, and stable KL and entropy throughout. Reasoning performance improved log-linearly from the SFT initialization all the way to the released checkpoint. A number of innovations under the hood made this possible, and the result is a strong testament to our post-training technology. We trained controllable thinking effort directly through RL. By varying the system message and per-token cost, the model learned to trade off tokens and performance on demand. We also saw an emergent shift in reasoning style: as RL progressed, the chain of thought became increasingly compressed, shedding grammatical overhead. Inkling reasons like a caveman mathematician—a distinctive style unlike that of other open-source models. We’re also previewing Inkling-small today and plan to release it very soon. It is exceptionally capable for its size, and we expect the community will find it broadly useful. Building a simple, stable, and scalable RL stack in such a short time was something few thought possible. This team proved otherwise.
Original Article
View Cached Full Text

Cached at: 07/16/26, 04:21 PM

Incredibly proud of the team. After countless late nights, Inkling is out, and I especially want to highlight the post-training stack and RL recipes behind it.

A few of my favorite details:

We scaled our largest RL run to 30M+ rollouts and thousands of continuous training steps—with no collapse, no restarts, and stable KL and entropy throughout. Reasoning performance improved log-linearly from the SFT initialization all the way to the released checkpoint. A number of innovations under the hood made this possible, and the result is a strong testament to our post-training technology.

We trained controllable thinking effort directly through RL. By varying the system message and per-token cost, the model learned to trade off tokens and performance on demand.

We also saw an emergent shift in reasoning style: as RL progressed, the chain of thought became increasingly compressed, shedding grammatical overhead. Inkling reasons like a caveman mathematician—a distinctive style unlike that of other open-source models.

We’re also previewing Inkling-small today and plan to release it very soon. It is exceptionally capable for its size, and we expect the community will find it broadly useful.

Building a simple, stable, and scalable RL stack in such a short time was something few thought possible. This team proved otherwise.

Frrrr

Thank you Clare!

Miss you

Similar Articles

Thinking Machines Lab Drops Its First Model

Wired

Thinking Machines Lab, founded by ex-OpenAI executives, releases its first open-weight AI model, Inkling, a 975-billion-parameter model capable of reasoning, coding, and processing audio, video, and text.