@MaziyarPanahi: I played Thinking Machines' new Inkling a 2-minute doctor's visit, on my Mac Studio The patient came in about her knee.…

X AI KOLs Timeline Models

Summary

Thinking Machines' Inkling, a 975B parameter model running locally on a Mac Studio via llama.cpp, listened to a doctor's visit audio and accurately diagnosed heart failure from subtle cues in small talk, demonstrating advanced medical reasoning without leaving the machine.

I played Thinking Machines' new Inkling a 2-minute doctor's visit, on my Mac Studio The patient came in about her knee. Inkling listened to the whole thing and flagged the heart failure instead. It heard the real audio, not a transcript. 975B params, the 1-bit GGUF (@UnslothAI), running on llama.cpp (@huggingface). Nothing left the machine. The knee was the easy part. Buried in the small talk: out of puff on the hill, sleeping on four pillows, ankles like balloons, and she'd quietly stopped her water tablet in spring. Five signals, one diagnosis, none of it why she booked. Her last line: "Should I have said something? It's only the knee I came about." A 975B model just did the hardest thing in medicine: it listened to what the patient didn't think was worth mentioning. What should it hear next?
Original Article
View Cached Full Text

Cached at: 07/17/26, 12:24 AM

I played Thinking Machines’ new Inkling a 2-minute doctor’s visit, on my Mac Studio

The patient came in about her knee. Inkling listened to the whole thing and flagged the heart failure instead.

It heard the real audio, not a transcript. 975B params, the 1-bit GGUF (@UnslothAI), running on llama.cpp (@huggingface). Nothing left the machine.

The knee was the easy part. Buried in the small talk: out of puff on the hill, sleeping on four pillows, ankles like balloons, and she’d quietly stopped her water tablet in spring. Five signals, one diagnosis, none of it why she booked.

Her last line: “Should I have said something? It’s only the knee I came about.”

A 975B model just did the hardest thing in medicine: it listened to what the patient didn’t think was worth mentioning.

What should it hear next?

Similar Articles

Welcome Inkling by Thinking Machines

Hugging Face Blog

Inkling by Thinking Machines is a large open multimodal LLM with ~1T parameters, 1M context, and native support for image, audio, and text. It uses a Mixture-of-Experts architecture and is available on Hugging Face with day-0 inference support.

thinkingmachines/Inkling

Hugging Face Models Trending

Inkling is a large open-weights multimodal model (975B total, 41B active parameters) using a sparse MoE architecture, accepting text, image, and audio inputs and generating text outputs, intended for agentic systems, coding assistants, and chatbots.

Thinking Machines Lab Drops Its First Model

Wired

Thinking Machines Lab, founded by ex-OpenAI executives, releases its first open-weight AI model, Inkling, a 975-billion-parameter model capable of reasoning, coding, and processing audio, video, and text.