@krandiash: Super cool to see we're #2 on this new benchmark out of the box (#1 soon) This makes Ink-2 a great fit for consumer app…
Summary
Ink-2 ranked #2 on the new VoiceCodeBench benchmark, demonstrating its suitability for real-time consumer apps, with GPT Live Transcribe taking the #1 spot.
View Cached Full Text
Cached at: 08/28/26, 03:53 PM
Super cool to see we’re #2 on this new benchmark out of the box (#1 soon)
This makes Ink-2 a great fit for consumer apps that are aimed at everyday interactions (and it runs real-time)
Vals AI (@ValsAI): VoiceCodeBench is now live on Vals, and GPT Live Transcribe takes the #1 spot.
@besimple_ai created VoiceCodeBench, which measures how well speech-to-text models handle structured values in English workplace speech.
Similar Articles
@beffjezos: Finally we have an American alternative to GLM 5.2 Awesome to see American Open Source catch up! Kudos to @thinkymachin…
Thinking Machines releases Inkling, an open-source multimodal AI model that reasons across text, image, and audio modalities, with full weights available for fine-tuning.
The Benchmarks of Thinking Machine's first open-source model Inkling
Thinking Machine has released its first open-source model, Inkling, alongside benchmark results demonstrating its performance.
@elonmusk: Grok Voice
Grok Voice Think Fast 2.0 reportedly surpasses GPT Realtime in both text and voice accuracy on VulcanBench, achieving 99% accuracy in text tasks.
@aaron_epstein: New model just released that beats sonnet 4.6, gemini 3 flash, and gpt 5.4 mini on OCR, vision, and STT tasks @interfaz…
A new AI model from interfaze_ai claims to outperform leading models (sonnet 4.6, gemini 3 flash, gpt 5.4 mini) on OCR, vision, and speech-to-text tasks.
@Lyubh22: Coding benchmarks are saturating. AI4Research is the next frontier. Thrilled to see our MLS-Bench (https://mls-bench.co…
Announcing MLS-Bench, the first AI4Research benchmark to gain broad community adoption, testing AI agents on 140 executable tasks across 12 domains to propose modular ML improvements. The post includes leaderboard scores for models like Claude Opus 4.6 and GPT-5.4.