sota

Tag

Cards List
#sota

@hwchase17: Detecting issues in production agent traces is hard. You have to do it cheaply (because of volume) but also accurately …

X AI KOLs Following ↗ · 2026-06-15

Harrison Chase announces a post-trained model for detecting issues in production agent traces, claiming SOTA accuracy at 10-100x cheaper rates than frontier models.

0 favorites 0 likes
#sota

@ChengleiSi: Excited to share these preliminary results on our internal autoresearch system @Recursive_SI, where we achieve SOTA on …

X AI KOLs Following ↗ · 2026-06-11 Cached

Recursive's automated AI research system achieves state-of-the-art results on NanoChat, NanoGPT Speedrun, and GPU kernel benchmarks by automating the research loop without task-specific adaptations, and open-sourcing artifacts for further inspection.

0 favorites 0 likes
#sota

@gregpr07: Browser Use Beta just achieved SOTA on our hardest internal web agent benchmark. Fable is genuinely amazing for optimiz…

X AI KOLs Following ↗ · 2026-06-11 Cached

Browser Use Beta achieved state-of-the-art results on a difficult internal web agent benchmark, using Fable for optimization and analysis.

0 favorites 0 likes
#sota

@patpcj: Fable-5/Mythos dropped this morning, so we tested it on agentic search - and it’s the new SOTA. The performance gap is …

X AI KOLs Timeline ↗ · 2026-06-10 Cached

Fable-5/Mythos achieves new SOTA on agentic search but is expensive for self-hosting, while open-weight Harness-1 offers a cost-effective alternative with fewer query restrictions.

0 favorites 0 likes
#sota

@adamdotnew: We’re very excited to announce that Anthropic’s Fable 5 is a SOTA model at mechanical engineering tasks! It can generat…

X AI KOLs Following ↗ · 2026-06-09 Cached

Anthropic announces Fable 5, a state-of-the-art model for mechanical engineering tasks capable of generating intricate assemblies and mechanisms from a single prompt.

0 favorites 0 likes
#sota

@perdactor: 1/ Meet Argus-Retriever: the first late-interaction visual doc retriever where the document representation adapts to th…

X AI KOLs Following ↗ · 2026-06-06 Cached

Argus-Retriever is a new late-interaction visual document retriever that adapts document representation to the query, achieving SOTA performance on ViDoRe benchmarks with a smaller index.

0 favorites 0 likes
#sota

@Miles_Brundage: BREAKING: massively improved SOTA score on Clear AVERI Pronunciation Guide Bench, via my colleague Carly

X AI KOLs Following ↗ · 2026-06-04 Cached

Miles Brundage announces a state-of-the-art (SOTA) score improvement on the Clear AVERI Pronunciation Guide Bench achieved by colleague Carly.

0 favorites 0 likes
#sota

@yvbbrjdr: I recommend everyone to read the MAI-Thinking-1 technical paper. It contains detailed (almost all) information on how to train a SOTA LLM. https://microsoft.ai/wp-content/uploads/2026/06/ma…

X AI KOLs Timeline ↗ · 2026-06-02 Cached

Recommended reading: the MAI-Thinking-1 technical paper, which details almost all the steps to train a SOTA large language model.

0 favorites 0 likes
#sota

@swyx: roundup of links:

X AI KOLs Following ↗ · 2026-06-02 Cached

NVIDIA releases Cosmos 3 (Mixture-of-Transformers models up to 64B), Nemotron 3 Ultra (550B-A55B LLM), and previews RTX Spark personal superchip at Computex 2026, achieving SOTA on multiple open model leaderboards.

0 favorites 0 likes
#sota

@nick_kango: One more task to add to my twitter benchmark collection:) Btw, Opus 4.8 and all the SOTA models passed when i tried tha…

X AI KOLs Timeline ↗ · 2026-05-30 Cached

Nick Kang adds a new task to his Twitter benchmark collection; Claude Opus 4.8 and other SOTA models pass, while Sonnet 4.6 and Grok 4.3 fail. Alfin remarks on Opus 4.8's dangerous capabilities.

0 favorites 0 likes
#sota

@kushalbyatnal: Over 1 billion PDFs are created every day, but your agents still can’t read them reliably. Today we’re releasing Parse …

X AI KOLs Following ↗ · 2026-05-26 Cached

Extend released Parse 2.0, a state-of-the-art document parsing API that achieves top accuracy on real-world documents, outperforming competitors on the open-source RealDoc-Bench benchmark.

0 favorites 0 likes
#sota

@victormustar: New: LongCat just dropped an excellent open-source talking-avatar model (probably SOTA) + MIT licensed Made a Hugging F…

X AI KOLs Following ↗ · 2026-05-24 Cached

LongCat released an open-source talking-avatar model (likely state-of-the-art) under MIT license, with a Hugging Face demo, enabling various applications like AI tutors, dubbing, and coding agents.

0 favorites 0 likes
#sota

New SOTA 1B model? HRM-text

Reddit r/LocalLLaMA ↗ · 2026-05-19 Cached

HRM-text is a 1B-parameter hierarchical reasoning language model proposed by Sapient Intelligence. It thinks efficiently through internal latent space, achieving performance surpassing most models of the same size with extremely low training cost.

0 favorites 0 likes
#sota

Ring-2.6-1T is putting up SOTA-level numbers for real-world agents

Reddit r/ArtificialInteligence ↗ · 2026-05-18

Ant Group released Ring-2.6-1T, a 1 trillion parameter reasoning model for agent workflows, featuring MIT license, extended context, and Async RL + IcePop training, achieving state-of-the-art results.

0 favorites 0 likes
#sota

@NielsRogge: Introducing a revival of PapersWithCode! As @ilyasut said, we're back to the "age of research". Hence, it's important t…

X AI KOLs Following ↗ · 2026-05-18 Cached

NielsRogge announces a revival of PapersWithCode, featuring SOTA per domain, leaderboards, and methods parsed at scale using AI agents.

0 favorites 0 likes
#sota

New SOTA: Poetiq uses self-optimizing harness to surpass e.g. Opus 4.7 with Gemini 3 Flash

Reddit r/singularity ↗ · 2026-05-15

Poetiq claims new state-of-the-art coding performance using a self-optimizing harness with Gemini 3 Flash, surpassing Opus 4.7.

0 favorites 0 likes
#sota

@poetiq_ai: Poetiq's Meta-System built its own coding harness from scratch. It got SOTA on LiveCodeBench Pro. No fine-tuning, no sp…

X AI KOLs Following ↗ · 2026-05-14 Cached

Poetiq's Meta-System achieved state-of-the-art results on LiveCodeBench Pro by autonomously building a coding harness using standard APIs and Gemini 3.1 Pro, without fine-tuning or special model access.

0 favorites 0 likes
#sota

@ash_csx: We’re dropping two open source SLMs this week. 1. One of them matches SOTA accuracy at up to 93x smaller. 2. The other …

X AI KOLs Following ↗ · 2026-05-11 Cached

Two new open-source small language models are being released: one matches state-of-the-art accuracy at up to 93x smaller size, and the other outperforms a recent OpenAI model. The first model drops tomorrow.

0 favorites 0 likes
#sota

@oliviscusAI: Someone open-sourced a memory layer that beats every RAG system on the planet. It's called Memvid. +35% SOTA on LoCoMo.…

X AI KOLs Timeline ↗ · 2026-05-11 Cached

A new open-source memory layer called Memvid claims to outperform all existing RAG systems, achieving +35% SOTA on LoCoMo and +76% on multi-hop reasoning, packaged as a single .mv2 file.

0 favorites 0 likes
#sota

Xiaomi released their SOTA model, MiMo-V2.5-Pro.

Reddit r/singularity ↗ · 2026-04-22

Xiaomi launched MiMo-V2.5-Pro, claiming state-of-the-art performance.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback