sota

Tag

Cards List
#sota

@intology: The models are improving the models. Locus, our automated AI research system, is SOTA on PostTrainBench and post-trains…

X AI KOLs Following · 6d ago Cached

Locus, an automated AI research system, achieves SOTA on PostTrainBench and post-trains Qwen3 base models that surpass human post-trained models. It also shows strong performance on Kaggle competitions.

0 favorites 0 likes
#sota

@TheAhmadOsman: Yet another thing where Kimi K3 is SoTA and beating the frontier

X AI KOLs Following · 6d ago Cached

Kimi K3 achieves state-of-the-art results on the new pmpp-hard benchmark, scoring 0.71 across 69 GPU kernel tasks and beating other frontier models.

0 favorites 0 likes
#sota

@NielsRogge: Qwen released a SOTA computer use agent with a technical report It's not on @arxiv, but it is on Papers with Code. You …

X AI KOLs Timeline · 6d ago Cached

Qwen released Qwen-CUA, a native computer-use agent with a 397B-A17B mixture-of-experts backbone, achieving state-of-the-art results on OSWorld-Verified and ranking #2 on WebArena. A technical report is available on Papers with Code.

0 favorites 0 likes
#sota

@MikeBradleyAI: TLDR on @deepseek_ai 0731 V4 Flash. It is comfortably the current SOTA for 190GB VRAM or unified memory based systems. …

X AI KOLs Following · 2026-08-02 Cached

Mike Bradley shares benchmark results claiming DeepSeek V4 Flash 0731 is the current state-of-the-art for 190GB VRAM systems, matching or exceeding an Unsloth 3-bit Qwen3.5-397B in quality while running about 3x faster.

0 favorites 0 likes
#sota

@VoidAsuka: there hasn't been a good open-source video generation model in a while, so we built one it's a sota video generation mo…

X AI KOLs Following · 2026-07-31 Cached

Announcement of an open-source SOTA video generation model, MiniMax H3, offering commercial-grade generation, unbeatable cost efficiency, and open weights.

0 favorites 0 likes
#sota

Kwaipilot/KAT-Coder-V2.5-Dev

Hugging Face Models Trending · 2026-07-23 Cached

KAT-Coder-V2.5-Dev is an open-weight MoE coding model with 35B total parameters (3B active), achieving state-of-the-art results on agentic coding benchmarks through SFT and RL training.

0 favorites 0 likes
#sota

@Azaliamirh: Check out LLM-as-a-Verifier: a simple, cheap, & general-purpose self-improvement technique that boosts performance on "…

X AI KOLs Timeline · 2026-07-10 Cached

LLM-as-a-Verifier is a simple, cheap, general-purpose self-improvement technique for agentic tasks, using fine-grained scoring and logprob-based ranking to achieve SOTA on multiple benchmarks like SWE-Bench Verified and Terminal-Bench V2.

0 favorites 0 likes
#sota

@mertunsal2020: Today, we are releasing Le Chaton L∃∀N, aka Leanstral 1.5. It achieves SOTA performance on graduate algebra benchmarks …

X AI KOLs Following · 2026-07-03 Cached

Today, we are releasing Le Chaton L∃∀N, aka Leanstral 1.5, which achieves SOTA performance on graduate algebra benchmarks FATE-H and FATE-X and improves the Pareto Frontier on PutnamBench, solving 587/672 problems with a x10 cheaper budget.

0 favorites 0 likes
#sota

Ornith-1.0 released on Hugging Face

Reddit r/LocalLLaMA · 2026-06-25

Ornith-1.0 has been released on Hugging Face, featuring a collection of models ranging from 9B to 397B parameters, including dense and MoE architectures, claiming state-of-the-art performance on various benchmarks.

0 favorites 0 likes
#sota

@NielsRogge: GLM-5.2 is the literal SOTA on PostTrainBench Beating GPT-5.5 and Opus 4.8 Learn more here https://paperswithcode.co/be…

X AI KOLs Following · 2026-06-20

GLM-5.2 achieves state-of-the-art results on PostTrainBench, outperforming GPT-5.5 and Opus 4.8.

0 favorites 0 likes
#sota

@teach_fireworks: https://x.com/teach_fireworks/status/2067243590447952212

X AI KOLs Timeline · 2026-06-17 Cached

SAG (SQL-Augmented Generation) is a novel SQL-based retrieval augmented generation method that converts data chunks into events and entities, enabling multi-hop reasoning via SQL join queries. On the MuSiQue dataset, recall increased from 65.13% to 80.04%. It supports second-level online retrieval of approximately 500 million data entries and has been open-sourced.

0 favorites 0 likes
#sota

@hwchase17: Detecting issues in production agent traces is hard. You have to do it cheaply (because of volume) but also accurately …

X AI KOLs Following · 2026-06-15

Harrison Chase announces a post-trained model for detecting issues in production agent traces, claiming SOTA accuracy at 10-100x cheaper rates than frontier models.

0 favorites 0 likes
#sota

@ChengleiSi: Excited to share these preliminary results on our internal autoresearch system @Recursive_SI, where we achieve SOTA on …

X AI KOLs Following · 2026-06-11 Cached

Recursive's automated AI research system achieves state-of-the-art results on NanoChat, NanoGPT Speedrun, and GPU kernel benchmarks by automating the research loop without task-specific adaptations, and open-sourcing artifacts for further inspection.

0 favorites 0 likes
#sota

@gregpr07: Browser Use Beta just achieved SOTA on our hardest internal web agent benchmark. Fable is genuinely amazing for optimiz…

X AI KOLs Following · 2026-06-11 Cached

Browser Use Beta achieved state-of-the-art results on a difficult internal web agent benchmark, using Fable for optimization and analysis.

0 favorites 0 likes
#sota

@patpcj: Fable-5/Mythos dropped this morning, so we tested it on agentic search - and it’s the new SOTA. The performance gap is …

X AI KOLs Timeline · 2026-06-10 Cached

Fable-5/Mythos achieves new SOTA on agentic search but is expensive for self-hosting, while open-weight Harness-1 offers a cost-effective alternative with fewer query restrictions.

0 favorites 0 likes
#sota

@adamdotnew: We’re very excited to announce that Anthropic’s Fable 5 is a SOTA model at mechanical engineering tasks! It can generat…

X AI KOLs Following · 2026-06-09 Cached

Anthropic announces Fable 5, a state-of-the-art model for mechanical engineering tasks capable of generating intricate assemblies and mechanisms from a single prompt.

0 favorites 0 likes
#sota

@perdactor: 1/ Meet Argus-Retriever: the first late-interaction visual doc retriever where the document representation adapts to th…

X AI KOLs Following · 2026-06-06 Cached

Argus-Retriever is a new late-interaction visual document retriever that adapts document representation to the query, achieving SOTA performance on ViDoRe benchmarks with a smaller index.

0 favorites 0 likes
#sota

@Miles_Brundage: BREAKING: massively improved SOTA score on Clear AVERI Pronunciation Guide Bench, via my colleague Carly

X AI KOLs Following · 2026-06-04 Cached

Miles Brundage announces a state-of-the-art (SOTA) score improvement on the Clear AVERI Pronunciation Guide Bench achieved by colleague Carly.

0 favorites 0 likes
#sota

@yvbbrjdr: I recommend everyone to read the MAI-Thinking-1 technical paper. It contains detailed (almost all) information on how to train a SOTA LLM. https://microsoft.ai/wp-content/uploads/2026/06/ma…

X AI KOLs Timeline · 2026-06-02 Cached

Recommended reading: the MAI-Thinking-1 technical paper, which details almost all the steps to train a SOTA large language model.

0 favorites 0 likes
#sota

@swyx: roundup of links:

X AI KOLs Following · 2026-06-02 Cached

NVIDIA releases Cosmos 3 (Mixture-of-Transformers models up to 64B), Nemotron 3 Ultra (550B-A55B LLM), and previews RTX Spark personal superchip at Computex 2026, achieving SOTA on multiple open model leaderboards.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback