leaderboard

Tag

Cards List
#leaderboard

@xiangxiang103: That's insanely impressive! Hugging Face unleashed a dark horse this week—one that's dominating two tracks at once. Spe…

X AI KOLs Timeline ↗ · 4d ago Cached

NetEase Youdao's open-source AI models R2T2 and T3PO have topped Hugging Face leaderboards for speech recognition and translation, outperforming major competitors with impressive real-time performance and stability.

0 favorites 0 likes
#leaderboard

@rauchg: We ran fresh Next.js evals. The tally: ① Opus 5.5 [𝟿𝟽%] ② GPT 6 Sol [𝟿𝟽%] ③ Fable 5.1 [𝟿𝟽%] ④ Grok 4.7 [𝟿𝟺%] No…

X AI KOLs Timeline ↗ · 4d ago Cached

Fresh evaluations of AI models on Next.js reveal Claude Opus 5.5, GPT-6 Sol, and Fable 5.1 tying at 97% success rates, with Grok 4.7 at 94% but offering lower costs.

0 favorites 0 likes
#leaderboard

@omarsar0: Recommended if you work on agent memory. Retrieval is the hardest memory problem in most harnesses because agents keep …

X AI KOLs Timeline ↗ · 5d ago Cached

The tweet from @omarsar0 recommends focusing on agent memory retrieval and announces the opening of the Agent Memory Challenge 2026 Cycle 2, which evaluates memory in AI agents using coding and text tracks.

0 favorites 0 likes
#leaderboard

@FiniYang: Made a little cat survival game Had the official Jev and local Laya (mlx) team up to protect the kitten Each step chose…

X AI KOLs Timeline ↗ · 5d ago Cached

The author created a cat survival game to test the performance of AI models Jev and Laya, finding that Jev is more accurate but slower, while Laya is fast but inaccurate, and plans to develop a leaderboard for such models.

0 favorites 0 likes
#leaderboard

GoBench: Evaluating LLMs on the game of Go [R]

Reddit r/MachineLearning ↗ · 2026-09-16

GoBench is a benchmark for evaluating large language models on 9x9 Go games, demonstrating strong correlation with ARC-AGI and featuring a leaderboard with KataGo opponents from random to superhuman levels.

0 favorites 0 likes
#leaderboard

Qwen3.8 Max (0902) scores 45 on the Artificial Analysis Intelligence Index, up 5 points in a month and back on top of China's leaderboard, nosing out GLM-5.3 (44.9) and Kimi K3 (43.8)

Reddit r/LocalLLaMA ↗ · 2026-09-16

Qwen3.8 Max has been upgraded and now leads China's AI leaderboard with a score of 45 on the Artificial Analysis Intelligence Index, surpassing GLM-5.3 and Kimi K3 after a 30-day improvement to the 2.4T MoE model.

0 favorites 0 likes
#leaderboard

@akshdeeps_001: New leaderboard just dropped

X AI KOLs Timeline ↗ · 2026-09-11 Cached

A new leaderboard related to AI or technology benchmarks has been announced in a tweet by @akshdeeps_001.

0 favorites 0 likes
#leaderboard

GPT-6 Astra 2,340 Elo on ChessBench - Ranked #11

Reddit r/singularity ↗ · 2026-09-11 Cached

The ChessBench AI chess leaderboard has been updated with new model performances, including GPT-6 Astra achieving 2,340 Elo and ranking #11, along with other models like Claude Fable 5.1 and Gemini 3.8 Flash.

0 favorites 0 likes
#leaderboard

@BenjaminDEKR: Astra understands spatial / 3D relationships better than other leading models. This is why it's so good at CAD, models,…

X AI KOLs Timeline ↗ · 2026-09-05 Cached

Astra reportedly surpasses other leading models in spatial and 3D understanding, achieving first place on the VoxelBench benchmark with an Elo rating exceeding 2600 and a lead of over 300 points.

0 favorites 0 likes
#leaderboard

@svpino: This is a cool benchmark that uses video games to test agents. Basically, they speedrun the game to test how well agent…

X AI KOLs Timeline ↗ · 2026-09-03 Cached

SpeedrunBench is a new benchmark that evaluates AI agents by having them speedrun video games to test planning, learning, and decision-making, with a public leaderboard for model comparison.

0 favorites 0 likes
#leaderboard

@ArtificialAnlys: Wan 3.0 debuts at #1 on the Artificial Analysis Video Editing Leaderboard, and is a close #2 in Text to Video with Audi…

X AI KOLs Timeline ↗ · 2026-09-03 Cached

Wan 3.0 is Alibaba's new all-in-one video generation and editing model that debuts at #1 on the Artificial Analysis Video Editing Leaderboard, featuring native audio and multimodal inputs, available in public preview via Alibaba Cloud.

0 favorites 0 likes
#leaderboard

@ArtificialAnlys: We have updated the Artificial Analysis Image Editing Arena to expand the range of editing tasks we test for, from Enha…

X AI KOLs Timeline ↗ · 2026-09-02 Cached

The Artificial Analysis Image Editing Arena has been updated to expand editing tasks and use cases, providing a comprehensive leaderboard that ranks AI models on human preference for various image editing scenarios.

0 favorites 0 likes
#leaderboard

@rohanpaul_ai: LLM rankings can be created by evaluation choices as much as model differences, so one setup should never decide the le…

X AI KOLs Timeline ↗ · 2026-08-31 Cached

The paper shows that LLM rankings are highly sensitive to evaluation setup choices, such as prompt format and scoring methods, leading to significant variability in model performance and rankings.

0 favorites 0 likes
#leaderboard

@freemanjiangg: I'm excited to share some of what i've been working on at Sesame! https://turnbench.sesame.com is our turn-taking bench…

X AI KOLs Following ↗ · 2026-08-28 Cached

Sesame releases TurnBench, a benchmark for evaluating real-time conversational turn-taking, with a leaderboard and labeled dataset to assess model performance in detecting speech events.

0 favorites 0 likes
#leaderboard

The Open ASR Leaderboard Adds Its First Global South Language

Hugging Face Blog ↗ · 2026-08-28 Cached

Voice Arena and Hugging Face have added Hindi and Indian English to the Open ASR Leaderboard, introducing new evaluation sets designed to capture diverse speaker attributes and address biases in automatic speech recognition.

0 favorites 0 likes
#leaderboard

SnakeRank

Product Hunt ↗ · 2026-08-27 Cached

SnakeRank is a pay-to-rank leaderboard where startups bid to climb positions in a snake-themed system, with bids stacking and rank based on total dollars committed.

0 favorites 0 likes
#leaderboard

We combined 3 public TTS leaderboards into one meta-ranking of 110 models

Reddit r/AI_Agents ↗ · 2026-08-26

A new meta-ranking combines three public TTS leaderboards into one unified ranking of 110 models across 46 providers, updated weekly with a fixed methodology.

0 favorites 0 likes
#leaderboard

A dataset with 52 Text to image model evaluation [P]

Reddit r/MachineLearning ↗ · 2026-08-26

A new benchmark dataset and evaluation methodology for 52 text-to-image models has been published, including results, a leaderboard, and a gallery to assess performance on challenging prompts.

0 favorites 0 likes
#leaderboard

@arcinstitute: Nearly 300 teams have already submitted to the Virtual Cell Challenge leaderboard to see how their initial models rank …

X AI KOLs Following ↗ · 2026-08-26 Cached

Nearly 300 teams have submitted initial models to the Virtual Cell Challenge leaderboard, with current standings based on six metrics this year.

0 favorites 0 likes
#leaderboard

@app_sail: https://outbid.lol This leaderboard hasn't changed in two days. Do you guys think there's still a chance, or is it just a flash in the pan? If it were your product, what would you do?

X AI KOLs Following ↗ · 2026-08-26 Cached

Discussing the performance of a creator monetization market named Tutti on X's leaderboard and questioning its sustainability.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback