Tag
NetEase Youdao's open-source AI models R2T2 and T3PO have topped Hugging Face leaderboards for speech recognition and translation, outperforming major competitors with impressive real-time performance and stability.
Fresh evaluations of AI models on Next.js reveal Claude Opus 5.5, GPT-6 Sol, and Fable 5.1 tying at 97% success rates, with Grok 4.7 at 94% but offering lower costs.
The tweet from @omarsar0 recommends focusing on agent memory retrieval and announces the opening of the Agent Memory Challenge 2026 Cycle 2, which evaluates memory in AI agents using coding and text tracks.
The author created a cat survival game to test the performance of AI models Jev and Laya, finding that Jev is more accurate but slower, while Laya is fast but inaccurate, and plans to develop a leaderboard for such models.
GoBench is a benchmark for evaluating large language models on 9x9 Go games, demonstrating strong correlation with ARC-AGI and featuring a leaderboard with KataGo opponents from random to superhuman levels.
Qwen3.8 Max has been upgraded and now leads China's AI leaderboard with a score of 45 on the Artificial Analysis Intelligence Index, surpassing GLM-5.3 and Kimi K3 after a 30-day improvement to the 2.4T MoE model.
A new leaderboard related to AI or technology benchmarks has been announced in a tweet by @akshdeeps_001.
The ChessBench AI chess leaderboard has been updated with new model performances, including GPT-6 Astra achieving 2,340 Elo and ranking #11, along with other models like Claude Fable 5.1 and Gemini 3.8 Flash.
Astra reportedly surpasses other leading models in spatial and 3D understanding, achieving first place on the VoxelBench benchmark with an Elo rating exceeding 2600 and a lead of over 300 points.
SpeedrunBench is a new benchmark that evaluates AI agents by having them speedrun video games to test planning, learning, and decision-making, with a public leaderboard for model comparison.
Wan 3.0 is Alibaba's new all-in-one video generation and editing model that debuts at #1 on the Artificial Analysis Video Editing Leaderboard, featuring native audio and multimodal inputs, available in public preview via Alibaba Cloud.
The Artificial Analysis Image Editing Arena has been updated to expand editing tasks and use cases, providing a comprehensive leaderboard that ranks AI models on human preference for various image editing scenarios.
The paper shows that LLM rankings are highly sensitive to evaluation setup choices, such as prompt format and scoring methods, leading to significant variability in model performance and rankings.
Sesame releases TurnBench, a benchmark for evaluating real-time conversational turn-taking, with a leaderboard and labeled dataset to assess model performance in detecting speech events.
Voice Arena and Hugging Face have added Hindi and Indian English to the Open ASR Leaderboard, introducing new evaluation sets designed to capture diverse speaker attributes and address biases in automatic speech recognition.
SnakeRank is a pay-to-rank leaderboard where startups bid to climb positions in a snake-themed system, with bids stacking and rank based on total dollars committed.
A new meta-ranking combines three public TTS leaderboards into one unified ranking of 110 models across 46 providers, updated weekly with a fixed methodology.
A new benchmark dataset and evaluation methodology for 52 text-to-image models has been published, including results, a leaderboard, and a gallery to assess performance on challenging prompts.
Nearly 300 teams have submitted initial models to the Virtual Cell Challenge leaderboard, with current standings based on six metrics this year.
Discussing the performance of a creator monetization market named Tutti on X's leaderboard and questioning its sustainability.