model-ranking

Tag

Cards List
#model-ranking

Closed models from Google & OpenAI currently take #2 & #3 on OpenRouter, which had traditionally a bias towards cheaper Chinese open weight models

Reddit r/ArtificialInteligence · 4d ago

Closed models from Google and OpenAI have taken the second and third positions on OpenRouter, a platform that traditionally favored cheaper Chinese open-weight models.

0 favorites 0 likes
#model-ranking

Top 10 most liked models on Qwen's HuggingFace page

Reddit r/LocalLLaMA · 2026-08-24

Qwen3.8-27B has rapidly become the most liked model on Qwen's HuggingFace page, surpassing QwQ-32B, with users preferring models around 27B and 9B parameters.

0 favorites 0 likes
#model-ranking

Qwen3.8 Max now ranked as the best overall model by agentic index

Hacker News Top · 2026-08-06 Cached

Qwen3.8 Max is now ranked as the best overall model on Artificial Analysis's agentic index, surpassing other leading AI models in independent evaluations.

0 favorites 0 likes
#model-ranking

Dynamically Allocating Evaluation Effort for Model Ranking

arXiv cs.CL · 2026-08-05 Cached

This paper formalizes multi-model human evaluation as a best-arm identification problem in a multi-armed bandit setup, adaptively allocating annotation effort to focus on competitive models and improve ranking discrimination.

0 favorites 0 likes
#model-ranking

Deepseek V4 flash 0731 ranks #21 on Agent Arena

Reddit r/LocalLLaMA · 2026-08-04

DeepSeek V4 flash 0731 ranks #21 on Agent Arena, below Sonnet 4.6 and Luna, but users appreciate its open-source nature for privacy and control.

0 favorites 0 likes
#model-ranking

@0xSero: Best models smallest to largest right now. - Gemma-4-12B - Qwen3.8-27B - Laguna-S2.1 - Deepseek-V4-Flash - Inkling-Smal…

X AI KOLs Timeline · 2026-08-03 Cached

A tweet ranking the best open-weight AI models from smallest to largest, highlighting the vibrant open-weights community.

0 favorites 0 likes
#model-ranking

Kimi K3 achieves 3rd Place on ArtificalAnalysis, beating out Claude Opus 4.8

Reddit r/singularity · 2026-07-16

Kimi K3 model ranks third on the ArtificialAnalysis benchmark, surpassing Claude Opus 4.8.

0 favorites 0 likes
#model-ranking

@mylifcc: http://Arena.ai just officially added 'Factuality' to model rankings. The leaderboard now supports weighted viewing of 'Human Preference + Factuality' (default 25% factuality weight). They have annotated over 2 million claims from real conversations (Text Ar…

X AI KOLs Timeline · 2026-07-15 Cached

Arena.ai has added Factuality to model rankings, supporting weighting of human preference and factuality, and showing changes in model rankings.

0 favorites 0 likes
#model-ranking

Spectral Signatures of Large Language Models

arXiv cs.CL · 2026-07-07 Cached

This paper introduces a spectral shape-based metric using Heavy-Tailed Self-Regularization theory to characterize, compare, and manage large language models. The approach is data-free, computationally efficient, and scale-invariant, enabling model lineage tracing, unsupervised clustering, and performance quantification across diverse model collections.

0 favorites 0 likes
#model-ranking

LLM rankings are not a ladder: experimental results from a transitive benchmark graph [D]

Reddit r/MachineLearning · 2026-05-09

The author introduces LLM Win, a tool that visualizes LLM benchmark results as a directed graph to analyze transitive relationships and ranking reversals. Experimental findings suggest that LLM rankings function more like a capability graph with high weak-to-strong reachability rather than a linear ladder.

0 favorites 0 likes
← Back to home

Submit Feedback