ranking

Tag

Cards List
#ranking

@BenjaminDEKR: Grok held the top rank on LMArena for just one month -- November 17, 2025, until December 22 No Grok model has been bac…

X AI KOLs Timeline · 3d ago Cached

Grok's latest model, grok-4.6-high, ranks #43 on LMArena text and 7th on webdev, failing to reclaim the top spot Grok last held in December 2025.

0 favorites 0 likes
#ranking

Ranking the Best Smart Glasses: Meta, Viture, & More (2026)

Wired · 2026-08-07 Cached

Wired's 2026 ranking of the best smart glasses, covering models from Meta, Viture, Even Realities, and RayNeo, with pros, cons, and specs for each.

0 favorites 0 likes
#ranking

Top 20 most visited AI tools by estimated web visits, May 2025 to Apr 2026

Reddit r/ArtificialInteligence · 2026-08-06

A chart ranking the top 20 most visited AI tools by estimated web visits from May 2025 to Apr 2026 shows ChatGPT dominating with 64.7B visits, far ahead of Canva, Gemini, and others, based on a study of 9,531 AI tools.

0 favorites 0 likes
#ranking

@efipm: The world's top fintech companies: 2026 - annual list by CNBC and Statista 40 Fintechs in 8 subverticals https://buff.l…

X AI KOLs Timeline · 2026-07-22 Cached

CNBC and Statista released their annual list of the world's top 40 fintech companies for 2026, covering 8 subverticals.

0 favorites 0 likes
#ranking

Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking

Hugging Face Daily Papers · 2026-07-22 Cached

This paper introduces a rubric-oriented framework for evaluating and optimizing document sets for large language models, including a new benchmark and a training-free method that improves downstream generation performance.

0 favorites 0 likes
#ranking

Preference-based Antibody Expression Ranking: Scaling with Large-scale Weak Supervision

arXiv cs.LG · 2026-07-21 Cached

This paper proposes a preference-based learning framework for antibody expression ranking, integrating scarce quantitative data with large-scale weak positive supervision from immunization sequences. The method adapts Direct Preference Optimization to protein language models using a union-masked log-likelihood approximation and IMGT-based alignment, achieving improved ranking performance on a diverse internal dataset.

0 favorites 0 likes
#ranking

Google has disappeared completely from the top 15

Reddit r/LocalLLaMA · 2026-07-20

Google has dropped out of the top 15 in a significant tech ranking, marking a notable decline in its competitive standing.

0 favorites 0 likes
#ranking

According to Agent Arena Kimi K3 ranks at same level as opus thinking

Reddit r/LocalLLaMA · 2026-07-20 Cached

Kimi K3 has achieved a ranking on the Agent Arena leaderboard equivalent to opus thinking.

0 favorites 0 likes
#ranking

@karminski3: Bring it on — we excel at competition. Chinese models have dominated OpenRouter's top 6 usage models for last week and this week. It's a complete reversal from July last year, when only DeepSeek and Qwen broke out. Now the top usage is Tencent Hunyuan3's free version (can be configured into Lobster…)

X AI KOLs Timeline · 2026-07-14 Cached

Chinese AI models dominated the top 6 usage on OpenRouter for last week and this week, with the free version of Tencent Hunyuan3 ranking first, Xiaomi MiMo-V2.5 and DeepSeek-V4-Flash in the lead, reflecting the significant increase in influence of domestic models in the open-source community.

0 favorites 0 likes
#ranking

RUBRIC: Realism--Utility Balanced Ranking for Imbalanced Classification

arXiv cs.LG · 2026-07-14 Cached

RUBRIC is a generator-agnostic filtering framework for imbalanced classification that selects synthetic samples by balancing realism (via a discriminator) and utility (margin-based scoring), improving F1-macro and recall on benchmarks like credit-card fraud detection.

0 favorites 0 likes
#ranking

Evaluating LLM Uncertainty in Long-Form Generation Using Deterministic Ground Truth

arXiv cs.AI · 2026-07-07 Cached

Introduces SALT, a benchmark with deterministic ground truth for evaluating LLM uncertainty at fine-grained atomic levels in long-form generation. Analysis of over 50 LLMs reveals insights into confidence functions, error propagation, and trade-offs with reasoning.

0 favorites 0 likes
#ranking

LineShine, a Chinese supercomputer, has topped the global supercomputer ranking

Reddit r/singularity · 2026-06-28 Cached

Chinese supercomputer LineShine has claimed the top spot in the global supercomputer rankings, marking a significant achievement in high-performance computing.

0 favorites 0 likes
#ranking

Finding the Best Dog Treat with Statistics

Hacker News Top · 2026-06-22 Cached

Uses the Bradley-Terry model and Elo rating system to statistically determine a dog's favorite treat through pairwise comparison experiments.

0 favorites 0 likes
#ranking

@onusoz: I created an LLM leaderboard based on Hugging Face download and like counts, grouped, filtered and time-averaged. Top 5…

X AI KOLs Following · 2026-06-20 Cached

An LLM leaderboard based on Hugging Face download and like counts, grouped, filtered, and time-averaged, highlighting the most popular models like Qwen and Gemma.

0 favorites 0 likes
#ranking

How the Peter Thiel-Linked Dialog Club Secretly Ranks Its Members

Wired · 2026-06-18 Cached

WIRED reveals that Peter Thiel's private Dialog club secretly ranks its members by wealth and fame using a hidden A, B, C grading system, tracking relationships and employing algorithms to manage attendance and seating, based on a leaked data trove of nearly 200 prominent individuals.

0 favorites 0 likes
#ranking

OneRank: Unified Transformer-Native Ranking Architecture for Multi-Task Recommendation

Hugging Face Daily Papers · 2026-06-15 Cached

OneRank proposes a Transformer-native multi-task ranking framework that integrates feature encoding and prediction to reduce inter-task interference and improve ranking performance in recommender systems.

0 favorites 0 likes
#ranking

Representation Curriculum: Stagewise Training for Robust Ranking and Allocation

arXiv cs.LG · 2026-06-10 Cached

This paper proposes Representation Curriculum (RC), a training-time intervention that stages feature utilization to reduce over-reliance on exposure-confounded historical signals and improve cold-start generalization in ranking systems. The method is theoretically analyzed and validated on public benchmarks and large-scale eBay search experiments.

0 favorites 0 likes
#ranking

TOPSIS-RAD: Ranking According to Desires

arXiv cs.AI · 2026-06-08 Cached

This paper proposes TOPSIS-RAD, a modified version of the TOPSIS method that incorporates decision-maker-defined reference levels (VPL and DPL) to address issues like misalignment with preferences, outlier sensitivity, and rank reversal.

0 favorites 0 likes
#ranking

Aligning Data-Driven Predictors with Allocation: A Decision-Focused Approach to Survival Analysis

arXiv cs.LG · 2026-06-03 Cached

This paper introduces a decision-focused learning approach for survival analysis that aligns predictive models with downstream allocation decisions, using NDCG optimization. Applied to US heart transplant data, it improves ranking performance by 50-100%, potentially yielding thousands of additional life-years annually.

0 favorites 0 likes
#ranking

@RuiTheBaker: GPT 5.5-level ranking but 27x faster?! @mixedbreadai

X AI KOLs Following · 2026-06-02 Cached

Mixedbread's reranker achieves GPT 5.5-level performance on OBLIQ-bench while being 27x faster, according to early results.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback