Tag
Grok's latest model, grok-4.6-high, ranks #43 on LMArena text and 7th on webdev, failing to reclaim the top spot Grok last held in December 2025.
Wired's 2026 ranking of the best smart glasses, covering models from Meta, Viture, Even Realities, and RayNeo, with pros, cons, and specs for each.
A chart ranking the top 20 most visited AI tools by estimated web visits from May 2025 to Apr 2026 shows ChatGPT dominating with 64.7B visits, far ahead of Canva, Gemini, and others, based on a study of 9,531 AI tools.
CNBC and Statista released their annual list of the world's top 40 fintech companies for 2026, covering 8 subverticals.
This paper introduces a rubric-oriented framework for evaluating and optimizing document sets for large language models, including a new benchmark and a training-free method that improves downstream generation performance.
This paper proposes a preference-based learning framework for antibody expression ranking, integrating scarce quantitative data with large-scale weak positive supervision from immunization sequences. The method adapts Direct Preference Optimization to protein language models using a union-masked log-likelihood approximation and IMGT-based alignment, achieving improved ranking performance on a diverse internal dataset.
Google has dropped out of the top 15 in a significant tech ranking, marking a notable decline in its competitive standing.
Kimi K3 has achieved a ranking on the Agent Arena leaderboard equivalent to opus thinking.
Chinese AI models dominated the top 6 usage on OpenRouter for last week and this week, with the free version of Tencent Hunyuan3 ranking first, Xiaomi MiMo-V2.5 and DeepSeek-V4-Flash in the lead, reflecting the significant increase in influence of domestic models in the open-source community.
RUBRIC is a generator-agnostic filtering framework for imbalanced classification that selects synthetic samples by balancing realism (via a discriminator) and utility (margin-based scoring), improving F1-macro and recall on benchmarks like credit-card fraud detection.
Introduces SALT, a benchmark with deterministic ground truth for evaluating LLM uncertainty at fine-grained atomic levels in long-form generation. Analysis of over 50 LLMs reveals insights into confidence functions, error propagation, and trade-offs with reasoning.
Chinese supercomputer LineShine has claimed the top spot in the global supercomputer rankings, marking a significant achievement in high-performance computing.
Uses the Bradley-Terry model and Elo rating system to statistically determine a dog's favorite treat through pairwise comparison experiments.
An LLM leaderboard based on Hugging Face download and like counts, grouped, filtered, and time-averaged, highlighting the most popular models like Qwen and Gemma.
WIRED reveals that Peter Thiel's private Dialog club secretly ranks its members by wealth and fame using a hidden A, B, C grading system, tracking relationships and employing algorithms to manage attendance and seating, based on a leaked data trove of nearly 200 prominent individuals.
OneRank proposes a Transformer-native multi-task ranking framework that integrates feature encoding and prediction to reduce inter-task interference and improve ranking performance in recommender systems.
This paper proposes Representation Curriculum (RC), a training-time intervention that stages feature utilization to reduce over-reliance on exposure-confounded historical signals and improve cold-start generalization in ranking systems. The method is theoretically analyzed and validated on public benchmarks and large-scale eBay search experiments.
This paper proposes TOPSIS-RAD, a modified version of the TOPSIS method that incorporates decision-maker-defined reference levels (VPL and DPL) to address issues like misalignment with preferences, outlier sensitivity, and rank reversal.
This paper introduces a decision-focused learning approach for survival analysis that aligns predictive models with downstream allocation decisions, using NDCG optimization. Applied to US heart transplant data, it improves ranking performance by 50-100%, potentially yielding thousands of additional life-years annually.
Mixedbread's reranker achieves GPT 5.5-level performance on OBLIQ-bench while being 27x faster, according to early results.