Tag
TasteBench introduces a multimodal benchmark and privacy-preserving competition for sensory prediction, spanning food-level ranking (21K+ human evaluations across 215 plant-based foods) and molecular-level taste classification (15K flavor molecules), with baselines reaching pairwise accuracy competitive with median human panelists. It aims to provide computational proxies for sustainable protein discovery, analogous to molecular docking in drug design.
An analysis of the top 50 most-cited AI researchers, highlighting how landmark papers like Attention is All You Need have shaped the field and whose work drives AI's influence.
The paper introduces reverse Item Response Theory to pharmacogenomics, treating cancer types as latent "subjects" and drugs as "items" to achieve sparsity-robust ranking recovery in fragmented drug-response matrices, outperforming simple averaging on GDSC2 data across multiple missingness regimes.
This paper empirically studies Jev, a decision-oriented "System One Model," for personalized recommendation reranking, finding it occupies a distinct quality–latency operating regime compared with recommendation-specific models and pointwise/listwise LLM rerankers across multiple domains and candidate-set sizes.
Sonnet 5.5 ranks second on the Artificial Analysis leaderboard, indicating its high performance in AI benchmarks.
This paper shows that LLM-based scorers producing equal ranking quality can still make different decisions when candidate order in the prompt changes, and introduces order-consistency SFT (OC-SFT) to penalize score disagreement across permutations. OC-SFT holds ranking quality while improving decision stability across reranking, response ranking, and multi-document QA tasks.
Meta's Muse app has topped the US App Store charts in the past 7 days, outperforming ChatGPT in rankings.
This paper introduces LIGE-GR, a method that uses large language models to smoothly transition from traditional ranking to generative recommendation systems, aiming to improve recommendation performance in the LLM era.
The article ranks individual barriers in the AI era, prioritizing self-brainwashing ability over mental resilience, credit, execution, judgment, filtering ability, and information asymmetry.
A tweet recommends viewing reviews of MapQuest's mobile app, which is currently the top app on the U.S. App Store, noting its nostalgic appeal.
This paper proposes MoPLEx, an algorithm for learning mixtures of Plackett-Luce models to handle heterogeneous preferences in AI alignment, showing improved clustering and ranking accuracy over baselines.
A new meta-ranking combines three public TTS leaderboards into one unified ranking of 110 models across 46 providers, updated weekly with a fixed methodology.
Stanford GSB professor Ilya Strebulaev releases a venture capital ranking based on 30 years and 230,000 investments, showing that 5% of VCs create 90% of profits, and it has a low correlation with the Midas List.
The 2026 US Top Venture Capital Firms Ranking has been officially released, listing the 100 leading firms in investments for the year.
Gary Marcus tweets that something has risen to #1 in technology, likely highlighting a notable trend or achievement in the tech field.
Cross-lingual Ranking Preference Optimization (CRPO) is a novel framework that enhances multilingual LLM alignment by transferring English preference knowledge to target languages through hierarchical ranking optimization, demonstrating improved performance in instruction-following and knowledge utilization across multiple languages.
This paper presents a method to rank neural operator models during deployment using shared physics responses, achieving high accuracy without ground-truth reference solutions for scientific computing applications.
UMER introduces a unified framework for multimodal retrieval that combines embedding and ranking via pair-aware discriminative reasoning, achieving state-of-the-art performance on the MMEB-V2 benchmark.
This paper analyzes selective prediction systems for rare-disease diagnosis, demonstrating that small open-weight LLMs have low recall on ultra-rare diseases and exploring the use of score margins for decision-making with limitations.
The paper identifies forecast collapse in time-series foundation models for hourly equity return prediction and introduces CalibRank to balance calibration and ranking, significantly improving cross-sectional correlation.