Tag
A user predicts DeepSeek V4 Flash 0731 will score 57±1 on Artificial Analysis, matching Kimi K3 level, based on linear regression. The post expresses excitement about the model's performance relative to its price.
The MS AI Frontiers team introduces BenchPress, a method that uses matrix completion to predict LLM benchmark scores from just five probes, showing the score matrix is effectively rank-2.
This research paper demonstrates that the scores of frontier AI models across 133 benchmarks are approximately rank-2, meaning only two latent factors explain over 90% of variation. The authors introduce BenchPress, a logit-space matrix completion method that predicts a model's full scorecard from just a few benchmarks, significantly reducing the cost of evaluation.