标签
This paper compares natively multimodal embedding models (Gemini Embedding 2, Amazon Nova 2) against frontier LLMs (GPT-4.1, Claude Sonnet 4.6) for hard-negative text-to-image retrieval, finding comparable accuracy but much lower latency for embedding-based ranking.
关于 Artificial Analysis 如何在 SciCode 基准测试中将 Gemma 4 排在 Qwen3.6 27b 之上的讨论,质疑该排名是否反映现实世界的编码能力,还是揭示了基准测试的问题。
马克斯·普朗克智能系统研究所推出了 Comparity AI,一个基于人类偏好进行 LLM 排名的研究平台,提供免费访问前沿模型和个人排行榜的服务。