Tag
This paper compares natively multimodal embedding models (Gemini Embedding 2, Amazon Nova 2) against frontier LLMs (GPT-4.1, Claude Sonnet 4.6) for hard-negative text-to-image retrieval, finding comparable accuracy but much lower latency for embedding-based ranking.
A discussion about how Artificial Analysis ranks Gemma 4 above Qwen3.6 27b on the SciCode benchmark, questioning whether the ranking reflects real-world coding ability or reveals a benchmarking issue.
Max Planck Institute for Intelligent Systems has launched Comparity AI, a research platform for human preference-based LLM rankings that provides free access to frontier models and a personal leaderboard.