llm-ranking

标签

Cards List
#llm-ranking

Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval

arXiv cs.AI ↗ · 2026-08-13 缓存

This paper compares natively multimodal embedding models (Gemini Embedding 2, Amazon Nova 2) against frontier LLMs (GPT-4.1, Claude Sonnet 4.6) for hard-negative text-to-image retrieval, finding comparable accuracy but much lower latency for embedding-based ranking.

0 人收藏 0 人点赞
#llm-ranking

为什么 artificialanalysis.ai 在 SciCode 上将 Gemma4 排在 Qwen3.6 27b 之上

Reddit r/LocalLLaMA ↗ · 2026-08-06

关于 Artificial Analysis 如何在 SciCode 基准测试中将 Gemma 4 排在 Qwen3.6 27b 之上的讨论,质疑该排名是否反映现实世界的编码能力,还是揭示了基准测试的问题。

0 人收藏 0 人点赞
#llm-ranking

语言模型的当前状态与基于人类偏好的排名 [R]

Reddit r/MachineLearning ↗ · 2026-08-06

马克斯·普朗克智能系统研究所推出了 Comparity AI,一个基于人类偏好进行 LLM 排名的研究平台,提供免费访问前沿模型和个人排行榜的服务。

0 人收藏 0 人点赞
← 返回首页

提交意见反馈