llm-benchmarks

Tag

Cards List
#llm-benchmarks

These LLMs are the best at resisting Russian propaganda

Ars Technica ↗ · 2026-06-04 Cached

A benchmark study by the Estonian Language Institute evaluates LLMs on their ability to resist Russian propaganda, finding that Nvidia's Nemotron, Alibaba's Qwen, and OpenAI's GPT-5.4 perform well, while Google's Gemini models show notable weaknesses, especially when prompted in Russian.

0 favorites 0 likes
#llm-benchmarks

Auditing LLM Benchmarks with Item Response Theory

arXiv cs.CL ↗ · 2026-06-01 Cached

This paper introduces an Item Response Theory-based method to detect mislabeled examples in LLM benchmarks at 95% precision, tracing errors to labeling heuristics and annotation issues.

0 favorites 0 likes
#llm-benchmarks

Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks

arXiv cs.AI ↗ · 2026-05-26 Cached

This paper identifies systemic measurement bias in production LLM inference benchmarks caused by single-process Python clients using asyncio, and proposes a multi-process evaluation framework and a new metric (NTPOT) to accurately profile serving engines at scale.

0 favorites 0 likes
#llm-benchmarks

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation

arXiv cs.AI ↗ · 2026-05-11 Cached

This paper introduces EnvSimBench, a benchmark for evaluating Large Language Models' ability to simulate environments for agent training. It identifies a 'state change cliff' in current LLMs and proposes a constraint-driven pipeline to reduce hallucinations and costs.

0 favorites 0 likes
#llm-benchmarks

@omarsar0: Cool paper from Apple. Most evaluation of tool-calling agents happens after the trajectory is over. By then the wrong c…

X AI KOLs Timeline ↗ · 2026-05-10 Cached

This Apple research paper introduces 'Reinforced Agent,' a method that moves evaluation into the execution loop using a specialized reviewer agent to correct tool-calling errors in real-time. It demonstrates significant accuracy improvements on benchmarks like BFCL and τ²-Bench without retraining the base agent.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback