@RemiCadene: Pretty cool :)
Summary
A tweet highlighting that benchmarking robot policies is broken and sharing results from thousands of evaluations over 12 manipulation tasks to determine which policy to use.
View Cached Full Text
Cached at: 07/29/26, 10:01 AM
Pretty cool :)
Ville🤖 (@VilleKuosmanen): Benchmarking robot policies is broken
We ran thousands of evaluations over 12 different manipulation tasks to evaluate which robot policy you should use.
Similar Articles
RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
RoboDojo is a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies, featuring 42 simulation tasks and 18 real-world tasks across multiple evaluation dimensions.
RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
RoboLab is a high-fidelity simulation benchmarking framework for evaluating task-generalist robotic policies, introducing the RoboLab-120 benchmark with 120 tasks across visual, procedural, and relational competency axes. It enables scalable, realistic task generation and systematic analysis of policy behavior under controlled perturbations to assess true generalization capabilities.
@_philschmid: https://x.com/_philschmid/status/2081744861829414977
EvoCode-Bench is a multi-turn coding benchmark with 26 tasks across 5 domains, designed to evaluate AI agents on evolving specifications and cumulative testing in a persistent workspace, revealing that single-turn scores dramatically overstate reliability.
@_lamaahmad: We (@CedricWhitney, @SandhiniAgarwal, @EstherTetruas, @OliviaGWatkins2, @dgrobinson) wrote about nuances we’ve observed…
OpenAI researchers share lessons learned from working with third parties on frontier model evaluations, highlighting the importance of considering the evaluation harness and potential validity issues like reward hacking, contamination, and sandbagging.
@KLieret: You can evaluate on ProgramBench yourself: https://github.com/facebookresearch/ProgramBench/… We will open the leaderbo…
ProgramBench is a new benchmark that tests AI agents' ability to reconstruct a complete codebase from a compiled binary and its documentation. The leaderboard will open for submissions soon.