All major robotics and VLA papers, ranked and benchmarked in a single place [P]

Reddit r/MachineLearning Tools

Summary

A new dedicated Robotics page on Papers with Code aggregates major benchmarks, trending papers with linked code, and open-source artifacts, tracking progress over time across benchmarks like LIBERO and SimplerEnv.

Hi folks, There is now a dedicated Robotics page on Papers with Code that lists the major benchmarks, trending papers with linked code, and open-source artifacts. Find it here: https://paperswithcode.co/tasks/robotics https://preview.redd.it/wvhcgu1q1fdh1.png?width=3034&format=png&auto=webp&s=9a43021f7b8e1b840b09f55bb8aee0c4b2d54215 Major benchmarks that most papers report evaluations on are: - LIBERO, as well as its subsets like LIBERO-Long and LIBERO-Spatial - SimplerEnv WidowX - RoboTwin and more. Currently, we have about 110 entries on each benchmark. Each benchmark's progress is visualized over time: https://preview.redd.it/myl3ysvv2fdh1.png?width=2880&format=png&auto=webp&s=3fae2b73d592f2c7ddd3118ec2c5a0d5c5e93b65 We also show which models are open source and which aren't. Let me know which others I missed. I'd be happy to add them. Also happy to hear any feedback, new tasks, or features to add! Kind regards, Niels ML Engineer @ HF
Original Article

Similar Articles

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research

Hugging Face Daily Papers

PRL-Bench is a comprehensive benchmark for evaluating LLMs' capabilities in frontier physics research, constructed from 100 curated Physical Review Letters papers across five physics subfields. The benchmark reveals significant gaps in current LLM performance (best scores below 50%), designed to test end-to-end research workflows, complex reasoning, and autonomous exploration.

@VincentLogic: Drowning in new Arxiv papers every day? Head spinning. Just discovered a treasure trove of a website that aggregates the latest AI papers and model benchmarks. Clean interface, just check Trending or filter by week/month. Best part: each paper directly links to the benchmarks and models it uses.

X AI KOLs Timeline

Recommend a free website sophon.at/papers that aggregates the latest AI papers and model benchmarks. Clean interface, supports Trending or weekly/monthly filtering. Each paper directly links to its benchmarks and models.

Reviving PapersWithCode (by Hugging Face) [P]

Reddit r/MachineLearning

Niels from Hugging Face announces the revival of PapersWithCode as paperswithcode.co, a platform that parses high-impact AI papers at scale and automatically generates leaderboards and benchmarks, incorporating features like trending papers, domain categorization, and external paper support.