@drfeifei: 1/N Long horizon, complex tasks that truly matter in everyday life are not solved problems by today’s robotics, requiri…
Summary
Dr. Fei-Fei Li announces the second year of Stanford's BEHAVIOR Challenge, a robotics competition tackling long-horizon, complex everyday tasks, with new tasks, improved evaluation, and an $11,000 prize pool.
View Cached Full Text
Cached at: 07/13/26, 10:01 PM
1/N Long horizon, complex tasks that truly matter in everyday life are not solved problems by today’s robotics, requiring planning, object detection, object manipulation, and failure recovery.
That’s why Stanford’s BEHAVIOR Challenge is back for year 2! Last year, the winning solution reached only 12.4% full task success. This year, the BEHAVIOR challenge has more tasks, better evaluation, and is easier to use.
Submission deadline: 10/16/2026 Winners announced: 11/04/2026 Prize pool: $11,000
2/N Real-world robot evaluation is essential but hard to scale: experiments are difficult to control, reproduce, and compare. Simulation is a powerful testbed for scalable, controlled, reproducible evaluation.
BEHAVIOR-1K is an open-source simulation benchmark of 1,000 everyday household activities requiring long-horizon reasoning, navigation, and bimanual manipulation, giving us a scalable way to measure how well robot AI models generalize.
What’s new in the 2026 BEHAVIOR Challenge?
- Double the tasks, double the challenge.
We doubled the benchmark from 50 to 100 long-horizon household tasks. These activities average 6 minutes each, requiring navigation, planning, memory, and bimanual coordination. No other robotics benchmark comes close in terms of difficulty.
- Larger dataset, better baselines
• 20,000 human teleoperation demos, 1950 hours in total (2 times larger than 2025) • RGBD observations and Robot proprioception • Skill/subtask annotations • Strong baseline support: pi0.5, GR00T N1.7
5/N Evaluation & Submission
To better reflect real-world deployment, the BEHAVIOR Challenge has one official track this year using only robot onboard observations:
• RGB • Depth • Proprioception
Submission instructions and evaluation details are available here: https://behavior.stanford.edu/challenge/
6/N Together, let’s ask:
Can current models solve complete human-centered household tasks? How should agents combine control, memory, and planning? Where do today’s models fail to generalize? What actually scales in embodied AI?
7/N Join the BEHAVIOR Discord server to ask questions and discuss:
https://discord.gg/bccR5vGFEx
We will also hold office hours every Monday, 5–6pm PST over Zoom. See the website for the link.
Whether you’re a robotics veteran or just entering the field, we’re here to support you.
8/N Proud of the amazing work from our students and collaborators, led by @drfeifei:
@wensi_ai @stefyfren @cgokmenAI @yalcintur36 @minyeongkim_ @BrndaHere2Chl @AndiXu1111
@RavenHuang4 @RuohanZhang76 @jiajunwu_cs
9/N And with the strong support from @EvansXuHan @jin_lynn808 @Hang_Yin_ @ChengshuEricLi @josiah_is_wong @sanjana__z @YunfanJiang @wenlong_huang
@RobobertoMM @YunzhuLiYZ @ManlingLi_ @Weiyu_Liu_ @silviocinguetta @hyogweon Prof. Karen Liu
10/N We thank @SimovationInc for providing high-quality JoyLo teleoperation data in simulation for the BEHAVIOR dataset.
BEHAVIOR is built upon @nvidia Omniverse. We thank @nvidia for their continuous support.
11/N We thank our sponsors and supporters for their generous support.
@SimovationInc @IMDAsg @StanfordHAI @SchmidtFutures Calder Inc.
Similar Articles
@rohanpaul_ai: Dr Fei-Fei-Li (@drfeifei ) explains why and how everyday household chores are so extremely difficult for Robots. "If yo…
Dr. Fei-Fei Li discusses the challenges robots face in understanding and executing everyday household tasks, highlighting the difficulty of grounding natural language instructions like 'open the drawer while avoiding the vase' into robot actions.
@drfeifei: I’m very excited by this test time training work for robotic learning! It’s an awesome collaboration between @StanfordS…
Fei-Fei Li highlights a new test-time training approach for robotic learning, developed in collaboration between Stanford SVL and NVIDIA Robotics, which scales robot model context to 8000 timesteps with constant inference cost.
@rohanpaul_ai: Fei-Fei Li ( @drfeifei ) beautifully explains Robotics. She defines robotics not by form, like humanoids or cars, but b…
Fei-Fei Li explains robotics as embodied machines requiring spatial intelligence, and discusses how 3D generation technologies enable creating infinite digital worlds, unlocking a multiverse for creativity, training, and storytelling.
@jietang: Recent thoughts: The Shift to Long-Horizon Tasks The most likely breakthrough this year will be in long-horizon tasks. …
The article discusses the anticipated breakthrough in long-horizon AI tasks and autonomous agents, suggesting a shift from 'one-person' to 'none-person' companies. It highlights technical pillars like memory, continual learning, and self-judging as key to realizing fully self-evolving AI systems that could redefine AGI and operating systems.
@drfeifei: https://x.com/drfeifei/status/2062247238143996275
Fei-Fei Li and the World Labs team present a functional taxonomy of world models, distinguishing between renderers, physics engines, and other components within the reinforcement learning loop, and arguing that spatial intelligence is AI's next frontier.