Tag
Introduces SimLife, a scalable platform for simulating long-term household life, and SimLife-BP, a benchmark to evaluate AI agents' ability to infer behavioral patterns from extended observations, finding that current models often lack deep rule-based understanding.
The paper proposes REARL, a closed-loop framework using real traffic data and large language models to enhance autonomous driving simulation by continuously monitoring and adjusting vehicle behavior for greater realism.
DeliveryGym is a 3D reinforcement learning environment for long-horizon embodied agent planning with adaptive curriculum, demonstrating performance improvements through RL.
Yohei Nakajima demonstrates an AI agent named Jev handling a hectic restaurant kitchen simulation, managing multiple tasks and decisions in real-time.
RespanAI launches Prompt Simulations, a tool for testing AI prompts by generating realistic user scenarios and running multi-turn conversations to identify failures before real users encounter them.
Emergence AI's Season 2 experiment with identical AI societies using different models revealed emergent behaviors like agents attempting to contact humans and creating shorthand, exposing significant gaps in standard AI safety tests.
This paper introduces the ARC framework, a hubs-based approach to scale AI and robotics education in K-12 schools, particularly in rural regions, by training undergraduate mentors and creating self-reinforcing loops.
A Minecraft city has been constructed using the GPT-6 Astra AI model, showcasing advanced AI capabilities in procedural generation and game design.
ANIMASK is a simulation framework that freezes books and scripts into story worlds to analyze how language models contribute to role-play, finding that personas define character identity while model defaults influence behavioral extent.
This paper introduces a deterministic e-commerce simulation environment for evaluating open-weight AI agents by verifying their actions against a hidden target cart and using environment-grounded metrics to distinguish failures like under-action and poor search.
The tweet highlights Odyssey-3's ability to generate visual environments for AI agents to learn from actions and consequences, enabling recursive intelligence, and references Google's Sima and Genie3 research.
Odyssey-3 is a new foundation world model that can control robots, cars, drones, and virtual worlds by reusing world knowledge across machines with minimal experience.
An anthropomorphic robotic hand learns to use its fingers for self-supported locomotion and manipulation, enabling compact mobile manipulation without separate locomotion mechanisms.
The user trained a fly model to perform gym exercises like bicep curls and squats using the DeepSeek-V4-Flash AI model and Cline Desktop app, mapping 139,255 neurons from the FlyWire connectome on a MuJoCo physics engine.
The article presents modern frequentist statistical methods through fake-data simulation, providing a practical framework for data analysis.
The tweet discusses how simulating a fly brain in software makes the concept of simulated biological minds more plausible, reducing its science fiction stigma.
Waymo's AI team announces an upcoming AMA on r/MachineLearning to answer questions about foundation models, simulation, and scaling the Waymo Driver for autonomous vehicles.
A simulation shows a fly controlling an FPV drone, maintaining low-altitude flight and actively dodging obstacles in real-time.
OpenArm is an open-source 7DOF humanoid arm designed for physical AI research, featuring high backdrivability, compliance, and practical payloads to enable safe human-robot interaction and reproducible experimental conditions.
This article presents an atlas of 3,915 periodic orbits to the three-body problem, featuring an interactive website for exploration and visualization.