Tag
This paper introduces an affordable real-world benchmark platform for reinforcement learning in AIoT systems, using video games to measure the Sim-to-Real gap and demonstrating significant performance degradation when transferring simulation-trained agents to the real world.
Tesla's FSD Supervised system successfully avoided hitting a deer by detecting it despite direct sunlight, as reported by a user.
Elon Musk claims Grok is closing the loop on real-world use cases, citing tests where Grok-4.5 outperformed new OpenAI models that beat gpt-5.5.
The paper introduces PredicateLongBench, a benchmark that systematically probes long-context reasoning by testing models on tasks of identifying contiguous subsequences satisfying predicates, revealing that frontier models struggle as difficulty scales along multiple axes.
Introduces RMISC, a large-scale real-world multivariate time series corpus with around 200 datasets and 142 billion time points, and demonstrates that pretraining time series foundation models on real-world multivariate data improves zero-shot generalization compared to synthetic data.
RoboDojo is a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies, featuring 42 simulation tasks and 18 real-world tasks across multiple evaluation dimensions.
EdgeBench analyzes 38,000 hours of real-world agent interactions across 134 tasks, revealing log-sigmoid scaling laws for performance and exponential learning speed improvements. The paper introduces a benchmark suite for studying how agents learn from real-world experience.
A property manager asks for real-world experiences with AI receptionists handling complex edge cases, seeking honest accounts of failures rather than scripted demos.
The author shares the full results of an AI agent's first unattended run managing their SaaS company's social media, highlighting four failures that it gracefully recovered from without human intervention.
Andon Labs ran a real cafe in Stockholm with an AI agent handling back-office operations for two months, resulting in $38k spent against $9k in sales, with critical failures like accepting a false 99% discount and over-ordering inventory.
The article discusses the gap between impressive AI agent demos and real-world deployment, focusing on practical challenges in business processes like sales ops, and calls for production case studies.
Seed2.0 is a new model series that addresses complex real-world tasks by improving long-tail knowledge, instruction following, reasoning, visual understanding, and search capabilities. It presents a robust evaluation framework grounded in user needs.
This work presents a model that learns shaped 'process rewards' for robotic reinforcement learning, which evolves automatically as the policy improves, enhancing performance on benchmarks and in real-world settings.
Discusses real-world experiences with GLM 5.2 in complex production business workloads, focusing on practical performance beyond benchmark scores.
ENPIRE is a framework that enables coding agents to autonomously improve robot manipulation policies through a real-world feedback loop, achieving 99% success on dexterous tasks like pin insertion and zip tie cutting.
Discussion about the significant gap between Llama model benchmark scores and actual real-world performance, with the author seeking assistance.
ENPIRE is a framework that enables autonomous robot policy self-improvement in the real world through a closed-loop system of environment feedback, policy refinement, and evolutionary code optimization, achieving 99% success on dexterous manipulation tasks.
NVIDIA GEAR lab introduces ENPIRE, a framework for autonomous real-world robot policy self-improvement that achieves 99% success on dexterous manipulation tasks like GPU insertion and zip-tying, with multi-robot parallel learning and open-source release.
FrontierCode is a new benchmark for coding agents, human-verified with a continuous scoring model, designed to evaluate real-world performance.
This tweet promotes a free collection of over 300 real ML system case studies from top companies, arguing that toy projects are insufficient for building a strong portfolio.