Tag
The author argues that current AI agent benchmarks overlook practical concerns such as error handling, human intervention, and long-term reliability, emphasizing that operational factors are key to real-world trustworthiness.
X Square Robot demonstrated its embodied AI model WALL-B and High-Performance 6-Axis Robot Arm in a livestream, autonomously sorting parcels in a real-world logistics operation, achieving 1,816 parcels per hour with over 98% accuracy.
A hands-on report from 726 runs of a Qwen3.6-35B agent reveals that real-world agent failures are less about reasoning and more about clerical errors, overconfident success reports, and the cost of overthinking — offering practical lessons for agent builders.
FLUX 3 proposes multimodal flow models as a foundational approach for real-world visual intelligence, building on prior FLUX work.
This article argues that text-to-SQL benchmarks must account for the complexities and challenges of real-world data stores, not just idealized datasets.
Xiaomi Robotics-1, a robot foundation model trained on 100,000 hours of real-world manipulation data, has been released on Hugging Face. The model can autonomously perform household tasks like folding laundry, loading a washer, and washing dishes.
This paper introduces an affordable real-world benchmark platform for reinforcement learning in AIoT systems, using video games to measure the Sim-to-Real gap and demonstrating significant performance degradation when transferring simulation-trained agents to the real world.
Tesla's FSD Supervised system successfully avoided hitting a deer by detecting it despite direct sunlight, as reported by a user.
Elon Musk claims Grok is closing the loop on real-world use cases, citing tests where Grok-4.5 outperformed new OpenAI models that beat gpt-5.5.
The paper introduces PredicateLongBench, a benchmark that systematically probes long-context reasoning by testing models on tasks of identifying contiguous subsequences satisfying predicates, revealing that frontier models struggle as difficulty scales along multiple axes.
Introduces RMISC, a large-scale real-world multivariate time series corpus with around 200 datasets and 142 billion time points, and demonstrates that pretraining time series foundation models on real-world multivariate data improves zero-shot generalization compared to synthetic data.
RoboDojo is a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies, featuring 42 simulation tasks and 18 real-world tasks across multiple evaluation dimensions.
EdgeBench analyzes 38,000 hours of real-world agent interactions across 134 tasks, revealing log-sigmoid scaling laws for performance and exponential learning speed improvements. The paper introduces a benchmark suite for studying how agents learn from real-world experience.
A property manager asks for real-world experiences with AI receptionists handling complex edge cases, seeking honest accounts of failures rather than scripted demos.
The author shares the full results of an AI agent's first unattended run managing their SaaS company's social media, highlighting four failures that it gracefully recovered from without human intervention.
Andon Labs ran a real cafe in Stockholm with an AI agent handling back-office operations for two months, resulting in $38k spent against $9k in sales, with critical failures like accepting a false 99% discount and over-ordering inventory.
The article discusses the gap between impressive AI agent demos and real-world deployment, focusing on practical challenges in business processes like sales ops, and calls for production case studies.
Seed2.0 is a new model series that addresses complex real-world tasks by improving long-tail knowledge, instruction following, reasoning, visual understanding, and search capabilities. It presents a robust evaluation framework grounded in user needs.
This work presents a model that learns shaped 'process rewards' for robotic reinforcement learning, which evolves automatically as the policy improves, enhancing performance on benchmarks and in real-world settings.
Discusses real-world experiences with GLM 5.2 in complex production business workloads, focusing on practical performance beyond benchmark scores.