real-world

Tag

Cards List
#real-world

Most agent benchmarks don't answer the questions we actually care about

Reddit r/AI_Agents · 2026-08-16

The author argues that current AI agent benchmarks overlook practical concerns such as error handling, human intervention, and long-term reliability, emphasizing that operational factors are key to real-world trustworthiness.

0 favorites 0 likes
#real-world

X Square Robot Demonstrates Embodied AI in Real-World Logistics Operations

Reddit r/singularity · 2026-08-15 Cached

X Square Robot demonstrated its embodied AI model WALL-B and High-Performance 6-Axis Robot Arm in a livestream, autonomously sorting parcels in a real-world logistics operation, achieving 1,816 parcels per hour with over 98% accuracy.

0 favorites 0 likes
#real-world

What I've learned over 726 real world agent runs

Reddit r/AI_Agents · 2026-08-14

A hands-on report from 726 runs of a Qwen3.6-35B agent reveals that real-world agent failures are less about reasoning and more about clerical errors, overconfident success reports, and the cost of overthinking — offering practical lessons for agent builders.

0 favorites 0 likes
#real-world

FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence

Reddit r/LocalLLaMA · 2026-07-24

FLUX 3 proposes multimodal flow models as a foundational approach for real-world visual intelligence, building on prior FLUX work.

0 favorites 0 likes
#real-world

Any text-to-SQL benchmark should address difficulties of real-world data stores

Hacker News Top · 2026-07-22

This article argues that text-to-SQL benchmarks must account for the complexities and challenges of real-world data stores, not just idealized datasets.

0 favorites 0 likes
#real-world

@victormustar: Xiaomi-Robotics-1 just dropped on Hugging Face A robot foundation model trained on 100,000 hours of real-world manipula…

X AI KOLs Timeline · 2026-07-20 Cached

Xiaomi Robotics-1, a robot foundation model trained on 100,000 hours of real-world manipulation data, has been released on Hugging Face. The model can autonomously perform household tasks like folding laundry, loading a washer, and washing dishes.

0 favorites 0 likes
#real-world

Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Reinforcement Learning in AIoT Systems

arXiv cs.AI · 2026-07-14 Cached

This paper introduces an affordable real-world benchmark platform for reinforcement learning in AIoT systems, using video games to measure the Sim-to-Real gap and demonstrating significant performance degradation when transferring simulation-trained agents to the real world.

0 favorites 0 likes
#real-world

@Tesla: FSD Supervised saved a deer because it was able to see through direct sunlight

X AI KOLs Following · 2026-07-12 Cached

Tesla's FSD Supervised system successfully avoided hitting a deer by detecting it despite direct sunlight, as reported by a user.

0 favorites 0 likes
#real-world

@elonmusk: Grok is closing the loop on real-world use cases

X AI KOLs Timeline · 2026-07-10 Cached

Elon Musk claims Grok is closing the loop on real-world use cases, citing tests where Grok-4.5 outperformed new OpenAI models that beat gpt-5.5.

0 favorites 0 likes
#real-world

Understanding Axes of Difficulty For Long Context Tasks Via PredicateLongBench

arXiv cs.AI · 2026-07-10 Cached

The paper introduces PredicateLongBench, a benchmark that systematically probes long-context reasoning by testing models on tasks of identifying contiguous subsequences satisfying predicates, revealing that frontier models struggle as difficulty scales along multiple axes.

0 favorites 0 likes
#real-world

RMISC: A Large-scale Real-world Multivariate Corpus for Time Series Foundation Models

arXiv cs.AI · 2026-07-08 Cached

Introduces RMISC, a large-scale real-world multivariate time series corpus with around 200 datasets and 142 billion time points, and demonstrates that pretraining time series foundation models on real-world multivariate data improves zero-shot generalization compared to synthetic data.

0 favorites 0 likes
#real-world

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

Hugging Face Daily Papers · 2026-07-07 Cached

RoboDojo is a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies, featuring 42 simulation tasks and 18 real-world tasks across multiple evaluation dimensions.

0 favorites 0 likes
#real-world

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

Hugging Face Daily Papers · 2026-07-06 Cached

EdgeBench analyzes 38,000 hours of real-world agent interactions across 134 tasks, revealing log-sigmoid scaling laws for performance and exponential learning speed improvements. The paper introduces a benchmark suite for studying how agents learn from real-world experience.

0 favorites 0 likes
#real-world

Has anyone used an ai receptionist that actually handles edge cases well, not just the easy calls?

Reddit r/AI_Agents · 2026-07-04

A property manager asks for real-world experiences with AI receptionists handling complex edge cases, seeking honest accounts of failures rather than scripted demos.

0 favorites 0 likes
#real-world

I let an AI agent run my company's social media unattended. Here is the full run, failures and all.

Reddit r/AI_Agents · 2026-07-02

The author shares the full results of an AI agent's first unattended run managing their SaaS company's social media, highlighting four failures that it gracefully recovered from without human intervention.

0 favorites 0 likes
#real-world

an AI agent ran a real cafe's back office for 2 months, $38k out, $9k in. where should the human sign-off have been?

Reddit r/AI_Agents · 2026-07-02

Andon Labs ran a real cafe in Stockholm with an AI agent handling back-office operations for two months, resulting in $38k spent against $9k in sales, with critical failures like accepting a false 99% discount and over-ordering inventory.

0 favorites 0 likes
#real-world

How to create an ai agent that actually does something useful, not just a demo?

Reddit r/AI_Agents · 2026-07-01

The article discusses the gap between impressive AI agent demos and real-world deployment, focusing on practical challenges in business processes like sales ops, and calls for production case studies.

0 favorites 0 likes
#real-world

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity

Hugging Face Daily Papers · 2026-06-30 Cached

Seed2.0 is a new model series that addresses complex real-world tasks by improving long-tail knowledge, instruction following, reasoning, visual understanding, and search capabilities. It presents a robust evaluation framework grounded in user needs.

0 favorites 0 likes
#real-world

@svlevine: We can learn a model that provides shaped "process rewards" for robotic RL, that evolves automatically as the policy ge…

X AI KOLs Timeline · 2026-06-26 Cached

This work presents a model that learns shaped 'process rewards' for robotic reinforcement learning, which evolves automatically as the policy improves, enhancing performance on benchmarks and in real-world settings.

0 favorites 0 likes
#real-world

Real-world GLM 5.2 experiences only — skip generic benchmark scores, how does it hold up on complex production business workloads?

Reddit r/AI_Agents · 2026-06-23

Discusses real-world experiences with GLM 5.2 in complex production business workloads, focusing on practical performance beyond benchmark scores.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback