real-world

Tag

Cards List
#real-world

Nvidia's Autonomous Robotics Research (6 minute read)

TLDR AI · 2026-06-22 Cached

ENPIRE is a framework that enables coding agents to autonomously improve robot manipulation policies through a real-world feedback loop, achieving 99% success on dexterous tasks like pin insertion and zip tie cutting.

0 favorites 0 likes
#real-world

Llama bench and real performance wayy different(Help)

Reddit r/LocalLLaMA · 2026-06-18

Discussion about the significant gap between Llama model benchmark scores and actual real-world performance, with the author seeking assistance.

0 favorites 0 likes
#real-world

ENPIRE: Agentic Robot Policy Self-Improvement in the Real World

Hugging Face Daily Papers · 2026-06-18 Cached

ENPIRE is a framework that enables autonomous robot policy self-improvement in the real world through a closed-loop system of environment feedback, policy refinement, and evolutionary code optimization, achieving 99% success on dexterous manipulation tasks.

0 favorites 0 likes
#real-world

@FinanceYF5: ENPIRE can now independently perform high-precision operations such as zip-tying, sorting fine needles, and installing GPUs, and has demonstrated a 'physical scaling' phenomenon: multiple robots exploring in parallel, with significantly faster progress. Part of the NVIDIA GEAR lab can now self-improve overnight, with humans only needing to review reports in the morning. The project will also be open-sourced. It...

X AI KOLs Following · 2026-06-17 Cached

NVIDIA GEAR lab introduces ENPIRE, a framework for autonomous real-world robot policy self-improvement that achieves 99% success on dexterous manipulation tasks like GPU insertion and zip-tying, with multi-robot parallel learning and open-source release.

0 favorites 0 likes
#real-world

@Murderlon: FrontierCode finally dropped, a coding agents benchmark for the real world. Human-verified through an extensive hardeni…

X AI KOLs Following · 2026-06-08 Cached

FrontierCode is a new benchmark for coding agents, human-verified with a continuous scoring model, designed to evaluate real-world performance.

0 favorites 0 likes
#real-world

@mdancho84: RIP toy projects. If your portfolio doesn’t touch real business problems, you’ll get filtered out. Here are 300+ real M…

X AI KOLs Timeline · 2026-06-08 Cached

This tweet promotes a free collection of over 300 real ML system case studies from top companies, arguing that toy projects are insufficient for building a strong portfolio.

0 favorites 0 likes
#real-world

What's the most useful AI agent you've seen in production?

Reddit r/AI_Agents · 2026-06-08

A discussion about the most useful AI agents actually deployed in production, highlighting simple, single-problem solutions like lead qualification and support triage.

0 favorites 0 likes
#real-world

@rohanpaul_ai: Arena just released a real-world agent leaderboard that ranks AI models by how well they complete actual user jobs, not…

X AI KOLs Following · 2026-06-05 Cached

Agent Arena is a new leaderboard that evaluates AI models on real-world agentic tasks such as coding, research, and file analysis, using signals like task success, steerability, and recovery, with GPT-5.5 High leading.

0 favorites 0 likes
#real-world

Are AI agents finally crossing the line from demos to real tools?

Reddit r/AI_Agents · 2026-06-05

Discussion on whether AI agents are transitioning from impressive demos to genuinely useful tools in research, coding, operations, and personal productivity.

0 favorites 0 likes
#real-world

6 weeks daily-driving an open-source desktop agent shell with a 3-model split (Haiku triager → Sonnet reviewer → Opus executor). Real cost numbers + what broke.

Reddit r/AI_Agents · 2026-06-05

A 6-week real-world experiment using an open-source desktop agent shell with a three-model split (Haiku triager, Sonnet reviewer, Opus executor) reports a 64% cost reduction and details failure modes like context bloat and runaway sub-agents.

0 favorites 0 likes
#real-world

Is anyone interested in seeing how advanced companies are actually running agents in production?

Reddit r/AI_Agents · 2026-05-26

The author, working at an AI infrastructure company, observes that running AI agents in production is less about the model and more about environment, access control, isolation, and safe state management, and asks if the community wants detailed architecture patterns.

0 favorites 0 likes
#real-world

I've built 50+ AI automations for clients, here's why most fail and what the working ones got right

Reddit r/AI_Agents · 2026-05-26

An agency founder shares lessons from 50+ AI automation implementations, highlighting that most fail due to broken underlying processes, lack of internal ownership, and over-engineering, while the most successful automations are simple, focused, and backed by a named client-side owner.

0 favorites 0 likes
#real-world

Apex-Testing: real-world, real repos, agentic coding benchmark (Update)

Reddit r/LocalLLaMA · 2026-05-23

Apex-Testing, a benchmark for evaluating agentic coding models using real private GitHub repositories, has been updated with recent models and detailed metrics including cost, time, and ELO-based leaderboard.

0 favorites 0 likes
#real-world

TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks

Hugging Face Daily Papers · 2026-05-21 Cached

This paper introduces TerminalWorld, a benchmark for evaluating AI agents on real-world terminal tasks, derived from 80,870 terminal recordings. Current systems achieve at most 62.5% pass rate, highlighting challenges in authentic terminal workflows.

0 favorites 0 likes
#real-world

Anyone else feel like AI agents are amazing right up until things get complicated?

Reddit r/AI_Agents · 2026-05-20

A reflection on the gap between impressive AI agent demos and dependable real-world execution, arguing that current agents excel at structured tasks but fail under unpredictable conditions, suggesting near-term AI roles will focus on narrow automation with human oversight.

0 favorites 0 likes
#real-world

AI agents feel impressive until the workflow gets messy

Reddit r/AI_Agents · 2026-05-19

A reflection on AI agents: impressive for narrow supervised tasks but fragile and unreliable in long-running, messy workflows due to issues like session expiration, context drift, and silent failures.

0 favorites 0 likes
#real-world

@cyrilXBT: ANTHROPIC JUST KILLED THE DEMO AGENT ERA. Their Agents team showed exactly what production grade looks like. Not theory…

X AI KOLs Timeline · 2026-05-19 Cached

Anthropic's Agents team unveiled a production-grade four-layer framework for multi-agent systems during a 30-minute presentation, marking a shift from demo to real-world applications.

0 favorites 0 likes
#real-world

Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

Hugging Face Daily Papers · 2026-05-19 Cached

Mega-ASR proposes scaling up real-world acoustic simulation to improve automatic speech recognition in challenging, wild conditions, aiming to narrow the performance gap between lab and real-world settings.

0 favorites 0 likes
#real-world

DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text Detection

arXiv cs.CL · 2026-05-18 Cached

DetectRL-X is a comprehensive multilingual benchmark for evaluating LLM-generated text detectors across 8 languages and 6 domains, including stress testing with AI-assisted writing operations and perturbations. It reveals strengths and limitations of current detectors in multilingual scenarios.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback