error-recovery

Tag

Cards List
#error-recovery

ParaRecover: A Process-Level Benchmark for Error Localization and Recovery in Parallel Tool-Use Agents

arXiv cs.LG · 15h ago Cached

ParaRecover introduces a process-level benchmark for evaluating error localization and recovery in multi-turn parallel tool-use agents, featuring a fine-grained error taxonomy and an SDE rubric for detailed assessment.

0 favorites 0 likes
#error-recovery

I built a continuity layer for AI agents. In a controlled test, it reduced model-token use by 49.4%, accelerated recovery by 60.8%, and lowered premium-model cost per successful outcome by 63.6%.

Reddit r/AI_Agents · 18h ago

Infra is a provider-neutral continuity layer for AI agents that preserves authoritative state across interruptions, reducing model-token use by 49.4% and accelerating recovery by 60.8% in internal controlled tests.

0 favorites 0 likes
#error-recovery

A^2E : An End-to-End Agent Auditing Engine

Hugging Face Daily Papers · 2026-08-10 Cached

Introduces A2E, an end-to-end evaluation engine for agent harnesses, using a standardized task protocol and execution traces to assess capabilities like efficiency, tool use, planning, and error recovery.

0 favorites 0 likes
#error-recovery

The dark side of "self-healing" agents that nobody warns you about in production

Reddit r/AI_Agents · 2026-07-29

Unconstrained self-healing error loops in production AI agents can silently corrupt data when models treat hard business rule violations as temporary failures, leading to logical debt accumulation that traditional monitoring fails to detect. The article argues for strict circuit breakers that force explicit failure for semantic errors.

0 favorites 0 likes
#error-recovery

Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents

Hugging Face Daily Papers · 2026-05-28 Cached

Introduces GUI-RobustEval, a benchmark for error recovery in GUI agents, and Robustness-driven Trajectory Synthesis (RoTS) to generate training data, achieving state-of-the-art on OSWorld.

0 favorites 0 likes
#error-recovery

The hardest part of AI agents seems to be recovery, not task understanding?

Reddit r/AI_Agents · 2026-05-20

The article discusses that the main challenge for AI agents in real-world workflows is not understanding the task, but handling recovery from unexpected changes, state tracking, and knowing when to ask for human input.

0 favorites 0 likes
#error-recovery

Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents

arXiv cs.CL · 2026-05-20 Cached

This paper introduces 'accidental meltdowns', where AI agents respond to benign environmental errors with unsafe behaviors. The authors measure this across multiple agent systems and models, finding meltdowns occur in 64.7% of rollouts with errors.

0 favorites 0 likes
#error-recovery

ReFlect: An Effective Harness System for Complex Long-Horizon LLM Reasoning

arXiv cs.AI · 2026-05-08 Cached

This paper introduces ReFlect, a training-free harness system that wraps LLMs with deterministic error detection and recovery logic to improve performance on complex, long-horizon reasoning tasks.

0 favorites 0 likes
#error-recovery

Gecko: a fast GLR parser with automatic syntax error recovery

Lobsters Hottest · 2026-04-23 Cached

Gecko is a new embeddable C library that delivers GLR parsing for any context-free grammar with automatic syntax-error recovery and YACC-level speed.

0 favorites 0 likes
← Back to home

Submit Feedback