failure-detection

Tag

Cards List
#failure-detection

Detecting Neural Network Failures through Spectral Analysis of Internal Activations

arXiv cs.LG · 2d ago Cached

This paper identifies spectral drift in internal activations of neural networks during misclassifications and introduces Self-Detecting Neural Networks (SDNN) that monitors spectral dynamics to detect failures, achieving 79% AUROC on CIFAR-10, outperforming confidence-based methods by 25-30 percentage points.

0 favorites 0 likes
#failure-detection

Adaptive Two-Stage Online Learning for Service-Affecting Failure Detection in Mobile Core Networks

arXiv cs.LG · 4d ago Cached

This paper proposes a two-stage online learning framework for detecting service-affecting failures in mobile core networks by modeling normal traffic dynamics and analyzing residuals, achieving improved precision-recall trade-off over static thresholds.

0 favorites 0 likes
#failure-detection

Why your agents "succeed" and then you find out three days later they didn't

Reddit r/AI_Agents · 2026-07-13

Discusses the phenomenon where AI agents appear to succeed at tasks but later reveal failures, highlighting challenges in agent evaluation and monitoring.

0 favorites 0 likes
#failure-detection

Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade

arXiv cs.AI · 2026-07-08 Cached

This paper proposes a recall-controlled abort cascade that uses lightweight probes on LLM agent internal representations to detect and abort doomed episodes early, saving up to 47% inference compute while maintaining high recall of successful episodes.

0 favorites 0 likes
#failure-detection

Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents

Hugging Face Daily Papers · 2026-06-22 Cached

Foresight is a failure detection framework for long-horizon robotic manipulation that uses action-conditioned world model latents and functional conformal prediction to monitor trajectories, trained only with final task labels. It demonstrates state-of-the-art performance across simulation and real robot tasks.

0 favorites 0 likes
#failure-detection

I asked how you all handle agent memory. Here's the pattern in the replies, and the one thing nobody's actually solved.

Reddit r/AI_Agents · 2026-06-09

A community discussion on agent memory reveals that while various patches exist for what to write down (e.g., plain files, layered memory, post-mortems), the unsolved problem is what to keep—detecting failures is tractable, but deciding which lessons persist still needs human judgment.

0 favorites 0 likes
#failure-detection

How Language Models Fail: Token-Level Signatures of Committed and Persistent Reasoning Failures

arXiv cs.CL · 2026-06-08 Cached

This paper characterizes two distinct processes by which language models fail in reasoning—committed failure and persistent uncertainty—using token-level uncertainty signals, and demonstrates implications for self-consistency and failure detection strategies.

0 favorites 0 likes
#failure-detection

AEGIS: A Backup Reflex for Physical AI

arXiv cs.AI · 2026-06-08 Cached

AEGIS uses activation-probe early warning to switch to a stronger policy before failures compound in long-horizon robot manipulation, recovering twice as many failures as budget-matched escalation.

0 favorites 0 likes
#failure-detection

Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring

Hugging Face Daily Papers · 2026-05-29 Cached

Hide-and-Seek is a framework that detects robot execution failures in VLA models by localizing failure-indicative actions through contrastive learning without step-level annotations, achieving state-of-the-art multi-task failure detection.

0 favorites 0 likes
#failure-detection

Lost in the Folds: When Cross-Validation Is Not a Deep Ensemble for Uncertainty Estimation

Hugging Face Daily Papers · 2026-05-18 Cached

This paper compares cross-validation ensembles to deep ensembles for uncertainty estimation in medical image segmentation. Deep ensembles outperform cross-validation ensembles in calibration and failure detection, while cross-validation ensembles better approximate inter-rater variability.

0 favorites 0 likes
← Back to home

Submit Feedback