Tag
The author reflects on evaluating outbound calling APIs, emphasizing operational failure modes and structured outcomes over basic demos, with mentions of projects like voygr and placeCall.
This paper applies Andrej Karpathy's AutoResearch paradigm to production-scale machine learning, identifying five recurring failure modes and proposing a multi-agent framework that achieves significant performance improvements over hand-tuned baselines.
This paper introduces a benchmark for evaluating LLM agents as forward-deployed engineers in post-training delivery, highlighting the critical 'trains but does not learn' failure mode where models optimize without actual learning.
This paper presents a risk-sensitive evaluation framework for LLM-generated contract clauses, focusing on legal failure modes and quality dimensions to assess risks beyond accuracy or fluency.
The article discusses the unique failure modes of AI agents compared to scripts, emphasizing the lack of accountability when they make mistakes and arguing that deploying AI in critical systems without proper diagnostics is indefensible.
This paper evaluates neural, workflow, and agentic systems for ICD-10-CM coding, identifies failure modes on rare and complex codes, and shows that tool-augmented agentic configurations can recover performance on specific subsets.
The article discusses potential failure modes in hierarchical multi-agent systems where AI agents manage other AI agents, asking about earliest breakdowns and concerns as systems grow more complex.
A developer discusses the lack of suitable observability tools for AI agents, expressing disappointment with existing solutions like Opik and hoping for a service that supports OpenTelemetry for analyzing agent sessions and failure modes.
CritICL is an inference-time framework that enhances large language model reasoning by leveraging structured failure patterns from smaller models as critique-based guidance, outperforming standard in-context learning with reduced generation and token costs.
This article compares how LangSmith, Langfuse, and Phoenix handle common AI agent failure modes, such as wrong tool calls and format drift, and introduces Future AGI as a tool with integrated guardrails and gateway for proactive blocking.
This paper studies criterion revision in language-model agents, identifying failure modes in current implementations like CMB-0.1 and proposing a new trace-anchored protocol, CMB-0.4, for more accurate future evaluation.
AI agents can fail silently without traditional errors, as illustrated by a public postmortem where a pipeline ran into loops and high costs without triggering alarms. The article suggests using tracing and per-agent spend monitoring to detect such issues.
The article catalogs ten documented failure modes in multimodal AI systems where models generate fluent answers that break correspondence with actual inputs, based on benchmark papers and research studies.
The article discusses the problem of AI agents that fail silently by not escalating when stuck, and suggests implementing explicit checks or circuit breakers to improve reliability in production.
An essay mapping AI agent failure modes—broken rollbacks, missing capability withdrawal, observability without enforcement, retrieval failures, hallucination, and prompt injection—to specific episodes in the Mahabharata, arguing the ancient epic already specified the risks.
OrchestraBench is a new benchmark that evaluates multi-agent orchestration frameworks on failure modes, recovery, and decomposition quality, using failure-injection and cascade-radius metrics to diagnose where and why pipelines fail.
This paper systematically investigates instability in reinforcement learning for small language model agents (70-500M parameters), identifying three failure modes and proposing robust techniques including a merge-and-reinitialize adapter approach and safety mechanisms; it achieves stable convergence and improved win rates.
This paper introduces a novel evaluation method called shadow evaluations to test whether AI agents can conduct open-ended AI research. In two case studies, agents completed all engineering without human help but could not make substantial progress on the research questions, revealing five recurring failure modes.
A discussion on how treating 'human rejection' as a separate failure mode from 'agent malfunction' significantly impacts the reliability and debugging of AI agents in production.
This paper systematically investigates failure modes in reinforcement learning for small language models (70-500M parameters) using PPO, identifies silent LoRA freezing, numerical overflow, and catastrophic policy collapse, and proposes a robust system with merge-and-reinitialize adapters, float32 precision, and a safety mechanism. The approach converges stably and outperforms baselines with less data.