ai-failures

Tag

Cards List
#ai-failures

@charliermarsh: Lots of good examples in these reports. For example: in DeepSWE v1.1, the verifier discards the agent's changes to some…

X AI KOLs Following · 6d ago Cached

Epoch AI introduces Benchmark Reviews to audit AI benchmarks, revealing flaws such as in DeepSWE v1.1 where the verifier discards changes without informing the model, leading to failures.

0 favorites 0 likes
#ai-failures

Retries can make AI failures worse

Reddit r/AI_Agents · 2026-09-01

This article discusses how retries in AI systems, particularly with LLMs and agents, can exacerbate failures when underlying issues are not addressed, leading to repeated mistakes with increased cost and latency.

0 favorites 0 likes
#ai-failures

The last two years I was trying to fix AI hallucinations now im dealing with a bigger problem

Reddit r/AI_Agents · 2026-07-11

Observations on the shift from addressing AI hallucinations to the more pressing problem of production AI failures, emphasizing the need for system reliability, tracking decisions, and limiting blast radius in enterprise deployments.

0 favorites 0 likes
#ai-failures

Are coding agents exposing how bad our specs actually are?

Reddit r/AI_Agents · 2026-06-21

The article argues that many failures of AI coding agents stem from vague specifications, not just model weaknesses. It suggests that writing clearer, more detailed work packets may be the next essential skill for developers using coding agents.

0 favorites 0 likes
#ai-failures

I think long context agents are failing in a very boring way

Reddit r/artificial · 2026-06-12

An opinion piece arguing that long context windows don't equate to memory and that agent failures are often mundane, like forgetting constraints or rereading files, emphasizing that reliability depends on context architecture decisions.

0 favorites 0 likes
#ai-failures

We built a public archive of AI failure patterns. The ones that keep coming back after changes.

Reddit r/artificial · 2026-05-29

An archive called Agent Fail Museum documents recurring AI failure patterns and provides regression test drafts for submitted failures, aiming to prevent repeat incidents.

0 favorites 0 likes
← Back to home

Submit Feedback