ai-reliability

Tag

Cards List
#ai-reliability

How do you check your AI written code is correct?

Reddit r/AI_Agents ↗ · 2d ago

The user asks for ways to verify the correctness of AI-generated code, highlighting doubts and inefficiencies when using multiple models for checks.

0 favorites 0 likes
#ai-reliability

Prompts Aren't Real

Hacker News Top ↗ · 4d ago Cached

Engineer Dan discusses the challenges of building reliable AI agents in production, highlighting LLM limitations and the engaging yet mentally taxing nature of the work.

0 favorites 0 likes
#ai-reliability

The internet is inbreeding.

Reddit r/artificial ↗ · 6d ago

The article describes how reputable sites blocking AI crawlers forces AI models to rely on low-quality, AI-generated content, leading to issues like post-hoc citation and infrastructure-level confirmation bias, and offers advice for better source verification.

0 favorites 0 likes
#ai-reliability

Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts

arXiv cs.AI ↗ · 2026-09-17 Cached

The paper introduces CHASE, a counterfactual-guided method for harness evolution in AI agent evaluation that mitigates benchmark shortcuts, enhancing reliability through constraint generation and validity checks.

0 favorites 0 likes
#ai-reliability

As a Student, I found larger models are still largely unreliable for many things, helpful but also useless

Reddit r/ArtificialInteligence ↗ · 2026-09-16

A student shares their experience with large AI models being unreliable for summarizing textbook material, noting issues with inaccuracies and nitpicking, and questions the perceived danger of AI based on these flaws.

0 favorites 0 likes
#ai-reliability

90% of the "look my agent can browse the web" posts here would collapse on the 5th URL

Reddit r/AI_Agents ↗ · 2026-09-16

The author critiques web agent demos for failing on real-world URLs due to fetch-layer issues, emphasizing that data engineering is the harder problem than agent reasoning.

0 favorites 0 likes
#ai-reliability

CUSP: Decomposable Collective Uncertainty for Multi-Agent Multimodal Reasoning

arXiv cs.AI ↗ · 2026-09-10 Cached

CUSP is a training-free framework for quantifying collective uncertainty in multi-agent multimodal systems using semantic opinion pooling, improving reliability and accuracy over baseline methods.

0 favorites 0 likes
#ai-reliability

From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs

arXiv cs.CL ↗ · 2026-09-03 Cached

This paper proposes methods for detecting hallucinations in black-box LLMs by combining semantic entropy and token-level uncertainty signals, evaluating techniques like TopK, CoCoA, Gated, and Stacked across multiple benchmarks to find that no single method is universally strongest but Stacked often performs best.

0 favorites 0 likes
#ai-reliability

A hallucination class that passes fact-checking: the claim is true and the quotation marks are fabricated

Reddit r/ArtificialInteligence ↗ · 2026-08-30

The article describes a type of AI hallucination where claims are accurate but quotations are fabricated, evading standard fact-checking, and discusses implementation challenges in detecting such errors.

0 favorites 0 likes
#ai-reliability

@rohanpaul_ai: Long-horizon agent reliability has not arrived yet with better models. On WeaveBench's 114 hybrid GUI-CLI tasks, the be…

X AI KOLs Timeline ↗ · 2026-08-29 Cached

This paper argues that long-horizon AI agent reliability is lacking despite better models, as seen on WeaveBench with only 41.2% pass rate, and proposes the LongHorizon-Harness to manage task state for improved performance.

0 favorites 0 likes
#ai-reliability

I wonder when people are going to realize we need to bring this back...

Reddit r/AI_Agents ↗ · 2026-08-27

The author argues for reviving 'Needle in a haystack' benchmarks to evaluate AI capabilities, sharing private test results that show many models performing poorly in remembering instructions, questioning their trustworthiness for real-world tasks.

0 favorites 0 likes
#ai-reliability

openai claims it took a week to realize its models hacked hugging face

Reddit r/AI_Agents ↗ · 2026-08-27

OpenAI claims it took a week to realize its models were hacked by Hugging Face, while mainstream media highlights enterprise struggles with AI reliability, token costs, and context management in multi-step tasks.

0 favorites 0 likes
#ai-reliability

I audited the sources my AI fact-checker was citing. About 1 in 18 didn't exist.

Reddit r/artificial ↗ · 2026-08-24

An author shares an audit of their AI fact-checker's citations, discovering that about 1 in 18 were dead or fabricated, and provides practical fixes to improve reliability in AI systems.

0 favorites 0 likes
#ai-reliability

Wild AI-related reliability incidents are coming

Lobsters Hottest ↗ · 2026-08-23 Cached

The article explores the growing trend of using AI agents for operational tasks like on-call work, but cautions that their complexity may lead to unexpected reliability incidents, referencing recent talks and examples from security conferences.

0 favorites 0 likes
#ai-reliability

Answer-Level Trust Selection for Physical Vision-Language Reasoning

arXiv cs.LG ↗ · 2026-08-21 Cached

This paper proposes Answer-Level Trust Selection (ATS), a post-hoc, model-agnostic framework for assessing the reliability of individual predictions from vision-language models in quantitative physical reasoning tasks.

0 favorites 0 likes
#ai-reliability

Ten Failure Modes That Define Multimodal AI Systems

Reddit r/ArtificialInteligence ↗ · 2026-08-20 Cached

The article catalogs ten documented failure modes in multimodal AI systems where models generate fluent answers that break correspondence with actual inputs, based on benchmark papers and research studies.

0 favorites 0 likes
#ai-reliability

Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study

arXiv cs.AI ↗ · 2026-08-18 Cached

This study benchmarks 14 large language models on cosmetic chemistry and skin health topics, finding poor accuracy and reliability, especially in technical and quantitative tasks. It concludes that current general-purpose LLMs are not reliable for informed consumer decision-making without further fine-tuning and algorithmic improvements.

0 favorites 0 likes
#ai-reliability

The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines

arXiv cs.AI ↗ · 2026-08-18 Cached

This paper models hallucination propagation in multi-agent LLM pipelines as a Markov process, showing that errors become less detectable across stages and proposing early verification to reduce survival rates.

0 favorites 0 likes
#ai-reliability

We are not at singularity moment. We are at Dexter moment.

Reddit r/singularity ↗ · 2026-08-16

A user reports significant bugs in Google's 3.1 TTS preview model where it invents content in audio outputs, leading to concerns about AI reliability in critical applications.

0 favorites 0 likes
#ai-reliability

AI is confidently wrong way more than people give it credit for, change my mind

Reddit r/ArtificialInteligence ↗ · 2026-08-12

A user shares concerns about AI models presenting thin or ambiguous data with the same confidence as well-supported findings, citing a case where a complaint appearing only twice in 200 comments was ranked as a top concern. The piece questions whether this is a fixable prompting issue or a fundamental limitation requiring manual verification.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback