reliability

Tag

Cards List
#reliability

@thsottiaux: The day we develop really good models. There will be signs. Reliability increasing despite load going up and up. Sudden…

X AI KOLs Timeline · 2026-07-31 Cached

The author speculates that truly advanced AI models will show signs such as improving reliability under increasing load, sudden efficiency gains, faster performance, and resets.

0 favorites 0 likes
#reliability

LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation

arXiv cs.CL · 2026-07-31 Cached

Introduces LayerRAG-Bench, a cross-layer reliability benchmark for agentic retrieval-augmented generation systems, covering 9 fault scenarios and 38,880 records across nine models, with findings that schema normalization fixes schema drift but not stale, unauthorized, or wrong-session evidence.

0 favorites 0 likes
#reliability

Recursive transformers for semiconductor thermo-mechanical reliability

arXiv cs.LG · 2026-07-31 Cached

This paper evaluates recursive weight-sharing transformer architectures as parameter-efficient surrogate models for semiconductor thermo-mechanical reliability prediction, comparing performance, parameter count, and computational cost on small engineering datasets.

0 favorites 0 likes
#reliability

Saturation: How Your Software Will Fail at Scale

Lobsters Hottest · 2026-07-30 Cached

This article discusses the phenomenon of software systems failing due to resource saturation as they scale, drawing an analogy with the limits of biological systems, and introduces the impact of physical resource (CPU, memory, disk, network) saturation and virtual limits on system reliability.

0 favorites 0 likes
#reliability

An AI agent without a stop policy is just an expensive loop

Reddit r/AI_Agents · 2026-07-30

A practical note on AI agent reliability, arguing that production agents need explicit gates for evidence thresholds, retry budgets, and impact assessment rather than relying on memory alone to determine task completion.

0 favorites 0 likes
#reliability

Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems

Hugging Face Daily Papers · 2026-07-30 Cached

The paper introduces Σ-Mem, an online reliability memory for LLM-based multi-agent systems that tracks historical competence of peers and peer relationships, enabling stable adaptation via spectral bounds and improving coordination through residual steering, routing, and weighted voting.

0 favorites 0 likes
#reliability

Why does an agent that nails every test case still go sideways after a few hundred real conversations?

Reddit r/AI_Agents · 2026-07-29

Explores why AI agents that perform perfectly on test cases often fail in real-world conversations, highlighting issues like distribution shift and overfitting.

0 favorites 0 likes
#reliability

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy

arXiv cs.AI · 2026-07-28 Cached

This paper investigates how LLMs' answers change under meaning-preserving paraphrases, finding that instance-level behavior is unstable (flip rates >23%) and that single-prompt accuracy masks substantial inconsistency, while a self-paraphrasing strategy can partially recover latent knowledge.

0 favorites 0 likes
#reliability

Treating "a human rejected this" as a different failure mode than "the agent broke" — turns out that distinction matters a lot in production

Reddit r/AI_Agents · 2026-07-27

A discussion on how treating 'human rejection' as a separate failure mode from 'agent malfunction' significantly impacts the reliability and debugging of AI agents in production.

0 favorites 0 likes
#reliability

AI agents are starting to look less like software and more like employees

Reddit r/artificial · 2026-07-26

The article argues that AI agents are evolving from being evaluated solely on intelligence to requiring operational reliability, governance, and team integration akin to human employees, highlighting the need for new infrastructure layers.

0 favorites 0 likes
#reliability

We're spending too much time building agents and not enough time thinking about production

Reddit r/AI_Agents · 2026-07-26

This article argues that the AI community focuses too much on building capable agents and not enough on the operational challenges of deploying them reliably in production, highlighting the need for better visibility, debugging, and system robustness.

0 favorites 0 likes
#reliability

@SpaceX: With flight tests, success comes from what we learn, and today’s test will help us advance Starship’s reliability ahead…

X AI KOLs Following · 2026-07-24

SpaceX states that success in flight tests comes from learning, and today's test will help improve Starship's reliability for future orbital missions and Starlink deployments.

0 favorites 0 likes
#reliability

Don't let the model write the audit log

Reddit r/AI_Agents · 2026-07-24

The article warns against using model-generated narration as the authoritative audit log for AI agents, advocating for persisting raw tool call data instead, and suggests a simple diff check to catch discrepancies.

0 favorites 0 likes
#reliability

The agent industry made stack overflow billable—but who owns the return?

Reddit r/AI_Agents · 2026-07-24

An analysis of the termination problem in recursive AI agent systems, highlighting how agents can lose verified position while still selecting valid actions, and questioning where termination logic should reside in the agent stack.

0 favorites 0 likes
#reliability

Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions

Hugging Face Daily Papers · 2026-07-23 Cached

This paper introduces MisKnow-Agent, a framework for generating misleading knowledge to test DeepResearch agents, showing that limited exposure to credible-looking false information can lead to false conclusions in final reports.

0 favorites 0 likes
#reliability

Operational Hallucination and Safety Drift in AI Agents

arXiv cs.AI · 2026-07-22 Cached

This paper identifies and characterizes two failure modes in LLM-based autonomous agents—Safety Drift and Operational Hallucination—and proposes a lightweight architectural layer to intercept violations without false positives.

0 favorites 0 likes
#reliability

Reliability Scales Inversely: Bigger Models Compound Mistakes Faster via a Hidden Auto-Regressive Risk Regime

arXiv cs.LG · 2026-07-22 Cached

This paper discovers that larger language models have a hidden auto-regressive risk regime where they commit to low-probability tokens and then snowball errors, causing reliability to degrade faster with scale. It shows that this failure mode is causal, dominant, and invisible to the model's own self-monitoring.

0 favorites 0 likes
#reliability

FYI, your agent can be "up" and completely broken at the same time

Reddit r/AI_Agents · 2026-07-21

A reminder that an AI agent can appear to be running ("up") while actually being broken or malfunctioning, highlighting the need for better monitoring and validation.

0 favorites 0 likes
#reliability

Multilingual Sentence Embeddings for Linguistic-Integrated Reliability Audit

arXiv cs.CL · 2026-07-21 Cached

This paper evaluates whether multilingual sentence embeddings can replace translation for linguistic-integrated reliability auditing across multiple languages in educational assessments, finding that native-language embeddings reproduce translation-based reliability estimates closely.

0 favorites 0 likes
#reliability

How Reliable Are Multimodal Signals of Conversational State? Evidence from Remote Dyadic Collaborative Tasks

arXiv cs.CL · 2026-07-21 Cached

This paper evaluates the predictive accuracy, cross-task generalizability, and test-retest reliability of multimodal features for measuring conversational states like cognitive load and power in dyadic remote collaborative tasks. Findings show that linguistic features predict well but generalize poorly, acoustic reliability degrades when controlling for speaker identity, and interaction features provide the most reliable signal.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback