diagnosis

Tag

Cards List
#diagnosis

What hidden states should an AI agent track when diagnosing CI failures?

Reddit r/AI_Agents · 2026-08-14

The article discusses hidden states that an AI agent should track when diagnosing CI failures, such as flaky tests, real bugs, and configuration errors, and seeks feedback on weaknesses and missing states.

0 favorites 0 likes
#diagnosis

SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification

Hugging Face Daily Papers · 2026-08-09 Cached

SymDiag is a neuro-symbolic framework that translates chain-of-thought reasoning into symbolic constraints and performs step-level satisfiability checks to localize failures in LLM reasoning, disentangling translation errors from reasoning errors.

0 favorites 0 likes
#diagnosis

Which area of healthcare will benefit most from AI in the next few years?

Reddit r/AI_Agents · 2026-07-27

Discussion on which area of healthcare will see the biggest transformation from AI in the next few years, including diagnosis, drug discovery, patient monitoring, and medical imaging.

0 favorites 0 likes
#diagnosis

CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data

arXiv cs.AI · 2026-07-20 Cached

CRAFT converts rubric-based evaluation into hierarchical capability diagnosis for LLMs, identifying specific weaknesses and generating targeted fine-tuning data, achieving stronger results on finance and legal benchmarks across four open-source models.

0 favorites 0 likes
#diagnosis

G-SHARE: A Guideline-Based Structured Reasoning Framework for Human-Factor Event Diagnosis

arXiv cs.CL · 2026-07-15 Cached

G-SHARE is a guideline-based structured reasoning framework for human-factor event diagnosis in nuclear power plants. It operationalizes a nine-step diagnostic guideline into a multi-stage pipeline with evidence extraction, stepwise reasoning, and consistency repair, outperforming one-shot LLM prompting and traditional baselines.

0 favorites 0 likes
#diagnosis

Imagined Rollouts are Kinematic, Not Dynamic: A Diagnosis of Long-Horizon World-Model Failure

Hugging Face Daily Papers · 2026-07-07 Cached

This paper diagnoses long-horizon failures in world models, attributing them to kinematic rather than dynamic imagination. The authors introduce a metric (iKCE) and show that imagined rollouts fail to capture dynamic regime changes even as policy rewards collapse.

0 favorites 0 likes
#diagnosis

Diagnosis is the missing skill in production agents

Reddit r/AI_Agents · 2026-07-03

The article argues that diagnosis—explaining why an agent failed in operational terms and what is safe to do next—is a missing first-class skill in production agent stacks, more critical than making agents sound smart.

0 favorites 0 likes
#diagnosis

RareDxR1: Autonomous Medical Reasoning for Rare Disease Diagnosis Beyond Human Annotation

arXiv cs.AI · 2026-07-02 Cached

Introduces RareDxR1, an end-to-end reasoning-centric large language model for open-domain rare disease diagnosis from unstructured clinical notes, using a progressive training framework and reflection-enhanced reasoning sampling, achieving state-of-the-art accuracy.

0 favorites 0 likes
#diagnosis

A specialized reasoning large language model for accelerating rare disease diagnosis: a randomized AI physician assistance trial

arXiv cs.AI · 2026-06-24 Cached

This paper presents RaDaR, a 32B open-source reasoning LLM trained on public and synthetic rare disease cases, which outperforms larger models like DeepSeek-R1 in diagnosis benchmarks and improves physician accuracy by 21.44 percentage points in a randomized trial.

0 favorites 0 likes
#diagnosis

When Agents Commit Too Soon: Diagnosing Premature Commitment in LLM Agents

Hugging Face Daily Papers · 2026-06-22 Cached

This paper introduces representational commitment, a cross-run hidden-state convergence that diagnoses when an LLM agent has locked onto a trajectory prematurely. It shows that commitment predicts trajectory consistency but not correctness, and proposes monitoring to detect when an agent is confidently settled rather than assuming consistency equals trust.

0 favorites 0 likes
#diagnosis

EBench: Elemental Diagnosis of Generalist Mobile Manipulation Policies

Hugging Face Daily Papers · 2026-06-20 Cached

EBench is a diagnostic benchmark for generalist mobile manipulation policies, providing a multi-dimensional profile across 26 tasks and 4 generalization axes, revealing structural strengths and weaknesses beyond aggregate success rates.

0 favorites 0 likes
#diagnosis

@OpenAI: Rare disease diagnosis is challenging, as sequencing can surface millions of variants, and medical knowledge changes co…

X AI KOLs · 2026-06-18 Cached

OpenAI highlights how o3 Deep Research can aid rare disease diagnosis by integrating clinical features, inheritance patterns, variant evidence, and scientific literature into actionable hypotheses for specialists.

0 favorites 0 likes
#diagnosis

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs

arXiv cs.AI · 2026-06-01 Cached

EHRBench is an automated and reliable benchmark for evaluating LLMs on clinical decision-making tasks using real-world electronic health records, covering nearly 1M QA items across diagnosis, treatment, and prognosis tasks.

0 favorites 0 likes
#diagnosis

Boston Children’s uses AI to unlock new diagnoses

OpenAI Blog · 2026-05-29 Cached

Boston Children's Hospital has integrated AI across its clinical and operational infrastructure, using a secure ChatGPT environment to diagnose over 40 rare conditions, reduce operational costs, and improve care delivery.

0 favorites 0 likes
#diagnosis

VBFDD-Agent for Electric Vehicle Battery Fault Detection and Diagnosis: Descriptive Text Modeling of Battery Digital Signals

arXiv cs.AI · 2026-05-22 Cached

This paper proposes VBFDD-Agent, a vehicle battery fault detection and diagnosis agent that uses descriptive text modeling of battery signals, large language models, and historical cases to generate interpretable diagnostic results and maintenance recommendations for electric vehicle batteries.

0 favorites 0 likes
#diagnosis

Elentaria

Product Hunt · 2026-05-18

Elentaria is a product launched on ProductHunt that helps with go-to-market strategy from diagnosis to execution.

0 favorites 0 likes
#diagnosis

Reconciling Consistency-Based Diagnosis with Actual-Causality-Based Explanations

arXiv cs.AI · 2026-05-12 Cached

This academic paper establishes connections between Consistency-Based Diagnosis and Actual Causality within the context of Explainable AI (XAI). It aims to integrate these two areas to improve explanations in AI and Explainable Data Management.

0 favorites 0 likes
#diagnosis

MEDSYN: Benchmarking Multi-Evidence Synthesis in Complex Clinical Cases for Multimodal Large Language Models

arXiv cs.CL · 2026-04-20 Cached

MEDSYN is a multilingual multimodal benchmark for evaluating MLLMs on complex clinical cases with up to 7 distinct visual evidence types per case. The study reveals that while frontier models match human experts on differential diagnosis generation, all MLLMs show significant gaps in final diagnosis selection due to poor synthesis of heterogeneous clinical evidence.

0 favorites 0 likes
#diagnosis

How AI Helps Solve Medical Mysteries at Boston Children’s Hospital | OpenAI Forum

YouTube AI Channels · 2026-08-05 Cached

OpenAI collaborated with Boston Children's Hospital and Harvard's Manton Center on a study covering 376 cases, using AI workflows to assist in diagnosing rare diseases and leading to 18 confirmed diagnoses. The core approach was to have AI perform literature retrieval, hypothesis ranking, and evidence summarization, which were then reviewed by human geneticists, rather than directly providing conclusions.

0 favorites 0 likes
← Back to home

Submit Feedback