selective-prediction

Tag

Cards List
#selective-prediction

One Score, Two Decisions: Selective Prediction on the Rare-Disease Tail

arXiv cs.LG · 2d ago Cached

This paper analyzes selective prediction systems for rare-disease diagnosis, demonstrating that small open-weight LLMs have low recall on ultra-rare diseases and exploring the use of score margins for decision-making with limitations.

0 favorites 0 likes
#selective-prediction

Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays

arXiv cs.LG · 2d ago Cached

The paper diagnoses three failure modes in per-field selective risk control for document extraction systems and introduces a validity ladder of fixes, demonstrating improvements through experiments on real-world data with frontier AI models.

0 favorites 0 likes
#selective-prediction

Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring

arXiv cs.CL · 6d ago Cached

The paper introduces RA-DPO, a reliability-aware direct preference optimization method that combines annotator agreement, model confidence, and token-level uncertainty for sexism detection, improving training efficiency and enabling selective prediction.

0 favorites 0 likes
#selective-prediction

Asymptotic Risk Calibration for Selective Question Answering

arXiv cs.CL · 2026-08-13 Cached

The paper proposes A-CRC-QA, a post-hoc calibration framework for selective question answering that controls error rates among accepted answers via asymptotic risk calibration, demonstrating improved reliability-retention trade-offs on CoQA and MedMCQA.

0 favorites 0 likes
#selective-prediction

Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents

arXiv cs.CL · 2026-08-13 Cached

This paper studies whether single-turn uncertainty quantification methods transfer to interactive LLM agent trajectories, evaluating white-box, black-box, and reflexive scorers across five LLMs and four tool-use datasets. Results show that transfer is uneven, with black-box self-consistency often strongest, and recommend revalidating UQ methods at the trajectory level.

0 favorites 0 likes
#selective-prediction

Confidence Estimation for Financial Vision-Language Models in Chart and Document Understanding

arXiv cs.CL · 2026-08-10 Cached

This paper studies confidence estimation for financial vision-language models in chart and document understanding, evaluating seven estimators across five LVLMs. It finds that calibration, not ranking, is the scarce property, and only trained probes produce thresholdable scores for safe deferral to human reviewers.

0 favorites 0 likes
#selective-prediction

Evidence-Grounded Constraint Checking in Construction Documents

arXiv cs.AI · 2026-08-03 Cached

Presents an evidence-grounded pipeline for constraint checking in construction documents, evaluating resolution-breadth tradeoffs in PDF evidence allocation on expert-referenced tasks.

0 favorites 0 likes
#selective-prediction

Calibrated Selective Fact-Checking via Evidence Chain Evaluation

arXiv cs.AI · 2026-07-22 Cached

This paper introduces Evidence Chain Evaluation (ECE), a selective fact-checking framework that allows LLM-based verification agents to abstain from giving verdicts when evidence is weak, sparse, or inconsistent. On ECE-Bench, ECE achieves 97.8% selective accuracy at 93.7% coverage, demonstrating a safety-oriented trade-off for handling epistemically weak evidence.

0 favorites 0 likes
#selective-prediction

Beyond Output-Space Calibration: Spectral Evidence Bundling for Selective Reliability Estimation in Time-Series Classification

arXiv cs.LG · 2026-07-22 Cached

This paper introduces SEB-Cal, a method that augments output-space calibration with spectral features (band energy, entropy, peak dominance, phase stability) to improve selective reliability estimation in time-series classification, achieving higher Corr-AUROC and lower [email protected] across multiple datasets.

0 favorites 0 likes
#selective-prediction

What Predicts Correctness in Text-to-SQL? A Selective-Prediction Study

arXiv cs.LG · 2026-07-09 Cached

This paper studies which signals best predict correctness in text-to-SQL for selective prediction. It finds that verification-based signals from LLM judges outperform black-box statistical signals like self-consistency, and that a two-provider ensemble achieves 0.82 AUROC with well-calibrated probabilities.

0 favorites 0 likes
#selective-prediction

When Can Conformal Risk Control Certify LLM Outputs? Bounds, Impossibility, and Adaptation for Structured Generation

arXiv cs.LG · 2026-06-30 Cached

This paper characterizes when conformal risk control can certify structured LLM outputs, proving impossibility bounds and analyzing certification hierarchies across different bounds. Empirical validation on six open-weight models shows that hard configurations are uncertifiable at low risk levels but practical certification is achievable at relaxed targets.

0 favorites 0 likes
#selective-prediction

Beyond Logprobs: A Multi-Signal Confidence Engine for LLM-Based Document Field Extraction

arXiv cs.CL · 2026-06-24 Cached

ExtractConf is a confidence estimation method for LLM-based document field extraction that uses two structurally different calls (field-guided and document-guided) to derive disagreement signals, achieving 0.928 ROC AUC on DocILE invoices and enabling reliable selective prediction for high-stakes automation.

0 favorites 0 likes
← Back to home

Submit Feedback