calibration

Tag

Cards List
#calibration

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

Hugging Face Daily Papers · yesterday Cached

This paper studies how instruction tuning affects model confidence and lexical diversity in question answering, finding that it alters confidence and reduces rationale diversity without improving calibration.

0 favorites 0 likes
#calibration

CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers

Hugging Face Daily Papers · yesterday Cached

This paper introduces CW-BASS v2, a saturation-aware pseudo-label selection method for semi-supervised semantic segmentation that adaptively switches between strict filtering and an adaptive confidence floor depending on the teacher's reliability. It shows improved results over baselines across several benchmarks with DINOv2 teachers.

0 favorites 0 likes
#calibration

Calibrating Post-Training Feature Shifts for LLM Data Contamination Detection

arXiv cs.CL · 2d ago Cached

This paper proposes CalibDCD, a calibration framework for feature-based LLM data contamination detection that mitigates feature shifts caused by post-training, improving detection performance by up to 7.0% AUC and 15.0% TPR@5%FPR.

0 favorites 0 likes
#calibration

MARCO: Click-Intent Decomposition for Calibrated Ads Conversion Prediction

arXiv cs.LG · 2d ago Cached

MARCO is a Meta AI framework that decomposes clicks by intent to improve ads conversion prediction, correcting per-intent calibration bias and lifting conversions per click by +2.80% and topline metrics by +0.98% in production.

0 favorites 0 likes
#calibration

Unsure but Certain: Uncovering the Representation-Confidence Gap in Diffusion Language Models

arXiv cs.CL · 3d ago Cached

This paper identifies a 'representation confidence gap' in diffusion language models: internal states detect input noise accurately but reported confidence stays high and answer ranking degrades under noise. It introduces a lightweight, training-free extraction tool that leverages hidden states to improve ranking without modifying the base model.

0 favorites 0 likes
#calibration

From token probabilities to calibrated confidence: An empirical study of mathematical question answering

arXiv cs.LG · 3d ago Cached

This paper empirically studies token-probability-based confidence estimation and calibration for LLMs in mathematical question answering, comparing single-pass and multi-pass estimators and evaluating post-hoc calibration methods.

0 favorites 0 likes
#calibration

Ask-E: An Environment for Calibrated Question Generation

arXiv cs.CL · 4d ago Cached

Ask-E is a new benchmark and training environment that evaluates and trains models on generating questions calibrated to specific skill levels, defined by the capabilities of two existing language models. Frontier models score below 50% on calibration, and training on Ask-E improves downstream math benchmarks without new math data or correctness-based rewards.

0 favorites 0 likes
#calibration

Dirichlet Follow-the-Leader Closes the Gap in Simultaneous Multiclass U-Calibration

arXiv cs.LG · 4d ago Cached

This paper introduces a simple Dirichlet-based forecaster that achieves optimal simultaneous multiclass U-calibration rates, closing the known dimension gap in regret bounds for bounded proper losses and removing extra additive terms for smooth losses.

0 favorites 0 likes
#calibration

Confidence Estimation for Financial Vision-Language Models in Chart and Document Understanding

arXiv cs.CL · 4d ago Cached

This paper studies confidence estimation for financial vision-language models in chart and document understanding, evaluating seven estimators across five LVLMs. It finds that calibration, not ranking, is the scarce property, and only trained probes produce thresholdable scores for safe deferral to human reviewers.

0 favorites 0 likes
#calibration

Mitigating Scoring Bias in LLM-as-a-Judge via Random Number Generation

arXiv cs.CL · 2026-08-07 Cached

This paper proposes a novel method to mitigate scoring bias in LLM-as-a-Judge by having LLMs randomly generate numbers to measure their latent numerical bias, then rectifying token generation probabilities accordingly. Experiments across four tasks show the method outperforms baselines and reveals that scoring bias varies across models, tasks, and score ranges.

0 favorites 0 likes
#calibration

Sample Complexity of Multicalibration for Multilevel Properties

arXiv cs.LG · 2026-08-06 Cached

This paper studies the sample complexity of multicalibration for a sequence of properties that are sequentially identifiable, establishing matching upper and lower bounds up to logarithmic factors.

0 favorites 0 likes
#calibration

The Calibration Floor: Format Repair Can Masquerade as Self-Correction at Small-to-Mid Scale

arXiv cs.CL · 2026-08-06 Cached

This paper shows that apparent LLM self-correction gains often stem from format repair rather than improved reasoning. Across multiple model scales, format effects dominate content effects, with content margins near zero on capable models, suggesting the field has misattributed a minority of measured self-correction to actual content improvement.

0 favorites 0 likes
#calibration

DUD: Decoupled Update Dynamics for Reliable Uncertainty Quantification in Large Language Models

arXiv cs.CL · 2026-08-05 Cached

The paper introduces DUD (Decoupled Update Dynamics), a framework that separates Feed-Forward Network and Attention contributions via causal interventions to improve uncertainty quantification and calibration in large language models, outperforming state-of-the-art baselines.

0 favorites 0 likes
#calibration

@chengyongru: One agent thinks A has a 51% probability of happening, another thinks 99%. But as long as they both end up submitting A, in FutureX's single-choice questions, a correct answer scores 1 point and a wrong answer scores 0. After reading the FutureX paper carefully, this example has stayed in my mind.…

X AI KOLs Timeline · 2026-08-04 Cached

The author comments on FutureX's prediction evaluation framework, pointing out that it only scores final answers and cannot distinguish the quality of probability calibration, and discusses finer-grained evaluation methods such as Brier score and log loss.

0 favorites 0 likes
#calibration

Subtype Robustness Is Not Just Accuracy: Calibration Under Unseen Subtype Shift

arXiv cs.LG · 2026-08-04 Cached

This paper presents the first systematic study of calibration under unseen subtype shift, showing that models become overconfident on novel subtypes within known coarse categories, and argues that subtype robustness should be evaluated with calibration metrics rather than accuracy alone.

0 favorites 0 likes
#calibration

The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models

arXiv cs.CL · 2026-08-03 Cached

This paper shows that knowledge distillation has asymmetric effects on bias in small language models: it improves context-following on unambiguous tasks but harms refusal calibration on ambiguous ones, and proposes PCCD, a protocol to diagnose such per-item harms that aggregate metrics miss.

0 favorites 0 likes
#calibration

Mitigating Class-Tail Undercoverage in Medical Vision-Language Models under Clinical Shift

arXiv cs.LG · 2026-08-03 Cached

Introduces CALCoDe, a post-hoc reliability layer for frozen medical vision-language models that mitigates class-tail undercoverage under clinical shift, achieving strong worst-class accepted coverage across multiple dermatology shifts and VLM backbones.

0 favorites 0 likes
#calibration

A Montage-Agnostic Encoder for Calibration-Light Cross-User Gesture Recognition from Surface Electromyography

arXiv cs.LG · 2026-07-31 Cached

Introduces a montage-agnostic encoder for calibration-light cross-user gesture recognition from surface EMG, using shared weights and electrode coordinates to handle variable channel counts and reduce per-user calibration. It outperforms per-user baselines on some datasets and analyzes factors affecting cross-user transfer.

0 favorites 0 likes
#calibration

How Often Should a Recommender Call an LLM? Value-Weighted Routing, Monitoring, and Seasonal Robustness

arXiv cs.AI · 2026-07-29 Cached

This paper introduces value-router, a simulation study for cost-aware routing between cheap heuristics and expensive LLM calls in recommender systems, showing that value-weighted routing improves precision and handles seasonal demand surges with adaptive budgets.

0 favorites 0 likes
#calibration

I built a tool to actually test which weights matter before quantizing, instead of guessing (Qwen3.6-27B, 3 builds: Bedrock/Tightrope/Gambit)

Reddit r/LocalLLaMA · 2026-07-28

A developer built a testing harness that measures KL divergence per weight group during quantization, leading to three custom quantized builds of Qwen3.6-27B (Bedrock, Tightrope, Gambit) with optimized compression. Tool calling is identified as the first capability to degrade under quantization.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback