llm-calibration

Tag

Cards List
#llm-calibration

Can a Model Catch Its Own Hallucinations for Free?: Label-Free Doubt Signals Hold Their Own Against a Labelled Dataset for Abstention

arXiv cs.CL · 2026-08-28 Cached

The paper introduces a label-free method for large language models to abstain when uncertain, using internal confidence signals, and shows it matches supervised abstention tuning in performance.

0 favorites 0 likes
#llm-calibration

Silico (3 minute read)

TLDR AI · 2026-07-16 Cached

GoodfireAI shares research on using probes to improve LLM calibration and detect reasoning faithfulness, and a hackathon using their product Silico for targeted fine-tuning with parameter decomposition.

0 favorites 0 likes
#llm-calibration

@dair_ai: New research from Google. LLMs hallucinate with high confidence, miss their own knowledge boundaries, and misreport unc…

X AI KOLs Timeline · 2026-07-02 Cached

A new research paper introduces RLMF (Reinforcement Learning with Metacognitive Feedback), a two-stage approach that uses the model's own self-judgments to calibrate confidence and express uncertainty faithfully, achieving state-of-the-art calibration across diverse tasks while preserving accuracy and surpassing standard RL by up to 63%.

0 favorites 0 likes
#llm-calibration

Counterfactual Graph for Multi-Agent LLM Calibration

arXiv cs.CL · 2026-06-01 Cached

This paper introduces CAGE, a counterfactual graph-based method for calibrating multi-agent LLM systems, evaluating on benchmarks like TriviaQA and MMLU-Pro across various communication topologies. The method outperforms existing post-hoc and LLM-elicited calibration approaches.

0 favorites 0 likes
#llm-calibration

Retrieval-Augmented Linguistic Calibration

arXiv cs.CL · 2026-05-20 Cached

This paper proposes Retrieval-Augmented Linguistic Calibration (RALC), a post-hoc pipeline for calibrating confidence signals in LLMs by modeling linguistic confidence as a distribution and using retrieval-augmented rewriting. It introduces Faithfulness Divergence metric and shows significant improvements across benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback