confidence

Tag

Cards List
#confidence

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

Hugging Face Daily Papers · 3d ago Cached

This paper studies how instruction tuning affects model confidence and lexical diversity in question answering, finding that it alters confidence and reduces rationale diversity without improving calibration.

0 favorites 0 likes
#confidence

"Uncensored" LLMs are measurably more optimistic than their base models

Reddit r/LocalLLaMA · 2026-07-29

A study on uncensored LLMs (Gemma and Qwen) shows that removing censorship makes them more optimistic in stock market predictions, but not more accurate. The effect varies by model family.

0 favorites 0 likes
#confidence

DeepLook: Deeper Thinking with Lookahead

arXiv cs.AI · 2026-07-28 Cached

DeepLook is a training-free framework that improves LLM reasoning by allocating compute at uncertainty bottlenecks, reducing token generation by 87.3% on average while improving accuracy on competition math benchmarks.

0 favorites 0 likes
#confidence

Dual-Confidence Contrastive Decoding for Retrieval-Augmented Generation

arXiv cs.CL · 2026-07-02 Cached

Proposes Dual-Confidence Contrastive Decoding (DCCD), a training-free method for retrieval-augmented generation that handles intra-context conflicts in multi-document settings by combining document-level and token-level confidence signals, and introduces the DRQA benchmark for factual-conflict QA.

0 favorites 0 likes
#confidence

Why do we trust AI answers simply because they sound confident?

Reddit r/artificial · 2026-07-01

The author reflects on why people trust confident AI answers, especially in finance, and introduces their project AutoFlow, a Credit Evidence Engine designed to verify financial claims against source evidence and highlight contradictions.

0 favorites 0 likes
#confidence

When Calibration Rankings Reverse: Accuracy-Controlled Evaluation for Fair Comparison of LLMs

arXiv cs.CL · 2026-07-01 Cached

This paper demonstrates that global calibration metrics like Expected Calibration Error are confounded by model accuracy, and proposes ACE, an accuracy-controlled evaluation framework for fair comparison of large language models.

0 favorites 0 likes
#confidence

TriageRA-CCF: Source-Side Clinical Confidence and Coverage Signals for Adaptive Rank Budgeting in Medical LLMs

arXiv cs.CL · 2026-06-30 Cached

This paper proposes TriageRA-CCF, a method for adaptive rank budgeting in LoRA for medical question answering. It uses source-side signals (base-model confidence, clinical coverage, counterfactual proxy) to dynamically choose rank budgets, achieving modest accuracy gains on Qwen3-8B and Llama3.1-8B.

0 favorites 0 likes
#confidence

An agent remembering everything sounds useful until it remembers the wrong crap

Reddit r/AI_Agents · 2026-06-17

The author critiques the idea of agents remembering everything and introduces TrueMemory, a system that converts memories into trait claims with confidence and evidence to better calibrate agent behavior.

0 favorites 0 likes
#confidence

If your agent makes a bad autonomous call, can you reconstruct why it decided that or just what it did?

Reddit r/AI_Agents · 2026-06-16

A developer building autonomous billing agents discusses the difficulty of reconstructing why an agent made a decision after the fact, and describes building a tool (Attova) that records decisions with evidence, alternatives, and confidence to improve debugging and human review.

0 favorites 0 likes
#confidence

LLMs Show No Signs Of Individuated Metacognition

arXiv cs.LG · 2026-05-26 Cached

This paper investigates whether frontier LLMs exhibit individuated metacognition—the ability to assess their own item-level capabilities beyond shared signals. Through factor analysis and pairwise calibration across 20 models and six benchmarks, the authors find no evidence of such metacognition; confidence differences reduce to a single shared difficulty factor, suggesting models rely on a common difficulty signal rather than model-specific self-knowledge.

0 favorites 0 likes
#confidence

Claude made me realize most AI models optimize for confidence, not truth

Reddit r/artificial · 2026-05-22

A reflection on how many AI models prioritize sounding confident over being truthful, using Claude as an example of a model that seems more focused on internal consistency and logical honesty.

0 favorites 0 likes
#confidence

Calibrating LLMs with Semantic-level Reward

arXiv cs.CL · 2026-05-18 Cached

Proposes CSR, a framework that calibrates LLMs directly in semantic space using a novel semantic calibration reward, reducing ECE by up to 40% and improving AUROC by up to 31% over verbalized-confidence baselines across multiple datasets.

0 favorites 0 likes
#confidence

@mitsuhiko: I think it would be great if people were upfront about declaring their own understanding of a topic / their pull reques…

X AI KOLs Timeline · 2026-05-16

Armin Ronacher (@mitsuhiko) suggests that people should be upfront about their actual understanding of a topic when making pull requests, as AI tools (referred to as 'clanker') make it easy to sound confident without real knowledge.

0 favorites 0 likes
#confidence

@WorldExecAI: Are second-generation rich clearly less confident than the first generation? At this banquet, several second-generation rich were seated next to Musk and Jensen Huang, but they had no interaction with the Silicon Valley giants, clearly not as confident as Jack Ma, Charles Zhang, and Robin Li. To the left of Tesla CEO Musk, smiling sheepishly, is Cao Hui, son of Fuyao Glass founder Cao Dewang. Next to Nvidia CEO Jensen Huang is Lu Weiding, from Wan...

X AI KOLs Timeline · 2026-05-14 Cached

The article discusses a banquet where second-generation rich were seated next to Musk and Jensen Huang but lacked interaction, contrasting with the confidence of first-generation entrepreneurs like Jack Ma and Charles Zhang, sparking discussion on the differences between the two generations of entrepreneurs.

0 favorites 0 likes
#confidence

The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation

Hugging Face Daily Papers · 2026-04-18 Cached

This paper identifies that on-policy distillation (OPD) in language models leads to severe overconfidence due to information mismatch between training and deployment, and proposes CaOPD, a calibration-aware framework that improves both performance and confidence reliability.

0 favorites 0 likes
← Back to home

Submit Feedback