priors

Tag

Cards List
#priors

Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors

arXiv cs.CL · 2026-07-02 Cached

This paper studies why language models hallucinate, proposing that hallucinations often stem from biased latent inference (inference misalignment) rather than missing knowledge. It introduces TrapQA, a controlled diagnostic testbed to test reasoning against priors, and demonstrates that hallucinations can arise from misleading latent associations.

0 favorites 0 likes
#priors

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators

arXiv cs.AI · 2026-06-09 Cached

This paper investigates the ability of LLMs-as-judges for safety to adapt to contextual information and varying safety definitions, finding that they are largely rigid and fail to adjust when the context contradicts their internal priors.

0 favorites 0 likes
#priors

On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance

arXiv cs.CL · 2026-06-02 Cached

This paper investigates how LLMs' internal priors affect zero-shot annotation performance, finding that nearly two-thirds of errors resist prompt-based correction and introducing Definition-Specific Familiarity as a better predictor than memorization metrics.

0 favorites 0 likes
← Back to home

Submit Feedback