Tag
This paper identifies a 'detectability gap' in hallucination detection: hallucinations split into high-agreement (Ghost) and low-agreement (Flickering) regimes with a 0.35–0.46 AUC gap, persisting across four models and three factual QA datasets even after freezing regime assignments and using stricter trajectory-based tests. The authors argue aggregate detection metrics hide model-dependent heterogeneity and call for regime-conditioned evaluation.
This paper proposes LUDI, a less uniform diffusion language modeling framework that fixes over-uniform training objectives and condition-target confusion in uniform diffusion LMs, enabling a 7B-scale UDLM with 3x-token-per-step speedup over autoregressive decoding and competitive complex reasoning performance.