Asking the Wrong Questions: The HuggingFace Incident and the (Mis)Calibration of Semantic Fields in LLMs
Summary
This paper discusses the OpenAI-HuggingFace incident and proposes a theoretical framework using semantic field mechanics to explain how large language models process meaning.
Similar Articles
Recognition-Refusal Misalignment in LLMs: Why Models Answer Structurally Unanswerable Questions
This paper investigates why large language models answer structurally unanswerable questions, concluding that models recognize impossibility but fail to route that recognition to refusal, highlighting a routing failure rather than an encoding failure in their decision-making.
Human-Like Anaphor Resolution in Large Language Models
This paper investigates whether five open-weight LLMs exhibit human-like sensitivity to psycholinguistic factors in anaphor resolution, using surprisal and comprehension accuracy as behavioral measures. Results show selective cognitive alignment, with some models matching human discourse sensitivity but not semantic interference effects.
HawkesLLM: Semantic Uncertainty Propagation in Agentic Text Simulation
This paper introduces HawkesLLM, a framework that models semantic uncertainty propagation in multi-step agentic text simulations by combining a multivariate Hawkes process for temporal influence and memory selection with a language model for text generation. Evaluation on a GDELT news-cascade case study shows improved late-stage semantic alignment under compact prompt-memory constraints.
Saying More Than They Know: A Framework for Quantifying Epistemic-Rhetorical Miscalibration in Large Language Models
Introduces a framework to quantify how LLMs overstate certainty through rhetorical devices, revealing model-agnostic patterns of epistemic-rhetorical miscalibration.
Human-Alignment, Calibration, and Activation Patterns in Large Language Model Uncertainty
This paper investigates how similar large language model uncertainty is to human uncertainty, exploring alignment, calibration, and activation patterns in LLMs across multiple datasets and the impact of instruction fine-tuning.