Tag
The paper identifies Harmfulness Propagation Dynamics in large language models and introduces Herald, a lightweight input moderator that uses cross-layer activation patterns to detect harmful prompts efficiently.
This paper proposes an adaptive spectral bandwidth control method for kernelized graph construction to align kernel spectral properties with intrinsic manifold dimensions, showing improvements in self-supervised learning embedding tasks on CIFAR-100.
This paper introduces a calibrated test of internal action maps in language models, showing that state signals can be decodable and causally usable without global affine closure, using an evidence lattice framework validated on finite worlds and the Qwen3-4B model.
This paper investigates why transformer intermediate representations are off-axis relative to the readout direction, showing that this off-axis subspace functionally insulates composition from the vocabulary and proposing methods to impose this geometry via rotation.
This paper introduces a framework using Jungian cognitive functions for activation steering in LLMs, demonstrating effective monotonic control over eight functions and revealing structured geometric relationships in activation space.
This paper proposes EffRank/n and D_act as low-overhead diagnostics to measure effects of reward attribution in cooperative multi-agent RL, and tests on SMACv2, finding that observation explains geometry while reward attribution mainly affects behavior.
This academic paper introduces finite-lag operator geometry for analyzing recurrent neural network hidden states, deriving a source-centered transport tensor and antisymmetric coordinate circulation to capture directed flow and deterministic recurrent motion beyond static snapshots.
This paper analyzes linear activation steering in language models by decomposing interventions into angular and radial components. It finds that concepts are primarily encoded in angular structure, but norm adjustments are crucial for stability, supporting spherical steering methods while showing that additive coefficients conflate geometry.
This paper systematically tests linear probes for deception detection in large language models, finding they fail under distributional shifts but style-augmented probes recover performance, and revealing that deception is encoded through distributed sub-threshold features.
Discusses that the mathematics used by AI is mainly linear algebra, calculus, etc., from before the 19th century, but emerging phenomena such as Scaling Law, emergent abilities, double descent, in-context learning, and representation geometry lack mathematical explanation. Analogizes to the clouds in physics in 1900, suggesting it may drive the development of 21st-century mathematics.
This paper introduces the concept of 'minimal cores' in overcomplete reasoning traces, showing that on average 46% of steps can be removed while preserving the final answer, and that minimal cores improve trace separation and reduce intrinsic dimensionality.
This paper introduces Minor Component Unlearning (MCU), a novel approach to LLM unlearning that targets minor components in representations to resist relearning attacks. It addresses the vulnerability of existing methods by focusing on robust directions within the model's spectral structure.