Tag
This paper presents TGO-IV, a topological framework using persistent homology to analyze how transformer representations evolve across layers, complementing prior spectral and geometric observatories.
This paper establishes theoretical bounds on the number of attention heads needed to produce vector representations that support multiple tasks, such as computing min/max and XOR, showing trade-offs between head count, embedding dimension, and precision.
This paper adapts Funder's personality triad framework to LLMs, using sparse autoencoders to discover and validate trait-like internal representations, and demonstrating controllable bidirectional behavioral shifts through feature-level interventions.
This paper demonstrates that large language models internally encode the strength of clinical evidence for claims, yet fail to accurately express this strength when asked, with stated evidence grades performing near chance.
This paper introduces 'fragility', a complementary metric to probe accuracy that measures activation-noise level at which probe accuracy collapses, enabling analysis of representation evolution during LLM pre-training even after accuracy saturates.
This paper investigates how lexical overlap, rather than semantic content, influences LLM representations across layers and architectures, and demonstrates that this lexical effect persists even in models trained for semantic similarity, leading to degraded performance on downstream tasks.
This paper investigates whether fMRI representations from different subjects' visual cortices can be aligned using unsupervised geometric methods, finding evidence for approximately isometric structure across individuals, extending the Platonic Representation Hypothesis to human brains.
A new paper shows that late-interaction retrieval model representations can effectively replace raw document text in RAG tasks, extending their utility beyond retrieval.