transformer-language-models

Tag

Cards List
#transformer-language-models

Representational Capacity: Geometric Limits on Feature Representation in Transformer Language Models

arXiv cs.LG · 2026-06-03 Cached

This paper introduces a quantitative framework for estimating how many near-orthogonal directions a transformer language model's latent space can support, based on the linear representation and superposition hypotheses. The authors define representational capacity as an upper bound on distinguishable features and show it is exponentially sensitive to the allowed deviation from orthogonality, with larger models favoring tighter constraints.

0 favorites 0 likes
#transformer-language-models

From Correlation to Cause: A Five-Stage Methodology for Feature Analysis in Transformer Language Models

arXiv cs.CL · 2026-05-22 Cached

This paper proposes a five-stage methodology for causal feature analysis in transformer language models, demonstrated on GPT-2 small for the IOI task. It finds that features are specifically causal but not necessary, and exposes a gap between detection and causal robustness.

0 favorites 0 likes
← Back to home

Submit Feedback