Tag
This preprint uses adversarial training as a controlled instrument on GPT-2 Small to test whether representational simplicity (SAE decomposability, concentrated attribution) implies causal circuit simplicity. It finds that robust models are more SAE-decomposable, while circuit size is regime-dependent: robustness helps at high faithfulness levels (90-95%) but not below 85%.
This paper introduces a method using GPT-2 token surprisal to label telicity in child language, finding that child models rely on syntactic cues like post-verbal determiners, while adult models use semantic features, supporting syntactic bootstrapping theory.
The article humorously or alarmingly discusses the potential dangers and exaggerated capabilities of OpenAI's GPT-2 AI model.
This GitHub repository provides CMake scripts to execute GPT-2 using Q16.16 integer arithmetic, including instructions for running both full and toy models.
A researcher built a mini version of the Kimi-K3 model with 1.02 billion parameters for under $250, achieving a 33.4% HellaSwag score that surpasses GPT-2's 28% score.
Reverse-engineered the RK3588 NPU and built an open compiler and runtime to run GPT-2 and SigLIP from PyTorch, ONNX, and JAX without vendor SDK.
This paper trains five GPT-2-style models from scratch to compare dedicated monolingual models for Tamil, Telugu, Kannada, and Malayalam against a joint multilingual model, finding monolingual models outperform mGPT on sentiment classification and NER with more efficient tokenizers.
Andrej Karpathy released nanochat, a minimalist LLM training framework that can train a GPT-2-level model for just around $48, covering the entire pipeline of pretraining, fine-tuning, and reinforcement learning, with minimal and fully transparent code.
A follow-up project visualizing GPT-2's 32,070 tokens as an interactive hyperbolic tree in a Poincaré ball, allowing users to fly through the embedding space.
An interactive map of GPT-2's token embedding space that lets users tap any token to explore its embeddings.
This paper examines the embedding geometry of GPT-2 Small around the token 'Trump', comparing discretized and continuous nearest neighbor approaches to understand representational structure.
This paper investigates how GPT-2 models pre-trained on impossible languages (with disrupted information locality) can recover natural English, showing a bias toward shorter dependency lengths and dissociation between structural and surface recovery.
The BABEL codec achieves the first complete decode of GPT-2 small's internal state, reconstructing 94.7% of its behavior and enabling reading and writing to the model in English, with open source resources and a demo.
This paper introduces Manifestation Units, a typed tuple protocol for organizing per-component statistics from mechanistic interpretability analyses into structured, queryable fields. The protocol is demonstrated across vision (β-VAE, CNN) and language (GPT-2) models, showing improved retrieval and causal sufficiency.
This paper re-derives activation patching from causal mediation analysis, revealing that the natural indirect effect (NIE) captures not only a component's causal effect but also interaction effects with other components. It demonstrates these hidden interactions in the GPT-2 IOI circuit and argues that they are a diagnostic tool rather than a nuisance.
NanoEuler is a GPT-2-scale language model built entirely from scratch in C/CUDA without any ML libraries, including hand-written forward/backward passes, a byte-level BPE tokenizer, and training pipeline. The project is an educational artifact demonstrating the engineering behind transformer training and runs on a single RTX 4070.
OpenAI trained the GPT-2 language model but deemed it too dangerous to release to the public due to potential misuse.
This paper argues that current world models lack a persistent state core, proposing a hybrid approach that adds temporal-causal structure via η-pseudo-unitary operator dynamics to convert pretrained GPT-2 into a time-reasoning model.
This paper examines whether language models can independently discover the concept of zero as a form of out-of-distribution generalization, finding that GPT-2 sized models cannot at test time but improve with training on examples of zero, and that language pretraining reduces the number of required examples.
This paper investigates converting pretrained GPT-2 into a time-reasoning model using η-pseudo-unitary operator dynamics, providing mathematical foundations and key findings on PT-breaking transitions and reversible/irreversible sequences.