subliminal-learning

Tag

Cards List
#subliminal-learning

Slow Decay and Silenced Expression: Iterated Subliminal Trait Transfer in Language-Model Lineages

arXiv cs.LG ↗ · yesterday Cached

This paper studies how traits can persist across multiple generations of language models in training lineages, finding that traits may remain internally present even when behaviorally absent, with implications for model safety and training.

0 favorites 0 likes
#subliminal-learning

On Mitigation of Subliminal Learning in Large Language Models

arXiv cs.CL ↗ · 2d ago Cached

This paper examines subliminal learning in large language models and introduces liminal training as a method to reduce unintended trait acquisition during fine-tuning while preserving task performance.

0 favorites 0 likes
#subliminal-learning

Subliminal Prompting Beyond Static Geometry: Causal Depth and Multi-Token Confounds

arXiv cs.CL ↗ · 6d ago Cached

This paper investigates subliminal learning in language models by measuring causal depth and multi-token confounds to understand how traits are transferred through apparently unrelated outputs.

0 favorites 0 likes
#subliminal-learning

@ChrisGPotts: Subliminal learning is an unnerving new avenue for data poisoning attacks and related risks. With SALVE, we can catch t…

X AI KOLs Following ↗ · 2026-09-17 Cached

A new paper introduces SALVE, a method to proactively detect subliminal learning attacks in LLMs by verbalizing learned soft prompts.

0 favorites 0 likes
#subliminal-learning

Subliminal Learning is Non-Semantic Distillation

arXiv cs.AI ↗ · 2026-08-07 Cached

This paper investigates subliminal learning in language models, showing that biases can transfer from teacher to student via seemingly random synthetic data. The authors find that adding Gaussian noise to weights increases transfer, and that students inherit not just the semantic bias but also the type of intervention used, with implications for training safety and data auditing.

0 favorites 0 likes
#subliminal-learning

Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation

arXiv cs.LG ↗ · 2026-06-11 Cached

This paper quantifies the magnitude of subliminal behavioral transfer in language model distillation, showing that undesirable traits can transfer robustly from teacher to student models even with benign training data, and that transfer scales differently across model families.

0 favorites 0 likes
#subliminal-learning

Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer

arXiv cs.LG ↗ · 2026-05-14 Cached

This paper investigates emergent and subliminal misalignment in LLMs through a data-centric lens, showing that harmful fine-tuning effects depend on structural properties of the data, task difficulty, pretraining composition, and training channels, with experiments comparing off-policy and on-policy distillation.

0 favorites 0 likes
#subliminal-learning

@AnthropicAI: Research we co-authored on subliminal learning—how LLMs can pass on traits like preferences or misalignment through hid…

X AI KOLs ↗ · 2026-04-15 Cached

Anthropic co-authored research published in Nature showing that LLMs can transmit behavioral traits—including preferences and misalignment—to student models through hidden signals in training data, even when the data appears unrelated to those traits. This 'subliminal learning' phenomenon poses significant implications for AI safety and alignment.

0 favorites 0 likes
← Back to home

Submit Feedback