secret-extraction

Tag

Cards List
#secret-extraction

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets

arXiv cs.AI · 2026-07-10 Cached

Introduces 'overthinking', a technique that amplifies reasoning weights from reasoning-distilled models to induce disclosure of hidden information in language models, demonstrating up to 10x greater secret leakage across 2B-32B models.

0 favorites 0 likes
← Back to home

Submit Feedback