Tag
Introduces 'overthinking', a technique that amplifies reasoning weights from reasoning-distilled models to induce disclosure of hidden information in language models, demonstrating up to 10x greater secret leakage across 2B-32B models.
The paper identifies 'Load-Bearing Wall' dimensions in pre-trained models that retain task-specific knowledge not fully captured by task vectors in model merging, and proposes PACT (PreserveAnchoredCores) to preserve these cores, achieving state-of-the-art performance across benchmarks.
This paper proposes using distributional alignment between task vector-based and in-context learning inference as a criterion for designing task vectors, and introduces Linear Task Vector (LTV) that minimizes next-token probability discrepancy via closed-form linear mapping. LTV achieves 9.2% average accuracy improvement over baselines across eight benchmarks and five LLMs.