Tag
This paper proposes SGRE, an Answer-then-Edit framework that modifies reasoning traces of LLMs to prevent unauthorized knowledge distillation while preserving answer accuracy and naturalness, drawing on cognitive load theory to increase extraneous load on student models.
Anthropic is removing hidden steganographic codes from its Claude Code tool that were used to detect and prevent unauthorized use and model distillation by Chinese competitors, citing the deployment of stronger mitigations.
This paper proposes Lossless Anti-Distillation Sampling (LADS), a novel sampling scheme that counters multi-account distillation by correlating responses across accounts while preserving exact statistical fidelity for individual benign users. Theoretical analysis and experiments show LADS degrades distilled student performance on image, math, and code generation.
Researchers propose trace rewriting methods to prevent unauthorized LLM knowledge distillation while preserving answer correctness and embedding detectable watermarks.