Tag
KItCAT introduces a lightweight training strategy for auto-regressive LLMs that uses input corruption to generate diverse training inputs, improving knowledge injection from niche documents without costly paraphrasing.
This paper investigates how neural networks maintain high accuracy even when over 90% of input features are corrupted, deriving a centroid-based decision rule in the high-noise limit using a mean-field approach.