Tag
The paper proposes Prior-Guided Tuning (PGT) and Contrastive Prior Steering (CPS) to use natural language priors as auxiliary learning signals, improving low-resource LLM training performance on tasks like AmbiMath, Jigsaw, and MNLI/HANS.
The paper challenges the assumption that cosine alignment between supervised latents and visual targets improves accuracy in vision-language models, finding a strong negative correlation. It introduces PRISM diagnostics revealing that answers are decoded downstream from latents, not within them, and that the auxiliary loss reshapes the language model via shared parameters.