Tag
This paper systematically evaluates LLM compression methods and reveals asymmetric harms, such as disproportionate loss of head knowledge and increased confidence in incorrect answers, which are masked by standard aggregate metrics.
This paper investigates whether language models can learn new facts in their weights through continual learning. Using invented facts and sequential writes into Qwen3 models, it finds that training data breadth determines knowledge type and retention: bare-statement facts are quickly forgotten (1% accuracy after 20 writes), while facts learned from diverse restatements retain 46% accuracy. Forgotten facts are not erased but become behaviorally inaccessible due to later writes redirecting questions, and context remains the reliable channel for fact composition and survival.
Anthropic master shared a prompt technique: let the AI first hide high-level concepts in a fable, then reveal and explain the metaphor to aid memory and understanding. This method leverages the human brain's preference for stories and images, making it easier to remember than directly giving definitions.
This paper introduces Act2Answer, a protocol to evaluate knowledge retention in Vision-Language-Action (VLA) models by requiring agents to answer questions through physical actions. It finds that VLAs retain basic knowledge but show gaps on richer semantic categories, and that VQA co-training helps.