Tag
The paper introduces MoFO, a momentum-filtered optimizer that mitigates forgetting in LLM fine-tuning by updating only parameters with large momentum magnitudes, preserving pre-trained knowledge without extra storage.