value-degradation

Tag

Cards List
#value-degradation

Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training

arXiv cs.CL ↗ · 2026-06-26 Cached

This paper investigates how post-training (SFT and RL) on helpfulness vs. coding data degrades compassion values in Llama 3.1 8B models that were mid-trained on compassion-oriented data. It finds that helpfulness training significantly reduces animal compassion and general moral reasoning compared to coding training, though the reasoning effect does not transfer cross-lingually.

0 favorites 0 likes
← Back to home

Submit Feedback