Tag
The paper examines how length penalties applied during chain-of-thought reasoning can reduce the ability to monitor the reasoning process, raising concerns for interpretability and alignment.