Tag
This paper investigates how the faithfulness of latent reasoning steps evolves during training, finding that it depends on training stage and answer format, rather than just final checkpoint performance.