reasoning-fine-tuning

Tag

Cards List
#reasoning-fine-tuning

Reasoning Fine-Tuning Induces Persistent Latent Policy States

arXiv cs.CL · 2026-07-22 Cached

This paper models Chain-of-Thought reasoning as a switching dynamical system, showing that reasoning fine-tuning globally reorganizes latent policy states, leading to improved multi-step reasoning. The proposed framework combines time-aware contrastive learning with discrete regime discovery, and experiments demonstrate that fine-tuned models exhibit richer latent-policy organization with functional specialization.

0 favorites 0 likes
← Back to home

Submit Feedback