Tag
This paper studies adaptive depth in looped Transformers by separating trajectory formation from exit readout, showing that fixed-prior depth supervision often outperforms jointly trained gates and that poor adaptive-compute performance stems more from the induced trajectory than from gate expressivity.