Tag
Latent Interface Training (LIT) addresses the issue of robot policies learning shortcuts from training images by first teaching action without visual input and then using a pose-supervised latent interface to preserve geometry for action.
The paper proposes Latent Interface Training (LIT), a two-stage strategy to enhance generalization in robotics foundation models by mitigating vision-action shortcuts through pose-supervised latent interfaces.