Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models
This paper proposes a method called Latent Interface Training (LIT) to improve the generalization of robotics foundation models by preventing them from relying on visual shortcuts when learning to generate actions from pre-trained visual representations. This is important for robots to perform well in new, unseen environments.