N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
This paper introduces a vision-tactile-language-action model that can perform fine-grained manipulation with tactile perception and control, and improve its policy offline from stored data. Practitioners may care because this model can be used to create more versatile and accurate tactile-driven manipulation policies.