Papers

Filtered to action generation · clear filter

Browse by term

continual learning 83reinforcement learning 50large language models 12benchmarking 11benchmarks 11language models 11vision-language models 11robotics 7natural language processing 6world models 6generative models 5recursive self-improvement 5attention mechanisms 4diffusion Transformers 4multi-agent systems 4multimodal learning 4multimodal models 4on-policy distillation 4self-distillation 4self-supervised learning 4transformers 4video generation 4vision-language-action models 4agent-based systems 3agentic models 3agentic search 3autonomous systems 3coding agents 3diffusion models 3image generation 3

Matching papers

Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

62 upvotes · 11 SEP 2026 · Jianman Lin, Shailesh Shailesh, Zhongyi Luo et al.

This paper proposes a method called Latent Interface Training (LIT) to improve the generalization of robotics foundation models by preventing them from relying on visual shortcuts when learning to generate actions from pre-trained visual representations. This is important for robots to perform well in new, unseen environments.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

59 upvotes · 16 JUL 2026 · Xiaomi Robotics Team, Jun Guo, Piaopiao Jin et al.

This paper introduces a vision-language-action model that can perform mobile manipulation tasks in unseen environments with minimal training data, and how it can be scaled up to achieve better performance. Practitioners might care about this model for building robots that can adapt to new tasks with minimal fine-tuning.