Papers

Filtered to vision-language-action models · clear filter

Browse by term

continual learning 83reinforcement learning 50large language models 12benchmarking 11benchmarks 11language models 11vision-language models 11robotics 7natural language processing 6world models 6generative models 5recursive self-improvement 5attention mechanisms 4diffusion Transformers 4multi-agent systems 4multimodal learning 4multimodal models 4on-policy distillation 4self-distillation 4self-supervised learning 4transformers 4video generation 4vision-language-action models 4agent-based systems 3agentic models 3agentic search 3autonomous systems 3coding agents 3diffusion models 3image generation 3

Matching papers

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

118 upvotes · 29 JUL 2026 · Hengyi Xie, Chenfei Yao, Xianjin Wu et al.

This paper introduces TurboVLA, a new vision-language-action model that reduces computation and memory overhead by directly exchanging information between visual observations and language instructions, allowing for faster and more efficient robotic manipulation. Practitioners might care about this approach for building more efficient and effective VLA models.

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

97 upvotes · 16 AUG 2026 · GigaBrain Team, Angen Ye, Axiang Sun et al.

This paper presents GigaBrain-0.7, a new embodied foundation model that achieves strong generalization across diverse robot embodiments and tasks, by improving the architecture and scaling it to large amounts of data. Practitioners might care about this research if they're working on developing generalist robots that can adapt to new tasks and environments.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

59 upvotes · 16 JUL 2026 · Xiaomi Robotics Team, Jun Guo, Piaopiao Jin et al.

This paper introduces a vision-language-action model that can perform mobile manipulation tasks in unseen environments with minimal training data, and how it can be scaled up to achieve better performance. Practitioners might care about this model for building robots that can adapt to new tasks with minimal fine-tuning.

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

55 upvotes · 27 AUG 2026 · Senqiao Yang, Chengyao Wang, Yuxin Chen et al.

This paper proposes a new approach to training Vision-Language-Action models by using a pre-trained backbone that captures generalizable visual-action knowledge from a large, diverse dataset of robot trajectories. This allows the model to perform well on new, unseen tasks without requiring a large amount of task-specific data.