118 upvotes · 29 JUL 2026 · Hengyi Xie, Chenfei Yao, Xianjin Wu et al.
This paper introduces TurboVLA, a new vision-language-action model that reduces computation and memory overhead by directly exchanging information between visual observations and language instructions, allowing for faster and more efficient robotic manipulation. Practitioners might care about this approach for building more efficient and effective VLA models.
59 upvotes · 16 JUL 2026 · Xiaomi Robotics Team, Jun Guo, Piaopiao Jin et al.
This paper introduces a vision-language-action model that can perform mobile manipulation tasks in unseen environments with minimal training data, and how it can be scaled up to achieve better performance. Practitioners might care about this model for building robots that can adapt to new tasks with minimal fine-tuning.
29 upvotes · 27 JUL 2026 · Yifan Ye, Yankai Fu, Yaoxu Lv et al.
This paper proposes a hierarchical structure for organizing data sources for embodied manipulation, aiming to balance scalability and robot alignment. Practitioners might care about this work if they're building or deploying embodied agents that require diverse and high-quality data.
4 upvotes · 17 JUL 2026 · Haoran Sun, Wentao Zhang, Junyang Hua et al.
This paper develops a service-oriented framework, JoyNexus, to efficiently train and deploy Vision-Language-Action models across multiple tenants, improving resource utilization and reducing costs. Practitioners may care about JoyNexus for its potential to streamline the training process and make VLA models more accessible.
4 upvotes · 13 JUL 2026 · Byungkun Lee, Dongyoon Hwang, Dongjin Kim et al.
This paper proposes a way to improve vision-language-action models so they can better understand the world from a robot's perspective, which is important for robots to make accurate decisions. By using robot-centric pointmaps, these models can generalize better across different camera setups and viewpoints.