Papers

Filtered to video world models · clear filter

Browse by term

continual learning 64reinforcement learning 34large language models 17benchmarking 12vision-language models 10generative models 8language models 8video generation 7multimodal models 6natural language processing 6robotics 6world models 6benchmarks 5diffusion models 5on-policy distillation 5policy optimization 5scalability 5self-distillation 5vision-language-action models 5autoregressive models 4computer vision 4diffusion transformers 4LLMs 4multimodal large language models 4verifiable rewards 4attention mechanisms 3embodied intelligence 3image editing 3long-term memory 3multimodal learning 3

Matching papers

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

48 upvotes · 20 JUL 2026 · AlayaWorld Team, Kaipeng Zhang, Chuanhao Li et al.

This paper develops a system called AlayaWorld that can generate interactive virtual worlds from text, images, or videos, allowing for customizable and evolving environments. Practitioners in areas like game development, virtual reality, or interactive storytelling might care about this research for its potential to streamline the creation of immersive experiences.

Wonder: Video World Model Done Better

17 upvotes · 28 JUL 2026 · Jiacong Xu, Hanwen Jiang, Zhixin Shu et al.

This paper introduces Wonder, a system that can generate videos of a virtual world in real-time, allowing users to interact with and explore the environment. Practitioners in fields like video game development or virtual reality might care about this technology for its potential to create more immersive experiences.

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

17 upvotes · 30 JUL 2026 · Jin Cao, Zian Meng, Kaipeng Zhang

This paper introduces ShadowDancer, a method for teaching video world models to perform any-action control by learning unified dynamics representations from a video and its shadow. Practitioners might care because it can improve action transfer and long-term control in complex environments.