Papers

Filtered to video generation · clear filter

Browse by term

continual learning 64reinforcement learning 34large language models 17benchmarking 12vision-language models 10generative models 8language models 8video generation 7multimodal models 6natural language processing 6robotics 6world models 6benchmarks 5diffusion models 5on-policy distillation 5policy optimization 5scalability 5self-distillation 5vision-language-action models 5autoregressive models 4computer vision 4diffusion transformers 4LLMs 4multimodal large language models 4verifiable rewards 4attention mechanisms 3embodied intelligence 3image editing 3long-term memory 3multimodal learning 3

Matching papers

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

52 upvotes · 20 JUL 2026 · Yiyang Cai, Nan Chen, Rongchang Xie et al.

This paper develops a video personalization method that focuses on human-object interactions, aiming to improve the accuracy of video generation by better understanding human-object relationships and incorporating intra-subject references. Practitioners may care about this research as it could lead to more realistic and engaging video content.

Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

40 upvotes · 17 JUL 2026 · Runmao Yao, Kairui Hu, Yukang Cao et al.

This paper introduces a benchmark to evaluate video generation models' ability to reason about physical laws, which is crucial for creating reliable world simulators. Practitioners caring about developing more realistic and physically intelligent AI models will find this research valuable.

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

32 upvotes · 27 JUL 2026 · Haopeng Li, Yitong Li, Junsong Chen et al.

This paper improves the efficiency of video generation by reducing the computational load of attention, a key bottleneck in diffusion transformers. Practitioners can benefit from the speedup achieved by Sol-Attn, which enables faster video generation and editing without compromising visual quality.

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

12 upvotes · 23 JUL 2026 · Sicheng Mo, Yuheng Li, Ziyang Leng et al.

This paper introduces a new method for generating videos in multi-agent environments, where each agent has its own view of the world. It's useful for applications like video games or simulations where multiple agents need to interact with each other and the environment.

Parallel Decoding Distillation for Fast Image and Video Generation

10 upvotes · 28 JUL 2026 · Neta Shaul, Chao Liu, Arash Vahdat et al.

This paper introduces Parallel Decoding Distillation (PDD), a new method to speed up image and video generation in diffusion and flow models, achieving state-of-the-art performance while improving video diversity. Practitioners might care about PDD's potential to accelerate video generation tasks.

FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation

8 upvotes · 17 JUL 2026 · Hao Liu, Chenghuan Huang, Ye Huang et al.

This paper develops a more efficient way to generate high-quality videos by balancing the workload across multiple GPUs during training, which can improve the performance of video generation models like those used in FVAttn. Practitioners in video generation and deep learning might care about this research because it can lead to faster and more efficient video generation models.