Papers

Filtered to generative models · clear filter

Browse by term

continual learning 64reinforcement learning 34large language models 17benchmarking 12vision-language models 10generative models 8language models 8video generation 7multimodal models 6natural language processing 6robotics 6world models 6benchmarks 5diffusion models 5on-policy distillation 5policy optimization 5scalability 5self-distillation 5vision-language-action models 5autoregressive models 4computer vision 4diffusion transformers 4LLMs 4multimodal large language models 4verifiable rewards 4attention mechanisms 3embodied intelligence 3image editing 3long-term memory 3multimodal learning 3

Matching papers

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry

31 upvotes · 20 JUL 2026 · Dingyun Zhang, Lixue Gong, Wei Liu

This paper creates a new AI model that can edit and generate videos without needing masks, and can also learn to mimic image editing capabilities. Practitioners might care about this because it could lead to more diverse and realistic video editing data, and enable AI models to understand and generate human-like video editing instructions.

Self Gradient Forcing: Native Long Video Extrapolation

30 upvotes · 22 JUL 2026 · Junhao Zhuang, Shiyi Zhang, Yuxuan Bian et al.

This paper proposes a new training method for autoregressive video diffusion models called Self Gradient Forcing, which helps them better remember and use past information to generate future frames. Practitioners might care about this because it could lead to more realistic and stable video extrapolation.

Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On

26 upvotes · 23 JUL 2026 · Yong Liu, Xiaolong Fu, Zihang Xu et al.

This paper introduces Oxygen-TryOn, a new AI model that can generate photorealistic images of people wearing any fashion item, in any setting. Practitioners in the fashion industry might care about this model because it can revolutionize virtual try-on, allowing for more realistic and diverse scenarios.

Wonder: Video World Model Done Better

17 upvotes · 28 JUL 2026 · Jiacong Xu, Hanwen Jiang, Zhixin Shu et al.

This paper introduces Wonder, a system that can generate videos of a virtual world in real-time, allowing users to interact with and explore the environment. Practitioners in fields like video game development or virtual reality might care about this technology for its potential to create more immersive experiences.

Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation

13 upvotes · 29 JUL 2026 · Alexi Gladstone, Heng Ji, Yilun Du

This paper introduces Explorative Modeling, a new approach to training generative models that allows for end-to-end generation by exploring multiple candidate matches between model generations and data. This can lead to improved performance and efficiency in various applications.

Three-Body Scattering for Generative Modeling

12 upvotes · 20 JUL 2026 · Peng Sun, Zhenglin Cheng, Deyuan Liu et al.

This paper introduces Three-Body Scattering Modeling (TBSM), a new approach to generative modeling that uses a distributional energy to induce motion and provide direct regression supervision for one-step generators. Practitioners might care because TBSM can achieve high-quality generation with reduced field noise and fewer sample-level losses.

Parallel Decoding Distillation for Fast Image and Video Generation

10 upvotes · 28 JUL 2026 · Neta Shaul, Chao Liu, Arash Vahdat et al.

This paper introduces Parallel Decoding Distillation (PDD), a new method to speed up image and video generation in diffusion and flow models, achieving state-of-the-art performance while improving video diversity. Practitioners might care about PDD's potential to accelerate video generation tasks.

Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling

7 upvotes · 26 JUL 2026 · Hengyuan Cao, Shizhuo Cheng, Mingxuan Liu et al.

This paper introduces Chamaileon, a new method for designing protein binders that can model multiple targets and conformations, which could help researchers and practitioners design more advanced proteins with specific functions. Practitioners might care about this approach because it could lead to breakthroughs in structural biology and protein engineering.