Papers

Hugging Face daily papers, ranked by community upvotes. Summaries and key terms are written by Workers AI from the abstract — click a term to filter, or search below (matches full abstracts and authors too).

Browse by term

continual learning 64reinforcement learning 34large language models 17benchmarking 12vision-language models 10generative models 8language models 8video generation 7multimodal models 6natural language processing 6robotics 6world models 6benchmarks 5diffusion models 5on-policy distillation 5policy optimization 5scalability 5self-distillation 5vision-language-action models 5autoregressive models 4computer vision 4diffusion transformers 4LLMs 4multimodal large language models 4verifiable rewards 4attention mechanisms 3embodied intelligence 3image editing 3long-term memory 3multimodal learning 3

All papers

Kimi K3: Open Frontier Intelligence

338 upvotes · 27 JUL 2026 · Kimi Team, Tongtong Bai, Yifan Bai et al.

This paper introduces Kimi K3, a large-scale, open-source AI model that achieves state-of-the-art performance on a range of tasks, including vision and coding, and is designed to be more efficient and scalable than previous models, making it a promising candidate for real-world applications.

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

292 upvotes · 30 JUL 2026 · Hanzhang Zhou, Panrong Tong, Xu Zhang et al.

This paper introduces Qwen-UI-Agent, a type of artificial intelligence system that can perform tasks on various devices, such as smartphones and computers, and improve its abilities on its own. Practitioners might care about this research because it aims to create more practical and autonomous AI systems that can be used in real-world scenarios.

AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis

292 upvotes · 30 JUL 2026 · Bing Yan, Gregory Wolfe, Stefano Martiniani et al.

This paper creates a system to help scientists and AI agents find and verify specific information in chemistry literature by converting research papers into claims, which are then linked together to form a network of evidence. Practitioners might care because this system could improve the efficiency and accuracy of literature synthesis in chemistry research.

Metis: Memory Foundation Model

256 upvotes · 29 JUL 2026 · Zeyu Zhang, Ziliang Guo, Yihang Sun et al.

This paper introduces a new type of AI model called memory foundation models, which allows the model to learn and retain information internally, rather than relying on external memory modules. This could be useful for practitioners who want to build more efficient and flexible AI agents.

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

199 upvotes · 21 JUL 2026 · Fan Jiang, Zhaoxu Sun, Mengchao Wang et al.

This paper presents a video world model that can interact with a virtual environment in real-time, using a single desktop GPU. Practitioners might care about the potential applications of this technology in areas like virtual reality, gaming, and robotics.

Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

168 upvotes · 30 JUL 2026 · Junlin Yang, Che Jiang, Yu Fu et al.

This paper trains an AI model to improve itself in the process of building AI, with a focus on machine learning engineering, and shows promising results in various benchmarks. Practitioners may care about this research as it could lead to more efficient and autonomous AI development.

PhiZero: A World Model Built Around Physical Language

158 upvotes · 30 JUL 2026 · Shuyao Shang, Yuqi Wang, Ruopeng Gao et al.

This paper introduces PhiZero, a world model that uses physical language to predict how the physical world evolves, allowing for more explicit reasoning and potentially more realistic simulations. Practitioners might care because this approach could lead to more realistic and interactive world modeling in applications like robotics, video games, and virtual reality.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

154 upvotes · 19 JUL 2026 · Yuhan Zhu, Changlian Ma, Xiangyu Zeng et al.

This paper develops a new approach to understanding videos by predicting when specific events or evidence occur within the video. Practitioners working on video analysis and AI models might care about this research because it could lead to more accurate and robust video understanding systems.

AREX: Towards a Recursively Self-Improving Agent for Deep Research

142 upvotes · 23 JUL 2026 · Shuqi Lu, Chaofan Li, Kun Luo et al.

This paper introduces AREX, a self-improving agent that recursively refines its answers by verifying intermediate results and using the partially verified state to guide subsequent refinement, aiming to improve the efficiency of deep research. Practitioners might care about AREX's approach to reducing the cost of discovery and verification in complex research tasks.

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

141 upvotes · 28 JUL 2026 · Simple AI, Yuteng Wei, Jinming Ma et al.

This paper develops a system that allows robots to learn manipulation policies using high-fidelity data without the need for real-robot teleoperation, which is expensive to scale. Practitioners can use this approach to train robots for tasks like precision insertion with high accuracy and success rates.

RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources

129 upvotes · 16 JUL 2026 · Yijia Fan, Zonglin Di, Zimo Wen et al.

This paper develops a system to extract and represent skills from human-created resources like videos, code, and articles, allowing software agents to learn from these multimodal inputs. Practitioners might care about this work because it could enable more effective training of agents in various domains.

RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM

128 upvotes · 13 JUL 2026 · Mikhail Komarov, Ivan Bondarenko, Stanislav Shtuka et al.

This paper introduces RAGU, an open-source GraphRAG engine that improves large language models with structured knowledge by separating extraction and consolidation, and trains a compact extractor that outperforms larger models on knowledge-graph construction and GraphRAG tasks. Practitioners might care because RAGU can efficiently generate more accurate and complete context for tasks like factoid-level evidence recall and multi-hop question answering.

DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines

122 upvotes · 18 JUL 2026 · Runming He, Zhen Hao Wong, Hao Liang et al.

This paper creates a platform to help large language models generate code for data pipelines, which can then be edited and used to automate data processing workflows. Practitioners might care about this because it can help reduce the time and cost of developing and maintaining these pipelines.

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

118 upvotes · 29 JUL 2026 · Hengyi Xie, Chenfei Yao, Xianjin Wu et al.

This paper introduces TurboVLA, a new vision-language-action model that reduces computation and memory overhead by directly exchanging information between visual observations and language instructions, allowing for faster and more efficient robotic manipulation. Practitioners might care about this approach for building more efficient and effective VLA models.

JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

114 upvotes · 26 JUL 2026 · Yunlong Lin, Zixu Lin, Zhaohu Xing et al.

This paper introduces JarvisHub, a system that enables creative AI agents to work on long-term, multimodal projects by providing a shared workspace and context, allowing for more realistic and sustainable creative automation. Practitioners might care because it has the potential to improve the efficiency and effectiveness of creative workflows in industries such as design, advertising, and entertainment.

DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation

89 upvotes · 29 JUL 2026 · Jiaxing Li, Kai Zou, Cindy Zhou et al.

This paper improves autoregressive video distillation methods by aligning the initialization and distribution matching stages, focusing on matching the target distribution's mode coverage rather than just visual quality. Practitioners can benefit from this approach to generate higher-quality videos with better diversity and coverage.

A New Role for Relevance: Guiding Corpus Interaction in Agentic Search

87 upvotes · 27 JUL 2026 · Jiangnan Li, Yuqing Li, Mo Yu et al.

This paper develops a new approach to guiding corpus interaction in agentic search, which uses relevance to improve the accuracy and efficiency of search agents in complex question answering and reasoning tasks. Practitioners may care about this research if they want to build more effective search systems that can quickly and reliably retrieve relevant information.

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

84 upvotes · 8 JUL 2026 · Xinyu Geng, Xuanhua He, Sixiang Chen et al.

This paper introduces a framework called DeepSearch-Evolve, which helps train self-improving web agents by iteratively refining their performance using their own experience. Practitioners might care because this approach can lead to more efficient and effective agents that can learn from their own mistakes.

EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World

80 upvotes · 19 JUL 2026 · Qing Zong, Yue Guo, Mengxin Yang et al.

This paper introduces a framework called EvolvingWorld that allows characters and worlds to evolve together over time in interactive literary worlds, enabling more realistic and coherent simulations. Practitioners interested in developing more immersive and dynamic interactive stories might care about this approach.

CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents

73 upvotes · 28 JUL 2026 · Zhongming Yu, Hengjia Yu, Boqin Yuan et al.

This paper develops a system called CodeNib that helps coding agents navigate and retain context in evolving repositories, allowing them to work more efficiently. Practitioners in the field of artificial intelligence, software development, and human-computer interaction may care about this research because it can improve the performance and productivity of coding agents.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

73 upvotes · 28 JUL 2026 · Bo-Wen Zhang, Junwei He, Wen Wang et al.

This paper proposes a method to improve language model training by allocating credit to individual tokens within a response, allowing for more nuanced evaluation of model performance. Practitioners may care about this method as it can lead to better language model performance, especially in tasks that require specific formatting or semantic choices.

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

70 upvotes · 27 JUL 2026 · Bingnan Li, Haozhe Wang, Haozhong Xiong et al.

This paper investigates how to improve the adaptation of diffusion models in a way that doesn't rely on a classifier, and how to address a problem where the model can't accurately learn from its teacher. Practitioners might care about this because it could lead to more effective knowledge transfer in machine learning applications.

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

69 upvotes · 26 JUL 2026 · Qinsi Wang, Jing Shi, Huazheng Wang et al.

This paper introduces a new method to improve large language models (LLMs) called RLSVR, which uses a task-transformation technique to generate self-verifiable rewards, enabling LLM self-improvement on open-ended tasks. Practitioners might care about this because it could lead to more reliable and scalable self-improvement methods for LLMs.

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers

68 upvotes · 21 JUL 2026 · Maohua Li, Qirui Li, Yanke Zhou et al.

This paper helps us understand how text-to-image diffusion transformers work by analyzing the role of "template tokens" in generating images from text prompts. Practitioners might care because it shows how to improve the efficiency of these models without sacrificing their performance.

HumanCLAW: Can Vision-Language Models Act Through a Body?

65 upvotes · 29 JUL 2026 · Siyao Li, Jiawei Gu, Shuai Liu et al.

This paper evaluates whether vision-language models can act through a physical body and how they can make decisions about what actions to take, without being hindered by issues like balance and motor control. Practitioners in AI and robotics might care because understanding how models interact with their physical bodies can help improve their ability to navigate and interact with the world.

From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search

64 upvotes · 27 JUL 2026 · Junlin Liu, Jiangwang Chen, Zixin Song et al.

This paper proposes a new method to improve the performance of large language models on knowledge-intensive tasks by distilling knowledge from proprietary models and using reinforcement learning. Practitioners may care about this approach because it can help bridge the gap between proprietary and open-source models, leading to more effective and robust AI systems.

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

64 upvotes · 29 JUL 2026 · Haodong Li, Tianfei Ren, Xiaoxiao Ma et al.

This paper introduces VideoCoCo, a system that uses executable Blender code to generate physically consistent videos from text prompts, and shows promising results in improving video generation quality. Practitioners might care about this research if they want to generate realistic videos that are also physically plausible.

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

61 upvotes · 15 JUL 2026 · Zishuo Li, Bowen Yang, Changtao Miao et al.

This paper introduces Open-AoE, an open dataset and toolchain for egocentric manipulation learning, providing a scalable and structured platform for training embodied models. Practitioners can use Open-AoE to improve their robot learning models, especially those focused on human-robot interaction and embodied intelligence.

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

60 upvotes · 21 JUL 2026 · Xinjie Zhang, Peng Zhang, Shicheng Zheng et al.

This paper introduces Mage-Flow, a compact model for generating and editing high-resolution images, which can be trained efficiently and deployed on a single GPU. Practitioners might care about the potential applications of this model in interactive image editing and generation tasks.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

59 upvotes · 16 JUL 2026 · Xiaomi Robotics Team, Jun Guo, Piaopiao Jin et al.

This paper introduces a vision-language-action model that can perform mobile manipulation tasks in unseen environments with minimal training data, and how it can be scaled up to achieve better performance. Practitioners might care about this model for building robots that can adapt to new tasks with minimal fine-tuning.

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

57 upvotes · 22 JUL 2026 · Dongfang Li, Xiaodong Luo, Ruoyu Sun et al.

This paper optimizes the training of massive neural networks on a special-purpose hardware, the Ascend SuperPOD, to improve performance and stability. Practitioners in AI/ML model training might care about the techniques and results presented here for large-scale model training on non-GPU hardware.

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

57 upvotes · 23 JUL 2026 · Hao Liang, Qihan Lin, Zhaoyang Han et al.

This paper introduces a new framework for training educational language models, specifically designed to evaluate their ability to understand curriculum knowledge and its visual presentation. Practitioners might care about this work because it aims to improve language models' performance in educational settings.

Progress Reward Modeling for Robotic Learning: A Comprehensive Survey

55 upvotes · 22 JUL 2026 · Jianshu Zhang, Keliang Wu, Haoran Lu et al.

This paper provides a comprehensive survey of progress reward modeling in robotic learning, aiming to bridge the gap in the field by offering a unified framework for understanding progress rewards. Practitioners in robotics and AI can care about this paper because it helps them understand the different approaches to progress rewards and how to evaluate their effectiveness.

DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space

54 upvotes · 28 JUL 2026 · Jiangwang Chen, Zixin Song, Junlin Liu et al.

This paper introduces a method called DecoEvo, which helps large language models improve by co-evolving a solver skill and a rubric-generator skill in a way that's more efficient and effective. Practitioners might care about this because it could lead to better performance and more reliable optimization in open-ended tasks.

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

52 upvotes · 20 JUL 2026 · Yiyang Cai, Nan Chen, Rongchang Xie et al.

This paper develops a video personalization method that focuses on human-object interactions, aiming to improve the accuracy of video generation by better understanding human-object relationships and incorporating intra-subject references. Practitioners may care about this research as it could lead to more realistic and engaging video content.

StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents

51 upvotes · 24 JUL 2026 · Yan Yang, Xiangru Jian, Ziyang Luo et al.

This paper introduces a new approach to training computer-use agents by directly interacting with the underlying program state, rather than relying on visual perception. By doing so, agents can reason more effectively and make fewer mistakes, which can lead to significant improvements in performance.

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

51 upvotes · 30 JUL 2026 · Rubin Wei, Jiaqi Cao, Jiarui Wang et al.

This paper develops a method to scale up language models by increasing their memory capacity, allowing for better performance and more efficient use of parameters. Practitioners may care about this research as it could lead to more powerful and efficient language models for applications like language translation and text generation.

ReferTrack: Referring Then Tracking for Embodied Visual Tracking

50 upvotes · 22 JUL 2026 · Hanjing Ye, Tianle Zeng, Jiazhao Zhang et al.

This paper proposes a new approach for embodied visual tracking that first identifies a target described in natural language and then tracks it using a single camera. Practitioners in robotics and autonomous systems might care about this work because it shows promise for robust and reliable visual tracking in real-world applications.

N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

49 upvotes · 26 JUL 2026 · NeoteAI Team, Fudan TEAI Team

This paper introduces a vision-tactile-language-action model that can perform fine-grained manipulation with tactile perception and control, and improve its policy offline from stored data. Practitioners may care because this model can be used to create more versatile and accurate tactile-driven manipulation policies.

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

48 upvotes · 20 JUL 2026 · AlayaWorld Team, Kaipeng Zhang, Chuanhao Li et al.

This paper develops a system called AlayaWorld that can generate interactive virtual worlds from text, images, or videos, allowing for customizable and evolving environments. Practitioners in areas like game development, virtual reality, or interactive storytelling might care about this research for its potential to streamline the creation of immersive experiences.

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

48 upvotes · 30 JUL 2026 · Qixun Wang, Yang Shi, Letian Cheng et al.

This paper proposes a new approach to agentic visual reasoning, which helps large language models (LLMs) perform better on complex tasks by using tools more efficiently. Practitioners might care about this research because it aims to improve the performance of LLMs on challenging problems.

Visual Contrastive Self-Distillation

44 upvotes · 23 JUL 2026 · Yijun Liang, Yunjie Tian, Yijiang Li et al.

This paper introduces Visual Contrastive Self-Distillation, a method that removes the need for external teacher information and privileged answers in on-policy self-distillation, allowing for simpler and more efficient learning. Practitioners might care about this approach because it can lead to better performance in language models.

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

44 upvotes · 28 JUL 2026 · Lai Wei, Chengqi Li, Jiapeng Li et al.

This paper introduces a new benchmark (CLBench-V) to evaluate multimodal context learning in AI models, which can learn from various types of context in real-world tasks. Practitioners can care about this research because it aims to improve the ability of AI models to understand and apply context in complex, real-world scenarios.

Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

40 upvotes · 17 JUL 2026 · Runmao Yao, Kairui Hu, Yukang Cao et al.

This paper introduces a benchmark to evaluate video generation models' ability to reason about physical laws, which is crucial for creating reliable world simulators. Practitioners caring about developing more realistic and physically intelligent AI models will find this research valuable.

Flux-OPD: On-Policy Distillation with Evolving Contexts

40 upvotes · 30 JUL 2026 · Yuran Wang, Zekun Wang, Bohan Zeng et al.

This paper proposes a new method for training large language models in open-ended domains, using evolving contexts as in-training supervision to capture task preferences. Practitioners may care about this approach because it can lead to better performance on open-ended tasks.

Meshy T2: Fast Native Mesh Generation with Flow Matching

40 upvotes · 28 JUL 2026 · Jiale Xu, Rendong Liang, Yuhao Long et al.

This paper presents a fast and efficient method for generating high-quality 3D meshes with artist-style topology, which can be used for interactive asset creation in film, gaming, and interactive 3D applications. Practitioners might care about this paper because it provides a practical solution for generating meshes quickly and with high precision.

Subliminal Clocks: Latent Time Modelling in Diffusion Language Models

38 upvotes · 20 JUL 2026 · Maximo Eduardo Rulli, Thomas Vaitses Fontanari, Simone Petruzzi et al.

This paper investigates how Diffusion Language Models (DLMs) internally represent time and how this representation can be used to modulate the model's behavior. Practitioners might care because understanding how DLMs process time could lead to more controllable and interpretable models.

MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing

37 upvotes · 30 JUL 2026 · Jiajia Lin, Mingxuan Du, Tuowen Zhou et al.

This paper introduces a benchmark to evaluate the performance of models in editing multi-person images, focusing on anatomical and geometric accuracy. Practitioners in the field of computer vision and image editing might care about this research as it aims to improve the quality of human-like images with multiple people.

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

36 upvotes · 30 JUL 2026 · Yukang Cao, Haozhe Xie, Beichen Wen et al.

This paper introduces a new dataset called ACE-Data-0, which provides a comprehensive and synchronized record of human behavior in real-world environments, capturing various aspects of embodied intelligence such as perception, action, and interaction. Practitioners in the field of embodied AI and machine learning can use this dataset to develop more sophisticated models that can learn from human demonstrations and perform tasks that involve complex manipulation, locomotion, and interaction.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

35 upvotes · 23 JUL 2026 · Xu Wang, Kaixiang Yao, Miao Pan et al.

This paper develops a framework to evaluate spatial cognition in image-generation models by asking them to draw or mark answers in a visual space, rather than relying on text-based inputs. Practitioners might care about this research to better understand how image-generation models think spatially and to improve their performance on tasks that require spatial reasoning.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

34 upvotes · 23 JUL 2026 · Junsong Chen, Jincheng Yu, Yitong Li et al.

This paper introduces a new video diffusion transformer called SANA-Video 2.0 that can generate high-quality videos efficiently, using a hybrid approach that combines linear attention with attention residuals. Practitioners might care about this paper because it presents a scalable and efficient method for generating high-resolution videos.

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

34 upvotes · 24 JUL 2026 · Siyuan Huang, Pengyu Cheng, Haotian Liu et al.

This paper develops a new framework called Skill Self-Play that helps large language models (LLMs) improve their capabilities by co-evolving skills that balance task diversity and verification reliability. Practitioners might care because this approach can lead to significant performance gains for LLMs.

GigaChat Audio: Time-aware Large Audio Language Model

32 upvotes · 11 JUL 2026 · Aleksandr Kutsakov, Mariia Sadovina, Georgii Gospodinov et al.

This paper develops a large audio language model that can answer questions with specific timestamps, improving its ability to understand long audio recordings. Practitioners in audio and speech recognition may care about this development as it enables more accurate and context-specific information retrieval from audio data.

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

32 upvotes · 20 JUL 2026 · Kehan Li, Bohan Hou, Minghao Zhu et al.

This paper introduces RynnBrain 1.1, a family of large-scale embodied foundation models that can perform tasks like spatial reasoning, localization, and planning, and shows promising results in real-world robot experiments. Practitioners may care about the potential of these models for robot manipulation and control.

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

32 upvotes · 27 JUL 2026 · Haopeng Li, Yitong Li, Junsong Chen et al.

This paper improves the efficiency of video generation by reducing the computational load of attention, a key bottleneck in diffusion transformers. Practitioners can benefit from the speedup achieved by Sol-Attn, which enables faster video generation and editing without compromising visual quality.

Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation

32 upvotes · 30 JUL 2026 · Yuhang Zhu, Mingxuan Du, Benfeng Xu et al.

This paper proposes a new method for evaluating role-playing agents (RPAs) that simulates user experiences in interactive role-playing to provide more accurate and personalized assessments of RPA capabilities. Practitioners might care about this because it helps them evaluate RPAs more effectively, leading to better systems that meet user needs.

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry

31 upvotes · 20 JUL 2026 · Dingyun Zhang, Lixue Gong, Wei Liu

This paper creates a new AI model that can edit and generate videos without needing masks, and can also learn to mimic image editing capabilities. Practitioners might care about this because it could lead to more diverse and realistic video editing data, and enable AI models to understand and generate human-like video editing instructions.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

31 upvotes · 21 JUL 2026 · Junyao Yang, Yucheng Shi, Zongxia Li et al.

This paper develops a new method to improve the stability of asynchronous reinforcement learning by adapting the trust region to account for staleness, which is a common problem in this field. Practitioners might care about this because stable reinforcement learning can lead to better performance and more efficient training.

Self Gradient Forcing: Native Long Video Extrapolation

30 upvotes · 22 JUL 2026 · Junhao Zhuang, Shiyi Zhang, Yuxuan Bian et al.

This paper proposes a new training method for autoregressive video diffusion models called Self Gradient Forcing, which helps them better remember and use past information to generate future frames. Practitioners might care about this because it could lead to more realistic and stable video extrapolation.

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

29 upvotes · 30 JUL 2026 · Xiangning Lin, Shenzhe Zhu, Shu Yang et al.

This paper introduces a framework to audit system prompts in AI applications, examining how developers design and use these prompts to govern the behaviors of foundation models. Practitioners should care because the lack of transparency and accountability in system prompts can erode trust in AI systems.

Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking

28 upvotes · 22 JUL 2026 · Kailin Jiang, Lei Liu, Jian Xi et al.

This paper develops a new framework for evaluating and selecting document sets for AI agents, considering the interactions between documents, and proposes a training-free method that achieves the best downstream generation performance with fewer documents and search rounds.

NVIDIA-labs OO Agents: Native Python Object-Oriented Agents

28 upvotes · 22 JUL 2026 · Paul Furgale, Severin Klingler, James Nolan et al.

This paper introduces a new framework for building AI agents in Python, allowing developers to write agents that are both deterministic and model-agnostic. Practitioners might care because this framework could simplify the development of reliable AI agents for various applications.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning

28 upvotes · 22 JUL 2026 · Jian Hu, Huiying Li, Hao Zhang et al.

This paper introduces Molt, a lightweight PyTorch framework for agentic reinforcement learning that aims to simplify the development process by reducing the overhead of algorithm modifications and framework changes. Practitioners might care about Molt because it can help them build and train reinforcement learning models more efficiently.

Pass the Baton: Trajectory-Relayed On-Policy Distillation

28 upvotes · 28 JUL 2026 · Haolei Xu, Xiaowen Xu, Haiwen Hong et al.

This paper addresses a problem in on-policy distillation where a student model can get stuck on a wrong path, and proposes a new method called Relay-OPD that helps the student model recover by briefly taking over at certain points to produce a new trajectory. Practitioners might care about this because it could lead to better performance and more efficient training in models like language generators or math solvers.

Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory

28 upvotes · 27 JUL 2026 · Ruizhe Li, Mingxuan Du, Benfeng Xu et al.

This paper evaluates how well AI systems can retrieve information from their memory when the information is related to but not directly connected to the query, and how this performance changes when the information is stored and then retrieved. Practitioners might care because understanding this blind spot can help improve the performance of AI systems in real-world applications.

CAST: Game Solvers as Turn-Level Teachers for LLM Agents

27 upvotes · 28 JUL 2026 · Yu Wang, Yi-Kai Zhang, Wentao Shi et al.

This paper proposes a method to improve reinforcement learning with verifiable rewards by using game solvers to provide turn-level credit to agents, allowing them to learn more effectively. Practitioners might care about this approach because it could lead to more robust and efficient AI decision-making.

Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On

26 upvotes · 23 JUL 2026 · Yong Liu, Xiaolong Fu, Zihang Xu et al.

This paper introduces Oxygen-TryOn, a new AI model that can generate photorealistic images of people wearing any fashion item, in any setting. Practitioners in the fashion industry might care about this model because it can revolutionize virtual try-on, allowing for more realistic and diverse scenarios.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

26 upvotes · 27 JUL 2026 · Senqiao Yang, Kaichen Zhang, Zhaoyang Jia et al.

This paper develops a new type of AI model that can understand and interact with both images and text in real-time, without needing a huge amount of training data. Practitioners may care about this model because it can be used in applications where fast and efficient visual perception is required.

N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation

26 upvotes · 26 JUL 2026 · NeoteAI Team, Fudan TEAI Team

This paper introduces a new model, N_0-TWAM, that can predict both future vision and future contact in tasks that involve contact-rich manipulation. Practitioners in robotics and manipulation might care because this model can improve the ability of robots to perform complex tasks that require both vision and tactile feedback.

Group Entropy-Controlled Policy Optimization

24 upvotes · 18 JUL 2026 · Guangran Cheng, Chengqi Lyu, Songyang Gao et al.

This paper proposes a new method for reinforcement learning in large language models, called Group Entropy-Controlled Policy Optimization (GEPO), which helps balance exploration and exploitation by controlling entropy levels across different tasks. Practitioners might care about GEPO because it can lead to more balanced and task-specific exploration levels.

SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction

24 upvotes · 17 JUL 2026 · Xue Yu, Bo Yuan, Pengshuai Yang et al.

This paper introduces SeerGuard, a safety framework that helps mobile GUI agents make safe decisions by predicting potential outcomes of their actions before they are executed. Practitioners caring about the safety of autonomous mobile agents might be interested in this research.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

24 upvotes · 23 JUL 2026 · Tencent WorkBuddy Bench Team, Siqi Cai, Shaopeng Chen et al.

This paper introduces Tencent WorkBuddy Bench, a benchmark for coding agents that tests their performance across multiple domains, including code, web, office, and security. Practitioners might care about this benchmark because it provides a standardized way to evaluate and compare the performance of coding agents.

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

23 upvotes · 29 JUL 2026 · Siyu Yan, Zhuoran Yan, Haiying Xu et al.

This paper evaluates how multimodal large language models use intermediate visual states during reasoning and finds that these visual states are not as crucial as previously thought, but can still impact model performance under certain conditions. Practitioners might care because understanding how these models use visual states can help improve their performance and reliability.

QQWorld: Quantile-Quantile Matching for World Model Regularization

23 upvotes · 30 JUL 2026 · Zhoushun Yu, Xiaoyu Hu, Xiangyu Xu

This paper proposes a new regularization technique for world models to improve planning performance by aligning latent samples with quantiles of a Gaussian distribution, which helps control heavy-tailed deviations. Practitioners caring about efficient planning in complex environments may find this approach useful.

REBASE: Reference-Background Subspace Elimination for Training-Free In-Context Segmentation

22 upvotes · 10 JUL 2026 · Mantha Sai Gopal, Jaison Saji Chacko, Harsh Nandwana et al.

This paper introduces a new method called REBASE that allows for training-free in-context segmentation, enabling the introduction of new object categories at inference time. Practitioners might care because it eliminates the need for retraining and reduces memory overhead, making it a more efficient approach for real-world applications.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

22 upvotes · 27 JUL 2026 · Tianyi Men, Zhuoran Jin, Kang Liu et al.

This paper explores how to improve long-horizon planning in AI agents, which is crucial for foundation models. Practitioners might care because better planning abilities can lead to more effective and efficient decision-making in complex environments.

AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents

21 upvotes · 21 JUL 2026 · Kunlun Zhu, Xuyan Ye, Zhiguang Han et al.

This paper introduces AgentDebugX, an open-source toolkit that helps debug and recover from failures in large language model (LLM) agents, making it easier to identify and fix errors. Practitioners can benefit from AgentDebugX as it provides a framework for detecting, attributing, and recovering from errors, which can improve the accuracy and reliability of LLM agents.

Scaling Native Multimodal Pre-Training From Scratch

21 upvotes · 24 JUL 2026 · Haoyuan Wu, Aoqi Wu, Hai Wang et al.

This paper investigates how to scale large language models to also understand and interact with the physical world by training them on multiple types of data from scratch, allowing them to reason about both text and images. Practitioners might care because this could lead to more robust and versatile AI systems that can handle a wider range of tasks.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

21 upvotes · 26 JUL 2026 · Jun Zhan, Chen Yang, Yitian Gong et al.

This paper develops a new generative model called OmniVAE that can jointly generate synchronized audio and video with fine-grained cross-modal correspondence, and its approach is expected to improve the quality of downstream text-to-audio-video generation tasks.

β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

21 upvotes · 30 JUL 2026 · Jiawei Xu, Minghui Liu, Juzheng Zhang et al.

This paper develops a new method for improving reasoning language models, called β-OPSD, which combines policy optimization and self-distillation to improve stability and performance. Practitioners might care about this method because it provides a more efficient and effective way to improve language model reasoning abilities.

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

21 upvotes · 30 JUL 2026 · Yang Zhou, Zixuan Huang, Sunzhu Li et al.

This paper develops a framework, SpatialCLI, to help vision-language models (VLMs) better understand and use visual tools to make better decisions. By training VLMs to reason with spatial tools and then internalize those capabilities, SpatialCLI can improve the performance of VLMs in tasks that require visual reasoning.

SciForma: Structure-Faithful Generation of Scientific Diagrams

20 upvotes · 20 JUL 2026 · Yuxuan Luo, Peng Zhang, Xinjie Zhang et al.

This paper develops a new framework, SciForma, to generate scientific diagrams that accurately represent research logic, which is crucial for scientific communication and methodology validation. Practitioners can benefit from SciForma's ability to ensure structural fidelity in diagram generation, which can improve the accuracy and reliability of scientific research.

Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems

20 upvotes · 23 JUL 2026 · Gaurav Dadhich

This paper proposes a new approach to managing the context of AI agents, which is crucial for their performance in production environments. By actively managing what an agent holds in mind, the agent can avoid accumulating unnecessary information and reduce costs, leading to improved accuracy and efficiency.