Papers

Filtered to reinforcement learning · clear filter

Browse by term

continual learning 83reinforcement learning 50large language models 12benchmarking 11benchmarks 11language models 11vision-language models 11robotics 7natural language processing 6world models 6generative models 5recursive self-improvement 5attention mechanisms 4diffusion Transformers 4multi-agent systems 4multimodal learning 4multimodal models 4on-policy distillation 4self-distillation 4self-supervised learning 4transformers 4video generation 4vision-language-action models 4agent-based systems 3agentic models 3agentic search 3autonomous systems 3coding agents 3diffusion models 3image generation 3

Matching papers

Scaling Automatic Research Agents via World Models

452 upvotes · 29 AUG 2026 · Xiyuan Yang, Sheikh Sarwar, Jingru Cheng et al.

This paper proposes a method to scale automatic research agents by replacing environment execution with a world model, which can reduce training costs and improve performance. Practitioners might care about this approach because it can accelerate training times and lead to better results for complex AI tasks.

StudentSim: Training LLM-based Student Simulators

399 upvotes · 1 SEP 2026 · Ke Yang, Chenglong Wang, Michel Galley et al.

This paper develops a method to create personalized AI tutors that can adapt to individual students' strengths and weaknesses, using a combination of pooled training and per-student fine-tuning. Practitioners in education and AI development may care about this work as it could lead to more effective and personalized learning experiences.

Kimi K3: Open Frontier Intelligence

338 upvotes · 27 JUL 2026 · Kimi Team, Tongtong Bai, Yifan Bai et al.

This paper introduces Kimi K3, a large-scale, open-source AI model that achieves state-of-the-art performance on a range of tasks, including vision and coding, and is designed to be more efficient and scalable than previous models, making it a promising candidate for real-world applications.

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

256 upvotes · 14 SEP 2026 · Tong Zheng, Xidong Wu, Zheng Zhang et al.

This paper introduces Dream-RSI, a framework for recursive self-improvement in exploration, which helps autonomous AI agents discover high-value solutions more efficiently by using a replay simulator to provide low-cost feedback. Practitioners might care because effective exploration is crucial for AI progress, and Dream-RSI can improve discovery quality and reduce costs.

Recursive Synthesis for Long-Horizon Terminal Tasks

223 upvotes · 5 AUG 2026 · Zhongzhi Li, Yucheng Shi, Zongxia Li et al.

This paper introduces a method to generate long-horizon training tasks for terminal agents at scale, which can be used to improve AI models' performance on tasks like navigation and decision-making. Practitioners might care about this because it could help train more advanced AI models with better performance on complex tasks.

Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models

185 upvotes · 26 AUG 2026 · Pengfei Zhou, Hexin Wang, Zhengfeiyang Zhang et al.

This paper proposes a way to improve the efficiency of training world models by using game development as a source of reward signals and trajectory data, allowing for more effective post-training of large language models using reinforcement learning. Practitioners might care about this approach because it could lead to more scalable and effective world models for applications like dialogue systems and visual question answering.

On-Policy Self-Distillation without Any Supervision

183 upvotes · 9 AUG 2026 · Yijiang Li, Bingyang Wang, Yijun Liang et al.

This paper shows how to make large language models improve themselves without needing external guidance or supervision, by using their own internal consistency to correct mistakes. Practitioners might care about this because it could lead to more robust and self-sufficient AI models.

Self-Supervised Visual On-Policy Distillation

158 upvotes · 14 AUG 2026 · Yijiang Li, Yijun Liang, Yunjie Tian et al.

This paper proposes a new method called Self-Supervised Visual On-Policy Distillation (S^2VOPD) that generates learning signals from asymmetric augmented views of images, allowing for effective on-policy learning without privileged information. Practitioners might care about this paper because it presents a simple yet effective way to improve performance on various perception benchmarks.

AREX: Towards a Recursively Self-Improving Agent for Deep Research

142 upvotes · 23 JUL 2026 · Shuqi Lu, Chaofan Li, Kun Luo et al.

This paper introduces AREX, a self-improving agent that recursively refines its answers by verifying intermediate results and using the partially verified state to guide subsequent refinement, aiming to improve the efficiency of deep research. Practitioners might care about AREX's approach to reducing the cost of discovery and verification in complex research tasks.

Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization

142 upvotes · 17 AUG 2026 · Yixuan Wang, Yifei Chen, Haichao Zhang et al.

This paper improves reinforcement learning for post-training language model reasoners by addressing two problems in existing methods: identical advantages for distinct reward profiles and fixed relative weights for all objectives. A new method, SA-MRPO, dynamically reallocates optimization effort toward under-optimized objectives while maintaining performance on well-satisfied objectives.

Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence

142 upvotes · 17 AUG 2026 · Xin Ding, Liang Mi, Mingzhe Huang et al.

This paper introduces Zetta, a system that enables embodied agents to learn and adapt in real-time while executing physical tasks, allowing for more efficient and effective learning. Practitioners in robotics and AI may care about Zetta's approach as it could lead to more reliable and scalable physical intelligence.

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

120 upvotes · 3 AUG 2026 · Ziyu Ma, Hailang Huang, Shun Zou et al.

This paper proposes a framework, LongHorizon-Harness, to help large language model agents tackle long-horizon tasks by explicitly tracking task states and verifying facts from the environment. Practitioners might care because it can improve the performance of these agents on real-world tasks.

EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction

110 upvotes · 2 SEP 2026 · Yuling Shi, Zhensu Sun, Junsen Dong et al.

This paper introduces EarlyEval, a method to reduce the cost of evaluating large language model (LLM) agents by predicting their outcomes early, allowing for earlier termination of agent runs. Practitioners may care because it can significantly reduce the computational cost of agent development.

DataFlex-RL: An Evaluation Platform for RLVR Data Policies

109 upvotes · 5 SEP 2026 · Hao Liang, Mingrui Chen, Hengyi Feng et al.

This paper introduces DataFlex-RL, an evaluation platform for comparing different data policies in reinforcement learning with verifiable rewards (RLVR), and finds that uniform data policy leads to better performance and is more reproducible than other methods. Practitioners might care about this because it provides a way to compare different data policies and ensure that their RLVR models are performing well.

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

107 upvotes · 12 AUG 2026 · Cheng Qian, Wenting Zhao, Liangwei Yang et al.

This paper explores a new approach to transferring capabilities from strong AI models to weaker ones at test time, rather than just during training. Practitioners might care about this because it could lead to more efficient and effective use of powerful models in real-world applications.

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs

98 upvotes · 20 AUG 2026 · Yunheng Li, Guohong Mu, Hao Li et al.

This paper introduces a method called OraRL to improve the efficiency and scalability of reinforcement learning for multimodal large language models (MLLMs) trained on video data. By leveraging annotations as a source of high-quality rollouts, OraRL can significantly reduce the number of required rollouts, leading to faster training times and better performance on video understanding tasks.

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

92 upvotes · 19 AUG 2026 · Yunhao Yang, Yuexin Bian, Yunjie Tian et al.

This paper introduces Co-RL, a framework for unsupervised multi-agent reinforcement learning that enables diverse and accurate reasoning in language and vision-language models. Practitioners can use Co-RL to improve their models' ability to reason and respond without relying on expensive ground-truth supervision.

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

90 upvotes · 18 AUG 2026 · Zhi Zheng, Rongsheng Chen, Yunpeng Ba et al.

This paper proposes a new method for fine-tuning large language models (LLMs) in reinforcement learning (RL) tasks with long horizons, using evolution strategies (ES) instead of traditional backpropagation-based training. Practitioners might care because it allows for more efficient and flexible fine-tuning of LLMs, enabling them to tackle complex tasks with larger models and longer interactions.

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

89 upvotes · 31 AUG 2026 · Yi Ding, Ruqi Zhang

This paper investigates whether on-policy distillation (OPD) truly improves student policies by analyzing the effects of noisy teacher supervision. It finds that OPD works by suppressing low-probability tokens, which can be achieved without a teacher, and introduces a new method called On-Policy Self-Adaptation (OPSA) that outperforms OPD and traditional reinforcement learning methods.

Show-Harness: Just a VLM Agent Can Play Robots

89 upvotes · 9 SEP 2026 · Yanzhe Chen, Zechen Bai, Zhijun Cao et al.

This paper shows how a vision-language model (VLM) can control robots without needing extensive pretraining or specialized hardware, by providing a compact interface that links the model's intentions to specific actions. Practitioners might care about this because it could make robots more accessible and user-friendly.

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

87 upvotes · 13 AUG 2026 · DreamX Team, Rui Chen, Xiangxiang Chu et al.

This paper introduces a new AI model called DreamX-Phi 1.0 that can predict what will happen in a robotic manipulation scenario, given an initial state and instructions. This model is useful for robotics developers because it can help them design more reliable and efficient robotic systems.

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

84 upvotes · 8 JUL 2026 · Xinyu Geng, Xuanhua He, Sixiang Chen et al.

This paper introduces a framework called DeepSearch-Evolve, which helps train self-improving web agents by iteratively refining their performance using their own experience. Practitioners might care because this approach can lead to more efficient and effective agents that can learn from their own mistakes.

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

78 upvotes · 8 SEP 2026 · Youngrok Park, Sangmin Bae, Hojung Jung et al.

This paper introduces a method called On-Policy Reverse Distillation that helps stronger models learn from weaker supervisors by selectively amplifying the parts of the weaker model's guidance that are most useful to the stronger model. This can lead to faster and more efficient learning, especially in situations where it's expensive to retrain models from scratch.

Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher

73 upvotes · 27 AUG 2026 · Shiyi Zhang, Mushui Liu, Yunze Tong et al.

This paper introduces a new method for training flow matching models called Self-OPD, which uses the model's own exploration to generate supervisory signals without needing a separate teacher model. Practitioners may care because it could lead to more efficient and effective training of flow matching models.

TTPO: Test-Time Policy Optimization

72 upvotes · 27 AUG 2026 · Aozhe Wang, Zhengxi Lu, Jianze Wang et al.

This paper proposes a new method called Test-Time Policy Optimization (TTPO) that allows large language models to be trained without ground-truth labels, enabling test-time training. Practitioners may care about this because it enables models to be trained without labels, which can be difficult or expensive to obtain.

RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments

70 upvotes · 14 SEP 2026 · Sibo Zhu, Shicheng Fan, Xinyue Wang et al.

This paper introduces a new framework called RSIAgent that helps digital agents adapt to new environments without needing to be retrained. A practitioner might care about this because it allows for more efficient and effective AI systems that can learn and improve on their own.

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

69 upvotes · 26 JUL 2026 · Qinsi Wang, Jing Shi, Huazheng Wang et al.

This paper introduces a new method to improve large language models (LLMs) called RLSVR, which uses a task-transformation technique to generate self-verifiable rewards, enabling LLM self-improvement on open-ended tasks. Practitioners might care about this because it could lead to more reliable and scalable self-improvement methods for LLMs.

ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments

69 upvotes · 16 SEP 2026 · Hejia Geng, Zesen Huang, Haoyang Li et al.

This paper creates a system called ScienceIDE that converts scientific code into environments that can be used to train artificial agents to perform scientific tasks. Practitioners might care because this could lead to more efficient and effective ways to develop scientific intelligence.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

68 upvotes · 30 JUL 2026 · Qiushi Sun, Kanzhi Cheng, Yian Wang et al.

This paper develops a standardized evaluation method for computer-using agents (CUAs) to ensure they fulfill task instructions, using vision-language models (VLMs) as judges. Practitioners can benefit from this work by using reliable and cost-effective reward signals for CUA training.

From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search

64 upvotes · 27 JUL 2026 · Junlin Liu, Jiangwang Chen, Zixin Song et al.

This paper proposes a new method to improve the performance of large language models on knowledge-intensive tasks by distilling knowledge from proprietary models and using reinforcement learning. Practitioners may care about this approach because it can help bridge the gap between proprietary and open-source models, leading to more effective and robust AI systems.

ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment

63 upvotes · 5 AUG 2026 · Yijun Lu, Rui Ye, Jiajun Wang et al.

This paper proposes a new method for training long-horizon search agents that can search, retrieve, and integrate evidence to reach a final answer. Practitioners in natural language processing and AI research might care about this paper because it shows a way to improve the performance of search agents, which can be used in applications such as question-answering systems.

It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning

62 upvotes · 1 SEP 2026 · Runpeng Dai, Kaili Huang, Changsung Kang et al.

This paper proposes a new retrieval framework called CoGR, which uses two generative models to co-evolve and optimize retrieval representations on both the query and item sides, leading to improved search and advertising performance. Practitioners might care because CoGR can potentially lead to better retrieval results and more efficient search systems.

Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening

60 upvotes · 16 SEP 2026 · Yizhuo Li, Jianhao Yan, Yun Luo et al.

This paper investigates a common problem in reinforcement learning for language models called Value Flattening, where critics fail to accurately estimate state values, and proposes a new method, SP^3O, to mitigate this issue by supervising only a few well-separated states per response.

UI-Venus-2 Technical Report

58 upvotes · 27 AUG 2026 · Venus Team, Zhuohan Cai, Haoxing Chen et al.

This paper introduces UI-Venus-2, a general-purpose GUI agent that can operate across different environments and perform various tasks, which could be useful for automating digital tasks in real-world applications. Practitioners may care about this work because it addresses the challenges of deploying multimodal GUI agents in practical settings.

Progress Reward Modeling for Robotic Learning: A Comprehensive Survey

55 upvotes · 22 JUL 2026 · Jianshu Zhang, Keliang Wu, Haoran Lu et al.

This paper provides a comprehensive survey of progress reward modeling in robotic learning, aiming to bridge the gap in the field by offering a unified framework for understanding progress rewards. Practitioners in robotics and AI can care about this paper because it helps them understand the different approaches to progress rewards and how to evaluate their effectiveness.

Iris: Climbing to the Search Frontier

53 upvotes · 3 SEP 2026 · Ziyuan Liu, Hengqi Liu, Zichuan Wang et al.

This paper presents two search agents, Iris-mini and Iris-pro, trained to solve complex search tasks using reinforcement learning and self-supervised learning. Practitioners might care because these models achieve state-of-the-art results on various benchmarks, demonstrating the potential of AI-powered search agents in real-world applications.

TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

52 upvotes · 21 AUG 2026 · Yibo Hu, Yu Qian, Mao Gu et al.

This paper proposes a model called TLive-Omni that can understand and process multiple types of data (image, video, audio, text) from e-commerce live streaming. A practitioner might care about this model because it can be used to analyze and respond to live-commerce streams in a more accurate and efficient way.

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience

52 upvotes · 3 SEP 2026 · Zixun Huang, Kishan Panaganti, Haitao Mi et al.

This paper introduces FlowBalance, a method that helps a reasoning model improve itself by learning from its own experiences, while avoiding overconfidence and focusing on the best solutions. Practitioners might care about this approach because it can lead to more accurate and diverse model performance.

DAPD: Dual-Anchored Policy Distillation

51 upvotes · 3 AUG 2026 · Jianyu Wu, Yizhou Wang, Encheng Su et al.

This paper addresses a problem in self-distillation, where a student model learns to mimic the behavior of a privileged teacher, but performs poorly at inference due to a "privilege illusion". The authors propose a new method, Dual-Anchored Policy Distillation, to resolve this issue and improve performance.

Intern-S2-Preview: Scientific Agentic Foundation Model

50 upvotes · 13 AUG 2026 · Lei Bai, Jiaqi Cao, Chiyu Chen et al.

This paper introduces Intern-S2-Preview, a type of AI model designed to support scientific discovery by reasoning over different types of evidence, interacting with scientific tools, and sustaining progress over long periods. Practitioners might care because it could lead to more accurate scientific understanding and forecasting.

DriveZero: End-to-End Driving Beyond Human Demonstrations

50 upvotes · 5 SEP 2026 · Hao He, Chengcheng Hu, Zirun Su et al.

This paper presents a method for end-to-end driving systems to learn beyond human demonstrations, using a combination of perception and action models that can learn from diverse data and provide goal-conditioned supervision. Practitioners might care about this approach because it allows for more robust and flexible autonomous driving systems.

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

49 upvotes · 24 AUG 2026 · Sungho Park, Wonjoong Kim, Rongyuan Tan et al.

This paper develops a system to automatically optimize the design of harnesses for LLM agents, which can improve their reliability on long-horizon tasks. Practitioners might care about this because it could lead to more robust and performant agent systems.

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

48 upvotes · 30 JUL 2026 · Qixun Wang, Yang Shi, Letian Cheng et al.

This paper proposes a new approach to agentic visual reasoning, which helps large language models (LLMs) perform better on complex tasks by using tools more efficiently. Practitioners might care about this research because it aims to improve the performance of LLMs on challenging problems.

SPADE: Self-Play in Adaptive Synthetic Executable Environments

47 upvotes · 19 AUG 2026 · Bo Liu, Simon Yu, Yiding Jiang et al.

This paper introduces SPADE, a self-play framework that enables language agents to learn from adaptive, self-generated environments, allowing them to improve continuously without fixed goal distributions. Practitioners might care because SPADE can lead to more robust and open-ended AI models.

Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report

47 upvotes · 22 AUG 2026 · TaoLive AIGC LLM Team, Yuhan Sun, Wenhao Lin et al.

This paper trains compact AI models to adapt quickly to changing digital avatar "harnesses" that define the tasks and tools available to the model, improving performance and reducing latency. Practitioners caring about real-time AI applications, such as chatbots or virtual assistants, might find this approach useful.