Papers

Filtered to world models · clear filter

Browse by term

continual learning 83reinforcement learning 50large language models 12benchmarking 11benchmarks 11language models 11vision-language models 11robotics 7natural language processing 6world models 6generative models 5recursive self-improvement 5attention mechanisms 4diffusion Transformers 4multi-agent systems 4multimodal learning 4multimodal models 4on-policy distillation 4self-distillation 4self-supervised learning 4transformers 4video generation 4vision-language-action models 4agent-based systems 3agentic models 3agentic search 3autonomous systems 3coding agents 3diffusion models 3image generation 3

Matching papers

Scaling Automatic Research Agents via World Models

452 upvotes · 29 AUG 2026 · Xiyuan Yang, Sheikh Sarwar, Jingru Cheng et al.

This paper proposes a method to scale automatic research agents by replacing environment execution with a world model, which can reduce training costs and improve performance. Practitioners might care about this approach because it can accelerate training times and lead to better results for complex AI tasks.

Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models

185 upvotes · 26 AUG 2026 · Pengfei Zhou, Hexin Wang, Zhengfeiyang Zhang et al.

This paper proposes a way to improve the efficiency of training world models by using game development as a source of reward signals and trajectory data, allowing for more effective post-training of large language models using reinforcement learning. Practitioners might care about this approach because it could lead to more scalable and effective world models for applications like dialogue systems and visual question answering.

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

108 upvotes · 17 AUG 2026 · Weiliang Chen, Haowen Sun, Jun Gao et al.

This paper develops a new method for evaluating world models, called HarnessEval-W, which provides more detailed and justifiable results than existing benchmarks. Practitioners might care about HarnessEval-W because it can help them build more trustworthy world models that better align with human preferences.