452 upvotes · 29 AUG 2026 · Xiyuan Yang, Sheikh Sarwar, Jingru Cheng et al.
This paper proposes a method to scale automatic research agents by replacing environment execution with a world model, which can reduce training costs and improve performance. Practitioners might care about this approach because it can accelerate training times and lead to better results for complex AI tasks.
171 upvotes · 26 AUG 2026 · Xuan He, Cong Wei, Yuhao Cheng et al.
This paper introduces VGI-bench, a new benchmark for evaluating video generation models' visual reasoning capabilities, and finds that current models can solve some visually grounded tasks but still struggle with reliability. Practitioners may care about developing more reliable video generation models that can perform better on this benchmark.
155 upvotes · 18 AUG 2026 · Keyu Tu, Zhuowei Chen, Mengqi Huang et al.
This paper introduces a new benchmark for video generation tasks that require both achieving a desired outcome and maintaining semantic consistency with a reference image. Practitioners might care about this research because it can help evaluate and improve the performance of video generation models in real-world applications.
120 upvotes · 3 AUG 2026 · Ziyu Ma, Hailang Huang, Shun Zou et al.
This paper proposes a framework, LongHorizon-Harness, to help large language model agents tackle long-horizon tasks by explicitly tracking task states and verifying facts from the environment. Practitioners might care because it can improve the performance of these agents on real-world tasks.
99 upvotes · 28 AUG 2026 · Yi Wang, Haopeng Zhang, Chengxiang Huang et al.
This paper introduces LoopArena, a benchmark for evaluating how well a model can guide a coding agent through a long-running task, and finds room for improvement in long-horizon loop control. Practitioners working on AI development might care because it could lead to more efficient and reliable development processes.
92 upvotes · 8 SEP 2026 · Jaewon Chu, Jinwoo Seo, Jaewon Cho et al.
This paper proposes a method to optimize prompts for multi-agent systems by identifying which agent's modification resolves a failure, and then using that agent's output as supervision to extract a fine-grained gradient. Practitioners might care because it could improve the performance of large language model-based multi-agent systems.
68 upvotes · 30 JUL 2026 · Qiushi Sun, Kanzhi Cheng, Yian Wang et al.
This paper develops a standardized evaluation method for computer-using agents (CUAs) to ensure they fulfill task instructions, using vision-language models (VLMs) as judges. Practitioners can benefit from this work by using reliable and cost-effective reward signals for CUA training.
64 upvotes · 3 SEP 2026 · Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham et al.
This paper proposes a new approach to Large Language Models (LLMs) that uses diffusion to speed up generation without sacrificing quality, allowing for faster and more efficient language processing. Practitioners may care about this research if they need to process large amounts of text quickly and efficiently.
55 upvotes · 18 AUG 2026 · Junwei Zhou, Zhen Sun, Binyu Li et al.
This paper introduces ASI-Bench, a new benchmark for evaluating AI systems' ability to explore new knowledge, create new ideas, and conduct scientific research autonomously. Practitioners might care about ASI-Bench because it aims to push the limits of current AI systems and accelerate the development of artificial superintelligence.
53 upvotes · 3 SEP 2026 · Ziyuan Liu, Hengqi Liu, Zichuan Wang et al.
This paper presents two search agents, Iris-mini and Iris-pro, trained to solve complex search tasks using reinforcement learning and self-supervised learning. Practitioners might care because these models achieve state-of-the-art results on various benchmarks, demonstrating the potential of AI-powered search agents in real-world applications.
50 upvotes · 15 AUG 2026 · Yansong Ning, Jingwen Ye, Zhongkai Wu et al.
This paper proposes a framework for training agents to create 3D open worlds based on user queries, and evaluates its performance using a large benchmark dataset. Practitioners might care about this research because it can help develop more capable multimodal agents that can understand user intent and generate realistic 3D environments.