Papers

Filtered to benchmarks · clear filter

Browse by term

continual learning 83reinforcement learning 50large language models 12benchmarking 11benchmarks 11language models 11vision-language models 11robotics 7natural language processing 6world models 6generative models 5recursive self-improvement 5attention mechanisms 4diffusion Transformers 4multi-agent systems 4multimodal learning 4multimodal models 4on-policy distillation 4self-distillation 4self-supervised learning 4transformers 4video generation 4vision-language-action models 4agent-based systems 3agentic models 3agentic search 3autonomous systems 3coding agents 3diffusion models 3image generation 3

Matching papers

Scaling Automatic Research Agents via World Models

452 upvotes · 29 AUG 2026 · Xiyuan Yang, Sheikh Sarwar, Jingru Cheng et al.

This paper proposes a method to scale automatic research agents by replacing environment execution with a world model, which can reduce training costs and improve performance. Practitioners might care about this approach because it can accelerate training times and lead to better results for complex AI tasks.

VGI-Bench: Probing Visual Intelligence in Video Generation Models

171 upvotes · 26 AUG 2026 · Xuan He, Cong Wei, Yuhao Cheng et al.

This paper introduces VGI-bench, a new benchmark for evaluating video generation models' visual reasoning capabilities, and finds that current models can solve some visually grounded tasks but still struggle with reliability. Practitioners may care about developing more reliable video generation models that can perform better on this benchmark.

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

155 upvotes · 18 AUG 2026 · Keyu Tu, Zhuowei Chen, Mengqi Huang et al.

This paper introduces a new benchmark for video generation tasks that require both achieving a desired outcome and maintaining semantic consistency with a reference image. Practitioners might care about this research because it can help evaluate and improve the performance of video generation models in real-world applications.

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

120 upvotes · 3 AUG 2026 · Ziyu Ma, Hailang Huang, Shun Zou et al.

This paper proposes a framework, LongHorizon-Harness, to help large language model agents tackle long-horizon tasks by explicitly tracking task states and verifying facts from the environment. Practitioners might care because it can improve the performance of these agents on real-world tasks.

LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering

99 upvotes · 28 AUG 2026 · Yi Wang, Haopeng Zhang, Chengxiang Huang et al.

This paper introduces LoopArena, a benchmark for evaluating how well a model can guide a coding agent through a long-running task, and finds room for improvement in long-horizon loop control. Practitioners working on AI development might care because it could lead to more efficient and reliable development processes.

AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems

92 upvotes · 8 SEP 2026 · Jaewon Chu, Jinwoo Seo, Jaewon Cho et al.

This paper proposes a method to optimize prompts for multi-agent systems by identifying which agent's modification resolves a failure, and then using that agent's output as supervision to extract a fine-grained gradient. Practitioners might care because it could improve the performance of large language model-based multi-agent systems.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

68 upvotes · 30 JUL 2026 · Qiushi Sun, Kanzhi Cheng, Yian Wang et al.

This paper develops a standardized evaluation method for computer-using agents (CUAs) to ensure they fulfill task instructions, using vision-language models (VLMs) as judges. Practitioners can benefit from this work by using reliable and cost-effective reward signals for CUA training.

Unlocking Lossless Speedups in LLMs via Discrete Diffusion

64 upvotes · 3 SEP 2026 · Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham et al.

This paper proposes a new approach to Large Language Models (LLMs) that uses diffusion to speed up generation without sacrificing quality, allowing for faster and more efficient language processing. Practitioners may care about this research if they need to process large amounts of text quickly and efficiently.

ASI-Bench: At the Dawn of Artificial Superintelligence

55 upvotes · 18 AUG 2026 · Junwei Zhou, Zhen Sun, Binyu Li et al.

This paper introduces ASI-Bench, a new benchmark for evaluating AI systems' ability to explore new knowledge, create new ideas, and conduct scientific research autonomously. Practitioners might care about ASI-Bench because it aims to push the limits of current AI systems and accelerate the development of artificial superintelligence.

Iris: Climbing to the Search Frontier

53 upvotes · 3 SEP 2026 · Ziyuan Liu, Hengqi Liu, Zichuan Wang et al.

This paper presents two search agents, Iris-mini and Iris-pro, trained to solve complex search tasks using reinforcement learning and self-supervised learning. Practitioners might care because these models achieve state-of-the-art results on various benchmarks, demonstrating the potential of AI-powered search agents in real-world applications.

VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

50 upvotes · 15 AUG 2026 · Yansong Ning, Jingwen Ye, Zhongkai Wu et al.

This paper proposes a framework for training agents to create 3D open worlds based on user queries, and evaluates its performance using a large benchmark dataset. Practitioners might care about this research because it can help develop more capable multimodal agents that can understand user intent and generate realistic 3D environments.