Papers

Filtered to benchmarking · clear filter

Browse by term

continual learning 64reinforcement learning 34large language models 17benchmarking 12vision-language models 10generative models 8language models 8video generation 7multimodal models 6natural language processing 6robotics 6world models 6benchmarks 5diffusion models 5on-policy distillation 5policy optimization 5scalability 5self-distillation 5vision-language-action models 5autoregressive models 4computer vision 4diffusion transformers 4LLMs 4multimodal large language models 4verifiable rewards 4attention mechanisms 3embodied intelligence 3image editing 3long-term memory 3multimodal learning 3

Matching papers

AREX: Towards a Recursively Self-Improving Agent for Deep Research

142 upvotes · 23 JUL 2026 · Shuqi Lu, Chaofan Li, Kun Luo et al.

This paper introduces AREX, a self-improving agent that recursively refines its answers by verifying intermediate results and using the partially verified state to guide subsequent refinement, aiming to improve the efficiency of deep research. Practitioners might care about AREX's approach to reducing the cost of discovery and verification in complex research tasks.

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

57 upvotes · 23 JUL 2026 · Hao Liang, Qihan Lin, Zhaoyang Han et al.

This paper introduces a new framework for training educational language models, specifically designed to evaluate their ability to understand curriculum knowledge and its visual presentation. Practitioners might care about this work because it aims to improve language models' performance in educational settings.

Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking

13 upvotes · 26 JUL 2026 · Haorui He, Xinwen Chen, Dacheng Wen et al.

This paper investigates the reliability of dynamic benchmarks for multimodal automated fact-checking by examining contamination risks and their impact on evaluation metrics. Practitioners should consider the potential for contamination in dynamic benchmarks to ensure accurate performance estimates.

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

13 upvotes · 27 JUL 2026 · Zichao Lin, Yifeng Xie, Bowen Qu et al.

This paper introduces a benchmark to evaluate the atomic visual perception capabilities of large language models, which are often unable to accurately perceive visual information. Practitioners may care about this research because it provides a standardized way to measure and diagnose the limitations of visual perception in MLLMs.

Can AI agents conduct open-ended AI research? Early evidence from two case studies

11 upvotes · 29 JUL 2026 · Peter Kirgis, Sayash Kapoor, Andrew Schwartz et al.

This paper investigates whether AI agents can conduct open-ended AI research and provides early evidence that they can perform the engineering aspects but struggle with critical parts of the research lifecycle, such as making progress on research questions and judgment about publishable research.

Reinforcement Learning for Code Optimization

9 upvotes · 28 JUL 2026 · Pierre Chambon, Kunhao Zheng, Juliette Decugis et al.

This paper uses reinforcement learning to optimize code for speed while maintaining correctness, and it finds that the approach can significantly improve performance, even when the reward function is noisy or sparse. Practitioners might care because optimizing code for speed can be crucial in many applications, and this research provides a promising method for achieving that goal.

FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents

7 upvotes · 21 JUL 2026 · Xianfu Cheng, Shiwei Zhang, Jiyu Zhao et al.

This paper creates a benchmark for testing the ability of AI agents to understand and analyze complex financial documents, and uses it to evaluate the performance of different agents in this task. Practitioners in finance and AI research can care about this work because it aims to improve the accuracy and reliability of financial document analysis.

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

7 upvotes · 29 JUL 2026 · Jingbo Zhou, Yusai Zhao, Qi Bao et al.

This paper introduces a benchmark to evaluate large language model (LLM) agents on long-horizon office-suite tasks, considering their cost-effectiveness and quality. Practitioners can care about this research because it aims to ensure LLM agents can assist users efficiently and effectively.

SceneActBench: Can Agents Act on the 3D Scenes They See?

6 upvotes · 24 JUL 2026 · Yifei Zhao, Xiangxin Zhou, Wenhao Yang et al.

This paper introduces SceneActBench, a benchmark for evaluating the ability of vision-language model agents to perform actions in 3D scenes, and analyzes the performance of 11 different VLM configurations on this benchmark. Practitioners may care about this research because it helps to identify the limitations of current VLM agents in performing complex tasks in 3D environments.

Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model

5 upvotes · 22 JUL 2026 · Markus J. Buehler

This paper investigates whether large language models, like Google's Gemma-4-E4B-it, represent scientific concepts and governing physics, and whether this representation affects their answers. Practitioners caring about the accuracy and reliability of language models in scientific domains might find this research valuable.