142 upvotes · 23 JUL 2026 · Shuqi Lu, Chaofan Li, Kun Luo et al.
This paper introduces AREX, a self-improving agent that recursively refines its answers by verifying intermediate results and using the partially verified state to guide subsequent refinement, aiming to improve the efficiency of deep research. Practitioners might care about AREX's approach to reducing the cost of discovery and verification in complex research tasks.
57 upvotes · 23 JUL 2026 · Hao Liang, Qihan Lin, Zhaoyang Han et al.
This paper introduces a new framework for training educational language models, specifically designed to evaluate their ability to understand curriculum knowledge and its visual presentation. Practitioners might care about this work because it aims to improve language models' performance in educational settings.
49 upvotes · 19 MAY 2026 · Hao Liang, Qifeng Cai, Yibo Lin et al.
This paper introduces a benchmark to measure how well large language models (LLMs) can prepare training data, and how well they evaluate the quality of that data. Practitioners might care because improving data preparation can lead to better model performance.
23 upvotes · 31 JUL 2026 · Yifan Ding, Xincheng Wei, Yoshua Y. Li et al.
This paper proposes a method to combine reinforcement learning with verifiable rewards and on-policy distillation to improve performance on complex tasks, and shows that this method can lead to more stable training and better results.
13 upvotes · 26 JUL 2026 · Haorui He, Xinwen Chen, Dacheng Wen et al.
This paper investigates the reliability of dynamic benchmarks for multimodal automated fact-checking by examining contamination risks and their impact on evaluation metrics. Practitioners should consider the potential for contamination in dynamic benchmarks to ensure accurate performance estimates.
13 upvotes · 27 JUL 2026 · Zichao Lin, Yifeng Xie, Bowen Qu et al.
This paper introduces a benchmark to evaluate the atomic visual perception capabilities of large language models, which are often unable to accurately perceive visual information. Practitioners may care about this research because it provides a standardized way to measure and diagnose the limitations of visual perception in MLLMs.
11 upvotes · 29 JUL 2026 · Peter Kirgis, Sayash Kapoor, Andrew Schwartz et al.
This paper investigates whether AI agents can conduct open-ended AI research and provides early evidence that they can perform the engineering aspects but struggle with critical parts of the research lifecycle, such as making progress on research questions and judgment about publishable research.
9 upvotes · 28 JUL 2026 · Pierre Chambon, Kunhao Zheng, Juliette Decugis et al.
This paper uses reinforcement learning to optimize code for speed while maintaining correctness, and it finds that the approach can significantly improve performance, even when the reward function is noisy or sparse. Practitioners might care because optimizing code for speed can be crucial in many applications, and this research provides a promising method for achieving that goal.
7 upvotes · 21 JUL 2026 · Xianfu Cheng, Shiwei Zhang, Jiyu Zhao et al.
This paper creates a benchmark for testing the ability of AI agents to understand and analyze complex financial documents, and uses it to evaluate the performance of different agents in this task. Practitioners in finance and AI research can care about this work because it aims to improve the accuracy and reliability of financial document analysis.
7 upvotes · 29 JUL 2026 · Jingbo Zhou, Yusai Zhao, Qi Bao et al.
This paper introduces a benchmark to evaluate large language model (LLM) agents on long-horizon office-suite tasks, considering their cost-effectiveness and quality. Practitioners can care about this research because it aims to ensure LLM agents can assist users efficiently and effectively.
6 upvotes · 24 JUL 2026 · Yifei Zhao, Xiangxin Zhou, Wenhao Yang et al.
This paper introduces SceneActBench, a benchmark for evaluating the ability of vision-language model agents to perform actions in 3D scenes, and analyzes the performance of 11 different VLM configurations on this benchmark. Practitioners may care about this research because it helps to identify the limitations of current VLM agents in performing complex tasks in 3D environments.
5 upvotes · 22 JUL 2026 · Markus J. Buehler
This paper investigates whether large language models, like Google's Gemma-4-E4B-it, represent scientific concepts and governing physics, and whether this representation affects their answers. Practitioners caring about the accuracy and reliability of language models in scientific domains might find this research valuable.