106 upvotes · 27 AUG 2026 · Tianjie Ju, Zheng Wu, Yueqing Sun et al.
This paper explores how large language models can turn local observations of a city into reliable actions, and whether these models can sustain goal-directed behavior in complex urban environments. Practitioners in AI/ML and urban planning might care about the limitations and potential of current models in navigating real-world cities.
98 upvotes · 20 AUG 2026 · Yunheng Li, Guohong Mu, Hao Li et al.
This paper introduces a method called OraRL to improve the efficiency and scalability of reinforcement learning for multimodal large language models (MLLMs) trained on video data. By leveraging annotations as a source of high-quality rollouts, OraRL can significantly reduce the number of required rollouts, leading to faster training times and better performance on video understanding tasks.
48 upvotes · 30 JUL 2026 · Qixun Wang, Yang Shi, Letian Cheng et al.
This paper proposes a new approach to agentic visual reasoning, which helps large language models (LLMs) perform better on complex tasks by using tools more efficiently. Practitioners might care about this research because it aims to improve the performance of LLMs on challenging problems.