Papers

Filtered to multimodal large language models · clear filter

Browse by term

continual learning 83reinforcement learning 50large language models 12benchmarking 11benchmarks 11language models 11vision-language models 11robotics 7natural language processing 6world models 6generative models 5recursive self-improvement 5attention mechanisms 4diffusion Transformers 4multi-agent systems 4multimodal learning 4multimodal models 4on-policy distillation 4self-distillation 4self-supervised learning 4transformers 4video generation 4vision-language-action models 4agent-based systems 3agentic models 3agentic search 3autonomous systems 3coding agents 3diffusion models 3image generation 3

Matching papers

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

106 upvotes · 27 AUG 2026 · Tianjie Ju, Zheng Wu, Yueqing Sun et al.

This paper explores how large language models can turn local observations of a city into reliable actions, and whether these models can sustain goal-directed behavior in complex urban environments. Practitioners in AI/ML and urban planning might care about the limitations and potential of current models in navigating real-world cities.

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs

98 upvotes · 20 AUG 2026 · Yunheng Li, Guohong Mu, Hao Li et al.

This paper introduces a method called OraRL to improve the efficiency and scalability of reinforcement learning for multimodal large language models (MLLMs) trained on video data. By leveraging annotations as a source of high-quality rollouts, OraRL can significantly reduce the number of required rollouts, leading to faster training times and better performance on video understanding tasks.

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

48 upvotes · 30 JUL 2026 · Qixun Wang, Yang Shi, Letian Cheng et al.

This paper proposes a new approach to agentic visual reasoning, which helps large language models (LLMs) perform better on complex tasks by using tools more efficiently. Practitioners might care about this research because it aims to improve the performance of LLMs on challenging problems.