292 upvotes · 30 JUL 2026 · Hanzhang Zhou, Panrong Tong, Xu Zhang et al.
This paper introduces Qwen-UI-Agent, a type of artificial intelligence system that can perform tasks on various devices, such as smartphones and computers, and improve its abilities on its own. Practitioners might care about this research because it aims to create more practical and autonomous AI systems that can be used in real-world scenarios.
256 upvotes · 29 JUL 2026 · Zeyu Zhang, Ziliang Guo, Yihang Sun et al.
This paper introduces a new type of AI model called memory foundation models, which allows the model to learn and retain information internally, rather than relying on external memory modules. This could be useful for practitioners who want to build more efficient and flexible AI agents.
154 upvotes · 19 JUL 2026 · Yuhan Zhu, Changlian Ma, Xiangyu Zeng et al.
This paper develops a new approach to understanding videos by predicting when specific events or evidence occur within the video. Practitioners working on video analysis and AI models might care about this research because it could lead to more accurate and robust video understanding systems.
128 upvotes · 13 JUL 2026 · Mikhail Komarov, Ivan Bondarenko, Stanislav Shtuka et al.
This paper introduces RAGU, an open-source GraphRAG engine that improves large language models with structured knowledge by separating extraction and consolidation, and trains a compact extractor that outperforms larger models on knowledge-graph construction and GraphRAG tasks. Practitioners might care because RAGU can efficiently generate more accurate and complete context for tasks like factoid-level evidence recall and multi-hop question answering.
122 upvotes · 18 JUL 2026 · Runming He, Zhen Hao Wong, Hao Liang et al.
This paper creates a platform to help large language models generate code for data pipelines, which can then be edited and used to automate data processing workflows. Practitioners might care about this because it can help reduce the time and cost of developing and maintaining these pipelines.
118 upvotes · 29 JUL 2026 · Hengyi Xie, Chenfei Yao, Xianjin Wu et al.
This paper introduces TurboVLA, a new vision-language-action model that reduces computation and memory overhead by directly exchanging information between visual observations and language instructions, allowing for faster and more efficient robotic manipulation. Practitioners might care about this approach for building more efficient and effective VLA models.
89 upvotes · 29 JUL 2026 · Jiaxing Li, Kai Zou, Cindy Zhou et al.
This paper improves autoregressive video distillation methods by aligning the initialization and distribution matching stages, focusing on matching the target distribution's mode coverage rather than just visual quality. Practitioners can benefit from this approach to generate higher-quality videos with better diversity and coverage.
84 upvotes · 8 JUL 2026 · Xinyu Geng, Xuanhua He, Sixiang Chen et al.
This paper introduces a framework called DeepSearch-Evolve, which helps train self-improving web agents by iteratively refining their performance using their own experience. Practitioners might care because this approach can lead to more efficient and effective agents that can learn from their own mistakes.
73 upvotes · 28 JUL 2026 · Bo-Wen Zhang, Junwei He, Wen Wang et al.
This paper proposes a method to improve language model training by allocating credit to individual tokens within a response, allowing for more nuanced evaluation of model performance. Practitioners may care about this method as it can lead to better language model performance, especially in tasks that require specific formatting or semantic choices.
70 upvotes · 27 JUL 2026 · Bingnan Li, Haozhe Wang, Haozhong Xiong et al.
This paper investigates how to improve the adaptation of diffusion models in a way that doesn't rely on a classifier, and how to address a problem where the model can't accurately learn from its teacher. Practitioners might care about this because it could lead to more effective knowledge transfer in machine learning applications.
64 upvotes · 20 JUL 2026 · Yuhang Wang, Yuling Shi, Shaoqiu Zhang et al.
This paper proposes a new pruning method for coding agents that prunes tool outputs directly inside the agent, rather than relying on a separate code classifier, and shows it can save up to 39% of tokens while preserving task quality.
60 upvotes · 21 JUL 2026 · Xinjie Zhang, Peng Zhang, Shicheng Zheng et al.
This paper introduces Mage-Flow, a compact model for generating and editing high-resolution images, which can be trained efficiently and deployed on a single GPU. Practitioners might care about the potential applications of this model in interactive image editing and generation tasks.
59 upvotes · 16 JUL 2026 · Xiaomi Robotics Team, Jun Guo, Piaopiao Jin et al.
This paper introduces a vision-language-action model that can perform mobile manipulation tasks in unseen environments with minimal training data, and how it can be scaled up to achieve better performance. Practitioners might care about this model for building robots that can adapt to new tasks with minimal fine-tuning.
54 upvotes · 28 JUL 2026 · Jiangwang Chen, Zixin Song, Junlin Liu et al.
This paper introduces a method called DecoEvo, which helps large language models improve by co-evolving a solver skill and a rubric-generator skill in a way that's more efficient and effective. Practitioners might care about this because it could lead to better performance and more reliable optimization in open-ended tasks.
51 upvotes · 30 JUL 2026 · Rubin Wei, Jiaqi Cao, Jiarui Wang et al.
This paper develops a method to scale up language models by increasing their memory capacity, allowing for better performance and more efficient use of parameters. Practitioners may care about this research as it could lead to more powerful and efficient language models for applications like language translation and text generation.
49 upvotes · 19 MAY 2026 · Hao Liang, Qifeng Cai, Yibo Lin et al.
This paper introduces a benchmark to measure how well large language models (LLMs) can prepare training data, and how well they evaluate the quality of that data. Practitioners might care because improving data preparation can lead to better model performance.
49 upvotes · 26 JUL 2026 · NeoteAI Team, Fudan TEAI Team
This paper introduces a vision-tactile-language-action model that can perform fine-grained manipulation with tactile perception and control, and improve its policy offline from stored data. Practitioners may care because this model can be used to create more versatile and accurate tactile-driven manipulation policies.
34 upvotes · 24 JUL 2026 · Siyuan Huang, Pengyu Cheng, Haotian Liu et al.
This paper develops a new framework called Skill Self-Play that helps large language models (LLMs) improve their capabilities by co-evolving skills that balance task diversity and verification reliability. Practitioners might care because this approach can lead to significant performance gains for LLMs.
32 upvotes · 11 JUL 2026 · Aleksandr Kutsakov, Mariia Sadovina, Georgii Gospodinov et al.
This paper develops a large audio language model that can answer questions with specific timestamps, improving its ability to understand long audio recordings. Practitioners in audio and speech recognition may care about this development as it enables more accurate and context-specific information retrieval from audio data.
30 upvotes · 11 JUL 2026 · Andrei Kuzmenko, Alexandr Maximenko, Aleksandr Kutsakov et al.
This paper develops a multilingual speech recognition model for underrepresented languages, addressing the issue of data scarcity and uneven performance, and shows promising results in controlled comparisons.
28 upvotes · 22 JUL 2026 · Kailin Jiang, Lei Liu, Jian Xi et al.
This paper develops a new framework for evaluating and selecting document sets for AI agents, considering the interactions between documents, and proposes a training-free method that achieves the best downstream generation performance with fewer documents and search rounds.
28 upvotes · 28 JUL 2026 · Haolei Xu, Xiaowen Xu, Haiwen Hong et al.
This paper addresses a problem in on-policy distillation where a student model can get stuck on a wrong path, and proposes a new method called Relay-OPD that helps the student model recover by briefly taking over at certain points to produce a new trajectory. Practitioners might care about this because it could lead to better performance and more efficient training in models like language generators or math solvers.
27 upvotes · 28 JUL 2026 · Yu Wang, Yi-Kai Zhang, Wentao Shi et al.
This paper proposes a method to improve reinforcement learning with verifiable rewards by using game solvers to provide turn-level credit to agents, allowing them to learn more effectively. Practitioners might care about this approach because it could lead to more robust and efficient AI decision-making.
24 upvotes · 23 JUL 2026 · Tencent WorkBuddy Bench Team, Siqi Cai, Shaopeng Chen et al.
This paper introduces Tencent WorkBuddy Bench, a benchmark for coding agents that tests their performance across multiple domains, including code, web, office, and security. Practitioners might care about this benchmark because it provides a standardized way to evaluate and compare the performance of coding agents.
23 upvotes · 29 JUL 2026 · Yihao Chen, Shi Chang, Khaled Chawa et al.
This paper teaches language models to synthesize complete software programs from scratch, which is a challenging task. Practitioners might care because this can improve the models' performance on software engineering tasks.
22 upvotes · 10 JUL 2026 · Mantha Sai Gopal, Jaison Saji Chacko, Harsh Nandwana et al.
This paper introduces a new method called REBASE that allows for training-free in-context segmentation, enabling the introduction of new object categories at inference time. Practitioners might care because it eliminates the need for retraining and reduces memory overhead, making it a more efficient approach for real-world applications.
22 upvotes · 27 JUL 2026 · Tianyi Men, Zhuoran Jin, Kang Liu et al.
This paper explores how to improve long-horizon planning in AI agents, which is crucial for foundation models. Practitioners might care because better planning abilities can lead to more effective and efficient decision-making in complex environments.
21 upvotes · 30 JUL 2026 · Jiawei Xu, Minghui Liu, Juzheng Zhang et al.
This paper develops a new method for improving reasoning language models, called β-OPSD, which combines policy optimization and self-distillation to improve stability and performance. Practitioners might care about this method because it provides a more efficient and effective way to improve language model reasoning abilities.
20 upvotes · 20 JUL 2026 · Yuxuan Luo, Peng Zhang, Xinjie Zhang et al.
This paper develops a new framework, SciForma, to generate scientific diagrams that accurately represent research logic, which is crucial for scientific communication and methodology validation. Practitioners can benefit from SciForma's ability to ensure structural fidelity in diagram generation, which can improve the accuracy and reliability of scientific research.
20 upvotes · 23 JUL 2026 · Gaurav Dadhich
This paper proposes a new approach to managing the context of AI agents, which is crucial for their performance in production environments. By actively managing what an agent holds in mind, the agent can avoid accumulating unnecessary information and reduce costs, leading to improved accuracy and efficiency.
18 upvotes · 29 JUL 2026 · Zhiyuan Yao, Yuxin Chen, Zhengxi Lu et al.
This paper introduces SkillRise, a framework that enables large language model agents to learn skills across related tasks, allowing them to reuse solution patterns and improve performance on multiple tasks. Practitioners can use SkillRise to train more efficient LLM agents that can adapt to new tasks and improve their performance over time.
17 upvotes · 23 JUL 2026 · Chenhui Gou, Haoqin Tu, Yunhao Fang et al.
This paper develops a method called Experience Distillation that allows agents to learn from their own interaction histories without needing additional environment interactions, making learning more sample-efficient. Practitioners might care about this because it can improve the performance of agents in complex environments with limited resources.
17 upvotes · 30 JUL 2026 · Chongjian Ge, Hanwen Jiang, Tianyu Wang et al.
This paper introduces Chimera, a hybrid visual diffusion transformer that efficiently processes text, image, and video tokens to generate high-resolution images, videos, and multimodal context. Practitioners might care about this paper because it provides a scalable solution for large-scale visual generation tasks.
16 upvotes · 21 JUL 2026 · Nischay Dhankhar, Dos Baha, Abulhair Saparov
This paper investigates using hypernetworks for large-scale knowledge injection into language models, a technique that can improve their ability to answer factual questions. Practitioners may care because it could lead to more accurate and scalable language models for applications like customer service or question-answering systems.
15 upvotes · 28 JUL 2026 · Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer et al.
This paper explores whether visual prompt engineering, which involves modifying the task image to improve video model performance, can be as effective as text-based prompt engineering for video models, and if so, why it's worth trying for practitioners looking to improve their visual reasoning tasks.
15 upvotes · 21 JUL 2026 · Mingxuan Xia, Yuhang Yang, Chao Ye et al.
This paper improves a type of reinforcement learning (RL) called rubric-based RL, which helps large language models (LLMs) perform well on open-ended tasks. A practitioner might care about this paper because it addresses a common problem in RL, where some criteria (or rules) are not explored properly, and it shows that its new method can improve performance on these tasks.
14 upvotes · 22 JUL 2026 · Jiarong Zhao, Zhikai Lei, Zhiheng Xi et al.
This paper develops a framework called NexForge that helps train more capable artificial agents by automatically generating a large number of tasks and training data, without requiring a lot of manual setup. Practitioners might care because it can improve the performance of their own agent models.
14 upvotes · 15 JUL 2026 · Dwip Dalal, Shivansh Patel, Chahit Jain et al.
This paper proposes a new method for fine-tuning vision-language models on robot demonstrations to improve their performance on real-world tasks, by preventing the overwrite of pre-trained representations and aligning language and action predictions. Practitioners may care about this work because it aims to improve the generalizability and robustness of vision-language-action policies in real-world applications.
14 upvotes · 30 JUL 2026 · Rong Wu, Daocheng Fu, Licheng Wen et al.
This paper proposes a new approach to memory-augmentation in large language model agents, allowing them to actively reconstruct and adapt past experiences to fit the current context, rather than simply replaying them. Practitioners might care because this approach can improve the robustness and intrinsic reasoning capabilities of agents in complex scenarios.
13 upvotes · 26 JUL 2026 · Haorui He, Xinwen Chen, Dacheng Wen et al.
This paper investigates the reliability of dynamic benchmarks for multimodal automated fact-checking by examining contamination risks and their impact on evaluation metrics. Practitioners should consider the potential for contamination in dynamic benchmarks to ensure accurate performance estimates.
13 upvotes · 30 JUL 2026 · Peilin Feng, Suorong Yang, Soujanya Poria
This paper introduces a new type of memory system for large language model (LLM) based multi-agent systems that tracks which agents can be trusted and under what conditions. Practitioners might care because it can help improve the reliability and coordination of these systems.
13 upvotes · 28 JUL 2026 · Junhan Sun, Hao Zhao, Guofeng Zhang
This paper introduces INTACT, a new method for training world models that can perform search-free actions without needing to test them. Practitioners might care because INTACT can improve the efficiency and effectiveness of world models in real-world applications.
13 upvotes · 29 JUL 2026 · Alexi Gladstone, Heng Ji, Yilun Du
This paper introduces Explorative Modeling, a new approach to training generative models that allows for end-to-end generation by exploring multiple candidate matches between model generations and data. This can lead to improved performance and efficiency in various applications.
11 upvotes · 16 JUL 2026 · Baohao Liao, Hanze Dong, Christof Monz et al.
This paper proposes a method to improve on-policy distillation by reusing pre-collected teacher data, allowing for faster training without interacting with the environment. Practitioners may care about this technique because it enables scalable and efficient distillation of complex agent models.
11 upvotes · 29 JUL 2026 · Peter Kirgis, Sayash Kapoor, Andrew Schwartz et al.
This paper investigates whether AI agents can conduct open-ended AI research and provides early evidence that they can perform the engineering aspects but struggle with critical parts of the research lifecycle, such as making progress on research questions and judgment about publishable research.
9 upvotes · 15 JUL 2026 · Zhihao Xie, Junfeng Wu, Xinting Hu et al.
This paper develops a method to transform video foundation models' representations into compact, reconstruction-capable, and generation-friendly video latents, which can be used in various generative modeling tasks. Practitioners can use VideoRAE to improve the performance of their models by leveraging the semantic and spatio-temporal structure captured by the frozen video foundation encoder.
9 upvotes · 21 JUL 2026 · Sam O'Nuallain, Nithya Rajkumar, Ramya Narayanasamy et al.
This paper introduces AutoIndex, a framework that learns to transform raw documents into representations for retrieval systems, allowing for more flexible and effective indexing. Practitioners may care about AutoIndex because it can improve the quality of search results in complex information retrieval tasks.
9 upvotes · 28 JUL 2026 · Pierre Chambon, Kunhao Zheng, Juliette Decugis et al.
This paper uses reinforcement learning to optimize code for speed while maintaining correctness, and it finds that the approach can significantly improve performance, even when the reward function is noisy or sparse. Practitioners might care because optimizing code for speed can be crucial in many applications, and this research provides a promising method for achieving that goal.
9 upvotes · 29 JUL 2026 · Chao-Han Huck Yang, Zih-Ching Chen, Piotr Zelasko et al.
This paper introduces a new approach to agentic speech recognition that uses a memory to help correct mistakes and improve accuracy. By limiting the corrections made, the system can avoid over-correcting and improve performance on challenging tasks.
9 upvotes · 29 JUL 2026 · Sizhe Zhou, Sheldon Yu, Hui Wei et al.
This paper investigates how Large Language Model (LLM) agents can use a file system to store and organize their memories, and whether this approach improves their performance. Practitioners might care because it shows that using a file system as memory can be beneficial for LLM agents, but there are limitations to this approach.
9 upvotes · 30 JUL 2026 · Yash Pandya, Sahil Gupta, Sarthak Harne et al.
This paper introduces a new method for training computer-use agents, called Echoverse, which generates evolving environments that mimic real-world applications. By using these environments, agents can learn more effectively and improve their performance on real-world tasks.
8 upvotes · 19 JUL 2026 · Chen Wang, Zhaochun Li, Jionghao Bai et al.
This paper proposes a new method called Distilled Reinforcement Learning that improves large language model post-training by providing fine-grained guidance to transfer new knowledge from a teacher model to a student model. Practitioners might care because it outperforms standard reinforcement learning and on-policy distillation methods in terms of knowledge transfer and model performance.
8 upvotes · 22 JUL 2026 · Arjun Majumdar, Avinash Sooriyarachchi, Benjamin Tibi et al.
This paper introduces Robostral Navigate, a vision-language model that enables robots to navigate using only a single monocular RGB camera, making it more scalable and cost-effective for deployment across various robotic platforms. Practitioners might care about this because it can simplify navigation tasks for robots in real-world environments.
8 upvotes · 20 JUL 2026 · Mei Yuan, Qi Long, Qifeng Wu et al.
This paper develops a new method for detecting anomalies in industrial video data, which is critical for quality control systems. Practitioners might care about this research because it aims to improve the accuracy and interpretability of anomaly detection in complex industrial settings.
8 upvotes · 28 JUL 2026 · Haoyang Huang, Wenjie Huang, Tianqi Xu et al.
This paper proposes a method to compress large language models by allocating a fixed budget to select the most important tokens, allowing for efficient inference and reduced memory usage. Practitioners may care about this research if they work on optimizing the performance of large language models for real-world applications.
7 upvotes · 23 JUL 2026 · Xiao Yu, Baolin Peng, Ruize Xu et al.
This paper creates a new framework, OpenForgeRL, that allows researchers to train AI agents in complex environments using real harnesses, rather than relying on simplified inference systems. Practitioners might care because it enables more realistic testing and training of agents in real-world settings.
7 upvotes · 30 JUL 2026 · Dongxiu Liu, Haoyi Niu, Peng Cheng et al.
This paper introduces ODEWorld, a new approach to modeling the physical world by learning a continuous latent velocity field that operates in physical time, allowing for more efficient and realistic predictions of future events. Practitioners in robotics and computer vision may care about ODEWorld's ability to provide rich planning-oriented information and high-quality image reconstruction.
6 upvotes · 24 JUL 2026 · Junlin Fang, Do Nguyen-Thanh, Xiaogang Xu et al.
This paper develops a method to clean up noisy signals in reasoning traces of large models to improve detection of hallucinations, which are false answers produced by models. Practitioners might care about this because it could lead to more accurate models that produce reliable answers.
6 upvotes · 29 MAY 2026 · Tzu-Heng Huang, Shengqi Qiu, Frederic Sala
This paper proposes a way to make automated evaluation systems more efficient, transparent, and reliable by distilling the decision logic of large language models into smaller, programmatic judges that can be easily inspected and edited. Practitioners might care because this approach could help reduce costs and improve the scalability of automated evaluation systems.
5 upvotes · 22 JUL 2026 · Hiskias Dingeto
This paper proposes a method to improve the explainability of natural-language autoencoders by making it harder for models to manipulate the explanations, and shows that this approach can increase the reliability of activation explanations and improve AI safety.
5 upvotes · 27 JUL 2026 · Jiahao Xie, Zhongbin Guo, Qianle Wang et al.
This paper introduces a systematic way to construct pretraining mixtures for Vision Language Models (VLMs) by breaking down the process into two parts: deciding which classes to combine and how to allocate data within each class. Practitioners can use this approach to improve the quality and diversity of their VLMs.
5 upvotes · 30 JUL 2026 · Yao Xiao, Reuben Tan, Zhen Zhu et al.
This paper proposes a new approach to improve vision-language models for visual retrieval, which can handle long visual contexts and large numbers of distractors. Practitioners might care because it can lead to better performance on image and video benchmarks.
4 upvotes · 20 JUL 2026 · Zhaokai Wang, Tianlin Gui, Jiayuan Rao et al.
This paper evaluates language models and deep-research agents at predicting football match outcomes before kickoff, using a dynamic benchmark that can be reused for future leagues. Practitioners can learn from the results to improve their own models' performance in similar tasks.
4 upvotes · 17 JUL 2026 · Haoran Sun, Wentao Zhang, Junyang Hua et al.
This paper develops a service-oriented framework, JoyNexus, to efficiently train and deploy Vision-Language-Action models across multiple tenants, improving resource utilization and reducing costs. Practitioners may care about JoyNexus for its potential to streamline the training process and make VLA models more accessible.