Papers

Filtered to continual learning · clear filter

Browse by term

continual learning 83reinforcement learning 50large language models 12benchmarking 11benchmarks 11language models 11vision-language models 11robotics 7natural language processing 6world models 6generative models 5recursive self-improvement 5attention mechanisms 4diffusion Transformers 4multi-agent systems 4multimodal learning 4multimodal models 4on-policy distillation 4self-distillation 4self-supervised learning 4transformers 4video generation 4vision-language-action models 4agent-based systems 3agentic models 3agentic search 3autonomous systems 3coding agents 3diffusion models 3image generation 3

Matching papers

Scaling Automatic Research Agents via World Models

452 upvotes · 29 AUG 2026 · Xiyuan Yang, Sheikh Sarwar, Jingru Cheng et al.

This paper proposes a method to scale automatic research agents by replacing environment execution with a world model, which can reduce training costs and improve performance. Practitioners might care about this approach because it can accelerate training times and lead to better results for complex AI tasks.

Atria Dawn: The Dawn of Agentic Superintelligence

424 upvotes · 14 SEP 2026 · Honglin Guo, Tao Gui, Yicheng Chen et al.

This paper introduces Atria Dawn Preview, a new type of AI model designed to work alongside humans in scientific research and engineering. Practitioners might care about this because it shows how AI can collaborate with humans more effectively, potentially leading to better research outcomes and more autonomous AI development.

StudentSim: Training LLM-based Student Simulators

399 upvotes · 1 SEP 2026 · Ke Yang, Chenglong Wang, Michel Galley et al.

This paper develops a method to create personalized AI tutors that can adapt to individual students' strengths and weaknesses, using a combination of pooled training and per-student fine-tuning. Practitioners in education and AI development may care about this work as it could lead to more effective and personalized learning experiences.

Compile by Training: Turning Natural-Language Specifications into Local Neural Functions

377 upvotes · 3 SEP 2026 · Yuntian Deng, Pengyu Nie, Stuart Shieber

This paper introduces a method to compile neural functions from natural-language specifications, allowing for faster and more reliable execution without relying on remote models. Practitioners may care about this approach for building efficient and flexible AI systems that can perform complex tasks without the need for expensive model calls.

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

360 upvotes · 31 AUG 2026 · Xin Zhou, Zongchuang Zhao, Zhibo Yang et al.

This paper develops a vision-language model for autonomous driving that combines 3D perception, question answering, and motion planning. A practitioner might care because it demonstrates a promising approach to integrating multiple tasks in autonomous driving.

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

331 upvotes · 8 SEP 2026 · NeoHorse Team, Guoliang Cao, Guohao Dai et al.

This paper proposes a method for recursive self-improvement in AI systems, where a model can learn from its own performance and use that knowledge to improve itself. Practitioners might care about this approach because it could lead to more efficient and effective AI systems.

NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

304 upvotes · 9 SEP 2026 · NCP Team, Jiaqi Cao, Chiyu Chen et al.

This paper introduces a new type of language model called NCP-ArchPreview that can generate text by predicting both individual tokens and larger concepts, and how this approach can improve performance. Practitioners might care about this because it could lead to better language models that can handle more complex tasks.

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

302 upvotes · 11 SEP 2026 · Jiyan He, Guang Liang, Hao Liu et al.

This paper introduces ZGCM-1, a highly efficient foundation model for math and agentic search that combines internal thinking with external tool use, and shows it can perform well on various benchmarks despite its compact size. Practitioners may care about the efficiency improvements and scalable architecture of ZGCM-1.

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

292 upvotes · 30 JUL 2026 · Hanzhang Zhou, Panrong Tong, Xu Zhang et al.

This paper introduces Qwen-UI-Agent, a type of artificial intelligence system that can perform tasks on various devices, such as smartphones and computers, and improve its abilities on its own. Practitioners might care about this research because it aims to create more practical and autonomous AI systems that can be used in real-world scenarios.

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

272 upvotes · 3 SEP 2026 · Jie Wu, Zhenru Zhang, Beichen Zhang et al.

This paper develops a method to turn agent trajectories into reusable environments, allowing for more efficient testing and interaction with the agent. Practitioners might care because this approach can help scale agent training and improve performance in various applications.

Metis: Memory Foundation Model

256 upvotes · 29 JUL 2026 · Zeyu Zhang, Ziliang Guo, Yihang Sun et al.

This paper introduces a new type of AI model called memory foundation models, which allows the model to learn and retain information internally, rather than relying on external memory modules. This could be useful for practitioners who want to build more efficient and flexible AI agents.

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

252 upvotes · 1 AUG 2026 · Yunhao Chen, Xin Wang, Yixu Wang et al.

This paper introduces OpenART, a platform for testing AI agents in open-ended environments, to evaluate their safety in complex and evolving scenarios. Practitioners can use OpenART to identify potential risks and improve the robustness of their AI systems.

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

236 upvotes · 1 SEP 2026 · Yuhao Wu, Jingyuan Zhang, Jiajun Shi et al.

This paper creates a benchmark (HarnessDev) to test whether large language models (LLMs) can design and improve their own execution infrastructure, called the agent harness, which affects their performance. Practitioners might care because it explores how models can adapt to changing environments and potentially improve efficiency.

On-Policy Self-Distillation without Any Supervision

183 upvotes · 9 AUG 2026 · Yijiang Li, Bingyang Wang, Yijun Liang et al.

This paper shows how to make large language models improve themselves without needing external guidance or supervision, by using their own internal consistency to correct mistakes. Practitioners might care about this because it could lead to more robust and self-sufficient AI models.

VGI-Bench: Probing Visual Intelligence in Video Generation Models

171 upvotes · 26 AUG 2026 · Xuan He, Cong Wei, Yuhao Cheng et al.

This paper introduces VGI-bench, a new benchmark for evaluating video generation models' visual reasoning capabilities, and finds that current models can solve some visually grounded tasks but still struggle with reliability. Practitioners may care about developing more reliable video generation models that can perform better on this benchmark.

Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

157 upvotes · 10 AUG 2026 · Mind Lab, Vin Bo, Asher Cai et al.

This paper introduces Macaron-V1, an open agent model family that enables continual learning and self-improvement in real-world environments, and explores its potential for collective intelligence. Practitioners might care about this work if they're interested in developing AI systems that can learn and adapt over time.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

154 upvotes · 19 JUL 2026 · Yuhan Zhu, Changlian Ma, Xiangyu Zeng et al.

This paper develops a new approach to understanding videos by predicting when specific events or evidence occur within the video. Practitioners working on video analysis and AI models might care about this research because it could lead to more accurate and robust video understanding systems.

Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training

149 upvotes · 27 AUG 2026 · Tingyun Li, Wenfeng Feng, Weiqing Li et al.

This paper proposes a method to determine which past update evidence in a large language model is still relevant and useful after subsequent training, to prevent wasting compute and potentially degrading the model's performance. Practitioners in the field of autonomous systems and language models might care about this problem because it can lead to better model performance and efficiency in adapting to changing domains and requirements.

SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

133 upvotes · 2 SEP 2026 · Junchao Huang, Guian Fang, Shengju Qian et al.

This paper introduces SolarWM, an open framework for training video world models from diverse datasets and using different video backbones, allowing for more consistent and reproducible results. Practitioners can use SolarWM to build interactive video world models that can be applied to various real-world applications.

RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM

128 upvotes · 13 JUL 2026 · Mikhail Komarov, Ivan Bondarenko, Stanislav Shtuka et al.

This paper introduces RAGU, an open-source GraphRAG engine that improves large language models with structured knowledge by separating extraction and consolidation, and trains a compact extractor that outperforms larger models on knowledge-graph construction and GraphRAG tasks. Practitioners might care because RAGU can efficiently generate more accurate and complete context for tasks like factoid-level evidence recall and multi-hop question answering.

FrontierChallenge: Evaluating Scientific Workflow Completion

127 upvotes · 25 AUG 2026 · Liangcai Su, Zhaopeng Feng, Zhuo Chen et al.

This paper introduces FrontierChallenge, a benchmark to evaluate the completion of end-to-end scientific workflows, and investigates the performance of various models in completing these workflows. Practitioners might care about this research to improve the accuracy of scientific workflow completion evaluations.

Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design

123 upvotes · 10 AUG 2026 · Qing Zong, Jiayu Liu, Junhao Shen et al.

This paper explores how agentic systems can improve on their own through co-evolution, where multiple agents and their environment adapt to each other, and discusses the challenges and benefits of building such systems that can learn beyond human design.

DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines

122 upvotes · 18 JUL 2026 · Runming He, Zhen Hao Wong, Hao Liang et al.

This paper creates a platform to help large language models generate code for data pipelines, which can then be edited and used to automate data processing workflows. Practitioners might care about this because it can help reduce the time and cost of developing and maintaining these pipelines.

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

120 upvotes · 3 AUG 2026 · Ziyu Ma, Hailang Huang, Shun Zou et al.

This paper proposes a framework, LongHorizon-Harness, to help large language model agents tackle long-horizon tasks by explicitly tracking task states and verifying facts from the environment. Practitioners might care because it can improve the performance of these agents on real-world tasks.

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

118 upvotes · 29 JUL 2026 · Hengyi Xie, Chenfei Yao, Xianjin Wu et al.

This paper introduces TurboVLA, a new vision-language-action model that reduces computation and memory overhead by directly exchanging information between visual observations and language instructions, allowing for faster and more efficient robotic manipulation. Practitioners might care about this approach for building more efficient and effective VLA models.

FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis

115 upvotes · 19 AUG 2026 · Kou Shi, Zun Wang, Qisheng Su et al.

This paper develops a method to generate high-quality terminal tasks for training agents, ensuring that the tasks accurately reflect the original instruction and environment, and providing a way to validate and improve the generated tasks. Practitioners in AI and robotics may care about this work because it addresses a common challenge in training terminal agents.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

114 upvotes · 3 AUG 2026 · Yu Zhang, Ruiqi Li, Changhao Pan et al.

This paper develops a system for generating speech and audio for various applications, including animation and video production, without reference recordings. Practitioners can use this system to create customized voices and control speaker styles, and to generate high-quality audio in complex scenarios.

EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction

110 upvotes · 2 SEP 2026 · Yuling Shi, Zhensu Sun, Junsen Dong et al.

This paper introduces EarlyEval, a method to reduce the cost of evaluating large language model (LLM) agents by predicting their outcomes early, allowing for earlier termination of agent runs. Practitioners may care because it can significantly reduce the computational cost of agent development.

LatentPress: Context Compression Beyond Text and Vision

109 upvotes · 1 SEP 2026 · Zhengze Zhou, Hejian Sang

This paper introduces LatentPress, a method to compress conversational histories and documents into a continuous memory token format that allows language models to directly read and process the context without text reconstruction. Practitioners might care about this because it could lead to faster and more efficient language model inference.

JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

108 upvotes · 26 AUG 2026 · Guibin Zhang, Leo Lu, Fangzhou Xie et al.

This paper develops a model that can automatically generate and adapt agent harnesses, which are crucial for the performance of AI models, to improve their ability to perform tasks. Practitioners might care about this because it could lead to more efficient and effective AI systems.

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

106 upvotes · 27 AUG 2026 · Tianjie Ju, Zheng Wu, Yueqing Sun et al.

This paper explores how large language models can turn local observations of a city into reliable actions, and whether these models can sustain goal-directed behavior in complex urban environments. Practitioners in AI/ML and urban planning might care about the limitations and potential of current models in navigating real-world cities.

LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering

99 upvotes · 28 AUG 2026 · Yi Wang, Haopeng Zhang, Chengxiang Huang et al.

This paper introduces LoopArena, a benchmark for evaluating how well a model can guide a coding agent through a long-running task, and finds room for improvement in long-horizon loop control. Practitioners working on AI development might care because it could lead to more efficient and reliable development processes.

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

97 upvotes · 16 AUG 2026 · GigaBrain Team, Angen Ye, Axiang Sun et al.

This paper presents GigaBrain-0.7, a new embodied foundation model that achieves strong generalization across diverse robot embodiments and tasks, by improving the architecture and scaling it to large amounts of data. Practitioners might care about this research if they're working on developing generalist robots that can adapt to new tasks and environments.

AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems

92 upvotes · 8 SEP 2026 · Jaewon Chu, Jinwoo Seo, Jaewon Cho et al.

This paper proposes a method to optimize prompts for multi-agent systems by identifying which agent's modification resolves a failure, and then using that agent's output as supervision to extract a fine-grained gradient. Practitioners might care because it could improve the performance of large language model-based multi-agent systems.

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

90 upvotes · 18 AUG 2026 · Zhi Zheng, Rongsheng Chen, Yunpeng Ba et al.

This paper proposes a new method for fine-tuning large language models (LLMs) in reinforcement learning (RL) tasks with long horizons, using evolution strategies (ES) instead of traditional backpropagation-based training. Practitioners might care because it allows for more efficient and flexible fine-tuning of LLMs, enabling them to tackle complex tasks with larger models and longer interactions.

DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation

89 upvotes · 29 JUL 2026 · Jiaxing Li, Kai Zou, Cindy Zhou et al.

This paper improves autoregressive video distillation methods by aligning the initialization and distribution matching stages, focusing on matching the target distribution's mode coverage rather than just visual quality. Practitioners can benefit from this approach to generate higher-quality videos with better diversity and coverage.

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

89 upvotes · 31 AUG 2026 · Yi Ding, Ruqi Zhang

This paper investigates whether on-policy distillation (OPD) truly improves student policies by analyzing the effects of noisy teacher supervision. It finds that OPD works by suppressing low-probability tokens, which can be achieved without a teacher, and introduces a new method called On-Policy Self-Adaptation (OPSA) that outperforms OPD and traditional reinforcement learning methods.

DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents

88 upvotes · 19 AUG 2026 · Hangrui Xu, Jiarui Wang, Yang Yang et al.

This paper proposes a new framework for training autonomous agents to perform multi-turn tool-calling tasks, addressing the challenge of dealing with vast solution spaces by using a diamond topology-aware approach. Practitioners may care about this research because it can help improve the performance of autonomous agents in complex tasks.

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

84 upvotes · 8 JUL 2026 · Xinyu Geng, Xuanhua He, Sixiang Chen et al.

This paper introduces a framework called DeepSearch-Evolve, which helps train self-improving web agents by iteratively refining their performance using their own experience. Practitioners might care because this approach can lead to more efficient and effective agents that can learn from their own mistakes.

SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries

74 upvotes · 6 AUG 2026 · Xingyu Tan, Xiaoyang Wang, Qing Liu et al.

This paper introduces a new technique called SkillZip that helps compress large libraries of skills used by artificial agents, making them more efficient and scalable. Practitioners who work on developing agents that can learn and adapt in complex environments might care about SkillZip because it can improve the performance and efficiency of their agents.

Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models

74 upvotes · 8 SEP 2026 · Zongjie Li, Alan Z. W, John Nicolas J et al.

This paper presents a data-centric framework to overcome challenges in training cyber agents, allowing for more efficient and effective training of open-weight models, and demonstrates the capability of a team of seven to train such models with leading agentic cyber capability.

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

73 upvotes · 28 JUL 2026 · Bo-Wen Zhang, Junwei He, Wen Wang et al.

This paper proposes a method to improve language model training by allocating credit to individual tokens within a response, allowing for more nuanced evaluation of model performance. Practitioners may care about this method as it can lead to better language model performance, especially in tasks that require specific formatting or semantic choices.

Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher

73 upvotes · 27 AUG 2026 · Shiyi Zhang, Mushui Liu, Yunze Tong et al.

This paper introduces a new method for training flow matching models called Self-OPD, which uses the model's own exploration to generate supervisory signals without needing a separate teacher model. Practitioners may care because it could lead to more efficient and effective training of flow matching models.

Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM

73 upvotes · 3 SEP 2026 · Sergii Kozyrev, Davyd Maiboroda

This paper investigates why a specific neural network architecture, Gated DeltaNet, can survive 4-bit quantization. The researchers found that a new quantization technique, NVFP4, can be used to quantize the recurrent state of GDN without significant loss of performance, and they provide a mechanistic explanation for why this is the case.

DarwinX: Evolving Agent Harnesses Through Natural Selection

72 upvotes · 31 JUL 2026 · Yifan Zhang, Yutong Dai, Juntao Tan et al.

This paper introduces DarwinX, a method for evolving agent capabilities through natural selection, which improves the agent's overall performance on various benchmarks without relying on task-specific patches. Practitioners might care about this approach as it allows for the creation of more general and durable agents that can adapt to changing environments.

TTPO: Test-Time Policy Optimization

72 upvotes · 27 AUG 2026 · Aozhe Wang, Zhengxi Lu, Jianze Wang et al.

This paper proposes a new method called Test-Time Policy Optimization (TTPO) that allows large language models to be trained without ground-truth labels, enabling test-time training. Practitioners may care about this because it enables models to be trained without labels, which can be difficult or expensive to obtain.

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

70 upvotes · 27 JUL 2026 · Bingnan Li, Haozhe Wang, Haozhong Xiong et al.

This paper investigates how to improve the adaptation of diffusion models in a way that doesn't rely on a classifier, and how to address a problem where the model can't accurately learn from its teacher. Practitioners might care about this because it could lead to more effective knowledge transfer in machine learning applications.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

68 upvotes · 30 JUL 2026 · Qiushi Sun, Kanzhi Cheng, Yian Wang et al.

This paper develops a standardized evaluation method for computer-using agents (CUAs) to ensure they fulfill task instructions, using vision-language models (VLMs) as judges. Practitioners can benefit from this work by using reliable and cost-effective reward signals for CUA training.

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution

65 upvotes · 8 AUG 2026 · Anton Razzhigaev, Andrei Gritsaev, Andrei Kaznacheev et al.

This paper introduces Ouroboros, a self-improving AI agent that develops its own tools and code through a process of reviewed commits, allowing it to learn and adapt over time. A practitioner might care about Ouroboros because it demonstrates a potential approach to creating autonomous AI systems that can improve themselves.

Unlocking Lossless Speedups in LLMs via Discrete Diffusion

64 upvotes · 3 SEP 2026 · Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham et al.

This paper proposes a new approach to Large Language Models (LLMs) that uses diffusion to speed up generation without sacrificing quality, allowing for faster and more efficient language processing. Practitioners may care about this research if they need to process large amounts of text quickly and efficiently.

Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction

64 upvotes · 8 SEP 2026 · Vishesh Tripathi, Abhay Kumar, Ramsha Khan

This paper introduces Grouped Value Attention (GVA), a technique to reduce the memory footprint of Transformer decoding by storing grouped values and reconstructing content keys. Practitioners might care about this because it could lead to faster and more efficient models.

It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning

62 upvotes · 1 SEP 2026 · Runpeng Dai, Kaili Huang, Changsung Kang et al.

This paper proposes a new retrieval framework called CoGR, which uses two generative models to co-evolve and optimize retrieval representations on both the query and item sides, leading to improved search and advertising performance. Practitioners might care because CoGR can potentially lead to better retrieval results and more efficient search systems.

Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

62 upvotes · 11 SEP 2026 · Jianman Lin, Shailesh Shailesh, Zhongyi Luo et al.

This paper proposes a method called Latent Interface Training (LIT) to improve the generalization of robotics foundation models by preventing them from relying on visual shortcuts when learning to generate actions from pre-trained visual representations. This is important for robots to perform well in new, unseen environments.

SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

61 upvotes · 20 AUG 2026 · Zhipeng Xu, Jiahao Lu, Yining Zheng et al.

This paper introduces SWE-bench Science, a benchmark for evaluating coding agents' performance in repairing scientific software, and identifies common failure mechanisms that hinder their success. Practitioners may care about this research as it can help improve the reliability and reproducibility of scientific findings by developing more effective coding agents.

What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents

61 upvotes · 27 AUG 2026 · Xingshan Zeng, Zishan Xu, Boju Zhang et al.

This paper develops a framework for generating agentic data that is consistent, useful, and sufficient for LLM agents to learn from. Practitioners might care because they want to ensure their agents learn from high-quality data that helps them adapt to changing environments.

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

60 upvotes · 21 JUL 2026 · Xinjie Zhang, Peng Zhang, Shicheng Zheng et al.

This paper introduces Mage-Flow, a compact model for generating and editing high-resolution images, which can be trained efficiently and deployed on a single GPU. Practitioners might care about the potential applications of this model in interactive image editing and generation tasks.

WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report

60 upvotes · 25 AUG 2026 · Junjie Zhou, Ke Mei, Lei Li et al.

This paper introduces WeMM-Embedding, a family of universal multimodal embedding models that can represent various types of data in a shared space, enabling applications like retrieval, recommendation, and classification. Practitioners might care because WeMM-Embedding achieves state-of-the-art performance on multiple benchmarks and has been deployed at scale in real-world applications.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

59 upvotes · 16 JUL 2026 · Xiaomi Robotics Team, Jun Guo, Piaopiao Jin et al.

This paper introduces a vision-language-action model that can perform mobile manipulation tasks in unseen environments with minimal training data, and how it can be scaled up to achieve better performance. Practitioners might care about this model for building robots that can adapt to new tasks with minimal fine-tuning.

EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents

59 upvotes · 5 SEP 2026 · Nanxi Li, Yingzi Ma, Yulong Cao et al.

This paper develops a framework to create customizable safety harnesses for AI agents, which can adapt to different models and domains to prevent harmful behavior. Practitioners may care about this research to improve the safety of AI systems in real-world applications.

Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search

57 upvotes · 16 AUG 2026 · Zhongwei Yu, Yan Song, Xue Yan et al.

This paper develops a new AI model that can efficiently search through vast spaces of possibilities, such as designing molecules or optimizing neural networks, by combining a generative model with a model that estimates the potential performance of each candidate. Practitioners might care about this approach because it can significantly speed up the discovery process and improve the quality of the results.

GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling

56 upvotes · 29 AUG 2026 · Guangting Zheng, Yiyuan Zhang, Tao Yang et al.

This paper introduces a new approach to training latent generative models, called GenFirst, which trains a generative model before reconstruction, and shows that this approach can achieve stable end-to-end training and state-of-the-art results on image generation tasks.

ASI-Bench: At the Dawn of Artificial Superintelligence

55 upvotes · 18 AUG 2026 · Junwei Zhou, Zhen Sun, Binyu Li et al.

This paper introduces ASI-Bench, a new benchmark for evaluating AI systems' ability to explore new knowledge, create new ideas, and conduct scientific research autonomously. Practitioners might care about ASI-Bench because it aims to push the limits of current AI systems and accelerate the development of artificial superintelligence.

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

55 upvotes · 27 AUG 2026 · Senqiao Yang, Chengyao Wang, Yuxin Chen et al.

This paper proposes a new approach to training Vision-Language-Action models by using a pre-trained backbone that captures generalizable visual-action knowledge from a large, diverse dataset of robot trajectories. This allows the model to perform well on new, unseen tasks without requiring a large amount of task-specific data.

DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space

54 upvotes · 28 JUL 2026 · Jiangwang Chen, Zixin Song, Junlin Liu et al.

This paper introduces a method called DecoEvo, which helps large language models improve by co-evolving a solver skill and a rubric-generator skill in a way that's more efficient and effective. Practitioners might care about this because it could lead to better performance and more reliable optimization in open-ended tasks.

Iris: Climbing to the Search Frontier

53 upvotes · 3 SEP 2026 · Ziyuan Liu, Hengqi Liu, Zichuan Wang et al.

This paper presents two search agents, Iris-mini and Iris-pro, trained to solve complex search tasks using reinforcement learning and self-supervised learning. Practitioners might care because these models achieve state-of-the-art results on various benchmarks, demonstrating the potential of AI-powered search agents in real-world applications.

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

51 upvotes · 30 JUL 2026 · Rubin Wei, Jiaqi Cao, Jiarui Wang et al.

This paper develops a method to scale up language models by increasing their memory capacity, allowing for better performance and more efficient use of parameters. Practitioners may care about this research as it could lead to more powerful and efficient language models for applications like language translation and text generation.

VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

50 upvotes · 15 AUG 2026 · Yansong Ning, Jingwen Ye, Zhongkai Wu et al.

This paper proposes a framework for training agents to create 3D open worlds based on user queries, and evaluates its performance using a large benchmark dataset. Practitioners might care about this research because it can help develop more capable multimodal agents that can understand user intent and generate realistic 3D environments.

SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking

50 upvotes · 11 SEP 2026 · Zhiwei Li, Lei Zhu, Hao Gu et al.

This paper proposes a new method to sparsify attention in Transformers, called Simple Attention Sparsification (SAS), which optimizes context ranking end-to-end with the language modeling loss. Practitioners might care because SAS can improve performance on downstream tasks by using attention budgets more effectively.

N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

49 upvotes · 26 JUL 2026 · NeoteAI Team, Fudan TEAI Team

This paper introduces a vision-tactile-language-action model that can perform fine-grained manipulation with tactile perception and control, and improve its policy offline from stored data. Practitioners may care because this model can be used to create more versatile and accurate tactile-driven manipulation policies.

On-Policy Self-Distillation in Diffusion Models

47 upvotes · 25 AUG 2026 · Wei Zhou, Xiongwei Zhu, Lingdong Kong et al.

This paper introduces a new method for improving diffusion models by using self-distillation to align them with human preferences and task-specific objectives. Practitioners might care about this approach because it can lead to more efficient and analyzable alignment of diffusion models with human goals.

Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report

47 upvotes · 22 AUG 2026 · TaoLive AIGC LLM Team, Yuhan Sun, Wenhao Lin et al.

This paper trains compact AI models to adapt quickly to changing digital avatar "harnesses" that define the tasks and tools available to the model, improving performance and reducing latency. Practitioners caring about real-time AI applications, such as chatbots or virtual assistants, might find this approach useful.

Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents

47 upvotes · 15 SEP 2026 · Caiqi Zhang, Xiaochen Zhu, Chengzu Li et al.

This paper proposes a new method for estimating confidence in language models, called XConf, which uses the model's past experiences to inform its confidence, rather than just relying on the current inference process. Practitioners might care about this because it could lead to more reliable and trustworthy deployment of language models.

Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory

46 upvotes · 7 AUG 2026 · Taeil Kim, Kangsan Kim, Sung Ju Hwang

This paper introduces Agent Memory Distillation, a technique that allows small language models to learn from a larger teacher model by transferring structured knowledge through hierarchical memory. Practitioners might care about this approach because it could improve the performance of small language models in tasks that require complex decision-making.

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

46 upvotes · 18 AUG 2026 · Hongyan Feng, Sunlai Chen, Xuanyu Liu et al.

This paper proposes a new framework for embodied navigation that overcomes limitations of existing methods by reformulating navigation into a 2D visual space, introducing selective reasoning and memory mechanisms, and designing an efficient alignment paradigm. Practitioners caring about efficient navigation for AI agents might care about this research.