22 upvotes · 22 JUL 2026 · Jihoon Tack, Philippe Laban, Jennifer Neville
This paper investigates how Large Language Models (LLMs) handle dynamic user intent in conversations, where users' goals change over time, and finds that current LLMs struggle with this capability, leading to significant performance drops when evaluated in a more realistic setting.
14 upvotes · 22 JUL 2026 · Jiarong Zhao, Zhikai Lei, Zhiheng Xi et al.
This paper develops a framework called NexForge that helps train more capable artificial agents by automatically generating a large number of tasks and training data, without requiring a lot of manual setup. Practitioners might care because it can improve the performance of their own agent models.
13 upvotes · 26 JUL 2026 · Haorui He, Xinwen Chen, Dacheng Wen et al.
This paper investigates the reliability of dynamic benchmarks for multimodal automated fact-checking by examining contamination risks and their impact on evaluation metrics. Practitioners should consider the potential for contamination in dynamic benchmarks to ensure accurate performance estimates.
13 upvotes · 30 JUL 2026 · Peilin Feng, Suorong Yang, Soujanya Poria
This paper introduces a new type of memory system for large language model (LLM) based multi-agent systems that tracks which agents can be trusted and under what conditions. Practitioners might care because it can help improve the reliability and coordination of these systems.