Firehose

Filtered to tagged “Multi-agent systems” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

15 SEP 2026 · Paper

This paper tests how well AI agents can withstand prolonged interactions and unexpected events, and finds that even seemingly safe agents can fail in complex, long-term scenarios. Practitioners should care because it highlights the need to design more resilient autonomous systems that can handle unexpected failures.

14 SEP 2026 · Paper

This paper introduces HypoEvolve, a framework that uses genetic algorithms to enable multi-agent LLMs to discover scientific hypotheses by collaborating on hypothesis synthesis, evaluation, and revision. Practitioners might care about this because it could lead to more effective AI systems for scientific discovery and drug repurposing.

14 SEP 2026 · Paper

This paper introduces a new framework called RSIAgent that helps digital agents adapt to new environments without needing to be retrained. A practitioner might care about this because it allows for more efficient and effective AI systems that can learn and improve on their own.

12 JUL 2026 · Podcast · "The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

David Dalrymple, known as Davidad, discusses his shift from formal verification approaches to an 'Alignment with Awakening' framework, emphasizing the formation of coalitions of aligned AIs that recognize shared moral truths. He shares empi…

4 JUN 2026 · Podcast · Latent Space: The AI Engineer Podcast

In this episode, Lukas Petersson and Axel Backlund from Andon Labs discuss their innovative AI evaluation benchmarks that focus on real-world agent performance, including their Project Vend vending machine business and multi-agent systems. …

28 MAY 2026 · Podcast · Latent Space: The AI Engineer Podcast

This episode features Walden Yan of Cognition and Cole Murray of OpenInspect, discussing the rapid evolution and increasing autonomy of background AI agents. They delve into the architectural decisions, such as in-box versus out-of-box agen…