This paper proposes a framework for generalizable recursive self-improvement (RSI) of agent harnesses, which can improve execution mechanisms without being specific to a particular task or benchmark. Practitioners can care about this work because it aims to create more adaptable and transferable AI agents.
Firehose
Filtered to tagged “recursive self-improvement” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives
Browse by tag
This paper introduces Dream-RSI, a framework for recursive self-improvement in exploration, which helps autonomous AI agents discover high-value solutions more efficiently by using a replay simulator to provide low-cost feedback. Practitioners might care because effective exploration is crucial for AI progress, and Dream-RSI can improve discovery quality and reduce costs.
This paper introduces a new framework called RSIAgent that helps digital agents adapt to new environments without needing to be retrained. A practitioner might care about this because it allows for more efficient and effective AI systems that can learn and improve on their own.
David Dalrymple, known as Davidad, discusses his shift from formal verification approaches to an 'Alignment with Awakening' framework, emphasizing the formation of coalitions of aligned AIs that recognize shared moral truths. He shares empi…
In this episode, Thomas Ahle discusses the development of thermodynamic computing chips and the challenges of chip design automation using AI agents. He explains how his team built an open-source Verilog simulator with AI collaboration to o…
OpenAI research scientist Noam Brown discusses how traditional AI benchmarks are failing to accurately evaluate modern models due to their increasing reliance on large-scale test-time compute. He argues that model capabilities are now a fun…
Carina Hong, CEO of Axiom Math, discusses the company's recent $200M Series A funding and their perfect Putnam exam score, highlighting their mission to scale "verified AI" through formal mathematics. She explains how formal verification, u…
The Model Eats the Scaffolding: DeepMind's Logan Kilpatrick & Tulsee Doshi on 3.5 Flash, Omni & More
This episode features Logan Kilpatrick and Tulsee Doshi of Google DeepMind discussing Google's AI strategy and new launches at Google I/O, including Gemini 3.5 Flash, the Omni video generation model, and the Gemini Spark agentic product. Th…
This episode features Beth Barnes and David Rein from METR discussing their 'Time Horizon' graph, a unified metric for measuring AI progress based on human task completion time. They explain how this benchmark addresses the limitations of t…
The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking
Kyle Corbitt, founder of OpenPipe and leader of CoreWeave's serverless training team, provides a master class on reinforcement learning (RL) and custom fine-tuning for AI models. He explains how RL differs from supervised fine-tuning (SFT) …