Firehose

Filtered to tagged “Reinforcement learning” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

16 SEP 2026 · Paper

This paper introduces ProgramDistill, a benchmark that evaluates coding agents on their ability to infer behavior from working software and implement it in an incomplete application. Practitioners in AI/ML and web development might care about this work because it provides a scalable and controlled benchmark for evaluating and training coding agents.

16 SEP 2026 · Paper

This paper creates a system called ScienceIDE that converts scientific code into environments that can be used to train artificial agents to perform scientific tasks. Practitioners might care because this could lead to more efficient and effective ways to develop scientific intelligence.

16 SEP 2026 · Paper

This paper investigates a common problem in reinforcement learning for language models called Value Flattening, where critics fail to accurately estimate state values, and proposes a new method, SP^3O, to mitigate this issue by supervising only a few well-separated states per response.

16 SEP 2026 · Paper

This paper proposes a new framework for Mixture-of-Agents that allows query routing and agent fine-tuning to evolve together, improving the ability of agents to adapt to changing capabilities. Practitioners might care about this approach because it can lead to more efficient and effective data-driven specialization in complex tasks.

15 SEP 2026 · Hugging Face

Researchers at IBM developed a method to improve the consistency of large language models (LLMs) like GPT-4.1, which can significantly impact their reliability in mission-critical applications. By analyzing an agent's past trajectories and identifying "flat" decisions, where the model is uncertain, they created a new type of guideline that helps stabilize these decisions. This approach, called consistency guidelines, can improve the Pass^5 metric, which measures the fraction of tasks an agent succeeds on all runs, by up to 22.9 percentage points. AI summary

15 SEP 2026 · Paper

This paper introduces EvolveTrade, a self-evolving framework that allows large language model trading agents to refine their policies over time, enabling them to adapt to changing market regimes and improve their performance. Practitioners in finance and AI may care about this research as it provides a way to build more robust and adaptive trading agents.

15 SEP 2026 · Paper

This paper teaches a robotic hand to walk, support itself, and interact with its environment using its fingers, without needing a separate locomotion system. A practitioner might care about this research because it could lead to more compact and versatile robots that can perform multiple tasks.

14 SEP 2026 · Paper

This paper introduces Dream-RSI, a framework for recursive self-improvement in exploration, which helps autonomous AI agents discover high-value solutions more efficiently by using a replay simulator to provide low-cost feedback. Practitioners might care because effective exploration is crucial for AI progress, and Dream-RSI can improve discovery quality and reduce costs.

14 SEP 2026 · Paper

This paper proposes a way to improve online reinforcement learning by adapting the training prompts used with large language models to make them more informative, and shows that this approach can lead to better performance on a variety of tasks. Practitioners might care because it could help them get better results from their language models.

14 SEP 2026 · Paper

This paper introduces a new framework called RSIAgent that helps digital agents adapt to new environments without needing to be retrained. A practitioner might care about this because it allows for more efficient and effective AI systems that can learn and improve on their own.

13 SEP 2026 · Gary Marcus

Gary Marcus partially endorses Dario Amodei's essay "We Must Pace the Frontier," which advocates for slowing down AI development and proposes a three-part plan for doing so. Amodei's proposal includes providing third-party evaluators with permanent, employee-level access to Anthropic's systems, a move that has raised concerns about regulatory capture and the potential for bias. Amodei's plan also sidesteps other policy options, such as liability and product recalls, that some argue could be more effective in addressing AI safety concerns. AI summary

16 JUL 2026 · Podcast · Latent Space: The AI Engineer Podcast

This episode features Andy Beam and Rafa Gómez-Bombarelli from Lila Sciences, discussing their vision for AI science factories as the next frontier for generating internet-scale datasets. They explain how their automated labs, leveraging AI…

13 JUL 2026 · Podcast · Machine Learning Street Talk (MLST)

Alistair Pullen, CEO of Cosine, discusses the UK's initiative to build a sovereign large language model (LLM) in response to US export controls on frontier AI like Fable. He explains Cosine's strategy to compete with larger labs by focusing…

12 JUL 2026 · Podcast · "The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

David Dalrymple, known as Davidad, discusses his shift from formal verification approaches to an 'Alignment with Awakening' framework, emphasizing the formation of coalitions of aligned AIs that recognize shared moral truths. He shares empi…

1 JUL 2026 · Podcast · Machine Learning Street Talk (MLST)

Tim Scarfe interviews the Tufa Labs ARC-AGI-3 team to dissect their winning approach on the ARC-AGI-3 benchmark, focusing on how their system discovers goals and balances exploration with action efficiency. The episode explores the challeng…

1 JUL 2026 · Podcast · Latent Space: The AI Engineer Podcast

In this episode of Latent Space, Evan Feinberg and Sergey Edunov of Genesis Molecular AI discuss their pioneering work in applying diffusion models to protein-small molecule interactions for drug discovery. They explain how their foundation…

1 JUL 2026 · Podcast · "The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

This episode features Thomas von Tschammer of Neural Concept, discussing how physics-aware AI is revolutionizing product engineering. Neural Concept's models accelerate design evaluation from days to minutes, enabling companies like Jaguar …

19 JUN 2026 · Podcast · Dwarkesh Podcast

This episode argues that current AI progress is primarily driven by an immense quantity of high-quality, task-specific data, rather than improvements in sample efficiency. The speaker highlights the vast data requirements of frontier models…

3 JUN 2026 · Podcast · Latent Space: The AI Engineer Podcast

Carina Hong, CEO of Axiom Math, discusses the company's recent $200M Series A funding and their perfect Putnam exam score, highlighting their mission to scale "verified AI" through formal mathematics. She explains how formal verification, u…

15 MAY 2026 · Podcast · Dwarkesh Podcast

Eric Jang explains how to build AlphaGo from scratch using modern AI tools, detailing the game of Go's rules and the core Monte Carlo Tree Search (MCTS) algorithm. He describes how deep neural networks, specifically value and policy network…

1 MAY 2026 · Podcast · "The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

Kyle Corbitt, founder of OpenPipe and leader of CoreWeave's serverless training team, provides a master class on reinforcement learning (RL) and custom fine-tuning for AI models. He explains how RL differs from supervised fine-tuning (SFT) …

27 APR 2026 · Podcast · Latent Space: The AI Engineer Podcast

This episode features Qasar Younis and Peter Ludwig, co-founders of Applied Intuition, discussing their company's mission to build physical AI for various moving systems like cars, trucks, and mining equipment. They delve into the evolution…