Firehose

Filtered to tagged “Reinforcement learning” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

31 JUL 2026 · Paper

This paper develops a new method to evaluate and verify image editing consistency across multiple references, addressing a challenge in reinforcement learning for multi-reference editing. Practitioners may care about this approach as it enables more accurate and reliable reinforcement learning for image editing tasks.

31 JUL 2026 · Paper

This paper proposes a method to combine reinforcement learning with verifiable rewards and on-policy distillation to improve performance on complex tasks, and shows that this method can lead to more stable training and better results.

30 JUL 2026 · Paper

This paper proposes a method to improve long-CoT reasoning in large language models by addressing the issue of unequal token contributions to the final outcome. It shows that current methods, such as GRPO, assign too much credit to highly sensitive tokens and proposes a new method, CSCR, that reduces credit for these tokens to improve performance.

30 JUL 2026 · Paper

This paper develops a method to improve the performance of Vision-Language-Action models by adapting their steering strategy at test time, allowing them to generalize better to new tasks and domains. Practitioners can benefit from this approach by improving the robustness of their VLA models in real-world applications.

30 JUL 2026 · Paper

This paper develops a framework, SpatialCLI, to help vision-language models (VLMs) better understand and use visual tools to make better decisions. By training VLMs to reason with spatial tools and then internalize those capabilities, SpatialCLI can improve the performance of VLMs in tasks that require visual reasoning.

30 JUL 2026 · Paper

This paper proposes a new approach to memory-augmentation in large language model agents, allowing them to actively reconstruct and adapt past experiences to fit the current context, rather than simply replaying them. Practitioners might care because this approach can improve the robustness and intrinsic reasoning capabilities of agents in complex scenarios.

30 JUL 2026 · Paper

This paper proposes a new approach to agentic visual reasoning, which helps large language models (LLMs) perform better on complex tasks by using tools more efficiently. Practitioners might care about this research because it aims to improve the performance of LLMs on challenging problems.

30 JUL 2026 · Paper

This paper develops a new method for improving reasoning language models, called β-OPSD, which combines policy optimization and self-distillation to improve stability and performance. Practitioners might care about this method because it provides a more efficient and effective way to improve language model reasoning abilities.

29 JUL 2026 · Paper

This paper develops a new framework for world modeling that takes into account the mental state of agents, which is essential for predicting human decisions. Practitioners caring about human decision-making and planning might find this research useful.

16 JUL 2026 · Podcast · Latent Space: The AI Engineer Podcast

This episode features Andy Beam and Rafa Gómez-Bombarelli from Lila Sciences, discussing their vision for AI science factories as the next frontier for generating internet-scale datasets. They explain how their automated labs, leveraging AI…

13 JUL 2026 · Podcast · Machine Learning Street Talk (MLST)

Alistair Pullen, CEO of Cosine, discusses the UK's initiative to build a sovereign large language model (LLM) in response to US export controls on frontier AI like Fable. He explains Cosine's strategy to compete with larger labs by focusing…

12 JUL 2026 · Podcast · "The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

David Dalrymple, known as Davidad, discusses his shift from formal verification approaches to an 'Alignment with Awakening' framework, emphasizing the formation of coalitions of aligned AIs that recognize shared moral truths. He shares empi…

1 JUL 2026 · Podcast · Machine Learning Street Talk (MLST)

Tim Scarfe interviews the Tufa Labs ARC-AGI-3 team to dissect their winning approach on the ARC-AGI-3 benchmark, focusing on how their system discovers goals and balances exploration with action efficiency. The episode explores the challeng…

1 JUL 2026 · Podcast · Latent Space: The AI Engineer Podcast

In this episode of Latent Space, Evan Feinberg and Sergey Edunov of Genesis Molecular AI discuss their pioneering work in applying diffusion models to protein-small molecule interactions for drug discovery. They explain how their foundation…

1 JUL 2026 · Podcast · "The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

This episode features Thomas von Tschammer of Neural Concept, discussing how physics-aware AI is revolutionizing product engineering. Neural Concept's models accelerate design evaluation from days to minutes, enabling companies like Jaguar …

19 JUN 2026 · Podcast · Dwarkesh Podcast

This episode argues that current AI progress is primarily driven by an immense quantity of high-quality, task-specific data, rather than improvements in sample efficiency. The speaker highlights the vast data requirements of frontier models…

3 JUN 2026 · Podcast · Latent Space: The AI Engineer Podcast

Carina Hong, CEO of Axiom Math, discusses the company's recent $200M Series A funding and their perfect Putnam exam score, highlighting their mission to scale "verified AI" through formal mathematics. She explains how formal verification, u…

15 MAY 2026 · Podcast · Dwarkesh Podcast

Eric Jang explains how to build AlphaGo from scratch using modern AI tools, detailing the game of Go's rules and the core Monte Carlo Tree Search (MCTS) algorithm. He describes how deep neural networks, specifically value and policy network…

1 MAY 2026 · Podcast · "The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

Kyle Corbitt, founder of OpenPipe and leader of CoreWeave's serverless training team, provides a master class on reinforcement learning (RL) and custom fine-tuning for AI models. He explains how RL differs from supervised fine-tuning (SFT) …

27 APR 2026 · Podcast · Latent Space: The AI Engineer Podcast

This episode features Qasar Younis and Peter Ludwig, co-founders of Applied Intuition, discussing their company's mission to build physical AI for various moving systems like cars, trucks, and mining equipment. They delve into the evolution…