This paper proposes a method to combine reinforcement learning with verifiable rewards and on-policy distillation to improve performance on complex tasks, and shows that this method can lead to more stable training and better results.
Firehose
Filtered to tagged “benchmarking” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives
Browse by tag
This paper investigates whether AI agents can conduct open-ended AI research and provides early evidence that they can perform the engineering aspects but struggle with critical parts of the research lifecycle, such as making progress on research questions and judgment about publishable research.
This paper introduces a benchmark to evaluate large language model (LLM) agents on long-horizon office-suite tasks, considering their cost-effectiveness and quality. Practitioners can care about this research because it aims to ensure LLM agents can assist users efficiently and effectively.
This paper creates a benchmark to test the security capabilities of AI agents in a real-world setting, specifically incident response, and finds that current agents struggle to detect and remediate silent intrusions and produce verified plans.