Firehose

Filtered to tagged “benchmarking” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

31 JUL 2026 · Paper

This paper proposes a method to combine reinforcement learning with verifiable rewards and on-policy distillation to improve performance on complex tasks, and shows that this method can lead to more stable training and better results.

29 JUL 2026 · Paper

This paper investigates whether AI agents can conduct open-ended AI research and provides early evidence that they can perform the engineering aspects but struggle with critical parts of the research lifecycle, such as making progress on research questions and judgment about publishable research.

29 JUL 2026 · Paper

This paper introduces a benchmark to evaluate large language model (LLM) agents on long-horizon office-suite tasks, considering their cost-effectiveness and quality. Practitioners can care about this research because it aims to ensure LLM agents can assist users efficiently and effectively.

29 JUL 2026 · Paper

This paper creates a benchmark to test the security capabilities of AI agents in a real-world setting, specifically incident response, and finds that current agents struggle to detect and remediate silent intrusions and produce verified plans.