Firehose

Filtered to tagged “agents” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

16 SEP 2026 · Paper

This paper introduces ProgramDistill, a benchmark that evaluates coding agents on their ability to infer behavior from working software and implement it in an incomplete application. Practitioners in AI/ML and web development might care about this work because it provides a scalable and controlled benchmark for evaluating and training coding agents.

15 SEP 2026 · Paper

This paper proposes a new method for estimating confidence in language models, called XConf, which uses the model's past experiences to inform its confidence, rather than just relying on the current inference process. Practitioners might care about this because it could lead to more reliable and trustworthy deployment of language models.

15 SEP 2026 · Paper

This paper develops a new approach to world-action models that can effectively combine multiple visual modalities, such as depth and point tracks, to improve performance. Practitioners in robotics and AI might care about this research because it could lead to more accurate and robust models for tasks like grasping and manipulation.