This paper investigates whether AI agents can conduct open-ended AI research and provides early evidence that they can perform the engineering aspects but struggle with critical parts of the research lifecycle, such as making progress on research questions and judgment about publishable research.
Firehose
Filtered to Papers, tagged “shadow evaluations” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives