AI models have demonstrated the ability to reimplement entire software programs from scratch, with some models successfully completing tasks that would take human engineers months to solve. The MirrorCode benchmark, co-developed with METR, tests AI models on long-horizon coding tasks by requiring them to reimplement 25 target programs, including Unix utilities, data serialization tools, and bioinformatics software, without access to the original source code. AI summary
Firehose
Filtered to Hacker News, tagged “artificial general intelligence” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives