This paper introduces a benchmark to evaluate large language model (LLM) agents on long-horizon office-suite tasks, considering their cost-effectiveness and quality. Practitioners can care about this research because it aims to ensure LLM agents can assist users efficiently and effectively.
Firehose
Filtered to Papers, tagged “human labor” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives