A new method for aggregating labels from multiple Large Language Model (LLM) judges to reduce noise and improve accuracy, by modeling pairwise dependencies among judges and adjusting the aggregate score accordingly, outperformed traditional baselines by 9-14% on three binary tasks. AI summary
Firehose
Filtered to Hacker News, tagged “judgment” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives