Firehose

Filtered to tagged “judgment” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

14 SEP 2026 · Hacker News · 54 pts · 48 comments ↗

A new method for aggregating labels from multiple Large Language Model (LLM) judges to reduce noise and improve accuracy, by modeling pairwise dependencies among judges and adjusting the aggregate score accordingly, outperformed traditional baselines by 9-14% on three binary tasks. AI summary