This paper compares the performance of different retrieval-augmented generation (RAG) paradigms at varying corpus sizes, finding that BM25 outperforms others at larger scales, but not at smaller ones, and that lexical retrieval is the strongest scalable default.
Firehose
Filtered to Papers, tagged “BM25” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives