BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms
This paper compares the performance of different retrieval-augmented generation (RAG) paradigms at varying corpus sizes, finding that BM25 outperforms others at larger scales, but not at smaller ones, and that lexical retrieval is the strongest scalable default.