This paper investigates the losslessness of a new language model architecture called Orthrus, which claims to achieve exact output sequences through a speculative decoding mechanism. The study finds that the architecture's performance depends on the numerical precision used, and that downstream task performance is not necessarily affected by the losslessness of the speculative decoding.
Firehose
Filtered to Papers, tagged “hybrid models” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives