Firehose

Filtered to Papers, tagged “long-CoT reasoning” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

30 JUL 2026 · Paper

This paper proposes a method to improve long-CoT reasoning in large language models by addressing the issue of unequal token contributions to the final outcome. It shows that current methods, such as GRPO, assign too much credit to highly sensitive tokens and proposes a new method, CSCR, that reduces credit for these tokens to improve performance.