This paper proposes a method to improve long-CoT reasoning in large language models by addressing the issue of unequal token contributions to the final outcome. It shows that current methods, such as GRPO, assign too much credit to highly sensitive tokens and proposes a new method, CSCR, that reduces credit for these tokens to improve performance.
Firehose
Filtered to Papers, tagged “token-level learning value” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives