This paper investigates a common problem in reinforcement learning for language models called Value Flattening, where critics fail to accurately estimate state values, and proposes a new method, SP^3O, to mitigate this issue by supervising only a few well-separated states per response.
Firehose
Filtered to Papers, tagged “value estimation” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives