Researchers at IBM developed a method to improve the consistency of large language models (LLMs) like GPT-4.1, which can significantly impact their reliability in mission-critical applications. By analyzing an agent's past trajectories and identifying "flat" decisions, where the model is uncertain, they created a new type of guideline that helps stabilize these decisions. This approach, called consistency guidelines, can improve the Pass^5 metric, which measures the fraction of tasks an agent succeeds on all runs, by up to 22.9 percentage points. AI summary
Firehose
Filtered to Companies, tagged “Reinforcement Learning” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives