Firehose

Filtered to Companies, tagged “Reinforcement Learning” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

15 SEP 2026 · Hugging Face

Researchers at IBM developed a method to improve the consistency of large language models (LLMs) like GPT-4.1, which can significantly impact their reliability in mission-critical applications. By analyzing an agent's past trajectories and identifying "flat" decisions, where the model is uncertain, they created a new type of guideline that helps stabilize these decisions. This approach, called consistency guidelines, can improve the Pass^5 metric, which measures the fraction of tasks an agent succeeds on all runs, by up to 22.9 percentage points. AI summary