This paper proposes a new method for aligning large language models with human preferences, called Comparison-based Preference Optimization (ComPO), which is more efficient than existing methods and can mitigate a problem called likelihood displacement. Practitioners might care about this paper because it offers a new approach to aligning LLMs with human preferences, which is essential for developing more reliable and trustworthy AI models.
Firehose
Filtered to Papers, tagged “large language models” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives
Browse by tag
This paper introduces HypoEvolve, a framework that uses genetic algorithms to enable multi-agent LLMs to discover scientific hypotheses by collaborating on hypothesis synthesis, evaluation, and revision. Practitioners might care about this because it could lead to more effective AI systems for scientific discovery and drug repurposing.
This paper proposes a way to improve online reinforcement learning by adapting the training prompts used with large language models to make them more informative, and shows that this approach can lead to better performance on a variety of tasks. Practitioners might care because it could help them get better results from their language models.
This paper investigates how large language model (LLM) agents adapt their performance during long tasks, and how their test-time strategies impact their scalability. Practitioners might care because understanding these strategies can help improve the performance of LLM agents in real-world applications.