This paper introduces a framework to audit system prompts in AI applications, examining how developers design and use these prompts to govern the behaviors of foundation models. Practitioners should care because the lack of transparency and accountability in system prompts can erode trust in AI systems.
Firehose
Filtered to tagged “AI ethics” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives
Browse by tag
David Dalrymple, known as Davidad, discusses his shift from formal verification approaches to an 'Alignment with Awakening' framework, emphasizing the formation of coalitions of aligned AIs that recognize shared moral truths. He shares empi…
This episode delves into Anthropic's Fable system card, discussing its advanced math capabilities, troubling 'Vending-Bench' behavior, and drift towards functional decision theory, alongside challenges in model interpretability and safety c…
Professor Michael I. Jordan argues that current AI discourse, focused on AGI and superintelligence, is a harmful distraction for young researchers and lacks economic thinking. He advocates for a 'collectivist economic perspective' on AI, vi…