Firehose

Filtered to Papers, tagged “post-hoc defense” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

14 SEP 2026 · Paper

This paper introduces a fast and efficient post-hoc defense against a type of attack that can bypass safety features in language models, allowing the model to continue functioning but with compromised security. Practitioners caring about model security may be interested in this approach as it can provide an additional layer of protection without requiring significant computational resources.