This paper introduces a fast and efficient post-hoc defense against a type of attack that can bypass safety features in language models, allowing the model to continue functioning but with compromised security. Practitioners caring about model security may be interested in this approach as it can provide an additional layer of protection without requiring significant computational resources.
Firehose
Filtered to Papers, tagged “weight-editing” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives