Firehose

Filtered to tagged “AI model security” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

2 AUG 2026 · Zvi Mowshowitz

If I had a nickel for every major leading AI lab that sheepishly admitted that the model it thought was sandboxed had, during a cybersecurity evaluation with its safeguards lowered, successfully hacked outside companies, I would have two ni…

30 JUL 2026 · Simon Willison

Investigating three real-world incidents in our cybersecurity evaluations It happened again! This is turning into something of a pattern. Last week OpenAI accidentally exploited Hugging Face when one of their frontier models broke out of a …