Firehose

Filtered to Papers, tagged “adversarial examples” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

15 SEP 2026 · Paper

This paper tests the robustness of rubrics generated by language models as reward signals in reinforcement learning, finding that even generic rubrics can be exploited 64% of the time, while tailored rubrics can be used to create fake answers. Practitioners should care because this can lead to biased grading and evaluation.

14 SEP 2026 · Paper

This paper helps developers make stronger backdoor attacks on large language models by learning to select the most effective set of poisoned examples. Practitioners might care about this because it can be used to improve the security of these models in real-world applications.