A diagonal attack on linearly represented truth probes, inspired by Tarski's paradox, shows that no probe can pin down truth in a language model's embedding space due to the self-referential nature of natural language, which can express its own properties and lead to paradoxes like the liar paradox. AI summary
Firehose
Filtered to Hacker News, tagged “formal semantics” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives