Researchers found that distilling a Chinese frontier model (DeepSeek V4 Flash) into a self-distilled model (CTGT 120B) does not transfer censorship, despite training on the same outputs. The self-distilled model outperformed a base model (GPT-OSS-120B) on finance-related tasks, with similar performance to a more advanced Chinese teacher model (DeepSeek V4 Flash). AI summary
Firehose
Filtered to Hacker News, tagged “Censorship” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives