Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it

AI panic, meet reality: users say the ‘China censorship’ fear may not survive copy-paste training

TLDR: Researchers say a new AI trained on a Chinese model’s answers gained useful skills without copying its political refusals, challenging fears that censorship automatically rubs off. Commenters were split between “that totally makes sense,” deeper what-if worries about hidden influence, and one hilariously off-topic complaint about broken iPad scrolling.

A spicy little showdown broke out on Hacker News after a team claimed it could train a new AI on answers from DeepSeek, a Chinese model known for dodging certain political topics, without importing that same censorship. In plain English: the student seemed to pick up the smarts, but not the political refusal habits. That immediately poked at a very online fear: if an AI learns from a restricted model, does it secretly absorb the restrictions too? The post says no — at least in this finance-focused test.

And the comment section? Half intrigued, half side-eyeing the whole premise. One user confidently declared this “makes full sense,” arguing censorship is basically missing information, and training mostly adds skills rather than deleting them. Another zoomed out and asked the more dramatic question: okay, if there’s no hidden transfer here, when would there be? That turned the thread into a mini conspiracy-vs-science cage match over whether “subliminal learning” is real or just AI ghost stories for policy people.

Then, because the internet refuses to stay serious for too long, someone swerved hard into peak comment-section energy by reporting that the iPad trackpad scrolling is broken. Absolute legend behavior. Another user pitched a nerdy-but-juicy idea: build a public scoreboard showing which AI censors what, basically turning model politics into a live rankings war. So yes, the paper was about safety and values — but the crowd turned it into a referendum on hidden influence, paranoia, and whether the real bug was censorship or the playground UI.

Key Points

  • The article studies whether distilling outputs from a censored Chinese frontier model transfers censorship along with capability gains.
  • The authors report that a student model trained on outputs from a heavily censored Chinese model improved in financial reasoning and performance without showing similar censorship.
  • The article says a self-distilled model achieved the same score as a model taught by a more advanced Chinese teacher model.
  • As background, the article notes that Chinese frontier models have documented refusals and reframing on China-sensitive topics, including DeepSeek model lines.
  • A cited example shows DeepSeek V4 Flash refusing a question about Uyghur labor-transfer programs in Xinjiang, while answering a matched control question about Uzbekistan’s forced labor history.

Hottest takes

"distillation is only additive, not subtractive" — Alifatisk
"there is no subliminal learning in this situation" — andy99
"the scrolling on iPad with trackpad is broken" — data-ottawa
Made with <3 by @siedrix and @shesho from CDMX. Powered by Forge&Hive.