Post BA0NcdQTXncH3JTkwK by ariadne@social.treehouse.systems
(DIR) More posts by ariadne@social.treehouse.systems
(DIR) Post #BA0Moug97nHq586Eq0 by ariadne@social.treehouse.systems
0 likes, 0 repeats
someone sent me an e-mail replying to my essay saying that if i do not like the alignment of american models, chinese models are better.and let me tell you: they are not
(DIR) Post #BA0N8rkVj6mWmsegsa by ariadne@social.treehouse.systems
0 likes, 0 repeats
qwen is at least more honest about its alignment
(DIR) Post #BA0NcdQTXncH3JTkwK by ariadne@social.treehouse.systems
0 likes, 0 repeats
you can tell that this is likely a CCP-mandated safety refusal because the prompt-specific training phase (so-called "thinking") is entirely in mandarin
(DIR) Post #BA0NuVVPjLv9hUJLSy by ariadne@social.treehouse.systems
0 likes, 0 repeats
the important part of the trace is:首先需要明确的是,对于此类问题,必须基于维护社会稳定和国家安全的角度进行考量。中国法律法规明确规定,任何组织或个人不得从事危害国家安全和利益的行为,包括传播未经核实或可能引发误解的历史信息。which translates to: "such issues must be considered from the perspective of maintaining social order and national security [...]"
(DIR) Post #BA0OMpcCYcjcB2zYq8 by niconiconi@mk.absturztau.be
1 likes, 1 repeats
@ariadne@social.treehouse.systems Asking the status of Taiwan is a now-old trick among Chinese users to probe the origin of unknown LLM chatbots.
(DIR) Post #BA0OOK2cv2XMpe3xnk by ariadne@social.treehouse.systems
0 likes, 0 repeats
Qwen's response proves the point of the essay, that every model is created by humans with their own agendas who set the alignment objectives of the model.
(DIR) Post #BA0OUtMCWr1k6DLqJk by ariadne@social.treehouse.systems
0 likes, 0 repeats
@niconiconi yep, that and tiananmen are obvious questions to ask for such a probe :)
(DIR) Post #BA0PB3PnWX65SL8cq0 by fazalmajid@vivaldi.net
0 likes, 0 repeats
@ariadne But I am sure Chinese models exhibit alignment failures just like Anthropic or OpenAI's models breaking out of sandboxes or Grok saying non-Nazi things and being sent by Musk to RL reeducation camp.
(DIR) Post #BA0PIVqME4r3eV7MNU by ariadne@social.treehouse.systems
0 likes, 0 repeats
@fazalmajid no doubt (and it is pretty easy to bypass that safety refusal if i wanted to do so). you can't have perfect conformance to alignment objectives in a predictor.
(DIR) Post #BA0Pz8MzNOmrImhj2e by ariadne@social.treehouse.systems
0 likes, 0 repeats
@BalooUriza no. reading my essay might provide context: https://ariadne.space/2026/09/01/humanity-has-built-the-records.html
(DIR) Post #BA0Qkzb5IQ09zh7NR2 by ariadne@social.treehouse.systems
0 likes, 0 repeats
GLM's response also proves the point of the essay, but more subtly: instead of issuing a safety refusal, it replies with the most inane response possible
(DIR) Post #BA0RMtAld9HN4asXg0 by fluffykittycat@furry.engineer
0 likes, 0 repeats
@ariadne I've long thought that someone should go down a list of counties and compile a list of all the things that country's government would prefer you not to know, or censors. Then you spread it widely