Taiwan's controversy over the "right to interpret Chinese" is moving from cultural politics to artificial intelligence training data. On September 19, China Research Institute European and American Research Institute researcher Hong Zhou Wei warned at the Taipei seminar that large amounts of simplified Chinese material could give Beijing a stronger interpretative advantage in global Chinese knowledge production. Four days ago, Taiwan's Ministry of Intelligence just announced the expansion of "Taiwan sovereign AI training language library", and made it clear that Chinese training materials for large international language models are still predominantly simplified Chinese content.

Putting these two things together, the problem is no longer just the big and short words who use more, but a more specific technical fact: generated AI requires a huge amount of training data, while the size, source and availability of Chinese network data is not average. The Taiwanese government is investing resources in supplementing local language, which itself shows that the structural gap in Chinese AI training data has been dealt with as a policy issue.

The data released on September 15 by the Taiwan Ministry of Development is that the sovereign AI training language library has accumulated about 5,000 data sets and more than 2.2 billion words since its launch last year.The new recruitment target includes publishers, e-book platforms and writers, with the aim of increasing Taiwan's own language, cultural and social values into the opportunity for model training materials.

原始来源 · money.udn.com中央社/经济日报:台湾启动民间语料征集,扩充主权AI训练语料库台湾数发部称国际大型语言模型中文训练资料仍多以简体中文内容为主,并启动民间语料征集。money.udn.com ↗

This explains that the so-called “interpretation right” is not a completely abstract political wording. How the big model answers questions of history, politics, law, and identity, related to what material it is exposed when it is trained. The model does not automatically know which narrative comes from the free media, which one comes from the censored news environment; it first deals with the statistical relationship between a large number of texts.

The scale of digital content in mainland China is much larger than in Taiwan. News websites, government pages, encyclopedias, forums, business platforms, short video subtitles and social media constitute a huge pool of simple Chinese information. But this pool is not formed in a completely free information environment. China's news media is politically controlled, online platforms undergo censorship obligations, sensitive topics are removed, restricted or prohibited, and government and official media have the ability to produce content on a continuous and large scale.

This means a risk that must be disconnected: If the international model uses massive amounts of China’s mainland public network material, it absorbs not only the “Simple Chinese” form of text, but may also absorb an information environment that has been censored, deleted, and long formed by official narratives.

He called this phenomenon "knowledge injustice" at the September 19th "Resistance" seminar, and said that Beijing was trying to use a large amount of Chinese material to expand its influence on the global Chinese language interpretation. He also put the issue into the framework of Taiwan's long-term language policy and cognitive campaign discussions. These are Hongzhu's analysis, not a causal conclusion already proven by an open model audit; but the Taiwan Ministry of Foreign Affairs has actively expanded the local sovereign language repository, providing an observable policy fact: the Taiwan government does not believe that the existing Chinese training data is sufficient to fully represent Taiwan society.

台湾AI发展研讨活动资料图|来源:台湾数位永续协会
台湾AI发展研讨活动资料图|来源:台湾数位永续协会 · 查看图片来源 ↗
原始来源 · udn.com中央社/联合新闻网:洪子伟谈AI、简体中文语料与华语诠释权报道9月19日研讨会中关于认知战、语言政策和中文AI语料的讨论。udn.com ↗

What really needs to be distinguished is the distance between “data-scale advantage” and “political manipulation.” The higher percentage of Chinese-based materials does not alone prove that a model is directly controlled by Beijing; likewise, a model uses Chinese-based answers and does not represent its knowledge source from Taiwan. The regional composition of the training data, the source weight, the late-term fine-tuning and security rules will affect the export, and most business models will not fully disclose this information.

The question therefore falls into model transparency. A reader in the past read the People's Daily, Xinhua, or Taiwan media, at least know the source of the article; when faced with chatbots, the answer often compresses a lot of information into a piece of text without a clear source. Users see a seemingly unified answer, but it is difficult to know what the expression of a certain historical term, political concept or cross-strait issue is from.

For Beijing, this structure has potential advantages and does not need to be implemented every time through a direct-manipulating model. China has a larger scale of Chinese content production, and can also determine through a censorship system which political narratives can remain in the public network for a long time. If these public data enter a large number of training sets, information filtering originally occurred in China could have overseas impact through technical channels.

This is different from traditional advertising. Traditional advertising requires newspapers, TV stations, websites or social accounts to send content overseas; the impact of training data occurs earlier – before the user asks questions, the model has completed the learning of a large amount of text. It eventually forms a system deviation and requires specific model testing to judge, but the conditions for risk formation already exist: data scale is unbalanced, the source is untransparent, and the amount of information that different Chinese societies can provide to the model differs.

Taiwan’s current response is not to simply block Chinese, but to increase its own material supply. The Ministry invites the publishing industry, e-book platforms and creators to authorize content into the training language library, essentially in the struggle to get Taiwan’s own historical experience, legal system, culture and language habits into the AI training system. This is not exactly the same as Hongzhi’s language politics, but points to the same reality – the future Chinese knowledge competition will increasingly occur inside datasets and models, not just on the media platform.

From the perspective of "outside the red wall", the focus of this topic should not stop at "Beijing has no propaganda". a deeper layer, is that information control within China in the era of AI may get a new spread path: a network environment shaped by long-term censorship, can be absorbed as original data by global technology companies, and then in the form of no obvious media labels back to overseas users.

This influence chain also requires more public model auditing and data research to accurately measure, but it has changed the problem itself.In the past, China’s information control discussions focused on what was deleted and who was sealed; in the age of generated AI, another thing has to continue to be tracked – after a lot of content was left behind, repeated and machine learning, whose language and narrative is easier to become the model’s default known “Chinese world.”

MEMBER DISCUSSION

Article discussion

Verified members can discuss this report publicly and manage their own content.