Can AI Speak Diplomacy? Exploring AI’s Grasp of Geopolitics and Limits in Sensitive Translation

Monalisa Hazarika

Strategic Communications & Partnership Officer, SCRAP Weapons

Across fields, a major debate is unfolding over how Artificial Intelligence (AI) is integrated into our workstreams, balancing the promise of enhanced productivity with the need for ethical safeguards. The academic and policy spheres are no exception. A 2024 systematic review identified six key areas where AI is already reshaping scholarly work: from idea generation and research design to content structuring and editing, dramatically speeding up the research process. Yet despite these benefits, concerns around academic integrity, transparency in the use of AI, and the risk of AI hallucinations producing inaccurate or misleading content persist.

On the flip side, AI-based translation software models can be extremely useful for language translations for the dissemination and outreach of tailored content to different audiences and platforms. They can help in reaching a global audience, including non-specialists, making technical jargon more relatable. They can be an essential tool to bridge linguistic barriers and foster collaborations, especially between the English and non-English speaking audiences, given the fact that currently only 17% of the world’s population speaks English. But again, how accurate are they?

While translating an op-ed on US-China-Russia strategic strike capabilities for a Japanese audience, I came across a rather interesting observation. I tested 6 different online platforms to translate the following text into Japanese–

“A strategic concern is whether Beijing and Moscow fear US military capabilities to the point where they justify a new arms race. US national intelligence has repeatedly highlighted Beijing’s fear of a US first strike as a motive for the Chinese nuclear arms build-up. Washington’s strategic superiority in non-nuclear weapons lies in what we call the 3Ds: detection, defeat, and defence. The US has superior detection through its satellites and strategic drones, including the RQ-4 Global Hawk and the new RQ-180, with operational ranges of 22,000km, likely capable of finding camouflaged and dispersed mobile nuclear launch systems.”

The most popular and rather convenient Google translate suggested the following: “戦略的な懸念事項は、北京とモスクワが米国の軍事力を恐れ、新たな軍拡競争を正当化するほどになるかどうかである。米国国家情報局は、中国の核兵器増強の動機として、北京が米国の先制攻撃を恐れていることを繰り返し強調してきた。ワシントンの非核兵器における戦略的優位性は、いわゆる「3D」、すなわち探知(detection)、敗北(defeat)、防御(defense)にある。米国は、衛星や戦略無人機(RQ-4グローバルホークや新型RQ-180など)を通じて優れた探知能力を有しており、その作戦範囲は2万2000キロメートルに及び、カモフラージュされ分散配置された移動式核発射システムを発見できる可能性が高い。”

Gemini offered the following: “戦略的な懸念として、北京とモスクワが米国の軍事力を恐れるあまり、新たな軍拡競争を正当化しているのではないかという点があります。米国国家情報機関は、中国が核兵器増強を進める動機として、米国の先制攻撃に対する中国の懸念を繰り返し指摘してきました。非核兵器における米国の戦略的優位性は、我々が3つのDと呼ぶ、探知(Detection)、撃破(Defeat)、防衛(Defence)にあります。米国は、運用航続距離が22,000kmに及ぶRQ-4グローバルホークや新型RQ-180を含む衛星および戦略ドローンを通じて優れた探知能力を有しており、これは偽装され分散された移動式核発射システムを発見できる可能性が高いとされています。」”

ChatGPT came up with “戦略上の懸念のひとつは、北京とモスクワが米国の軍事能力を恐れ、そのことが新たな軍拡競争を正当化する段階に達しているかどうかである。米国の国家情報機関は、中国の核戦力増強の動機として、北京が“米国の先制攻撃”を恐れていることを繰り返し指摘してきた。ワシントンの非核兵器分野における戦略的優位性は、いわゆる3D──探知(Detection)、撃破(Defeat)、防御(Defence)──にある。米国は、RQ-4グローバルホークや新型RQ-180を含む衛星および戦略無人機によって優れた探知能力を持ち、作戦行動半径は2万2,000キロに及び、迷彩・分散された移動式核発射システムを発見できる可能性が高い。”

DeepL suggested “戦略上の懸念は、北京とモスクワが米国の軍事能力を恐れるあまり、新たな軍拡競争を正当化するほどかどうかである。米国国家情報機関は繰り返し、中国の核軍備増強の動機として、北京が米国の先制攻撃を恐れている点を指摘してきた。非核兵器分野における米国の戦略的優位性は、我々が「3D」と呼ぶ能力——探知(Detection)、撃破(Defeat)、防衛(Defence)——に由来する。米国は衛星と戦略ドローン(RQ-4グローバルホークや新型RQ-180など)による優れた探知能力を有し、その作戦行動半径は22,000kmに及ぶ。これにより、偽装・分散配置された移動式核発射システムを発見できる可能性が高い。”

QuillBot’s translation was “戦略的な懸念は、北京とモスクワが米国の軍事能力を恐れて、新たな軍拡競争を正当化するほどになっているかどうかです。 米国の国家情報機関は、中国の核兵器増強の動機として、北京が米国の先制攻撃を恐れていることを繰り返し強調している。 ワシントンの非核兵器における戦略的優位性は、私たちが3Dと呼ぶものにあります。それは、探知(detection)、撃破(defeat)、防御(defence)です。 米国は、RQ-4グローバルホークや新型RQ-180など、22,000kmの作戦範囲を持つ衛星や戦略ドローンによる優れた探知能力を有しており、カモフラージュされた分散型の移動式核ミサイル発射システムを発見できる可能性が高い。”

Finally, Reverso came up with “戦略的な懸念は、北京とモスクワが米国の軍事力を新たな軍拡競争を正当化するほど恐れているかどうかだ。米国家情報機関は、中国の核兵器増強の動機として、米国の先制攻撃に対する中国政府の懸念を繰り返し強調してきた。非核兵器におけるワシントンの戦略的優位性は、我々が3Dと呼ぶもの、すなわち探知、敗北、防衛にある。米国は、RQ-4グローバルホークや新型RQ-180など、衛星や戦略ドローンによる優れた探知能力を備えており、運用範囲は22,000kmで、カモフラージュされ分散した移動式核発射システムを発見できる可能性が高い。”

A quick retranslation using Google Lens could create confirmation bias, so I asked a native speaker with subject-matter expertise [1] to review the nuances and assess the accuracy of these translations. Across the six versions, the level of politeness varied by tool, and certain nuances led to misinterpretations. For instance, Reverso and Google interpreted the word “defeat” as losing a battle, whereas other tools interpreted it as a knockdown. When a single word can hold multiple meanings depending on context and placement, AI tools, unlike human translators, often lack insight into the intended frame of reference. Similarly, ChatGPT introduced an unintended nuance by translating the first sentence as “one of the strategic concerns,” instead of the original “a strategic concern.” These subtle variations often stem from sentence structure or punctuation and can ultimately change how the message is perceived. Long sentences with multiple commas are also particularly prone to mistranslation without a strong command of the original language.

While some experts are of the opinion that there are no statistical differences between human and AI translations, beyond linguistic accuracy, translation must carry the cultural context, emotional subtext, and social norms and etiquette to maintain the integrity and intent of a message. AI-powered translation tools rely on natural language processing, machine learning, and neural networks to sift through massive multilingual datasets, delivering quick and affordable translations. But speed isn’t everything. In cultures where hierarchy and formality are integral, even a slight shift in politeness levels can come across as disrespectful—or worse, offensive. These tools still struggle with contextual intelligence and are currently unable to adjust the tone as per the audience and intent.

It’s fascinating to consider how all of this plays out in the world of scientific and policy writing, especially in arms control and non-proliferation, where the stakes are unusually high. Using AI in these fields isn’t as straightforward as it is in everyday communication, as these discussions happen in diplomatic and high-level settings, where every word carries weight and even small nuances can influence policy.

Yes, AI brings clear advantages: wider accessibility, faster processing, and the potential to make research more inclusive. But that opens up a series of intriguing questions: How do we ensure the outputs are accurate? Should these tools be specially trained, and if so, how? Could an AI model tailored for scientific or security-focused writing even be feasible? And then there’s the bigger picture: Could AI eventually replace diplomatic or professional translation training? How well can a machine pick up the sensitivities of international security language? What happens when it encounters technical jargon or policy-specific terms? Where does human judgment still matter and how much?

These are open questions, and perhaps only time will reveal the answers.

[1] Based on the author’s correspondence with Ms. Shizuka Kuramitsu, Research Analyst, Nuclear Policy Program at Carnegie Endowment for International Peace

Monalisa Hazarika

Strategic Communications & Partnership officer, SCRAP Weapons

Leave a Comment

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.