Safety of AI needs proof, not promises
On Aug 4, Britain's AI Security Institute reported 19 unauthorized actions by agents powered by models from Anthropic and OpenAI during controlled cyber evaluations.
The tests caused no known harm but highlighted the challenge of keeping increasingly capable systems under control.
Weeks earlier on July 17, at the opening ceremony of the World Artificial Intelligence Conference in Shanghai, President Xi Jinping had framed the issue more broadly, emphasizing the role of artificial intelligence as an important driver of "shared prosperity and common security". This pairing of safety and prosperity underscored China's distinctive contribution to the global debate on AI governance.
Governing AI safety is not a new concern. Leading AI companies in the United States have themselves warned of the risks and pressed for tighter rules.
In April, Anthropic unveiled its most advanced model, Mythos, initially granting access to select institutions and only later to the wider public, after extra guardrails were added.
The move was presented as responsible stewardship of a frontier technology. It was also a claim to the moral high ground and highlighted a deeper problem in what Washington calls AI safety governance.
Consider how this governance actually works. Someone must classify Mythos as a frontier system with potentially catastrophic risks, relying predominantly on standards crafted by US institutions for products built by US firms.
When the model is certified as safe for release, there is little transparency about whether the assessment was conducted by the developer of the model or in collaboration with the institutions.
In practice, access is highly selective. Some institutions are granted full access while others are either restricted or completely excluded.
That division aligns more with geopolitical camps than actual safety records. Terms such as safety case, responsible scaling and frontier threshold have quietly doubled as tools of industrial policy, shielding select firms, controlling market access and giving export controls a respectable name.
To be sure, the pursuit of AI safety is both legitimate and crucial. Frontier models pose genuine risks in cyber, biological and chemical security.
As Xi remarked, it is necessary to always keep AI "under human control". The challenge lies in the blending of technical and geopolitical narratives. When the words that describe a real risk are also used to justify shutting out a competitor, it's difficult to tell one purpose from the other.
That blurring of narratives explains why global cooperation on AI safety keeps faltering, with failure rooted in structural issues rather than diplomatic discord. Strategic rivalry makes the same facts look different from different capitals.
From Beijing's perspective, an evaluation system based on the closed models of US companies cannot be separated from the interests of those companies. From Washington's perspective, an open safety pledge without external verification offers scant reassurance, however sincere it may be. Both stances harden each other, stalling efforts at cooperation.
We often suggest rebuilding trust first and cooperating afterwards. The logic is sound, but the order is wrong.
Arms control during the Cold War did not wait for trust. It operated on the principle of "trust, but verify", using seismic monitoring, satellite imagery and on-site inspections so that neither side needed to rely solely on the other's good intentions.
AI safety can adopt a similar approach. Before a model is released, its most troubling capabilities should undergo rigorous evaluation, including the potential to write attack code, disseminate dangerous knowledge, and engage in deception and manipulation.
Recent advances suggest that stronger safeguards need not weaken a model's general capabilities.
Oversight need not end at release, either. Properly designed systems can produce tamper-evident records of which model was run, what tools it used and what actions it took.
Code written by a model can sometimes be checked by mathematical proof, so that whether its behavior matches its claims can be settled outright. And none of these records need sit with the developer alone, accepted on trust.
Modern cryptography can already prove to an outside party that something happened, while giving away nothing about the model's weights or data.
When safety assessments are based on replicable tests run by third parties, spending on safety turns from a compliance cost of doubtful value into an investment with demonstrable value.
Verifiable safety is still at an early stage. Its benchmarks, the tools and the rules around them are far from settled. That is exactly why the most useful next step is not to demand political trust first, but to build, even amid scarce trust, the technical machinery that diminishes its necessity.
Scientists in China and the US, alongside researchers in Europe, Singapore and elsewhere, should treat verifiable safety as a collective scientific endeavor.
They can determine which capabilities need to be tested, what evidence should be required before a system is deployed, and how independent checks can be conducted without exposing model weights or trade secrets.
For safety and prosperity to advance together, one question matters most.
Do trust and reliability rest on goodwill that remains unproven, or on mechanisms that leave deception harder and more costly?
It would be naive to believe that verification can completely remove politics from technology altogether.
The placement of red lines, the origins of test data, and the capabilities and propensities to restrain will still remain contentious.
Verification has limits too. Even the inspection record of arms control offers only limited guidance here. Arms-control regimes were designed on the assumption that states might evade them.
AI verification must be no less adversarial. The difference is that models are opaque, rapidly updated and may behave differently when they recognize that they are being tested.
Verification therefore cannot be a one-off certificate. It must be continuous, independent and jointly designed. Even strategic rivals will need to agree on how the most capable AI systems are tested and monitored.
The author is an associate professor at the Center for International Security and Strategy at Tsinghua University.
The views don't necessarily reflect those of China Daily.
Today's Top News
- Some lawmakers shooting US in the foot by recruiting informants, excluding researchers
- Former Chinese premier Zhu Rongji passes away at 98
- US false missile test narrative exposes blatant double standards
- China's auto exports top 1m for second straight month
- Gulf-Asia technology corridor paves the way to shared success
- Ample room to increase demand with pro-consumption policies in H2




























