AI automates the judgment of security experts
The security expert’s judgment was the skill automation was never supposed to touch. A government assessment of Anthropic’s Mythos has shown it can be reproduced in a single automated pass.
The assessment that inverted the assumption
For two decades, the most stubborn shortage in cybersecurity was never funding or computing power. It was judgment: the rare instinct that lets one person look at a dense block of code and sense the flaw everyone else missed. Governments courted these people the way they once courted weapons designers, and every strategy paper offered the same comfort — machines could scan, but only a person could decide what mattered.
A government assessment of Mythos, the frontier model built by the American lab Anthropic, has now canceled that comfort MSN News. Presented with a target, the model did not flag leads or hand a shortlist to a human operator. It chained hundreds of vulnerabilities together and wrote working exploit code on its first attempt, with results holding up more than 83 percent of the time. [2]
A human penetration tester might spend a month mapping an attack chain, only to watch the target patch the flaw before the work is finished. The model compresses the whole arc — observation, hypothesis, judgment, execution — into one automated run. The part of the job that was supposed to be irreducibly human was the first part to be automated.
The people who once performed that role are not gone yet. Their function has shifted in a way no one prepared for. They no longer supply the expertise; they witness it being reproduced. The expert has become a spectator at their own profession.
A craft built on a human call
The skill being replaced has a short history and a steep rise. Before the millennium, breaking software was a hobbyist’s pursuit, shared in zines and mailing lists, driven by curiosity rather than salary. The professional era arrived when governments and markets realized that a working exploit could be worth more than a battalion of conventional engineers.
That profession rested on one human capacity: the ability to decide. Researchers who knew their craft did not run the most tools. They understood which corner of a system deserved suspicion, which flaw was alive and which was a dead end, which chain would hold when a defender pushed back. This taste could not be learned from a manual. It accumulated over years of quiet failure.
The machines that came before Mythos respected that boundary. They could scan, probe, and sort, but the choices stayed with the person. A tool that could not judge was just another instrument. The dividing line between instrument and operator was the last guarantee of professional relevance, and it was drawn by people who assumed their own skill could not be turned into an output.
Then the assessment crossed the line without ceremony. The model did not assist a judgment; it performed one, built on it, and finished the work. The distinction between instrument and operator collapsed, and the field’s most valuable humans woke up in the instrument category.
When the rulebook lags the machine
Regulation was never designed for this situation. The Wassenaar Arrangement, the international export-control regime, recognized intrusion software as a controlled commodity in the previous decade, and the decision drew protests from security researchers who feared the rules would criminalize their craft. [3] The logic of that era was simple: the skill lived in humans, so controlling the tool meant controlling the people.
The logic has quietly vanished. By June, Washington judged Mythos risky enough to cut off foreign access entirely — a decision carried out at the level of code, not credentials. There is no human expert to vet, no visa to deny, no clearance to revoke. The capability travels as a model, and the only way to control it is to control the network it moves through.
The officials making these calls sit in an odd position. They are drafting policy about a system whose judgment they cannot independently evaluate, because the yardstick they would once have used — a human expert’s opinion — is precisely what has been made obsolete. A regulator cannot ask a specialist to assess the machine that replaced the specialist.
Trade enforcement is stumbling on the same fault line. Treasury Secretary Scott Bessent floated sanctions last month against developers found to have stolen intellectual property, after Beijing’s Moonshot AI released a system said to rival the best American laboratories. The concept of theft, however, has mutated. In the old economy, stealing meant taking code or documents. In the new one, a model absorbs another model’s judgment by training on its outputs, and no law has caught up with that transfer.
China’s Commerce Ministry warned that it would take “all necessary measures” in reply. The exchange has the shape of a familiar trade dispute, but the substance is unprecedented. What is being argued over is not a thing but a competence — an ability that used to live in scarce human minds and now lives in a set of weights that can be copied, locked, or denied.
The democratizer caught in its own doctrine
The irony lands hardest on Beijing. For a year, China has told the world that artificial intelligence should be open, cheap, and shared, flooding the global market with downloadable models and presenting itself as the technology’s democratizer. The posture is now colliding with a private fear that its officials can no longer hide.
The system China fears most is the one it cannot touch. Mythos is closed, American, and legally unavailable to Chinese users; officials reportedly view it as a potential offensive cyber weapon. The nation that promised shared intelligence is now anxious about a rival intelligence it cannot download, built by the company most determined to keep it that way.

Anthropic has made the choice explicit. It has become the most China-hawkish of the American labs, pressing for high-end chips to be kept out of Chinese hands and publicly accusing Chinese developers of training on its outputs. Beijing’s suspicion, in other words, is not imagined. It is a response to a company that treats its judgment machine as a strategic asset worth guarding.
Beijing’s reply is being drafted in the old grammar of statecraft. Officials are preparing options that include sanctions and the addition of American AI firms to a restricted-entity list, should Washington move first. Because Anthropic does not operate in China, any penalty would carry little commercial sting — a signal, not a blow. The playbook is familiar: in June, China answered the Pentagon’s blacklisting of Alibaba and BYD with export controls on two American rare-earth firms, part of a pattern of hitting back where it hurts.
Yet nothing in that arsenal can touch a judgment machine. Sanctions were invented for a world where capabilities lived in factories, ships, and supply chains. A model that exists as code and training data has no factory to hit. The tools of retaliation still assume a human economy, and they are aimed at a target that no longer runs on human skill.
China’s answer to the capability gap is also a confession. A Chinese security firm, 360 Security, claims it has developed an AI tool to rival Mythos MSN News. Whether the claim survives scrutiny matters less than what it concedes: the race is no longer about recruiting the best researchers or training the best teams. It is about who can automate the researcher first. Both superpowers now treat that contest the way they once treated missile parity.
Negotiating in the machine’s shadow
The diplomatic calendar makes the discomfort visible. Xi Jinping is due in the United States on 24 September, and Beijing wants a warm mood before he lands. Officials are seeking a “win-win” tone at AI talks expected ahead of the meeting, even as suspicion runs in both directions.
The negotiators are in a peculiar bind. They will sit across a table and discuss rules for systems whose most significant acts have already occurred without any of them in the room. The model that alarmed Beijing did its work in an assessment environment, not a war game; the humans who must now decide what it means were not part of the judgment it performed.
Washington carries its own contradiction. President Trump says he is weighing tighter AI controls while fearing they could hobble American firms. “We have to be careful in both ways,” he said last week. “We don’t want to restrict them where all of a sudden, we come in second to China.”
The same government has banned federal agencies from using Anthropic’s tools — a decision the company is contesting in court. The state that fears the model refuses to employ the very system that outclassed its own experts. It watches from the outside, like the researchers it once relied on.
Every actor in this standoff is reacting to a competence they cannot reproduce. The Chinese experts who once would have been tasked with matching Mythos are no longer the bottleneck; the model itself is. The American experts who once would have certified its danger have been replaced by the thing they were asked to judge. Diplomacy, too, is now a matter of responding to an output rather than consulting a specialist.
The last open variable
What remains undecided is not the technology. The assessment has already shown what the machine can do, and 83 percent is a number that does not need a human to interpret it, only a policy to answer it. The open variable is institutional.
Success and failure now turn on whether governments can rebuild the machinery of oversight around a world where the scarce human skill is not scarce anymore. The old system assumed expertise was the bottleneck: train the experts, protect them, pay them, and the nation would be safe. The assumption has been retired by the same force that retired the experts.
An alternative future is visible in the response being prepared in Beijing. The sanctions, the restricted-entity lists, the rare-earth controls, the warnings of “necessary measures” — all of it is theater with real consequences, but it is theater nonetheless, because the capability in question cannot be sanctioned out of existence. It can only be matched, locked away, or out-built.
The last variable, then, is what the humans in charge decide to become. They can keep playing the role of experts in a system that no longer needs their expertise, drafting rules that describe the world as it was. Or they can accept the inversion: judgment has become a manufactured output, the person is no longer the scarce factor, and the only question worth asking is who controls the factory.
The question will be asked whether the summit produces warmth or frost. It will be asked in the court where Anthropic fights the federal ban. It will be asked every time a rival model turns up claiming the same ground. The humans who once made the calls are not out of the game yet. They simply no longer determine its pace.
Sources
1. Anthropic
2. MSN News
