At 5:21 p.m. Eastern on Friday, June 12, Anthropic received a United States government directive. By nightfall, two of the most capable artificial-intelligence models in the world had gone dark after a Commerce Department export-control directive prompted Anthropic to take them offline worldwide.
The directive required Anthropic to prevent access to Claude Fable 5 and Claude Mythos 5 by any foreign national, whether that person was inside or outside the United States. It extended even to Anthropic’s own foreign-national employees. Because the company had no reliable way to verify every user’s nationality in real time, it suspended both models for everyone.
The controls remained in place until June 30. On July 1, Anthropic restored Fable globally. Mythos took a narrower path: after government approval on June 26, access returned to a set of United States organizations while Anthropic continued coordinating with the government over broader domestic and international access.
That sounds like a short outage with a happy ending. It was neither.
I have spent years building systems where a dependency going unavailable is not an abstraction. Once software is woven into a product or an operating process, access to it becomes part of the architecture. A model that can be withdrawn immediately under a company-specific determination and factual rationale no customer could have read is no longer only a technical dependency. It is a bet on unpublished policy.
The United States needs to govern frontier A.I. It also needs companies, researchers, investors, and customers to know what the rules are before the government enforces them. June showed how far apart those two needs remain.
The security concern was real
The government’s concern was not invented. Fable and Mythos shared the same underlying model, but Fable was released with strong safeguards for general use while Mythos, with fewer safeguards, went only to a small group of defensive-cybersecurity partners. Anthropic says Mythos can find and exploit software vulnerabilities more effectively than any other model and all but the most skilled human security experts. That is a serious capability to govern.
The specific report behind the June directive, however, concerned a bypass of Fable’s safeguards. According to Anthropic’s account after the controls were lifted, Amazon researchers prompted Fable to identify several software vulnerabilities and, in one case, produce code demonstrating how a vulnerability could be exploited.
That deserves serious attention. A model that materially lowers the cost of finding and exploiting security flaws can affect national security even when its ordinary users have harmless intentions.
But the details also matter. Anthropic said less capable models could identify the same vulnerabilities and that every other model it tested could reproduce the single exploit demonstration. It characterized the task as routine defensive-security work that its safeguards had blocked out of caution, not the exposure of a distinct Mythos capability. Anthropic is an interested party, so its interpretation should not be the final word. It should be tested against an independent technical standard.
That is precisely what was missing.
The initial directive arrived without a published severity test, a disclosed threshold, or a graduated response. The models went from released to globally unavailable in a few hours. This came despite thousands of hours of pre-release red-teaming by Anthropic, the United States government, the United Kingdom’s A.I. Security Institute, and outside organizations.
Anthropic subsequently trained a more targeted classifier, which the Commerce Department’s Center for A.I. Standards and Innovation tested along with the earlier safeguards. The company also proposed a common framework for judging jailbreaks and committed to expanding early government access, dedicating staff to government evaluations, and providing resources for joint research. Those are useful responses. They arrived after the shutdown because no shared process had told either side what response a finding of this kind should trigger.
A reproducible security problem can justify emergency action. It does not justify governing indefinitely by surprise.
The earlier blacklist changes how June must be read
The June export directive did not occur in isolation.
For months, Anthropic and the Pentagon had been negotiating over military access to Claude. The Pentagon wanted contractual permission for “all lawful uses” and argued that a private vendor should not have veto power over military operations. Anthropic supported government use of its models but insisted on two exclusions: mass domestic surveillance of Americans and fully autonomous weapons. The company’s position was not that it should choose military objectives. It was that current models were not reliable enough to select and engage targets without human control, and that A.I.-enabled mass surveillance would threaten fundamental rights.
Those limits were contractual, not a technical kill switch. Once Claude Gov was deployed, Anthropic could not enforce them through the model and had no direct visibility into how the government used it.
On February 27, after Anthropic refused to remove those limits, President Trump ordered every federal agency to cease using its technology. Defense Secretary Pete Hegseth then directed the Pentagon to designate Anthropic a supply-chain risk and announced that no military contractor, supplier, or partner could conduct any commercial activity with it. The public language was not restrained procurement language. The president called Anthropic a “radical left, woke company,” while Hegseth accused it of arrogance and betrayal.
The authority invoked for the designation was designed for a different problem. Section 3252 of Title 10 defines a supply-chain risk around the danger that an adversary will sabotage or covertly subvert a national-security system. It lets the Pentagon exclude a source from sensitive procurements when that source creates such a risk. It does not turn a public disagreement over contract terms into sabotage.
Anthropic challenged the actions in federal court. On March 26, Judge Rita Lin issued a preliminary injunction. On August 27, after reviewing the government’s administrative record, she granted Anthropic summary judgment on its First Amendment, due-process, and core administrative-law claims.
The summary-judgment opinion is more damning than the preliminary one. The record supporting the designation consisted principally of a four-page memorandum written after two of the three challenged actions. The government retreated from its claim that Anthropic retained backdoor access to models once they were deployed in national-security systems. It conceded that Anthropic’s models were no riskier in that respect than comparable black-box models. What remained was the department’s claim that it could not trust Anthropic because the company had criticized it publicly.
Judge Lin concluded that the designation violated the statute and that the broader measures were “illegal and baseless.” She also emphasized what the ruling did not do: the Pentagon remains free to select a different vendor. Procurement discretion was never the problem. Using national-security authority to punish a company for public criticism was.
That ruling resolved the district-court challenge to the Pentagon measures under Section 3252. A parallel proceeding in the D.C. Circuit challenges the department’s procurement determination under Section 4713 of Title 41. The appeals court denied interim relief in April without deciding the merits, later consolidated a second petition, and had not resolved the merits as of August 31.
The August ruling does not establish that the separate June export directive was retaliation. The two actions came from different legal processes, addressed different claimed risks, and must be judged on their own records. There is no public evidence directly tying the June suspension to retaliation.
But the sequence matters. Eleven weeks before an export-control directive prompted Anthropic to suspend two frontier models worldwide, a federal judge had already found serious evidence that the administration used a national-security rationale pretextually against the same American developer. Anyone asked to trust an undisclosed rationale in June was entitled to remember what the court had found in March. The August merits decision makes that concern stronger, not weaker.
This is not only about Anthropic
It would be easy to reduce the story to whether one likes Anthropic’s politics or trusts its safety claims. That misses the larger risk.
OpenAI’s release of GPT-5.6 Sol followed a different path, but it reflected the same unsettled boundary between government oversight and product deployment. OpenAI said it had previewed its plans and the models’ capabilities to the government and, at the government’s request, initially limited access to trusted partners whose participation was disclosed to Washington. OpenAI also said that arrangement should not become the long-term default.
That limit was temporary. On July 9, OpenAI made the GPT-5.6 family, including Sol, generally available across ChatGPT, Codex, and its A.P.I.
That was not a forced shutdown, and it should not be described as one. Pre-release evaluation and a staged launch can be responsible ways to manage an unusually capable model. The important point is that release conditions for frontier systems are already being shaped through direct, case-by-case government involvement while the durable public framework is still being built.
Case-by-case coordination is unavoidable at the edge of a fast-moving technology. As a governing model, it is fragile. It rewards access, leaves outsiders guessing, and can make materially similar decisions turn on which company is in the room or which official signs the letter.
Leadership is a system, not a benchmark
When people ask whether the United States will remain the A.I. leader, the conversation often collapses into a horse race: Which company owns the best model this month? Which country tops the latest benchmark?
That is not what durable leadership means. Leadership is the ability to create, finance, deploy, improve, and commercialize successive generations of systems. It rests on at least four inputs: talent, capital, compute, and a predictable legal environment.
America is extraordinarily strong in the first three. Stanford’s 2026 A.I. Index counts 5,427 data-center facilities in the United States, more than ten times the count in any other country, although facility count does not measure size, computing capacity, or utilization. American private A.I. investment remains far ahead of every competitor. On Stanford’s open-source measure—cumulative GitHub stars for qualifying projects—U.S.-based projects lead, though the measure excludes activity on Chinese platforms such as Gitee and GitCode.
The same report contains a warning. Its 12-month rolling net flow of selected top A.I. authors and inventors into the United States—arrivals minus departures—has fallen 89 percent since 2017, including an 80 percent decline in the latest year measured. The measure remained positive at 26.0 in 2025, and the United States still has more A.I. talent than any other country. This is a collapse in the recruitment margin, not 89 percent fewer people arriving and not an empty domestic workforce. The decline began long before the events described here. It nevertheless shows that America’s ability to attract people is not automatic.
Neither is its ability to attract deployment.
Capital can price a known regulatory burden. Engineers can design around a published restriction. Customers can diversify around a disclosed availability risk. What none of them can reliably model is a rule discovered only when a product is switched off, a company is publicly blacklisted, or an official later supplies the legal theory.
The immediate cost is disruption. The larger cost is that teams quietly decide the United States is a superb place to train a model but an unpredictable place to release one.
What durable frontier-model governance should look like
The answer is not to leave model companies alone. Frontier developers should not get to grade their own security work, define every acceptable use, and declare every launch safe. Government has both a legitimate role and capabilities private labs do not have.
That role needs structure.
First, restrictions should attach to capabilities and evidence, not to a company name. If a reproducible jailbreak crosses a national-security threshold, the same threshold should apply to every model that can perform the relevant task. Anthropic’s claim that older and competing models reproduced the underlying vulnerability findings and exploit demonstration is exactly the kind of assertion an independent evaluator should confirm or reject.
Second, emergency orders should contain their own expiration and review. Government sometimes must act before it can publish everything it knows. An immediate restriction can still identify its legal authority, the category of risk, the people or systems covered, the official responsible for review, and the date on which the order expires unless renewed with additional findings.
Third, the response should be proportional. A technical finding might justify a classifier update, monitored access, a trusted-researcher program, a delayed release, or a temporary limit on one capability. Global suspension should be available when the risk warrants it, but it should be the end of the ladder rather than the only rung.
Fourth, compliance must be technically possible. An order based on a foreign-person distinction cannot simply assume that a public A.P.I. provider reliably collects and can operationalize that status in real time. If the government requires an identity distinction that the service does not collect, it must define a workable mechanism or choose a narrower control. Otherwise an apparently targeted rule predictably becomes a worldwide ban.
Fifth, decisions need due process. A company should receive the evidence it can lawfully see, an opportunity to respond, and a path to prompt independent review. The government must remain free to stop buying a product. A procurement decision and a punitive declaration that reaches unrelated private business are not the same thing.
Finally, the public deserves an after-action standard. Classified details can remain classified. The governing rule cannot. Once an emergency passes, government should explain what category of event triggered action, which mitigation ended it, and how a future developer can avoid repeating the same failure.
None of this is a plea for slow government. It is a design for fast government whose decisions survive the emergency.
The reversal is not the resolution
The June controls were lifted. Fable returned globally, while Mythos returned first to a limited set of approved United States organizations. Anthropic improved its safeguards and expanded government testing. The federal district court has now ruled against the Pentagon’s blacklist on a full record, while the parallel D.C. Circuit proceeding remains unresolved.
A skeptic can call that the system working. In part, it is. The government’s position changed after technical review, safeguard changes, and negotiated commitments, and judicial review constrained unlawful action in the Pentagon case.
But speed of reversal does not cure arbitrariness of imposition. A rule learned by having a product switched off is not a rule. It is a decision. A company vindicated in the Pentagon case six months after being branded a national-security threat still spent those months under the threat.
The country best positioned to nurture the next generation of A.I. is still this one. That is why these episodes should alarm anyone who wants it to remain so. American leadership is likelier to be squandered at home than taken in a single dramatic victory by a foreign challenger. It can be lost one directive, one after-the-fact restriction, and one unpublished judgment at a time.
Can America remain the A.I. leader? Yes. Whether it will is now a decision the government makes every time it reaches for the pen.
