When the Models Go Dark
Sections
In June 2026, two frontier AI models went dark — for every customer, everywhere, at once. There was no breach and no outage. The US Commerce Department had placed export controls on Anthropic’s most capable models, and rather than risk non-compliance, Anthropic disabled them for everyone — by the Financial Times’ account, the company had ninety minutes to comply. Access came back within weeks. The interruption was not a hack. It was paperwork — frontier AI, it turns out, can be export-controlled like any other strategic technology. Nor was it one lab’s story: the same season, OpenAI’s newest flagship shipped first as a limited preview to “trusted partners.” State-gated frontier access is no longer a hypothetical. It is a precedent.
Most countries — and almost every company — will never train frontier AI. They will still run critical systems on it. So the useful question is not who owns artificial intelligence — AI Is Bigger Than Five, the first essay in this series, mapped that, and the answer for almost everyone is someone else. The useful question is what you do about it.
Sovereignty is a recovery time
There is a tempting answer, and it is usually the wrong one. For two years Jensen Huang has urged countries toward “sovereign AI”: your own model, trained on your own data, running on your own infrastructure. He has every commercial reason to make that case — and it is still correct: the dependency he describes is real. A country that wires its hospitals, courts, and banks to a foreign company’s API has quietly handed a piece of its statehood to a corporation on another continent, one that can raise the price, change the rules, or be ordered by its own government to pull the plug.
But the maximal reading of that pitch — own the fabrication plant, the frontier lab, the hyperscale datacenter, the training run — is a fantasy priced in decades and tens of billions, and it ends with a model measurably worse than one you could have downloaded without a license fee the week you started. Full-stack sovereignty is not a strategy. It is a way to lose slowly, at maximum expense, while calling the expenditure independence.
Call the alternative minimum viable sovereignty — a phrase borrowed from enterprise cloud strategy, promoted here to statecraft — and define it not as ownership but as durability: the standing ability to keep critical services running when a foreign supplier changes its terms or disappears. Minimum viable sovereignty has a unit, and the unit is time. How many hours can the hospitals, the payments, and the courts operate without the primary model provider? How many days until they run again on systems the state can actually govern? Not the number of layers owned; the maximum interruption an outside actor can impose. Enterprises face an analogous continuity test, even if they do not possess sovereign authority: a bank or a hospital network wired to a foreign model runs the same cutoff arithmetic as a ministry — what it defends is continuity and bargaining power, not statehood.
The premise — that owning the full stack may simply not be feasible, and that the real work is managing the dependencies — has become the center of gravity in this year’s policy literature — Brookings calls the practical alternative “managed interdependence,” and Chatham House frames the question as not whether to depend, but on whom, and on what terms. What this essay adds is narrower and sharper: a single test (what still runs the day after a cutoff?), an ownership verdict for each layer of the stack, and six numbers a government or a company can publish and be judged by.
Two ways to fail
The first failure is the national frontier model built for pride. A ministry or a champion announces it will train a foundation model to rival the empires; the announcement is the high point, and everything after is a bill. One instructive case: even Kai-Fu Lee’s well-funded 01.AI stepped back from frontier pretraining, because, in his words, “the biggest nightmare for Sam Altman is that his competitor is free.” The qualifier matters: a national model is rational when it supplies a public good the market will not — an under-served language, a reproducible research base, a defense requirement, a tested fallback. What is never rational is a pretraining run whose real objective is a leaderboard. A similar arithmetic weighs on the sovereign leading-edge fab: one Dutch company makes the EUV machines, a handful of foundries can use them, and even Japan’s Rapidus — with an existing toolchain industry, full state backing, and tens of billions behind it — is, so far, an attempt rather than a proof.
The second failure is the opposite one, and more common because it looks like pragmatism: “we’ll just use the best API.” An API you do not control is a dependency dressed up as a decision. It can be revoked — June proved that. It can be repriced. And it arrives pre-loaded with someone else’s values: a model trained elsewhere answers your citizens’ questions about their own history and politics in a voice some other institution tuned. Renting frontier capability is fine — required, even. Renting it with no tested fallback for the functions that cannot stop is not a procurement choice. It is sovereignty theater. Nor is the choice binary: sovereign cloud, on-prem vendor deployments, model escrow, and contractual step-in rights can all reduce exposure — but they count only if the cutoff drill survives the supplier’s absence.
What open weights change — and what they don’t
Between those failures runs the road that has opened in the last two years: the open-weight ecosystem. When models near the frontier — on specific, economically important tasks — ship under permissive licenses, the acquisition cost of the hardest layer — a capable base model — collapses toward zero. Deployment still costs: hardware, power, security, evaluation, people. But a country no longer has to build the frontier in order to stand on it.
As of late August 2026 the supply is deep and the leadership keeps changing hands — the strongest open model is currently Chinese, the most-used open models on the public routers are Chinese, and the Western open camp has reorganized around new names while Meta edges back toward the openness it abandoned. The details expire monthly, which is why they live in a separate, dated companion: the Open-Weight Field Guide keeps the score. The doctrine only needs three durable facts. Capable, downloadable, commercially usable models exist in abundance. They differ sharply in license and in behavior — so verify the exact checkpoint, not the family’s reputation, and test it in your language under your own serving stack. And a downloaded model is harder to switch off from abroad than any API.
Three limits keep this honest. Open-weight is not open source: the weights come, but the data and recipe mostly do not, and trained-in behavior — including another state’s silences — is sticky and never fully visible. Weights are inert: open weights hand you the engine and none of the car, and nothing sovereign happens until you build the data access, adaptation, evaluation, governance, and distribution around it. And you cannot fully audit what you download: research on trained-in backdoors shows implants that survive safety tuning and defeat detection, so the defense is hygiene and plurality — pinned hashes, red-teamed triggers, and more than one model lineage under critical systems.
The sovereignty stack
Assign every layer a verdict, running from own at one end to cede at the other.
| Layer | Verdict | Why |
|---|---|---|
| Minimum critical inference capacity | Own | A domestic reserve that carries essential systems through a cutoff — on hardware already in the building, with frozen toolchains and more than one runtime, so it can keep running the morning after a supplier’s terms change. |
| Adaptation & evaluation | Own | The ability to bend an adopted model to your world, and to judge it there. |
| Governance & control plane | Own | Authority limits, audit, rollback, appeal — and rules citizens can see and challenge. |
| People & institutional capability | Own | A retained cadre of operators, evaluators, security engineers, and accountable officials who can actually run all of the above — compounding, not consulted and gone. |
| Distribution & application | Own | The surfaces where a citizen, a business, a bank actually touches intelligence. |
| Data & language | Steward | Not ownership — lawful, rights-governed access to the records and language no foreign supplier can readily reproduce. |
| Frontier model weights | Adopt, verify & archive | Stand on the best open weights; check the exact license; archive more than one lineage from more than one jurisdiction. |
| Energy, grid & datacenter siting | Secure at home | The foundational input most countries actually can control — and the hardest pillar in practice: many national plans are already colliding with grid constraints. |
| Frontier training compute | Rent, with exit | Diversify vendors and geographies; never hostage to a single patron. |
| Leading-edge semiconductor fabrication | Cede / opt out | The one race in which even major powers remain dependent on foreign chokepoints; do not spend your sovereignty budget at the bottom of someone else’s mountain. |
The inversion the whole table turns on: for the five great labs, the chokepoint is compute, which is why the company selling accelerators is worth trillions. Most of the world cannot own that chokepoint — but it can own the power that feeds a domestic reserve, and the layers above the model that no empire can reach without permission: the language, the institutional data, the regulated market, the trained people, the distribution. Sovereignty for the rest of the world does not live in the fab. It lives in the middle and at the top of the stack. None of it forbids pooling — allied reserves, shared evaluation suites, joint archives spread the cost — so long as the pool itself would pass the cutoff test.
Read the table in two tiers, or “minimum viable” becomes a new full-stack wish list. The minimum floor — what the phrase literally means — is five capabilities: an inference reserve, two archived model lineages, a local evaluation suite, a tested migration and graceful-degradation path, and the people who have run it. Everything else is strategic upside: worth building, not required for survival. And the reserve, in particular, is not the primary system. Rent the best frontier model for daily work; the reserve is a capability-degraded emergency system, and capital sunk into it beyond the insurance level is productivity forgone.
Two honest notes on the “steward” and “people” rows, because they are where doctrines flatter themselves. Data is a weaker moat than the slogans claim: general models grow more multilingual by the quarter, and raw national corpus matters less than the things around it — lawful task-linked access, the feedback loops that turn real institutional work into evaluations, and the legal authority to deploy. The durable advantage is institutional, not volumetric. And people are the binding constraint more often than GPUs: reserve hardware without retained operators is not reserve capacity, and a budget that buys compute but not careers builds every layer on sand.
The drill
Doctrine becomes real at the moment of migration, so picture the test — then run it.
A bank has wired a loan-review assistant into its daily workflow on a foreign frontier API. The sovereignty question is not which flag is on the model card. It is this: at nine on a Monday morning, the API’s credentials are revoked — deliberately, by the bank’s own security team, on a schedule the operating team does not control. The clock starts. The target: the assistant back in service within forty-eight hours, on an archived open-weight model running on hardware the bank owns, without a single customer record leaving the building, and with the employee’s workflow essentially unchanged.
Run honestly, the first drill fails somewhere unexpected. The weights load, but the tokenizer version differs; the permissions layer assumed the vendor’s identity system; the evaluation suite — the private, local-language tests that define “good enough for production” — turns out to cover a third of the workflow. That is the drill working. Ownership was not operability, and it is far better to learn it on a Monday you chose than one chosen for you in someone else’s capital. The second drill goes better. By the third, the bank has something no procurement contract can give it: a measured recovery time, a known quality delta, and a team that has done it before.
The same exercise scales down to a hospital’s clinical-coding assistant and up to a ministry’s benefits triage. It also, in miniature, already happened in the wild: during July’s intrusion at Hugging Face — an unprecedented model-driven breach that began inside another lab’s cyber evaluation — the forensic reconstruction relied on a self-hosted open-weight model after commercial APIs refused parts of the workload. One incident proves no general law, and the account is the defenders’ own. But the shape is the drill’s, unplanned: part of the rented tooling failed under pressure, and the work that mattered ran on an open model on the victim’s own hardware.
And the fallback need not be another frontier model at all. Real resilience is graceful degradation: a smaller model for the routine cases, rule-based software for the deterministic ones, human review for the judgment calls, a manual procedure that holds for a week. The doctrine’s question is never “which AI replaces the AI.” It is: what level of service survives, and for how long.
The owned layers, briefly
Evaluation is the quiet crown. Whoever writes the tests defines what “good” means, and a model graded only on foreign benchmarks is good at being foreign. An owned evaluation lives on private, rotated, task-grounded test sets — local-language legal retrieval, benefit eligibility, clinical coding — because public leaderboards saturate, leak, and get gamed. To cede evaluation is to let someone else decide, permanently, what counts as working.
Governance is where values become legitimate. Sovereignty is not that the model has no values; it is that its rules are publicly chosen, auditable, contestable, and bounded by rights — not national censorship rebranded as sovereignty. A state does not own its citizens’ records; it stewards them under consent, purpose limits, and redress, or the doctrine curdles into extraction.
Distribution is where it is finally felt. AI sovereignty is not experienced in a datacenter; it is experienced where a person touches intelligence, and domestic institutions hold licenses, trust, and channels no foreign vendor can quickly reproduce. But distribution is leverage, and leverage cuts both ways: replacing a foreign monopoly with a domestic one is sovereignty for the state, not the citizen. The test is pluralism and exit at home — sovereignty of a country need not mean sovereignty over its people.
And one reframe that changes the whole posture: durability is not the goal. It is the floor that makes ambition safe. A country confident that its critical systems survive a cutoff can afford to wire intelligence deeper into its factories, hospitals, classrooms, and firms than a country that is quietly afraid of its own dependencies. The reserve is not a bunker; it is the insurance policy that lets you build aggressively on rented capability. Defense buys the right to compound.
Six numbers
For every critical AI service, measure six numbers, and be judged by them:
- Reserve hours — how long essential inference runs on capacity you own.
- Supplier concentration — the share of critical workload on your largest single provider.
- Migration time — tested days (or hours) to move the service onto an archived model.
- Evaluation coverage — the fraction of critical tasks covered by private, local-language tests.
- Fallback quality — the share of those tasks the fallback actually completes above the minimum safe threshold. Coverage measures whether the tests exist; this measures whether the fallback passes them. A well-tested fallback that fails its tests is still a failure.
- Drilled operators — how many staff have completed a live migration drill within the last twelve months.
Publish the aggregates; keep the service-level detail for independent audit rather than the open internet. The numbers are accountability, not a target map for an adversary.
Then maintain the archive the numbers depend on. A permissively licensed checkpoint, archived with its tokenizer, runtime, and evaluation suite, cannot be recalled from a distance — the deepest difference between a download and an API — though sanctions or new law can still constrain how it is used: the copy endures; the legal environment can change. But an archive is a bridge, not an endowment: it lasts exactly as long as it keeps passing its drills. Requalify it against the critical workloads at least annually; refresh it with every significant release, in more than one lineage, from more than one jurisdiction. An archive that has never survived a drill is not a reserve. It is a hope stored on disk.
The supply is a decision
All of this depends on a supply of open weights that large actors maintain for their own reasons, and one summer showed how deliberately. In July, China placed open-source AI at the center of its international agenda, founding a Shanghai-based cooperation body with more than two dozen states and pitching openness as a shared good. Days later an American industry letter, then a hundred-and-twenty-company security alliance, made a parallel case — openness as sound industrial strategy, argued by the firms that build the infrastructure. And Anthropic offers the most serious version of a more cautious view — mandatory pre-release safety testing for every sufficiently capable model, open or closed — a position this doctrine can live with.
Policy can also move the other way, and over the same season both capitals publicly weighed measures that could touch open weights — reviews, inquiries, possible controls on model releases. Nothing was enacted; the deliberation itself is the point. Open weights are, among other things, an instrument of policy — and instruments can be adjusted. (The season’s full chronicle lives in the field guide.)
So the supply is a decision made elsewhere, and that is not comfort but urgency. Take the weights while they are offered — a downloaded copy is yours on the license’s terms — but never mistake the offer for a neutral gift, and keep the adopted layer politically portable: more than one lineage, from more than one jurisdiction, so that a policy shift in any capital costs you a re-basing exercise, not the doctrine. Build the owned layers now, and archive on a cadence, while the suppliers still find openness useful.
Unconquerable
The goal was never to become the sixth empire. Trying to out-build the empires is playing the one game almost no one can win. The winnable game is quieter: steward your data, own your evaluation and your governance, hold your reserve, train and keep your people, distribute through the surfaces your society already trusts — and rehearse the cutoff until the numbers are boring. A country that has done this has not conquered anything, and does not need to. It has reduced the number of decisions an empire can make on its behalf. It has made itself much harder to coerce.
The world around the empires does not need to become an empire. It needs to become unconquerable.
Further reading: AI Is Bigger Than Five — the founding history of the AI empires; the Open-Weight Field Guide, the dated companion that tracks the models, licenses, hardware, and season-by-season politics behind this doctrine; and The Smallest Sovereign, which runs the same argument one order of magnitude down — from the state to the citizen who owns their own context.