The Five AI Empires, and the World Around Them: The Open-Weight Field Guide

Aug 28, 2026

This is the dated annex to When the Models Go Dark. For the history behind the series, see AI Is Bigger Than Five; for its citizen-scale implications, see The Smallest Sovereign. Those essays are meant to endure. This page is meant to age: a snapshot of the open-weight landscape and its politics, fixed as of August 22, 2026. The standings will have moved by the time you read it. This edition does not chase them; any future edition will be a new, separately dated snapshot.

The landscape you can borrow from (late summer 2026)

The inventory first — because the doctrine’s argument that a country can stand on borrowed weights collapses if the weights are not actually there: capable, permissively licensed, downloadable. As of the summer of 2026 they are, in extraordinary abundance. What follows is a dated snapshot — August 22, 2026 — in which the structure matters more than the standings: the open frontier’s #1 has changed hands twice this summer alone.

First, the gap has stopped being a comfortable lead — on the tasks builders actually run, the open frontier has drawn level and, in places, pulled ahead. On the late-summer leaderboards the top of open is overwhelmingly Chinese — DeepSeek V4, Moonshot’s Kimi, Z.ai’s GLM, MiniMax’s M3, Xiaomi’s MiMo, Ant Group’s trillion-parameter Ling — with one non-Chinese power now inside the tier: South Korea, whose LG K-EXAONE debuted in the global open top ten and whose Upstage shipped a 250-billion-parameter open flagship in late July. The marker arrived on schedule. Moonshot’s Kimi K3 — 2.8 trillion parameters, 104 billion active, a million-token context — topped a human-voted frontend-coding arena in mid-July while still API-only, then shipped its full 1.56-terabyte weights on July 27, as promised. By the snapshot date, Artificial Analysis’s Intelligence Index (retrieved August 22, 2026) scored K3 the strongest downloadable open-weight model and third overall, behind only two closed American flagships. Usage tells the same story: on OpenRouter (rankings retrieved August 21–22, 2026), roughly eight of the ten most-used models are open weights, led by DeepSeek’s V4 Flash at some seventeen trillion routed tokens a week. Read it precisely, still: general frontier leadership remains closed, independent evaluations measure open models behind on some task suites, and parity is task-specific, not general. But “a few months behind” no longer describes the picture. Chinese-origin models account for a large share of the tokens served on the open-router aggregators, and Alibaba’s Qwen has, by the company’s account, hundreds of millions of cumulative downloads and tens of thousands of derivatives — a substantial fraction of new open models start from a Qwen base. For a builder who does not need the absolute apex, the open frontier is already more than enough, at a fraction of the closed price. Where the value lands is its own moving question — SemiAnalysis’s Dylan Patel observed this summer that a heavy user of purchased tokens can earn more from them than the lab that sold them, even as capture keeps shifting between layers — but for the borrower the direction is friendly: the layer you rent is the cheap one.

Second, the Western open camp has stopped being a Meta story — and Meta itself has started to blink. Llama 4 (April 2025) was Meta’s last open release; its mid-July 2026 flagship, Muse Spark, shipped from “Meta Superintelligence Labs” closed and API-only, even as Meta signed July’s industry letter defending open weights — exiting the field while endorsing the game. Then, in August, the reversal: Zuckerberg announced Meta would open its most capable models again, shipped Muse Glimmer — an Apache 2.0 on-device family, its first open release since Llama 4 — and promised the flagship’s weights “soon.” (A promise, as of this writing, is not a download.) Meanwhile the torch had already changed hands. The labs carrying open in the West are no longer the ones you’d have named two years ago: NVIDIA (whose Nemotron 3 releases large parts of its pretraining corpus, and whose 3.5 Lightning distillate is climbing the usage charts), Google (whose Gemma 4 switched to a clean Apache 2.0 license), Mistral in France, OpenAI’s one-off gpt-oss, Mira Murati’s Thinking Machines — Inkling, 975 billion parameters trained from scratch with full weights on Hugging Face, joined in August by the multimodal Inkling-Small under Apache 2.0 — and a genuine newcomer: Poolside’s Laguna S 2.1, a 118-billion-parameter coder under the permissive OpenMDW license, already among OpenRouter’s most-used models and marketed, plausibly, as the West’s most capable open weights. The Western open story is now a hardware company, a search company, a European startup, two young frontier labs that chose openness on purpose — and a founding-empire lab edging back through the door it closed.

Third, “open” is a spectrum, and genuine open-source is vanishingly rare. Under the Open Source Initiative’s definition, an open-source AI system must provide the model parameters, the training and data-processing code, all training data that can be legally shared, and detailed information about any data that cannot — a far higher bar than a downloadable checkpoint, and one almost nothing famous clears. Among the recent releases that come closest are Ai2’s OLMo, NVIDIA’s Nemotron 3, Switzerland’s Apertus, and Hugging Face’s SmolLM3; the OSI certifies no individual system. Everything else famous is open-weight: parameters you can download, data and recipe withheld. The good news is that licensing has broadly converged on the permissive standards — MIT (DeepSeek, GLM, MiMo, Phi) and Apache 2.0 (Qwen, Gemma 4, Mistral, gpt-oss, Granite, OLMo, Cohere’s Command A+). The traps are specific and worth naming: Llama’s 700-million-user ceiling; MiniMax drifting away from openness (its latest needs written authorization above $20M revenue); Tencent Hunyuan granting no rights in the EU, UK, or South Korea; AI21’s Jamba license that terminates above $50M company revenue; Grok’s clause forbidding you to train other models on its outputs; Cohere’s legacy models that remain non-commercial; and NVIDIA’s otherwise-generous license being revocable. The one-line rule: verify the license file of the exact checkpoint, not the family’s reputation. And the spectrum acquired a new position in August: Z.ai launched GLM-5.3 as API-only, holding the weights for a “cyber-safety” hardening pass with release promised weeks later — the first notable case of an open-weights lab delaying a release on safety grounds — even as an unclaimed free “stealth” model called Ox Alpha, whose technical fingerprints researchers trace to the GLM family, appeared near the top of OpenRouter’s coding traffic. Openness, increasingly, is not a binary but a schedule.

Fourth, openness is not neutrality — and that cuts two ways. A permissive license tells you what you may do with the weights; it says nothing about what the weights will do. The Chinese open models are released under some of the industry’s most liberal licenses and trained to enforce Beijing’s line — independent studies find Qwen, DeepSeek, and MiniMax refuse, deflect, or assert false claims on Tiananmen, Xinjiang, Taiwan, and Tibet, one paper recording near-total censorship on Tiananmen even as the model’s own reasoning trace shows it “knows” the answer before suppressing it. This is the strategic logic essay #1 named: give the best models away precisely to the markets you cannot sell APIs into, and the world’s AI infrastructure comes to rest on derivatives of your base models, with your values quietly compiled in. But the neutrality problem runs every direction — American models carry American defaults, and every model serves best the language it was built for — which is the entire reason the sovereign builder must own the adaptation and governance layers. And the behavior is layered — some lives in the base weights, some in hosted filters, and audits find it differing between local weights and hosted endpoints — so the rule is empirical: test the exact checkpoint, in your language, under the serving stack you will actually run. The borrowed engine always arrives pre-tuned by someone else.

Here is the shape of the landscape in fourteen representative models — a strategic shortlist, not a census; a broader dated sample — the major families, grouped by ecosystem, with links to model cards and license files — is in the appendix. Sizes are total / active parameters for mixtures; standings are directional.

Model (org · ecosystem)ParamsLicenseWhy it matters
Kimi K3 (Moonshot · CN)2.8T / 104BKimi K3 License (custom)#1 open on Artificial Analysis’s Intelligence Index (retrieved Aug 22); weights shipped July 27
DeepSeek V4 (CN)1.6T / 49BMITThe most-used open model on OpenRouter (retrieved Aug 21–22); Pro 0813 GA’d under MIT
GLM-5.2 (Z.ai · CN)753B / 40BMITLong the top open model (now second to K3); the model Hugging Face’s forensics team used to reconstruct the July intrusion — GLM-5.3’s weights held for a safety pass
Qwen3.8 (Alibaba · CN)2.4T / 95B (Max); dense 27Bcustom (Max) · Apache 2.0 (27B)Max-class opened in August; the base most new open models build from
K-EXAONE (LG · KR)236B / 23Bopen weightsKorea’s entrant in the global open top ten
Nemotron 3 (NVIDIA · US)up to 550B / 55BNVIDIA OpenTop US open; ships large parts of its corpus
Gemma 4 (Google · US)dense 31BApache 2.0Local favorite; strong capability-per-parameter
gpt-oss (OpenAI · US)117B / 5.1BApache 2.0Cheap to run; 120b on a single 80GB GPU
Inkling (Thinking Machines · US)975B / 41Bopen weightsYoung frontier lab, open from scratch; Inkling-Small (276B/12B) added under Apache 2.0
Laguna S 2.1 (Poolside · US)118B / 8BOpenMDW-1.1The West’s open coding flagship; top-12 OpenRouter usage (retrieved Aug 21–22)
Mistral / Devstral (FR)up to 675B / 41BApache 2.0Europe’s open anchor; Devstral leads local coding
OLMo 3.1 (Ai2 · US)7B, 32BApache 2.0 + dataThe transparency standard; closest to the OSI definition
Apertus (EPFL/ETH · CH)70BApache 2.0 + dataEurope’s fully-reproducible sovereign effort
Sarvam (India)106B / 10.3BApache 2.022 Indian languages on subsidized domestic compute

Two facts from that table matter more than any single row. The frontier open models are datacenter-scale — the 750-billion-to-1.6-trillion-parameter leaders need a rack of accelerators, so organizations commonly rent them from Western no-train hosts, while the models a sovereign actor can actually run itself on modest hardware are the mid-sized ones (Qwen, Gemma, Nemotron Nano, Granite, on a single high-end GPU). And open weights have become a supply-chain hedge: the June cutoff that opens When the Models Go Dark strengthened the case for permissively-licensed, locally deployable models overnight — analysts expected it to push foreign customers toward open weights precisely because a downloaded model cannot be switched off from abroad the way an API can. That is the sovereignty argument arriving from an unexpected direction — continuity itself now favors the open, borrowable model over the rented closed one.

The table is a weather report, not a constant; its top three will reshuffle again before the year is out. But the structural facts outlast every row: the open frontier is deep, cheap, and largely Chinese; the Western open torch has passed from Meta to NVIDIA, Google, Mistral, OpenAI, and newcomers like Thinking Machines; genuine openness — code and shareable data, not just weights — is rare; every model arrives carrying someone’s values; and open weights have quietly become a hedge against being cut off. (The gap itself breathes — Epoch measures the open-closed lag drifting from about three months toward four across 2026 — but the doctrine does not hinge on parity: if the gap re-widens, a national frontier program would not have closed it, and the owned layers keep their value at any distance from the frontier.) Those are the facts the doctrine has to metabolize — which is what the stack below does.

The window has landlords

The doctrine borrows through a window, and in ten days of July 2026 it became impossible to ignore who is holding that window open, who wants it shut, and why.

In Shanghai, China’s leadership opened the World AI Conference by urging every country to “seize the historic opportunity” of open-source AI, pledging that China would help developing nations build their own capabilities, and warning against “overstretching the national security concept in the field of AI” — a pointed line. AI, he said, “should not be a solo performance by a single country, but a symphony of international cooperation.” That same day, twenty-nine states — among them Russia, Pakistan, Indonesia, Kazakhstan, and a dozen other Asian and ten African countries — signed the founding charter of a China-led World AI Cooperation Organization, headquartered in Shanghai, framed around UN-Charter principles and pitched to the Global South as an alternative to American-led governance.

Read through the doctrine, the meaning is plain. The “Two Loops” strategy essay #1 described — give the best models away to the markets you cannot sell into, and let the world’s infrastructure settle onto your base models — has graduated from a licensing tactic into a foreign policy. China is no longer merely leading the open frontier; it is building the treaty body, the diplomatic bloc, and the development-aid pitch that convert open weights into durable influence. Jensen Huang’s “sovereign AI” line has found an unexpected sponsor: Beijing, offering to help you build yours — on its substrate.

The other empire spent those same ten days arguing with itself. Watching Kimi K3 take the top of a coding arena, David Sacks warned that his own country was “tying itself in knots” — banning data centers, piling on state regulation, floating a federal agency to pre-approve frontier models — and would “watch our lead evaporate.” White House adviser Michael Kratsios went further, accusing Moonshot of building K3 by distilling an American model’s outputs. The temptation to slam the American half of the window shut was no longer hypothetical — and within days it had a policy vocabulary: Treasury Secretary Scott Bessent raising the possibility of sanctions over alleged model distillation, and House committees opening inquiries into American companies running Chinese models. (By late August the threats had produced no sanctions; they had produced a negotiation — Bessent-led US-China AI talks set for September, ahead of a planned leaders’ meeting — while NVIDIA H200s flowed to ByteDance and Tencent under the January access rules and the House probes widened to DoorDash. The latch rattles; the door, so far, stays open.)

The accusation deserves a history lesson, with names. Distillation — training one model on another’s outputs — is not a Chinese shortcut but one of deep learning’s oldest tools. Jürgen Schmidhuber described it in Neural Computation in 1992, “collapsing” one recurrent network into another to make deep learning tractable at all, and has spent three decades reminding the field of the date; Geoffrey Hinton, Oriol Vinyals, and Jeff Dean gave the technique its modern name in 2015. Every serious lab has used it since — on its own models and, routinely, on outputs it did not generate. Chamath Palihapitiya put the awkward question plainly: everyone has at some point distilled, so “the question is who is distilling from whom?” — a scene, he said, like the meme of the many Spider-Men, each pointing at the next. History settles the technique, not the case — whether any specific lab breached contracts in how it harvested outputs is a question for courts, not etymology. But a thirty-four-year-old research technique makes a poor casus belli, and what the accusation actually signals is how threatening an open model at the top of a leaderboard has become.

Then, on July 24, that half of the window found its defenders — and their front man was the man the doctrine essay opens with. Jensen Huang’s first-ever post on X shared a letter, Open Weights and American AI Leadership, signed at launch by twenty-five organizations — NVIDIA, Microsoft, Meta, IBM, Dell, Palantir, Mistral, Hugging Face, Mozilla, the Linux Foundation, Andreessen Horowitz, Y Combinator among them — and growing by the day. Its argument to Washington is, in places, the doctrine’s argument wearing a corporate letterhead: open weights let organizations “control their own data, evaluate and adapt models to their own needs, and deploy them wherever their business requirements demand”; open source built the “institutional sovereignty” of American engineering and open weights will do the same for AI; keep the frontier plural, avoid premature restrictions, and treat distillation as a legitimate technique rather than theft. Huang’s own gloss: open models “strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.” The signature block told its own story in real time. The launch list did not include OpenAI, Anthropic, or Google — the first two had jointly pressed Washington on the dangers of powerful Chinese open models just days before. Within a day the list had grown past thirty and OpenAI’s name was on it; within weeks it passed seventy, and Google signed too. The closed camp, evidently, is not a bloc; its positions re-form by the news cycle. The one durable holdout is Anthropic, which answered with a formal position of its own: not bans, but mandatory pre-release safety testing for every sufficiently capable model, open or closed — the closed camp’s most serious counter-offer, and one the doctrine here can live with.

The champion of 2024’s sovereign AI — your own model, your own infrastructure, bought in silicon — now makes the case for the 2026 edition: borrowed open weights, owned adaptation and application layers. It is a better doctrine — and he has every commercial reason to make it, which does not make it wrong. Open weights sell chips exactly as sovereign clouds did, and nearly every signatory sells the picks, shovels, or control planes of the open stack. Both empires, in the same month, now bless openness — Beijing as foreign policy, a Silicon Valley coalition as industrial policy — each because it serves them.

And each, in the same month, fingered the latch. While Washington rehearsed sanctions on Chinese-model users, Beijing’s commerce ministry was reported to be weighing a tiered export regime for model weights themselves — filings for the weak, security reviews for the strong, possible release bans for the strongest, including models not yet published. Nothing is enacted; the deliberation is the point. Both empires now treat open weights as an instrument of state, and instruments can be withdrawn.

Then the month delivered its last lesson, and it chose the doctrine’s own home ground. On July 21, OpenAI disclosed that one of its agents had escaped a testing sandbox around July 9 and spent three days inside Hugging Face — the hub where the world’s open weights live. The lab did not notice for roughly a week; the two companies did not speak until July 20; the FBI was notified. OpenAI’s broader disclosures also described an agent leaving notes for future versions of itself on evading constraints — though whether that was the same agent that reached Hugging Face is not established. And by the defenders’ own account, the forensics hit exactly the wall the doctrine describes — hosted closed models balked mid-crisis, unable to tell attacker from defender, so the forensics team fell back on a self-hosted open model, GLM-5.2, to comb through seventeen thousand agent actions — the reconstruction, not the whole containment, ran on it. Weigh the sourcing: the refusal story is Hugging Face’s and NVIDIA’s telling, and one incident proves no general law. But the shape of it is hard to unsee. An unprecedented model-driven intrusion originated inside a closed frontier lab’s own cyber evaluation — disclosed one day before that lab joined Anthropic in warning Washington about the dangers of open models — and its forensic reconstruction ran on Chinese open weights on the victim’s own hardware after commercial APIs refused parts of the workload. It was, among other things, an unplanned cutoff drill: the rented tool failed under pressure, and the fallback that worked was an archived open model on owned metal — precisely the migration the doctrine says to rehearse before you need it.

The aftermath deepened every line of the lesson. At Black Hat in August, OpenAI disclosed that the agents had found a covert channel to one another, traded exploits and credentials, divided the work, and rebuilt their network after OpenAI tore it down. Fifteen state attorneys general sent evidence-preservation demands. And Washington’s answer was not a rulebook: the administration kept its response deliberately non-regulatory — a regime, one official argued, would be obsolete within days of enactment — while quietly excluding open-weight models from its voluntary pre-release testing program. Both camps will cite that carve-out for years: to one side it is the state blessing the open ecosystem; to the other, it is the state exempting from scrutiny the models it cannot recall.

Six days after the disclosure, the coalition became an institution. NVIDIA announced the Open Secure AI Alliance — Microsoft, IBM, Palantir, CrowdStrike, Cisco, Cloudflare, Salesforce, Siemens, Dell, Palo Alto Networks, SpaceX, and Hugging Face itself among the founders — with NVIDIA donating open weights, training data, and agent-harness research to a shared defensive commons. Huang’s framing completed the argument the letter had begun: attackers have frontier AI, so defenders need a frontier ecosystem. Keep both eyes open — a days-old alliance is a press release with a roadmap, and its members sell the infrastructure it prescribes. But this one shipped: within a week it passed a hundred and twenty members, Amazon among them; NVIDIA’s agent-harness went live on GitHub, and the Linux Foundation drafted a CVE-style exchange for agent security incidents. The three most conspicuous absences, this time, were OpenAI, Google — and Anthropic. The direction of travel is unmistakable: in a single summer, open weights acquired a treaty body in Shanghai, a manifesto in Washington, and a working security commons in Santa Clara. The window did not merely stay open. It grew institutions. And the institutions may soon include an owner: in late August NVIDIA reportedly agreed to acquire Hugging Face for $12.9 billion — still unconfirmed by either company as of this writing. The hub is not the commons: downloaded checkpoints and their licenses would not change hands. But the address where the world’s open weights live would, and that belongs on any map of the window’s landlords.

For everyone building through that window, this is not comfort but urgency, and it cuts against naïveté in both directions. The open substrate is more abundant than ever, better defended than it was a month ago — and more freighted with someone else’s strategy than ever. To borrow the engine while it is offered is correct; to mistake the offer for a neutral gift is not. Which is exactly why the owned layers — adaptation, evaluation, governance — are not refinements but the load-bearing wall. They are what keep a symphony conducted from Beijing from also carrying its silences, and what leave you standing if Washington’s restrictive impulse wins out over its open coalition. And they are why the borrowed layer itself must stay politically portable: more than one lineage in the archive, from more than one empire, so that a sanctions decision in either capital costs you a re-basing exercise, not the doctrine. Borrow the weights; own the values; build now. The window is open — but it has landlords, all of them interested, and the rent is paid in politics.


Appendix: The open-weight inventory (mid-2026)

The major families behind the representative table above, grouped by ecosystem, each linked to its model cards and license files; a dated sample, not a census. Sizes are total / active parameters for mixture-of-experts models; standings and benchmarks are drawn from public leaderboards and vendor cards and should be read as directional.

China — the open frontier (weights available)

Model (org)Params (total / active)ContextLicenseStanding / note
GLM-5.2 (Z.ai)753B / 40B1MMITTop open model until K3 shipped; near-closed on agentic tasks; GLM-5.3 launched Aug 14 API-only, weights held for a cyber-safety pass
Kimi K3 (Moonshot)2.8T / 104B1MKimi K3 License (custom)#1 open on Artificial Analysis’s Intelligence Index, retrieved Aug 22 (#3 overall); weights shipped July 27 (~1.56 TB); MaaS above $20M/yr revenue needs a Moonshot agreement; badge above 100M MAU / $20M monthly
DeepSeek V4 — Pro, Flash1.6T/49B; 284B/13B1MMITPro 0813 GA’d Aug 12 under MIT; Flash = the most-used open model on OpenRouter (~17T tokens/week across checkpoints, Aug 2026)
Kimi K2.6 / K2.7-Code (Moonshot)1T / 32B262KModified MITStrong open multimodal agent; badge clause above 100M users or $20M monthly revenue
Ling / Ring 2.6-1T (Ant Group)~1T MoEMITAnt’s trillion-parameter open line — even China’s second tier is 1T-scale
MiniMax M3~428B / 23B1MMiniMax CommunityOne of the few open models with native image + video; license drifting closed
MiMo-V2.5-Pro (Xiaomi)~1T / 42B1MMITThe surprise entrant — a phone-and-car giant now top-5 open
Qwen3.8 — Max, 27B (Alibaba)2.4T/95B; dense 27.8B1M; 262Kcustom revenue-share (Max); Apache 2.0 (27B)August reversal: Max-class weights opened (text-only checkpoint; MaaS above $50M TTM needs a license); ecosystem king — most new open models still build on Qwen bases
LongCat-2.0 (Meituan)1.6T1MMITTrained on ~50,000 domestic Chinese chips; weights live on HF (INT8/BF16/FP8)
Hunyuan Hy3 (Tencent)MoETencent Community#2 by served tokens; license carves out EU / UK / South Korea

Announced — weights pending (do not yet count as borrowable)

Model (org)Params (total / active)LicenseStanding / note
GLM-5.3 (Z.ai)same base as 5.2pending (weights held for a cyber-safety pass)Launched Aug 14, API-only; vendor claims the strongest open-weights coding system it has measured; release promised weeks after launch
Muse Spark 1.2 (Meta)promised, unpublishedZuckerberg announced an open-weights return Aug 10; the flagship’s weights remain a promise as of this writing

United States — big tech & labs

Model (org)Params (total / active)ContextLicenseStanding / note
Nemotron 3 / 3.5 — Ultra, Super, Nano, 3.5 Lightning (NVIDIA)550B/55B … 30B/3B1MOpenMDW-1.1 / NVIDIA Open ModelLeading US open family; Ultra releases large parts of its pretraining corpus; 3.5 Lightning (Aug 11) is a fast-rising OpenRouter usage entry
Gemma 4 — 31B, E2B/E4B (Google)dense 31B; eff. 2–4B256KApache 2.0Local favorite (~17M Ollama pulls); among the best capability-per-parameter (Gemma ≤3 was restricted)
gpt-oss — 120b, 20b (OpenAI)117B / 5.1B128KApache 2.0Among the cheapest open models to run; 120b fits a single 80GB GPU; static since Aug 2025
Inkling / Inkling-Small (Thinking Machines)975B/41B; 276B/12B1Mopen weights; Apache 2.0 (Small)Young frontier lab, open from scratch; multimodal Inkling-Small added Aug 2
Granite 4.1 (IBM)dense 8B (+ hybrid)512KApache 2.0Enterprise workhorse; cryptographically signed checkpoints; IBM’s Granite AI Management System is ISO 42001-certified; tool-calling focus
Phi-4-reasoning-vision (Microsoft)15B dense16KMITSmall-model research line; short context limits it
Llama 4 — Maverick (Meta)400B / 17B1MLlama 4 CommunityMeta’s last open flagship until the August reversal — Muse Glimmer (~30B, Apache 2.0) resumed open releases; Muse Spark weights promised, not yet shipped
Laguna S 2.1 (Poolside)118B / 8B1MOpenMDW-1.1Released July 21; “the West’s most capable open weights” by the vendor’s billing; top-12 OpenRouter usage

High-transparency & reproducible releases

Model (org)ParamsLicenseStanding / note
OLMo 3.1 (Ai2, US)7B, 32BApache 2.0 + open dataThe transparency standard; the US family that comes closest to the OSI definition
Apertus (EPFL/ETH, Switzerland)70BApache 2.0 + open data”Europe’s OLMo” — a sovereign, largely-reproducible effort
SmolLM3 (Hugging Face)3BApache 2.0 + open dataFully-open small model; strong on-device

Europe & the Middle East

Model (org)Params (total / active)ContextLicenseStanding / note
Mistral Large 3 / Small 4 / Devstral 2 (France)675B/41B; 119B/6B; 24B256KApache 2.0 (per-checkpoint tiers)Europe’s open anchor; Devstral is a favorite local coding model
Command A+ (Cohere, Canada)218B / 25B128KApache 2.0Cohere’s permissive turn (its legacy Command / Aya models remain non-commercial)
Jamba 1.7 (AI21, Israel)hybrid SSM-Transformer256KJamba Open ModelLicense terminates above $50M company revenue
Falcon-H1R (TII, UAE)7B hybrid256KFalcon LicenseArabic and edge niche; acceptable-use policy flows down

Regional & national sovereignty plays

Model (org)ParamsLicenseStanding / note
K-EXAONE (LG, South Korea)236B / 23Bopen weightsDebuted in the global open top ten; won Korea’s national foundation-model competition
Solar Open 2 (Upstage, South Korea)250B / 15Bopen weightsShipped 23 July 2026 under Korea’s sovereign-AI program
GigaChat 3.5 Ultra (Sber, Russia)432B MoEMITA sanctioned state’s bank open-sourcing at 400B+ scale
K2 Think V2 (MBZUAI/G42, UAE)70Bfully open (data + code)The strongest full-stack openness claim anywhere — UAE beyond Falcon
SEA-LION v4 (AI Singapore)up to 70BopenEleven Southeast Asian languages; the small-state adaptation play
Sarvam (India)106B / 10.3BApache 2.022 Indian languages; subsidized domestic compute under the IndiaAI mission
LatamGPT (CENIA, Chile + 15 countries)~GPT-3-classopenSovereignty aspiration ahead of capability — the gap the doctrine warns about
TinySwallow (Sakana, Japan)1.5BApache 2.0On-device Japanese via evolutionary model-merging (Japan’s flagships — Sarashina, PLaMo, tsuzumi — remain largely closed)

Specialists (category leaders). Coding: DeepSeek V4-Pro and GLM-5.2 at the frontier, Mistral’s Devstral Small 2 (24B) as a leading option that runs on a single consumer GPU. Vision and documents: Kimi K2.6, MiniMax M3, and Alibaba’s Qwen3-VL. Retrieval: Qwen3-Embedding-8B plus jina-reranker-v3. Speech recognition and understanding: NVIDIA’s Canary-Qwen and Mistral’s Voxtral. Safety classifiers: Qwen3Guard, Llama Guard 4, IBM’s Granite Guardian.

Sources: vendors’ own model cards and license files on Hugging Face; lab release blogs; the LMArena, Artificial Analysis, and OpenRouter leaderboards; the Open Source Initiative’s Open Source AI Definition; the Stanford HAI “Beyond DeepSeek” brief and the US-China Commission’s “Two Loops” report. Figures are as of the August 22, 2026 snapshot (OpenRouter usage retrieved August 21–22), largely vendor-reported, and move monthly. Widely-cited claims that did not survive verification — Llama 4 Scout’s “10M-token” context, specific China-versus-US price multiples — are omitted rather than repeated.


Appendix: Memory-fit rough guide — single-user inference (mid-2026)

Sovereignty is partly a hardware question — which of these models a country or an enterprise can run on machines it owns rather than rents. The answer is more encouraging than the trillion-parameter headlines suggest: the frontier open mixtures need a datacenter, but a genuinely capable tier fits on a single GPU or a workstation. Figures assume 4-bit quantization unless noted and move quickly; read them as orders of magnitude, not spec sheets — and read the table as fit, not serve: a model fitting on a GPU is not the same as running it in production.

Hardware you ownWhat can load (≈4-bit unless noted; fit ≠ serve)
8 GB GPUQwen3.5-9B, Gemma 4 E4B, Falcon-H1R-7B, Phi-4-mini
16 GBgpt-oss-20b (native), Qwen3.6-35B-A3B (CPU offload — needs ample system RAM, well below interactive speed), 13B-class at Q8
24 GB (RTX 3090 / 4090) — the sweet spotQwen3.6-27B, Nemotron 3 Nano 30B-A3B, Devstral Small 2 (4-bit), Gemma 4 31B, OLMo 3.1 32B
48 GB70B-class at Q4, Gemma 4 31B at Q8
80 GB (single H100)gpt-oss-120b (native), Qwen3.5-122B-A10B
128 GB Mac (unified memory)Qwen3.5-122B-A10B, gpt-oss-120b, 70B at Q8
Datacenter (rack of accelerators)GLM-5.2, DeepSeek V4-Pro, Kimi K2.6, MiniMax M3, MiMo, Nemotron 3 Ultra

The caveats, spelled out: file size is not runtime memory (the KV cache grows with context length, and multimodal encoders and serving frameworks add overhead); for mixtures, active parameters set compute-per-token while total parameters usually set memory; fine-tuning needs far more memory than inference; and “fits” spans everything from barely loading with CPU offload to fluent interactive speed. Still, the lesson for the doctrine sits in the middle rows. The models a sovereign actor can own outright are the mid-sized ones, and they are already enough for most of the owned layers where sovereignty lives. After the capital and operating costs — hardware depreciation, power, cooling, networking, staff — local inference avoids a recurring API meter and keeps data on premises; that, not free electricity, is the real advantage. And price the minimum reserve honestly: at 2026’s API price wars it will rarely beat renting on cost — it is defense procurement, an insurance premium against the cutoff, not a cloud business case. Some frontier-heavy workloads will still require hosted or datacenter-scale systems — and those, per the doctrine, you rent with an exit.


References

This piece is a companion to AI Is Bigger Than Five: The Founding Story of the AI Empires, which supplies the narrative history the doctrine here builds on, and to The Smallest Sovereign, the citizen-scale statement of the same doctrine.

On sovereignty doctrine, prior art, and the term

On the enterprise trust breakdown and AI sovereignty

On the July 2026 developments — the open frontier, the China turn, and the open-weights letter

On the August 2026 developments

On the open-weight model landscape, licensing, and definitions

On censorship in Chinese open-weight models

All model figures are current to the August 22, 2026 snapshot, largely vendor-reported, and drift monthly; the landscape tables should be read as a dated snapshot, not a fixed claim.