Closed at the summit, open at the base


· 20 min read
The biggest AI story of the last quarter was not a better model but the moment buyers realised they did not need one for most work, and it arrived as a cost disclosure rather than a model release. In late June, Coinbase chief executive Brian Armstrong revealed the company had cut its AI spending by nearly half while token consumption kept growing, and was specific about how: "Not with friction and spend alerts. With better defaults, routing, and caching." Coinbase now defaults its engineers to open-weight models for routine work, reserving frontier models for the tasks that need them. Uber reportedly exhausted its 2026 AI coding budget in four months and capped usage, as the market shifted suitable work to cheaper open-weight models.
Neither is a story about companies using less AI, since usage rose in both, but about demand decoupling from centralised pricing: intelligence is shifting from metered distant racks to wherever the economics, the energy and the data say it should run. That shift creates a competitive question with a specific answer. The companies best positioned to win in the West will not choose between open and closed models; they will orchestrate both, routing each task to the cheapest adequate intelligence and wrapping the whole in the governance, support and assurance that enterprises actually buy. Not the purest frontier lab nor the rawest open-weight download, but the managed middle: closed at the summit, open at the base, with guardrails governing the space between. Electricity, computing, storage and networking each became distributed once economics allowed; intelligence now follows the same path.
The transition is forced by hard limits, and the first is energy. The largest hyperscalers are expected to spend hundreds of billions on AI and data-centre capacity this year, with recent estimates clustering around $650 to $710 billion, and every dollar must be energised through grids not built for it, with multi-year interconnection queues across Western markets. I argued in Power Before Performance that the societies combining abundant firm power with dense local compute will shape the century, and the capex numbers sharpen the point. Energy sets the clock: new centralised capacity waits years in queues while the distributed estate scales on shipping silicon. Satya Nadella put it plainly in January: "The choices we make about where we apply our scarce energy, compute, and talent resources will matter." When energy is the binding constraint, moving computation to where power already exists, in devices, gateways, vehicles and factories that are powered anyway, stops being an optimisation and becomes a strategy.
The second limit is silicon, and specifically memory. DRAM and NAND prices have risen sharply since late 2025, with Counterpoint reporting an 80 to 90 per cent jump in the first quarter of 2026 and TrendForce warning that high-bandwidth-memory demand is tightening conventional DRAM supply, because the new capacity Samsung and SK Hynix are building skews overwhelmingly towards the data centre. The squeeze cuts both ways: in the near term it makes fleet-scale on-device inference expensive and pushes workloads back towards the cloud, and by mid-2026 the rises had cooled only because device makers hit the ceiling of what they could absorb, not because supply eased, while the same investment cycle delivers commodity relief between late 2027 and 2029, at almost exactly the moment the next generation of edge silicon and radio standards arrives. The window in which distributed intelligence becomes broadly economic sits on the memory makers' capacity schedules today.
The third force is not a limit but a pull. The workloads growing fastest, the agentic systems that act on the world rather than answer questions about it, are the ones that most benefit from being local. A gateway diagnosing its own Wi-Fi or a hospital device processing patient audio needs low latency, generates data that should not leave the premises by default, and must keep working when the connection does not. Open models that are a year behind the frontier now score in the high seventies on graduate-level reasoning benchmarks while running on hardware people own, and the best open single-GPU models sit in the high eighties. Compression carries capability down the stack: in July 2026 PrismML released Alibaba's Qwen cut by over ninety per cent, all 27 billion parameters running on an iPhone, Apple in reported talks. The frontier remains ahead on the hardest long-horizon tasks, worth paying for on a shrinking class of work, but for everything else the intelligence has left the building. The speed of that departure is easier to see than to describe, and Figure 1 maps it directly, plotting capability against the power envelope each model deploys into.
Figure 1. The open continuum is closing in on the frontier: GPQA Diamond scores by indicative deployment power envelope, July 2026. Growth Under Hard Limits, 2/4 · lishawa.com
It matters that the intelligence which left is open. Open weights are to distributed intelligence what open protocols were to the internet: the property that lets capability be embedded, not merely accessed. A closed model can be called; an open one can be carried into a gateway, a robot or a hospital device, inspected before it is trusted, fine-tuned on data that need not leave the building, and run when the connection fails or the vendor changes terms. Distribution without openness is just a longer wire back to the same meter, which is why the fastest-diffusing layer of the stack is open by construction.
Underneath the economics sits a distinction the industry talks around and this argument depends on: the difference between general and focused intelligence, and the fact that the capability a task requires scales with its open-endedness, not with its economic value. General intelligence, the open-ended reasoning that can plan across long horizons and attack novel problems, is rare in the economy's daily workload and extravagant in what it demands, living in the largest models and the heaviest serving infrastructure. Focused intelligence is everything else, and everything else is almost everything: perceiving, classifying, monitoring, translating, summarising, controlling. A model that disaggregates a household's power signal into its appliances runs in a few hundred megabytes on a gateway processor at single-digit watts, pairing a specialised network for the signal with a sub-billion-parameter model for the meaning. A quality-inspection camera, a network firewall and a robot arm's perception stack are all of the same kind, enormously valuable precisely because they are bounded.
The mapping to physics is direct, in that focused intelligence fits small models, small models fit the power envelopes of devices and gateways, so it is distributed by nature; general intelligence fits large models, large models fit racks, so the summit is centralised by nature. The signature error of the centralised era has been serving focused tasks with general infrastructure, paying summit energy for base work, and the buyers' routing is the market correcting it. The economy does not run on genius but on competence, repeated, and competence has become small enough to live where the work is.
The Coinbase and Uber episodes proved that buyers have found the mechanism, right-sizing: matching each task to the cheapest model capable of completing it. Armstrong's cost levers are a routing policy, not a retreat. The Chinese open-weight models Coinbase now defaults to cost a fraction of the equivalent Western frontier rates per token, and the leading open models now sit within a few points of the best closed ones on coding benchmarks.
Nor is the shift anecdotal, since the volume has moved. On OpenRouter, a large neutral marketplace for model traffic, the combined US share of Google, OpenAI and Anthropic fell sharply over the year to mid-2026 while total volume rose several-fold, and DeepSeek became the single most-used model on the platform, with close to a fifth of token share by early June. The premium survives at the top of the same dataset, where the leading closed lab holds a fraction of the volume but far more of the revenue: the structure before anyone designed it. The distance between base and summit is closing faster than the consensus allowed. Moonshot, the Beijing lab behind Kimi, released K3 on 16 July, at two to three trillion parameters China's largest model to date, published as open weights, reported to beat Anthropic's Opus 4.8, short of the restricted class above it. If the results hold under independent testing the assumed eight-to-twelve-month Chinese lag becomes hard to sustain, and Mistral's chief executive has argued that the gap between the most advanced models is rarely more than six months, a claim the capability tables keep supporting.
Pricing moves against the convergence: market shorthand prices a barrel of intelligence at $56 from Anthropic and fifty cents from the Chinese labs, and Anthropic has scheduled a fifty per cent rise for September, the premium raised as the base climbs. Not all of that gap is a market signal, with distillation and state-backed capital behind part of the Chinese price, but the direction is unmistakable. Commoditisation cuts the base too, Z.ai and MiniMax falling 27 and 16 per cent on K3's debut: margin thins among the open labs, the durable value sitting in the layer above. Valuations price the difference, Moonshot raising at around $31.5 billion and DeepSeek at about $71 billion against Anthropic's $965 billion and OpenAI's $852 billion; part of the gap may reflect distribution and perceived assurance advantages.
The pattern now runs beyond the technology buyers into infrastructure itself: in December 2025 Deutsche Telekom, Europe's largest operator, announced a multi-year collaboration with OpenAI taking frontier models by contract, months after committing with Nvidia to a sovereign industrial AI cloud on German soil. Two separate initiatives form one hybrid strategy, the summit rented and the infrastructure kept sovereign, at national-champion scale.
The same episodes exposed why raw open weights alone will not carry the Western enterprise market. Within days of Armstrong's disclosure, the model makers he had adopted were named in a congressional security probe, and self-hosting improves data residency, but not the assurance question. Who evaluated this model for the failure modes that matter in this industry, who indemnifies and documents it for the regulator, and who answers the phone when an agent built on it does something expensive? Under the EU AI Act's obligations on general-purpose and systemic-risk models and the UK's growing assurance expectations, those are procurement criteria, not paranoia. Mistral's chief executive concedes regulated enterprises hesitate over Chinese models on security grounds, yet cost already pulls others across, Siemens and Airbnb among the adopters. A model with no licence fee still carries operating cost and liability, and most enterprises can price the first but not the second.
This is the gap the managed hybrid layer fills. It provides both open and closed models, because the frontier still earns its price on the hardest work while open models take the volume. It routes between them automatically, because the routing layer is where the savings live, and the reflex of throwing the most powerful model at every task is where budgets die. It wraps the whole in guardrails and the compliance documentation that converts cheap intelligence into deployable intelligence, and it comes with support, because enterprises buy accountability as much as capability. None of these layers is glamorous, but together they are the product.
The distributed shift is also a sovereignty shift, and it operates at three altitudes at once. At the personal level, the household's data becomes something stewarded at the edge under revocable consent, since the sensing that makes a home intelligent, the power signal and the daily routine, is the data most people would not choose to export. On-device inference makes that promise real, since the intelligence acts on the data without the data ever leaving, a stronger guarantee than any privacy policy.
At the enterprise level, sovereignty lives in the weights. A company that fine-tunes an open model on its own operations holds that capability as intellectual property it controls, inspectable and portable, and beyond a vendor's terms. That is a different asset from a metered call to a model somebody else owns, which is why the enterprises thinking clearly are treating their fine-tuned models as owned infrastructure.
At the national level, sovereignty is about the supply channel. Access to a foreign frontier model is a policy variable, not a fixed input. June's US export controls on Anthropic's Mythos and Fable made that literal, forcing European businesses to confront dependence on American systems. A nation that can run capable intelligence on infrastructure it controls has an option a nation dependent on a distant provider does not, and the distributed base is where that option is exercised. Intelligence you can run yourself, on data you hold, is sovereignty made operational; intelligence you can only rent is a dependency waiting to be priced.
Local inference does not always minimise cost: devices bring fragmented hardware and low utilisation, while racks batch work on specialist accelerators. The distributed case is strongest where privacy, resilience, latency or data gravity outweigh those economies of scale; the winning architecture is hybrid, not purely local. The argument is sometimes read as anti-frontier, and it is the opposite. Intelligence is a general-purpose input, like electricity or literacy, and general-purpose inputs create most value when cheapest and most widely available. The base going open and cheap is not a threat but the precondition for the distributed economy to exist at all.
The summit is different, and the difference justifies a carve-out. Self-improving frontier research, models that advance the capability of the next models, is the one part of the stack where the case for concentrated resources and controlled development, even government supervision, is strong, because the work is dangerous to get wrong and dependent on scale. The system cost of a wholly closed model is that it consumes the scarce inputs, memory, power, grid capacity, that the distributed base needs. Capacity ahead of internal demand is becoming visible: in mid-2026 Meta was reported to be building a cloud business to sell excess AI compute after raising 2026 capex guidance to $145 billion. Markets are pricing the question: mid-July delivered the worst semiconductor week since the April 2025 rout, memory makers falling hardest and investors asking when the build-out pays back. For holders of frontier equity and hyperscaler debt, capacity ahead of demand is unproductive capital until the base absorbs it. The fair configuration and the efficient one are the same: a closed, supervised summit doing the research only scale can do, and an open, distributed base doing the work only proximity can do.
Distributed intelligence is a network property, and the network's decisive metric is no longer speed but latency. Physics sets the terms: light in fibre travels roughly two hundred kilometres in a millisecond, so a round trip to a distant cloud region costs tens of milliseconds before a single token is generated, and no capex changes it. The workloads that carry the next decade of value have latency budgets the geometry forbids. A robot arm's control loop, a vehicle interpreting its surroundings, a grid balancing demand response in sub-second granularity, an assisted-living system deciding whether a fall just happened: each needs its intelligence within milliseconds of the task, which means on the device, the gateway or the network edge, and never across an ocean.
This reframes what national connectivity policy is for. Gigabit coverage is the headline metric across advanced economies, yet speed alone does not monetise the network or deliver the distributed era, whereas latency, jitter, reliability and the quality of the last metre do. Around three quarters of broadband traffic terminates over Wi-Fi, so the delivered experience of every local service is set one hop past the fibre. In-home wireless pushes tail latency down every generation, and operators have begun to expose latency and quality as capabilities the agentic economy will buy.
The societal ledger follows, because low-latency local intelligence makes care at home dignified not institutional, a fall detected in the moment with sensing staying on the premises, and it gives the energy transition a missing half. The ledger has a cost side: capital sunk in stranded central capacity is capital not building the edge, and the economies adjusting fastest will let the base absorb it. Intelligence at the meter and the gateway can shift household demand to when the grid is long on power and aggregate homes into dispatchable flexibility, scaled by the GB system operator to 10 gigawatts by 2030, delivered where the electrons are used rather than by cloud round trips, so the grid strain the centralised build creates is partly answerable by the very distributed intelligence it competes with for silicon. The countries that treat low-latency connectivity as productive infrastructure, as earlier generations treated electrification, will convert the distributed transition into measured productivity, and the operators who own that connective tissue hold a position no platform can occupy from outside.
The physical carriers of the shift are moving. Nvidia, whose economics still favour the centralised story, has moved towards the network edge with a billion-dollar stake in Nokia to put its silicon into the radio network, while Qualcomm, MediaTek and Broadcom turn phones, gateways and home equipment into inference hardware, and Ericsson and Nokia make network latency something a developer can buy. Huawei, sanctioned into self-reliance, has assembled the most complete vertical expression of the distributed stack anywhere, its own silicon, networks, devices and openly released models.
For the hyperscalers and the frontier labs alike, the shift is directly disruptive: every token that runs on a device or a gateway is inference that never reaches a metered rack, and every open weight adopted is pricing power surrendered at the summit. The summit is also reaching for the base: OpenAI's hardware push, now the subject of an Apple trade-secrets suit, is a frontier lab trying to own the device layer.
The hyperscalers' positioning shows they have read it the same way. Microsoft, Google and Amazon are each building the managed middle in their own idiom, less as frontier labs than as the layer that industrialises access to intelligence. Amazon has been the most explicit: its Bedrock service made leading closed models available alongside open ones through 2026 behind one routing, governance and evaluation layer, and Andy Jassy's framing, that most AI value accrues to the companies that make models usable rather than the ones that make models, states the platform view plainly. Google pairs its frontier line with a model-neutral cloud, and Microsoft, built on a frontier partnership, is diversifying towards a portfolio-and-tooling posture, because a platform that routes across models is worth more than one locked to a single supplier. All three still build frontier models; each also bets on being the layer where models are selected, governed, evaluated and served.
Reliability strengthens the case where hybrids are designed with genuine local fallback, since a service failure took Gemini down for roughly seven hours on 10 June 2026, and Downdetector's disruption-day count rose from six in the first quarter of 2025 to fifty-one in the same quarter of 2026. A hybrid that degrades gracefully where a rented estate stops is a design requirement, not a hedge.
America builds the world's smartest models while China embeds intelligence into the world's largest physical system, and China is pursuing the same distributed logic by a different route, through the state coordination and open release that its industrial structure allows. The scale is the headline: China installed roughly 295,000 industrial robots in 2024, more than half the world total. Separately, Counterpoint estimates it accounted for more than 80 per cent of humanoid installations in 2025. The direction is more telling than the count. UBTech is deploying humanoids running DeepSeek-derived reasoning onto the assembly lines of carmakers, a deployment no Western firm has matched at scale, although industrial deployment remains early everywhere. Beijing converts capability into position: nearly thirty countries joined a Shanghai-based World AI Cooperation Organization aimed at the global South.
The connectivity substrate is the part most easily missed, since China has close to five million 5G base stations and the world's densest fibre footprint, the physical precondition for distributed intelligence at national scale. A state-backed venture fund on the order of $138 billion, aimed at robotics and other advanced industries over twenty years, signals where national priority sits. The strategic reading is not that China leads on frontier models, where it does not, but that it has understood the value of AI is realised where it is embedded. The ageing that drives this is shared across Japan, Korea, Germany, Italy and the United Kingdom, making the distributed answer China is building a template every ageing advanced economy will need.
The two strategies converge on the same insight from opposite directions: the value of AI is realised where it is embedded, not where it is served, and the economics now agree. The same diffusion that keeps a factory productive keeps an ageing citizen at home with dignity, provided the data is stewarded at the edge under consent. The difference is that China's diffusion runs through state coordination and open weights while the Western version must run through markets, and markets need what raw weights do not provide, which is assurance. That is why the managed hybrid framework is not a compromise but the Western strategy properly understood: it takes the cost curve the open ecosystem created and the capability the frontier still provides, and makes both deployable inside the legal, security and support expectations Western enterprises, citizens and regulators actually have.
The practical test follows: ask of any supplier whether it can route a task across open and closed models on cost and capability rather than loyalty. The closing weeks of this quarter put dates and prices on the argument. An open-weight model tuned to beat the strongest public closed system ships in the season its target raises prices by half, and a compression startup puts a 27-billion-parameter model on an iPhone while Apple negotiates for the technique. Anthropic's February distillation complaint conceded the point: capability leaks downhill and only position does not. The positions that hold are physical and contractual rather than clever: the silicon and power that run intelligence at scale, and the networks and consented data that carry and feed it at the edge. The companies assembling the managed middle around those assets are doing so with published prices and named customers rather than promises, because the summit will continue inventing intelligence, and the base will decide who creates value from it.
Benchmark figures are reasoning-mode and directional; per-task costs indicative.
illuminem Voices is a democratic space presenting the opinions of leading Sustainability Thought Leaders, their views do not necessarily represent those of illuminem.
The world needs sustainability knowledge. At illuminem, no interest group or shareholder can influence our work. Thank you for supporting our mission to make high-quality and independent sustainability information free for all. Every contribution helps. Thank you for donating today.
illuminem briefings

AI · Corporate Governance
Andrea Bonime-Blanc

Green Tech · AI
illuminem briefings

Power Grid · AI
The Wall Street Journal

AI · Ethical Governance
The Washington Post

AI · Public Governance
energynews

Green Tech · Sustainable Business