The AI Market Is Not One Market
Buyers enter linked markets for models, infrastructure, coordination, and institutional assurance.
The Bill Reveals the Market Structure
In June 2026, Microsoft made Copilot Cowork generally available and priced its use through Copilot Credits, a usage-based billing layer that sits on top of the per-user subscriptions enterprises already pay.
Microsoft’s own explanation says that a Cowork task consumes credits according to model use, context retrieval, tool calls, and runtime. Administrators receive budgets, alerts, allocation controls, prepaid options, pay-as-you-go billing, reporting, and hard caps. A workplace assistant now needs cost governance because the unit being sold is no longer simply a software seat.
One Cowork task can draw on markets for model use, context retrieval, tool execution, runtime and hosting, and enterprise cost governance. The interface presents one assistant; the bill records several economic relationships underneath it.
That distinction is the first clue that the phrase “the AI market” hides more than it explains. AI is better understood as a market web, a set of linked markets organized around four broad functions: making models, running them, coordinating their work, and providing institutional assurance. Buyers usually enter several at once. The bill keeps the categories concrete: a seat is not a token, a token is not a task, and a task is not a data-center contract.
Because those control points operate on different economic logic, fragmentation does not distribute power evenly. Model choice can widen while infrastructure supply remains narrow, and open weights can pressure model prices while platforms capture routing, billing, and procurement. The interface unifies the experience. The market structure does not.
Figure 1: The Bill Reveals the Market Structure
For a startup, the purchase may stop at an API and the cloud capacity behind it. A bank buys a longer operating chain because identity controls, audit records, evaluation, and compliance have to travel with the model. Government procurement lengthens it again by adding language coverage, residency, and a supplier able to work inside public institutions. Different buyers enter different combinations of the same markets.
Here, “market web” simply describes a purchase path that intersects several markets with different sellers and pricing units. Model choice, model production and inference, infrastructure, coordination, and sovereign-AI demand each have different bottlenecks and control points.
Model Choice Becomes More Modular
The model market is the most visible part of AI. It is where OpenAI, Anthropic, Google DeepMind, Meta, Mistral, DeepSeek, Sarvam, and other labs release, update, price, and distribute their systems through proprietary services, open weights, and cloud catalogs.
For bounded tasks, the model-selection layer is becoming more modular. A buyer can compare frontier proprietary systems, lower-cost models, regional models, and open-weight alternatives through catalogs such as Microsoft Foundry, Amazon Bedrock, and Google’s Vertex AI. In April 2026, Microsoft announced DeepSeek V4 Flash and DeepSeek V4 Pro in Microsoft Foundry, presenting them as choices that could be matched to tasks by quality, speed, and cost without changing the surrounding platform.
Catalogs make it easier to choose a model. They do not remove the cost of changing one. A change can require prompt revisions, evaluation reruns, latency checks, output-format changes, security review, and workflow testing. The cost is small for a casual chat and larger when the model sits inside document systems, permissions, tools, logs, and user routines; selection can become modular even while switching remains a paid activity performed by evaluators, integrators, security teams, and platform operators.
Chinese open-weight models add pressure at the model layer. CSIS describes recent Chinese models as close enough to compete in many practical tasks, with the gap framed as months rather than years and one US-government evaluation placing DeepSeek V4 Pro roughly eight months behind leading US models. This does not establish parity across every capability. It establishes that many buyers now have credible alternatives for some workloads.
These capable alternatives put pressure on model-layer prices. A support workflow, classification step, or routing task may not require the most expensive frontier model on every call. Open weights intensify that pressure because they can be adapted, hosted, and distributed by firms other than the originating lab.
Some open-weight releases, especially in China’s diffusion-first strategy, are better read as ecosystem-building or strategic projects than as products that must recover their costs through API margins. Proprietary frontier labs, by contrast, sell access against the background of large training and infrastructure commitments. An open-weight supplier pursuing diffusion and a proprietary frontier lab selling current access can therefore meet in the same model market without carrying the same revenue obligations.
Modular model choice is therefore only the first layer of the market web. An available model still needs to be run, supplied with context, evaluated, secured, billed, and governed. Behind and around that model sit three additional economies: training and inference, the infrastructure beneath both, and the coordination layers that turn a selected model into usable work.
Making Models and Running Them Diverge
The next economic boundary appears once a model has been selected. Using a model and making one are different economic activities. They require different firms, capital commitments, planning horizons, and revenue models.
A model is the learned system that produces outputs such as text, code, predictions, or action suggestions from a prompt and surrounding context. Training is the expensive process of producing or improving that system. Inference comes later, when the trained model is run for each query, summary, tool call, or agent step.
Training a frontier model is episodic and capital-intensive. It requires large clusters, specialized engineering teams, data pipelines, and long planning cycles. The expense is incurred before the resulting capability can be sold across APIs, subscriptions, open-weight releases, or enterprise contracts. The producer bears this large fixed and speculative cost before demand is known; enterprise buyers compare finished models and purchase variable usage after a much smaller set of organizations has absorbed that risk.
Inference is recurring. Every query, summary, code edit, tool call, or agent step consumes compute, memory, power, cooling, and network capacity. Its economics turn on volume, latency, reliability, and utilization. At small scale, the cost of a call can look trivial. At enterprise scale, repeated execution turns into a budget line, which is why routing rules, usage caps, and cheaper models become valuable.
Data creates another set of markets across both phases. Training may require licensing, annotation, synthetic-data production, provenance checks, and benchmark access. Deployment may require retrieval systems, permissioned enterprise data, audit records, and contractual rules about how data is stored or reused. Calling all of this “the model market” erases suppliers whose product is not a model at all.
The training-inference distinction matters because it prevents “compute” from becoming an empty catch-all. Training scales with the ambition of the next model. Inference scales with the volume and duration of use. They pull on some of the same physical inputs, but on different schedules and through different contracts.
This economic divergence changes how the market should be read. Catalogs can make trained models look abundant because buyers encounter them after the production cost has already been absorbed. The underlying production market remains narrower, and every act of serving those models creates recurring demand for a different set of inputs.
Infrastructure Concentrates Beneath Both
AI infrastructure is often compressed into a shortage of GPUs. In practice, accelerators become usable capacity only when they are combined with memory, packaging, fabrication, networking, cooling, land, and power. Each has its own suppliers, constraints, and bargaining structure.
The OECD’s work on AI infrastructure competition makes the concentration visible. It reports examples in which one firm is said to hold more than 80 percent of three AI-infrastructure segments, while the top three firms collectively hold more than 60 percent in three others. The OECD cautions that these are not formal market-definition findings. The direction remains useful: model competition can stay dynamic while infrastructure and complementary assets remain concentrated.
The OECD table also points to the variety inside infrastructure concentration: advanced lithography, advanced AI-chip fabrication, GPUs, high-bandwidth memory, cloud provision, and electronic-design automation are different markets, not one generic chip shortage. ASML is a concentration point in advanced lithography tools. TSMC is a concentration point in leading-edge fabrication. CoWoS and other advanced-packaging systems are constrained by construction and qualification timelines.
High-bandwidth memory illustrates a different bottleneck. It is supplied by a small group of firms, so limited qualified supply can become a planning constraint rather than a spot purchase. Networking depends on switches, interconnects, optical components, and systems integration that determine whether many accelerators can work as one system.
Power introduces a different economy again. A data center can have capital and chips but still wait for grid interconnection. Its viability can turn on tariff design, take-or-pay commitments, land, water, and whether local institutions accept the distribution of costs. Transformers, batteries, cooling equipment, and power-supply systems are mundane beside a frontier model, but they can decide when capacity is available and at what price.
These infrastructure layers concentrate through different mechanisms. A lithography-tool supplier gains power through technological concentration, while a foundry gains it through process capability and capital intensity. A memory producer gains it through limited qualified supply, while a utility or regulator gains it through control over interconnection and cost allocation. Concentration is shared. Its sources are not.
The current AI investment cycle amplifies the link between infrastructure, capital, and demand. Takshashila’s analysis of the AI investment cycle makes a useful distinction: a buildout can be financially overextended and still leave behind valuable infrastructure. Data centers, power contracts, engineering teams, procurement relationships, and operational knowledge do not disappear simply because some investors earn disappointing returns.
For infrastructure, the familiar choice between “bubble” and “transformation” is too crude. Financial returns and infrastructure formation can diverge. Whether the capital is eventually justified depends partly on whether organizations absorb AI quickly enough to create sustained demand for what has been built.
Enterprise demand, in turn, authorizes investment upstream. When a buyer signs a multi-year capacity contract, commits cloud spend, or moves a workflow into a platform, that commitment helps suppliers finance data centers, power agreements, chips, and memory. The buyer’s purchase commitment becomes the supplier’s capex authorization.
The physical substrate and the cloud platforms built on top of it concentrate for different reasons. The physical layers are constrained by hard supply limits: capital intensity, long lead times, inelastic inputs, and geographic chokepoints. Hyperscalers concentrate through accumulated software, identity, procurement, billing, data gravity, and workflow switching costs. Concentrated infrastructure also shapes which coordination platforms can operate at scale, and on what terms.
Coordination Becomes a Market Layer
Coordination turns a selected model into reliable work inside an organization. It means choosing which model handles which task, connecting it to data and tools, enforcing permissions, billing usage, evaluating outputs, and routing failures. A consumer may experience this as one chat box; an enterprise has to buy and govern the machinery around the chat box.
When model choice gets easier, economic value often moves into the assets required to make models useful. David Teece described these as complementary assets: the distribution channels, specialized systems, and commercial infrastructure through which an innovation reaches users. In AI, those assets include cloud capacity, model routing, enterprise identity, security controls, evaluation, billing, and workflow integration.
Complementary assets become valuable because technological improvements change where scarcity resides. Joshua Gans’s microeconomic account of AI helps explain why this happens. If AI reduces the cost of prediction, the scarce complements move elsewhere: judgment, data, decision authority, feedback, coordination, and organizational redesign.
Put differently, the bottleneck moves: cheaper model access makes coordination, judgment, and assurance more valuable. A market forms around those complements as much as around the prediction engine itself.
Wider model choice therefore expands the coordination problem rather than removing it. Firms that operate the systems around model selection can capture value even when they do not produce the selected model.
The coordination pattern is familiar from other modular systems. Manufacturing offers the older version: when components become easier to source, value can move toward system design, quality control, logistics, distribution, and customer relationships. Modularity does not remove coordination. It moves value toward whoever can assemble the modules into a reliable product.
The personal-computer era demonstrated the software equivalent of this dynamic. As components became easier to combine, operating systems became more valuable. The operating system did not manufacture every component. It coordinated access to them, provided a stable environment for applications, and became the point through which users and developers entered the system.
In AI, this coordination market sells model routing, hosting, catalogs, identity, policy controls, billing, evaluation, procurement, and compliance. Azure, Bedrock, Foundry, and their peers are more than neutral shelves: they turn models into purchasable enterprise objects by placing them inside existing security systems, cloud contracts, data-residency arrangements, and administrative controls.
Microsoft’s multi-model strategy is valuable because the platform coordinates model choice. The company does not need every available model to originate inside Microsoft; it needs buyers to select, route, meter, and govern those models through its environment. A model-agnostic platform can still be a concentration layer.
This supplier-platform overlap extends beyond Microsoft. OpenAI models, Codex, and managed agents entering AWS and Bedrock procurement paths show how a model supplier can expand distribution while the cloud provider retains billing, security, procurement, and infrastructure relationships. Anthropic’s compute and distribution relationships across AWS, Google, and Nvidia-based capacity show how the same firms can occupy several positions at once. In AI, a company can be a rival in one market, a customer in another, and a supplier or distributor in a third.
Coordination is easiest to see wherever models that appear interchangeable at selection time must execute multi-step work under organizational controls. The clearest current example is the agent execution environment.
An agent does more than generate an answer. It constructs context, calls tools, uses credentials, takes actions, checks results, retries failures, and decides when to stop. The surrounding system prepares the task, supplies tools and permissions, records intermediate steps, checks results, and handles failure. Builders often call that surrounding system the harness.
Benchmark evidence shows that this surrounding system changes practical capability. The paper “Stop Comparing LLM Agents Without Disclosing the Harness” reports that the same model can score differently under different scaffolds, including a Claude Opus 4.5 result on SWE-bench Pro that varied from 45.9 percent under one standardized scaffold to 55.4 percent under Claude Code. The paper’s broader point is methodological: agent performance is jointly produced by the model and the execution environment.
The benchmark finding does not prove a permanent moat for any native harness. It establishes that harnesses can improve, compete, and sometimes change model rankings. The execution environment is therefore a distinct market layer. Model choice grows more modular; the work environment re-bundles model, tools, context, verification, and control.
The value shift does not end with generation. Generation gets cheaper while evaluation stays scarce. Agent systems intensify the imbalance because each additional action may need validation, error handling, auditability, or human review. The coordination layer is where that scarcity is priced.
Figure 2: Different Market Layers, Different Competitive Conditions
Sovereign AI Is a Market Segment
Sovereign AI is the same market web seen through a public buyer. When governments use the phrase, they are usually not buying sovereignty as one thing. They are buying attributes spread across the linked markets: language coverage, local hosting, data residency, auditability, procurement compatibility, and institutional support. The label names a purchasing bundle rather than a product category.
Mistral offers a European-rooted route across models, enterprise tooling, deployment, and compute. Sarvam provides an Indian route built around Indian-language models, enterprise and government deployment, and participation in the wider IndiaAI ecosystem. These firms are not equivalent national champions, and neither demonstrates full-stack sovereignty. Their market significance is more concrete: they package attributes that global model catalogs do not supply on their own.
India makes sovereign AI’s market segmentation visible. An Indian public agency may want language coverage, data-residency assurances, domestic hosting, predictable cost, and an implementation partner able to work with public-sector procurement. No single technical layer supplies all of those attributes. A model firm, cloud provider, chip supplier, systems integrator, evaluation provider, and government program may all participate in the purchase.
India’s present AI strength is closer to application, adaptation, and deployment than to full-stack control. That makes sovereign AI less a claim of self-sufficiency than a question of which market layers Indian institutions can shape, learn from, and govern.
Institutional assurance becomes a market alongside technical performance. Evaluation reports, audit logs, deployment controls, local support, procurement eligibility, and compliance documentation become part of what is sold. Local-language capability and domain adaptation are one purchasable attribute; regional hosting, residency, auditability, and public-sector assurance are another. The same model may therefore become a different product when offered through a regional cloud, a public-sector contract, or an enterprise platform with identity and logging controls.
National AI programs also generate demand for infrastructure suppliers: accelerators, memory, cloud capacity, software tools, and integration services. Sovereign-AI purchasing can therefore buy language and deployment capability at one layer while revenue still flows to compute, cloud, or hardware suppliers.
National ambition operates inside the AI market as one source of demand, shaping which bundles suppliers build and which layers attract capital. The useful test for a sovereign-AI offering is therefore a purchasing test: who is the buyer, what attributes are being purchased, which firms supply each attribute, and where does the resulting revenue accumulate?
Buyers Enter a Market Web
Return to the administrator setting a monthly credit cap on a workplace assistant. The cap governs tasks that draw on model use, context retrieval, tool execution, runtime and hosting, and enterprise cost governance. One interface coordinates them; one invoice may bundle them; neither fact turns them into one economy.
Different competitive conditions can exist across AI at the same time. Ask how competitive AI is, and the honest answer is: at which layer?
Model selection can become more competitive while training remains capital-intensive; physical inputs can stay concentrated while several platforms compete to aggregate them. Coordination can remain contested even as billing, identity, procurement, and workflow integration pull buyers toward a small set of environments, while sovereign-AI demand adds political differentiation to markets already divided by cost structure and scale.
One pattern runs through all of them. As components become easier to substitute, economic power settles with whoever coordinates them.
The question “who has the best model?” can describe only one part of this structure. Market analysis has to ask who trains, who serves, who supplies inputs, who coordinates work, who validates performance, who owns procurement, and who captures recurring revenue. The history of modular technologies is full of buyers mistaking an interface for an industry. AI reproduces that error when an assistant conceals the linked markets required to make it usable.
Once the linked markets are visible, another question follows. Each layer creates a different relationship between buyer and supplier, and it is those relationships, not the technologies alone, that pose the next strategic question: which of them can later be changed, and at what cost? Sorting those relationships is the next task.
Buyers do not buy AI. They enter a market structure.
Visual note: The diagrams in this essay are original Yukti visuals, designed from the author’s briefs and produced with AI-assisted code generation, then reviewed before publication.
Earlier Essays
The Stack Beneath the Interface - why AI capability has to be read layer by layer; this essay adds the market-structure overlay.
Operational Capacity Is AI Capability - why evaluation, procurement, and cost governance are institutional capabilities, not administrative afterthoughts.
Digital Public Infrastructure and Its Limits - where coordination through shared rails can steer markets, and where substrate constraints bind.
The Dynamic Trilemma of Technology Strategy - the strategic trade-offs produced by this market structure.
Further Reading
David Teece, “Profiting from Technological Innovation” - the classic formulation of complementary assets and value capture around innovation.
Joshua Gans, The Microeconomics of Artificial Intelligence - useful economics foundation for AI as cheaper prediction, with value moving into judgment, complements, pricing, and market structure.
Arvind Narayanan and Akash Kapur, “Up the Stack: How AI’s Escape From the Commodity Trap Risks Enterprise Lock-in” - an adjacent economics argument that model inference may face price pressure while value capture moves into enterprise products, orchestration, switching costs, and lock-in.
Sources and Case Materials
Microsoft, Copilot Cowork is now generally available - the opening case for task-level AI cost governance.
Microsoft Learn on Copilot Credits and usage-based billing - cost-management controls for usage-billed Copilot services.
Microsoft Foundry announcement for DeepSeek V4 Flash and V4 Pro - model-catalog example for task-based selection.
CSIS, What to Know About Chinese AI Models - grounding for Chinese open-weight model pressure and the “months, not years” framing.
OECD, Competition in Artificial Intelligence Infrastructure - competition-policy support for concentration in infrastructure and complementary assets.
Takshashila, The AI Investment Cycle - why AI investment can run ahead of revenue while still leaving durable infrastructure behind.
OpenAI, OpenAI models, Codex, and Managed Agents come to AWS - evidence for model distribution through AWS security, governance, procurement, and Bedrock paths.
Anthropic and Amazon expand collaboration for up to 5 gigawatts of new compute - evidence for compute, cloud, billing, and enterprise-distribution entanglement.
Anthropic, Expanding our use of Google Cloud TPUs and Services - evidence for multi-provider compute relationships.
Anthropic, Higher usage limits for Claude and a compute deal with SpaceX - evidence for Nvidia-GPU capacity entering model-company infrastructure strategy.
Zhang et al., Stop Comparing LLM Agents Without Disclosing the Harness - evidence that agent performance is jointly produced by the model and the execution environment.
Mistral Compute and Le Chat Enterprise - case material for regional model, enterprise, deployment, and compute packaging.
Economic Times on Sarvam’s 2026 financing - case material for Sarvam’s sovereign-AI positioning, Indian-language models, enterprise and government deployments, and HCLTech partnership.
Government of India, Cabinet approval for the IndiaAI Mission - official source for IndiaAI compute, indigenous foundation-model, datasets, startup, and safe-AI pillars.




