Export PDF

A working theory of AI value accrual over time

AI gets cheaper.
What stays
valuable?

As capable models get cheaper,1 pricing power shifts toward companies that run agents reliably, connect them to customer data, and own the trusted interface where work gets done. Leading model labs won’t automatically control those businesses.

THE THESIS

Pricing power follows what customers need and can’t easily replace. In AI, our hypothesis is that the main source of that power shifts from intelligence to deployment, then to context and ultimately distribution. These advantages coexist. Later ones gain bargaining power only as earlier inputs become more commoditized. Adding compute can be faster than getting permission to use customer data or earning trust to act on it.

Explore the four shifts

The idea behind the curve

AI can spread while its suppliers lose pricing power. When customers can switch to an adequate alternative, the premium on a scarce input comes under pressure. That can happen even as the frontier advances.

Even if frontier development slows, compute demand may keep rising as more work goes into validating outputs, running evals and auditing agent actions. That work may favor different chips and systems from those used for frontier training. Coordinating multiple models and harnesses can also squeeze more useful work from existing models by choosing the right setup for each task. Lower costs per task can coexist with higher total compute use.

Carlota Perez’s framework ↗ separates infrastructure build-out from widespread use. The curve tracks adoption, not profits. Applied to AI, it suggests that pricing power can move from models and compute to customer context and distribution, with the same company competing across several layers.

Source of pricing power
Capability sets the price

Intelligence scarcity

At the beginning of the cycle, only a handful of companies could produce the strongest models. That gap in capability, and the hardware required to sustain it, supported exceptional pricing power.

How AI is built and usedIntelligence
Typical setup

Closed-cloud frontier; centralized by necessity

The model is the product. Frontier capability is reached through a proprietary cloud API, while open and local models remain useful but materially behind on the tasks that command the highest willingness to pay.

Where AI runs

Large GPU clusters, HBM, interconnect and power; consumer hardware is largely an endpoint, not an inference substrate.

Model deploymentRole in this source of pricing power

Closed cloud

Dominant

The only practical route to frontier capability. The lab owns the weights, the serving stack and the meter.

Open cloud

Emerging

Open weights broaden access, but most serious workloads still require rented cloud hardware and lag the frontier.

Hybrid

Specialist

Some data remains private while difficult reasoning is sent to the cloud; integration is bespoke rather than a default architecture.

Local / edge

Constrained

Small models run on customer hardware, but memory, thermals and capability limit the economically important use cases.

Harness / agent layer

Thin wrappers around scarce models

Market shape: Thousands of applications; very few own the underlying capability

Where the power law sits: The power law sits below the harness: a handful of frontier labs and compute suppliers collect most of the leverage.

Who holds pricing power?

Many wrappers; power stays with model and chip suppliers.

Each bar = a provider · taller = more power

App / harness providers

Illustrative market structure · not measured data

“Harness” means the layer that turns a model into a usable agent: interface, instructions, tools, memory and policy.

Few companies had the research expertise or hardware to train frontier models. Training depended on scarce accelerators, memory, interconnect, power and engineering talent. Model labs charged for access to a capability few others possessed; NVIDIA, TSMC and the hyperscalers charged for the means of producing and serving it.

Those two forms of scarcity need not clear together. The premium attached to model quality can contract while power, memory and data-center capacity remain constrained.2 Intelligence may become substitutable before the infrastructure built to provide it has finished depreciating. Treating all of this as ‘compute’ obscures the most important timing difference in the cycle.

Frontier labs have so far captured most enterprise model spending. Coding shows why: better models can take on more of a large, expensive labor pool’s work. Our hypothesis is that the premium for each new capability gain grows more slowly as cheaper models catch up. Demand for the frontier can still grow wherever better performance makes new work possible, including in defense, cybersecurity and scientific research.

Open models can put downward pressure on prices wherever their capability is adequate. A cheaper alternative makes a closed lab’s premium harder to sustain in that use case, but frontier advantages can reopen.1 Distillation can accelerate the transfer of capability. The commercial territory reserved for frontier access may therefore shrink even while the frontier itself advances.

ScarceFrontier models, chips, interconnect, power
Controlled byLabs + compute substrate
Value accrues toFrontier access and the physical capacity behind it
Where pricing power may settleHover / focus for timing, mechanism and risk

Frontier model labs

OpenAI · Anthropic · Google DeepMind · xAI · Meta · Mistral · DeepSeek · Cohere

Why it matters

Frontier training remains concentrated among organizations with exceptional research talent, capital and cluster access.

How value accrues

They can charge for capabilities that customers cannot reproduce or obtain from a cheaper model. The rent is the performance gap—not intelligence in the abstract.

What compresses the rent

Open models, distillation and diminishing willingness to pay narrow that gap use case by use case.

AI accelerators + systems

NVIDIA · AMD · Google TPU · AWS Trainium · Intel Gaudi · Cerebras · SambaNova

Why it matters

Accelerator supply and integrated systems remain difficult to add quickly, even as model access becomes less scarce.

How value accrues

Training and high-volume inference are bounded by usable compute, not chip counts alone. Vendors that combine silicon, networking and software capture more of the system economics.

What compresses the rent

Custom silicon, better utilization and a shift from training to cheaper inference weaken the premium on general-purpose frontier hardware.

Foundry, packaging + memory

TSMC · Samsung Foundry · SK hynix · Micron · Samsung Memory · ASE Technology

Why it matters

Leading-edge fabrication, advanced packaging and high-bandwidth memory expand on industrial rather than software timelines.

How value accrues

These suppliers control capacity that cannot be replicated with code or purchased at short notice. Their bottlenecks can persist after model quality begins to converge.

What compresses the rent

Committed capacity eventually arrives; when it does, volume may remain high while scarcity margins fall.

Networking + interconnect

Broadcom · Arista Networks · Marvell · NVIDIA Networking · Astera Labs · Cisco

Why it matters

Ever-larger clusters require more bandwidth within and between racks, making network architecture part of model economics.

How value accrues

A costly accelerator is useful only when data can reach it. Suppliers capture value by removing communication bottlenecks that otherwise leave scarce chips idle.

What compresses the rent

Standards, merchant silicon and slower growth in frontier cluster size can transfer value back to buyers.

Power, cooling + electrical systems

Vertiv · Eaton · Schneider Electric · GE Vernova · Siemens Energy · Caterpillar

Why it matters

Grid connections, generation, cooling and power-management equipment have lead times measured in years.

How value accrues

The constraint is increasingly energized capacity rather than server procurement. Suppliers sell into a build-out that cannot proceed without them and is difficult to accelerate.

What compresses the rent

Their order books can outlast peak scarcity, but returns normalize when data-center construction slows or designs become more efficient.

Hyperscale + specialist capacity

AWS · Microsoft Azure · Google Cloud · Oracle Cloud · CoreWeave · Crusoe · Nebius

Why it matters

Few operators can finance, provision and run large clusters across regions with enterprise-grade availability.

How value accrues

They monetize balance sheet, procurement access and operating expertise by turning scarce physical infrastructure into capacity customers can consume immediately.

What compresses the rent

Large fixed costs make this category vulnerable once capacity is interchangeable and utilization becomes the central problem.

Signals of transition

Compare capability gains with changes in willingness to pay, rather than benchmarks alone. Further signs of transition include labs selling outcomes instead of tokens, sovereign buyers supporting a high-cost frontier, and chipmakers backing open ecosystems to protect demand for their hardware.

Source of pricing power
Execution becomes the product

Deployment scarcity

As models become more available, the practical constraint shifts from producing intelligence to delivering it reliably, at the required speed and cost.

How AI is built and usedDeployment
Typical setup

Multi-model cloud; routing becomes the architecture

Models multiply faster than production systems can absorb them. Closed and open models share the cloud; the scarce product is a dependable execution layer that chooses where a request runs and keeps it running.

Where AI runs

Hyperscale and specialist inference fleets dominate, while open models begin to run on developer workstations and private servers.

Model deploymentRole in this source of pricing power

Closed cloud

Strong

Frontier APIs remain the easiest route to difficult tasks, but applications can switch among several credible suppliers.

Open cloud

Rapid growth

Open models are served by competing inference clouds, turning weights into a portable workload rather than a proprietary destination.

Hybrid

Forming

Routers send routine work to cheaper or private models and escalate harder requests to frontier clouds.

Local / edge

Practical

Quantization and better runtimes make local inference credible for development, privacy-sensitive work and bounded tasks.

Harness / agent layer

Framework proliferation; platform consolidation

Market shape: Hundreds of frameworks and libraries; a much smaller set of production runtimes, gateways and sandboxes

Where the power law sits: Moderate and unsettled. Order flow can concentrate in a few gateways, but open standards and low switching costs keep the contest open.

Who holds pricing power?

A few gateways may pull ahead; the market is still open.

Each bar = a provider · taller = more power

Gateways / agent runtimes

Illustrative market structure · not measured data

“Harness” means the layer that turns a model into a usable agent: interface, instructions, tools, memory and policy.

The center of activity moves from training to inference. Customers must choose among models, chips and clouds while managing latency, availability and cost.3 Inference providers, gateways and agent runtimes turn that complexity into a service. Developer distribution matters because the default route can capture order flow even when the underlying suppliers change.

Deployment scarcity may erode especially quickly. Routing makes inference services more commoditized; optimization extracts more work from each chip; open software lowers switching costs. The economics still depend on how much of the rented capacity is put to use.4 The firms solving deployment scarcity are also helping to eliminate it. That does not make them poor businesses, but it does place a burden on them to build workflow, distribution or scale advantages before basic inference becomes a utility.

The physical layer presents a separate risk. Data centers are financed and built in parallel, with each operator planning against similar demand. Once the capital is committed, the incentive is to bid aggressively for utilization. The result has been familiar across railways, canals and fiber: a useful network, followed by disappointing returns for a portion of those who financed it.

The risk is that economic returns deteriorate before GPU assets finish depreciating. Contracted revenue can cushion that adjustment, but renewal prices and uncontracted capacity remain exposed.5 If inference efficiency continues to improve while new capacity arrives, a shift in bargaining power from deployment to context could combine falling unit prices with an overbuilt asset base. Consolidation would then be a consequence of the build-out, not evidence that demand had disappeared.

ScarceServing capacity, routing, latency, reliable execution
Controlled byInference + orchestration layers
Value accrues toReliable throughput across a fragmented supply base
Where pricing power may settleHover / focus for timing, mechanism and risk

Inference + model clouds

Together AI · Fireworks AI · Baseten · Modal · Groq · Cerebras · DeepInfra · Replicate · Cloudflare Workers AI

Why it matters

Supply is fragmented across chips, clouds and model architectures while production demand is rising faster than internal serving expertise.

How value accrues

They capture the spread between raw accelerator time and reliable, model-specific throughput. Optimized kernels, batching and capacity management let them deliver more useful tokens per dollar.

What compresses the rent

That spread compresses as optimization diffuses, capacity grows and customers treat inference providers as interchangeable.

Routing + model gateways

OpenRouter · Portkey · LiteLLM · Cloudflare AI Gateway · Kong · Helicone · Martian

Why it matters

Rapid model releases create meaningful differences in price, latency, capability and availability from one request to the next.

How value accrues

The gateway sees the order flow and can choose the supplier. That position compounds through usage data, policy controls, failover and one integration across many models.

What compresses the rent

Routing succeeds by commoditizing what it routes. Clouds and application platforms can absorb the feature once the decision logic becomes standard.

Open + local inference runtimes

Ollama · vLLM · llama.cpp · SGLang · LM Studio · Open WebUI · BentoML

Why it matters

Capable open models and smaller hardware footprints make private and on-device deployment practical for a growing set of workloads.

How value accrues

These projects set the deployment standard and reduce dependence on proprietary clouds. Commercial value can accrue through managed hosting, enterprise support and control of the local model workflow.

What compresses the rent

Adoption alone does not create a business; agents also reduce the labor premium customers once paid to avoid self-deployment.

Model + developer distribution

Hugging Face · GitHub Models · AWS Bedrock · Vertex AI Model Garden · Azure AI Foundry · Replicate

Why it matters

Developers need discovery, testing, governance and deployment across an unusually fast-changing supplier base.

How value accrues

Platforms capture value by becoming the place a model is first found, evaluated and integrated. Distribution can persist after any individual model loses its premium.

What compresses the rent

The largest clouds can bundle the catalog, while open standards keep migration costs lower than in traditional platform markets.

Agent execution + sandboxes

Daytona · E2B · Temporal · Modal · Browserbase · Fly.io · Cloudflare

Why it matters

Agents are moving from single responses to long-running jobs that use browsers, code, files and external systems.

How value accrues

They sell the controlled environment in which unreliable model output becomes resumable, observable work. Security isolation and state management matter before agents can receive broader permissions.

What compresses the rent

Basic sandboxing and orchestration may consolidate into clouds, operating systems or agent frameworks.

Inference optimization

NVIDIA TensorRT-LLM · Red Hat vLLM · Modular · SambaNova · CentML · BentoML

Why it matters

Unit economics depend heavily on compiler choices, quantization, caching and hardware-aware scheduling.

How value accrues

Small efficiency gains multiply across enormous token volumes. Optimizers capture value while expertise is scarce and serving stacks remain heterogeneous.

What compresses the rent

The economics are self-eroding: successful techniques enter open runtimes, chip libraries and cloud platforms.

Signals of transition

The hand-off is under way when inference capacity is marketed as interchangeable, routing is absorbed into larger platforms, management teams emphasize utilization rather than expansion, and distressed assets begin to consolidate.

Source of pricing power
Context becomes capability

Context scarcity

Once generic capability is widely available, performance depends increasingly on what a system knows about a particular organization, what it is permitted to see and what it is authorized to do.

How AI is built and usedContext
Typical setup

Hybrid by policy; models are selected inside the workflow

No single model topology wins. Workload placement follows context: sensitivity, latency, policy and the value of the task. Hybrid systems become normal because the enterprise boundary matters more than allegiance to one model.

Where AI runs

Public cloud, private cloud, on-premise servers and edge devices operate as one governed estate; locality becomes a permission decision.

Model deploymentRole in this source of pricing power

Closed cloud

Selective

Used where frontier performance justifies cost and data policy permits an external service.

Open cloud

Established

Open weights provide control and portability without requiring the customer to own every piece of serving infrastructure.

Hybrid

Default

Policy-aware routing keeps sensitive context private, uses local or open models for routine work and buys frontier capability when needed.

Local / edge

Strategic

Private inference supports continuous memory, low latency and regulated workloads; hardware ownership becomes part of the trust model.

Harness / agent layer

Domain-specific agents around systems of record

Market shape: Many enterprise and vertical harnesses; a small number can dominate each workflow or industry

Where the power law sits: The power law fragments by domain. The winners are attached to live data, identity and outcomes—not necessarily to the most popular base model.

Who holds pricing power?

Leaders emerge within each workflow, not across all agents.

Each bar = a provider · taller = more power

Providers in one workflow

Illustrative market structure · not measured data

“Harness” means the layer that turns a model into a usable agent: interface, instructions, tools, memory and policy.

Context is more than data retrieval. It includes current state, institutional memory, identity, policy, secrets and the authority to take a consequential action. A model may know how to approve an invoice; it still needs to know which invoice, under whose policy and with whose liability.

This constraint clears at institutional speed. Security reviews, procurement, data migration and regulation do not improve at the rate of model inference. Context scarcity should therefore last longer than deployment scarcity and favor businesses already embedded in systems of record and operational workflows.

The adjective ‘proprietary’ is not enough to make a dataset defensible. Static information can be exported, reproduced or lose relevance. A stronger position is a live feedback loop: software participates in the work, observes the result, improves from it and earns broader permission over time. The scarce asset is not simply knowledge. It is authorized participation.

The next shift is conditional. If customers can carry their records and permissions into another service, a trusted interface may choose among competing workflows. If context cannot travel safely, the workflow remains hard to replace. Ownership of the customer relationship only becomes the stronger advantage when it brings a credible ability to redirect work.

This also explains why open-source software can support valuable companies. Enterprises do not pay only for access to code. They pay for integration, accountability and the transfer of operational risk. Agents may reduce the effort required to deploy software, but they do not eliminate the need for someone to be responsible when it fails.

ScarceCurrent data, identity, permissions, evaluation
Controlled bySystems embedded in proprietary workflows
Value accrues toAccess to live context and permission to act on it
Where pricing power may settleHover / focus for timing, mechanism and risk

Systems of record + workflow

Salesforce · ServiceNow · SAP · Oracle · Workday · Epic · Guidewire · Atlassian

Why it matters

Enterprises can access capable models before they can safely connect them to the systems where consequential work is recorded.

How value accrues

These incumbents already hold current state, permissions and workflow. They can place agents inside an approved operating environment rather than asking customers to reconstruct one elsewhere.

What compresses the rent

Possession of data is not enough if a new interface owns the user and treats the system of record as a replaceable back end.

Data + retrieval platforms

Databricks · Snowflake · MongoDB · Confluent · Elastic · Pinecone · Weaviate · Redis

Why it matters

Generic intelligence is abundant, but useful answers depend on fresh, permission-aware access to proprietary state.

How value accrues

They sit between models and changing enterprise data. Value comes from governed access, lineage and real-time retrieval—not from attaching the word ‘AI’ to stored data.

What compresses the rent

Static datasets move easily; durable value requires staying inside the live read-and-write path of the workflow.

Connectors + tool protocols

Model Context Protocol · Zapier · Workato · Boomi · MuleSoft · Merge · Pipedream · Composio

Why it matters

The immediate obstacle is not reasoning but the fragmented work required to connect agents to thousands of software tools and data sources.

How value accrues

Connection layers can become the authorized route through which agents discover and use enterprise tools. Breadth, maintenance and policy enforcement create the commercial advantage.

What compresses the rent

Open protocols reduce lock-in, and systems of record may expose their own first-party agent interfaces.

Identity, secrets + permission

Okta · Auth0 · Microsoft Entra · CyberArk · HashiCorp Vault · Infisical · Aembit

Why it matters

Agents need identities distinct from their human sponsors and permissions bounded by task, time and resource.

How value accrues

Autonomy cannot expand without a control plane that can authenticate an agent, limit its authority and revoke it. Existing trust relationships and integrations are difficult to reproduce quickly.

What compresses the rent

Identity suites can bundle agent controls; point solutions must own a genuinely new policy layer rather than a renamed service account.

Evaluation + observability

LangSmith · Braintrust · Arize · Langfuse · Galileo · Patronus AI · Humanloop

Why it matters

Non-deterministic agents are reaching production before organizations have adequate evidence of how they behave.

How value accrues

These platforms turn traces and outcomes into tests, audit records and feedback. That evidence lets customers grant more autonomy without accepting unknowable risk.

What compresses the rent

Tracing is easy to commoditize; durable value depends on proprietary evaluation data, workflow integration or becoming part of the approval process.

Agent security + governance

Credo AI · HiddenLayer · Protect AI · Noma Security · Lakera · Robust Intelligence

Why it matters

Prompt injection, tool misuse and data leakage become budgeted problems once agents can take actions rather than merely generate text.

How value accrues

They capture value by reducing the expected cost of autonomy and by producing the controls regulators, insurers and security teams require before deployment.

What compresses the rent

Security platforms can absorb these capabilities, and fear alone will not sustain a category without measurable risk reduction.

Vertical workflow owners

Harvey · Abridge · Glean · Sierra · Decagon · Hebbia · Cursor · Ambience Healthcare

Why it matters

Narrow domains can support agents earlier because the workflow, vocabulary, data and acceptable outcome are more clearly defined.

How value accrues

They capture value by owning the completed task and the feedback it generates. Deep integration and domain liability matter more than access to any one model.

What compresses the rent

A vertical product is vulnerable if customers can reproduce the workflow with a general assistant and a small number of connectors.

Data + environment suppliers

Scale AI · Turing · Mercor · Surge AI · Handshake · AfterQuery · Invisible Technologies

Why it matters

Continuous learning requires specialized human feedback, task environments and outcome data after the easy public corpus has been consumed.

How value accrues

They organize scarce expertise and generate the environments in which models learn to perform economically useful work. The value lies in repeatable production, not a one-off dataset.

What compresses the rent

Margins fall when data production is undifferentiated or when customers internalize the workflow; privileged supply and quality control must persist.

Signals of transition

Look for agent identity to separate from human identity, permissions to become specific to a task and duration, and evaluation to move from demonstrations to replayable evidence. For a shift toward distribution, watch whether customers can change service providers while retaining context, permissions and audit history. If those remain locked to one workflow, its bargaining power persists.

Source of pricing power
From attention to delegation

Distribution scarcity

Distribution matters from the start. It becomes more powerful when a trusted interface can choose among several capable suppliers on the customer’s behalf. That shift requires alternatives and permission to act. It does not require free intelligence or local computing.

How AI is built and usedDistribution
Typical setup

Customer-led service, flexible computing

A trusted interface receives the request and selects services behind it. It may use paid cloud models, open models or local inference. The defining feature is authority over supplier choice, not a particular computing architecture.

Where AI runs

Cloud, private servers or customer devices, depending on capability, cost and privacy. Local execution is optional.

Model deploymentRole in this source of pricing power

Closed cloud

Premium input

Frontier models remain valuable for difficult work, but increasingly appear as an invisible supplier inside someone else’s service.

Open cloud

Alternative supply

Open models can give the interface more supplier choice. They need not be free or replace the frontier for every task.

Hybrid

Possible mix

The agent decides whether work stays on the device, moves to a private service or purchases frontier reasoning. The user need not see the choice.

Local / edge

Possible route

Capable local models could support private memory and low-latency tasks. Distribution can remain valuable even if most work stays in the cloud.

Harness / agent layer

A few trusted agents above a vast market of skills

Market shape: Single-digit daily agent relationships per user; potentially millions of downstream services, tools and specialist agents

Where the power law sits: A barbell power law: winner-take-most at the customer interface, extreme fragmentation among suppliers behind it.

Who holds pricing power?

A few agents own demand; many suppliers compete behind them.

Each bar = a provider · taller = more power

Customer-facing agents

Downstream tools / services

Illustrative market structure · not measured data

“Harness” means the layer that turns a model into a usable agent: interface, instructions, tools, memory and policy.

A workflow owner knows how to complete a task inside its system. A customer interface receives the broader request and may decide which systems do the work. The balance changes when reliable connections, portable context and compatible permissions let that interface replace a supplier without forcing the customer to rebuild the process. The workflow is still necessary, but it has less control over the next purchase.

Three conditions make that hand-off plausible: several suppliers can perform the task well enough; the interface can access the context and authority needed to use them; and customers trust it to choose. Without those conditions, distribution may deliver leads while the workflow owner still sets the terms. A large audience alone does not transfer bargaining power.

Context and distribution often belong to the same company. An accounting platform could keep both the records and the assistant that receives the request. In that case, the advantage deepens inside the incumbent rather than migrating to a new firm. The test is who can replace whom: can the assistant switch workflows while keeping the customer, or can the workflow replace the assistant without losing demand?

Expensive cloud inference is compatible with this argument. An assistant can own the customer relationship while purchasing every model call. If it can switch model suppliers, it has negotiating power; if only one model can do the job, that supplier can still claim a large share of the economics. Distribution can gain power without eliminating intelligence scarcity.

Local-first computing is one possible route, not a necessary condition. Capable models on customer hardware could lower serving costs, keep more data private and support continuous use. Paid cloud models or a hybrid architecture could support the same customer relationship. Where inference runs affects cost and privacy; it does not by itself determine who owns demand.

Incumbents have formidable advantages here: installed distribution, capital, data and default placement. Their victory is not automatic. Rebuilding a product around a new interface is harder than adding AI features to an existing one; Microsoft owned GitHub yet did not monopolize coding agents. Platform shifts have a long record of making secure positions look less secure in retrospect.

Consumer demand is also likely to divide. Convenience products will accept extensive automation because the desired outcome is the service itself. Leisure, culture and status operate differently. Authorship and human participation remain part of what is being consumed. As synthetic production becomes abundant, verified human origin may command a premium precisely because it is scarce.

The likely result is a barbell: inexpensive, automated utility at one end and scarce human craft or prestige at the other. The middle—competent but undifferentiated production—faces the greatest pressure. Enterprise buyers are less sentimental, except when the reputation of the provider is itself part of the risk decision.

ScarceCustomer trust and authority to choose suppliers
Controlled byDefault interfaces + direct relationships
Value accrues toControl of the customer relationship and the allocation of demand
Where pricing power may settleHover / focus for timing, mechanism and risk

Operating systems + devices

Apple · Google · Microsoft · Meta · Samsung · Amazon · emerging AI-native hardware

Why it matters

Capability differences become less visible to mainstream users while assistants are embedded into phones, PCs, glasses and home devices.

How value accrues

The operating surface receives intent before a model is selected. Defaults, hardware integration and installed base let the platform route demand and make upstream suppliers invisible.

What compresses the rent

A default is valuable only if people use it. Weak products can still cede the relationship to an application that users actively choose.

General-purpose assistants

ChatGPT · Claude · Gemini · Microsoft Copilot · Meta AI · Perplexity

Why it matters

Assistants are building direct consumer habits before the underlying models become fully substitutable.

How value accrues

Repeated use creates memory, preference data, identity and trust. If users delegate decisions to the assistant, it becomes the buyer of downstream models and services rather than a reseller of answers.

What compresses the rent

The relationship can be displaced by the operating system, browser or specialist product where the user’s intent actually originates.

Work surfaces

Microsoft 365 + Teams · Google Workspace · Slack · Notion · Zoom · Atlassian · Canva

Why it matters

Enterprise agents first appear inside products employees already open each day and that already carry organizational identity.

How value accrues

These products own both attention and workflow context. They can choose the model, surface the agent at the moment of intent and retain the feedback generated by completed work.

What compresses the rent

They lose leverage if the agent becomes the primary interface and reduces the underlying application to a tool call.

Search + browsers

Google Search + Chrome · Safari · Microsoft Edge · Perplexity · Dia · Brave

Why it matters

Research, navigation and commercial discovery are being compressed into conversational and agentic interfaces.

How value accrues

Browsers and search products observe high-intent queries at the point where users choose what to read, buy or do. That position can direct traffic, transactions and model selection.

What compresses the rent

Answer interfaces can weaken the open-web supply that makes them useful, while operating systems can move discovery one layer higher.

Commerce + payment rails

Amazon · Shopify · Stripe · Visa · Mastercard · Instacart · DoorDash · Booking Holdings

Why it matters

Delegation becomes economically consequential when agents can compare, purchase and settle on a user’s behalf.

How value accrues

These networks combine merchant access, identity, payment credentials and dispute handling. An agent needs all four to convert a recommendation into a trusted transaction.

What compresses the rent

Standards may separate the agent relationship from fulfillment and payment, forcing each layer to compete on price.

Vertical consumer relationships

Abridge + Epic · Duolingo + Khan Academy · Intuit · Harvey · WhatsApp · travel and health incumbents

Why it matters

Users delegate sooner in domains where a product already understands the task, history and acceptable boundary of action.

How value accrues

Domain trust and longitudinal context make the product a plausible representative rather than a generic interface. The strongest positions own a recurring decision, not a single query.

What compresses the rent

The relationship is less durable where a general assistant can import the history and reproduce the service without regulatory or workflow friction.

Local + personal agent ecosystems

Ollama · LM Studio · Open WebUI · on-device Apple + Android models · personal-agent entrants

Why it matters

Capable local models make it possible to keep private context and continuous memory on hardware the user controls.

How value accrues

If local models become capable and economical enough, applications could use private context without a metered cloud call for every task. A personal agent could combine local inference with paid cloud services. This is one possible architecture for owning the customer relationship, not a prerequisite.

What compresses the rent

Convenience may outweigh sovereignty for most consumers, leaving the category influential technically but difficult to monetize directly.

AI-native services + software

Cursor · Harvey · Abridge · Sierra · Decagon · Replit · new consumer and small-business services

Why it matters

Falling inference costs allow products to deliver completed work in markets that could not support the cost of a human service or repeated frontier-model calls.

How value accrues

These companies can turn labor into a repeatable software product: research completed, claims processed, customers served or code shipped. Cheap intelligence creates the supply; workflow ownership, brand and distribution capture the margin by packaging it into an outcome someone will pay for.

What compresses the rent

Cheaper intelligence can lower the barrier for competitors too. Without proprietary workflow, trust, data or a low-cost route to customers, the service becomes as interchangeable as the model beneath it.

Authored media + scarce attention

YouTube · TikTok · Spotify · Substack · Patreon · Live Nation · creator and luxury brands

Why it matters

Synthetic production becomes abundant while human attention remains fixed and verified authorship becomes easier to value by contrast.

How value accrues

Platforms and brands capture value by aggregating scarce audiences, identity and cultural signal. Utility content can automate; participation, status and human provenance cannot be manufactured at the same marginal cost.

What compresses the rent

The undifferentiated middle is exposed. Only products with genuine audience ownership, trusted curation or scarce human origin retain the premium.

Signals of transition

The transition becomes visible when model brands recede behind product brands, agents transact with other agents, and changes in defaults move demand more than changes in benchmarks. In consumer markets, watch automated utility separate from work whose value depends on human authorship.

The investment question

What remains scarce after the build-out?

The shifts in bargaining power are conditional. When models become more commoditized, reliable execution can command a premium. When execution becomes easier to buy, authorized context can matter more. A trusted interface gains power over those workflows only if it can access their context, choose alternatives and keep the customer. These advantages may accumulate in one company, and the hand-off may never happen in some markets.

This is why the hand-off from deployment to context deserves particular attention. It may coincide with the largest financial risk in the cycle: inference prices falling, infrastructure still arriving and enterprise demand constrained by permissions rather than capacity. The assets will remain useful. Their owners may not all earn the returns they expected.

WHAT WOULD MAKE THIS WRONG?
  • Model quality keeps widening willingness to pay faster than open alternatives narrow it.
  • Context and permissions remain tied to indispensable workflows, so interfaces cannot redirect demand.
  • Agents never earn enough trust to receive meaningful delegated authority.
  • Incumbents absorb the interface shift without surrendering their old economics or their distribution.
  • Demand for intelligence proves effectively unbounded, keeping compute scarce despite the build-out.

Look beyond the best model. Ask what customers pay extra for today, how long that advantage will last, and what will keep them from switching.

Footnotes

Sources and calculations behind the economic claims. Prices are dated snapshots in USD; worked examples use stated assumptions.

  1. Cheaper at a given capability

    Stanford reports that inference prices at GPT-3.5-level performance fell over 280× between November 2022 and October 2024. That measures the price of reaching a capability threshold, not the expense of every frontier workload. Its 2026 report also finds that the open–closed model performance gap reopened in 2025, with closed models ahead as of March 2026. Cheaper adequate intelligence can put pressure on prices even while a frontier premium survives; these observations do not establish an irreversible loss of pricing power.

  2. Two different build-out clocks

    The IEA’s Electricity 2026 report puts planning, permitting and building new grid infrastructure at 5–15 years, compared with 1–3 years for a data center. This mismatch helps explain why physical constraints can outlast changes in model availability. The ranges describe infrastructure projects, not a universal wait for every grid connection. They support the timing distinction; they do not establish that all regions will face shortages, or that an eventual capacity glut is inevitable.

    Verified physical-scarcity discussion
  3. What one model call costs

    Anthropic lists Claude Sonnet 5 at $2 per million input tokens and $10 per million output tokens on its standard API. For an illustrative invoice attempt using 5,000 input and 1,000 output tokens:

    (5,000 × $2 + 1,000 × $10) / 1,000,000 = $0.02 per attempt
    $20 per 1,000 attempts

    This is our calculation using assumed token counts, not an invoice benchmark or a cost per successful payment. It excludes caching and batch discounts, taxes, retries, tools and human review. The relevant service metric is total expense divided by successfully completed tasks; API prices do not reveal the model provider’s production costs.

    Verified model-call cost discussion
  4. Idle time changes the unit cost

    Lambda lists a single H100 SXM, 80 GB, at $4.29/hour. Suppose it remains rented continuously, including idle time. Treat utilization here as the share of rented hours spent productively serving work:

    $4.29 / 25% = $17.16
    $4.29 / 75% = $5.72
    Rental expense per productive GPU-hour

    The utilization levels are illustrative assumptions, not measured fleet statistics. This calculation shows why filling rented capacity matters; it is not a measure of provider margins. It excludes taxes and additional operating costs. Translating GPU-hours into token costs would require measured throughput for a specific model, workload and latency target.

    Verified deployment economics
  5. Accounting life and economic life

    CoreWeave’s 2025 annual filing assigns technology equipment a six-year accounting life. It also describes committed contracts as take-or-pay: customers owe payment regardless of utilization. Those commitments can protect contracted revenue from low customer usage. The risk is that achievable prices and returns deteriorate before assets finish depreciating, particularly on uncontracted capacity or at renewal. An accounting estimate does not establish when a GPU becomes uneconomic, and this filing does not prove that an industry-wide write-down is coming.

    Verified GPU depreciation discussion
  6. Efficiency, demand and a possible slowdown

    The IEA’s 2026 assessment links AI energy demand to “improvements in efficiency, surging uptake, and changing model capabilities.” It also identifies power, chip and financing constraints on expansion. Cheaper tasks need not mean lower total demand; energy consumption does not establish industry revenue or profitability. The effects of a coordinated frontier slowdown described here are our scenario analysis, assuming deployment and optimization continue. The report does not model that agreement.

    Verified frontier slowdown discussion
Framework session / 24 July · Draft 0.4A working thesis, not investment adviceTop ↑