The Great Open Source AI Acceleration: Commoditised Intelligence, Cost Collapses, Demand Inflection, and Portfolio Value Creation
Earlier this year we argued, from first principles, that foundation models would be commoditised over time, and that value would accrue to the layers around them.
At the time, this was largely a theoretical framework. The prevailing narrative suggested that a handful of closed frontier model labs would maintain durable, tollbooth style pricing power. Yet over recent quarters, the capability and rapid adoption of high performance open source (open weight) models has accelerated faster than many expected.
To understand what this means for investors, we need to reframe how we look at open source AI. This is not about declaring winners and losers in a zero sum battle, nor is it a prediction that proprietary pioneers like OpenAI or Anthropic will simply disappear. Those businesses may well still capture enormous inference workloads and margin dollars, especially in premium enterprise use cases. The more useful question is not who loses at the frontier, but where the margin pool moves as the cost of a unit of intelligence falls and demand inflects.
Our answer, in plain terms: open source and cheaper models do not shrink demand for hardware or infrastructure. They shift the increasing pool of margin dollars away from the frontier model layer, where gross margins have been estimated at around 80 to 95%, and down into hardware, infrastructure, and the SaaS platforms that turn those cheaper tokens into workflows.
Elasticity and Compute Demand
A common misconception in markets, vividly illustrated during the "DeepSeek moment" when low cost reasoning models triggered a temporary sell off across tech stocks, is that cheaper, open source models will reduce the overall need for AI infrastructure.
At a physical level, generating a token on an open weight model requires the same underlying compute as generating a token on a closed model of comparable parameter scale and architecture. Compute demand does not vanish simply because the model weights are free.
What changes fundamentally is the end user pricing and the resulting economic elasticity:
-
Margin Re-allocation: Closed frontier APIs historically carried estimated gross margins of 80% to 95% at the model software layer. As open source architectures offer comparable reasoning at near zero software markup, those gross margin dollars are effectively redistributed down the stack to the infrastructure layer (hardware accelerators and cloud hyperscalers) and up to the application layer (SaaS).
- Jevons Paradox in Action: When the unit cost per token plummets, total consumption does not stay flat, it grows. Lower costs transform experimental chatbots into continuous, autonomous agentic workflows where a single enterprise task can consume millions of background tokens.
Real World Validation: Dynamic Routing in Enterprise Software and Platforms
We are already seeing enterprise software platforms adapt to this reality by becoming strictly model agnostic. Rather than locking into a single proprietary vendor, forward thinking platforms are deploying intelligent routers that direct queries to the lowest cost model capable of completing the task.
Recent commentary from enterprise platforms highlights this transition:
Block: Highlighted on recent earnings calls that its platform is built to dynamically route workloads to the most cost effective intelligence available. This architecture ensures that as underlying model costs fall, the margin benefits flow directly to Block rather than being captured by a third party model vendor.
Uber: Has detailed how it pairs open source models with domain specific retrieval augmented generation (RAG) and internal fine tuning. For high volume workloads like dispatch support, eater recommendations, and code assistance, open models deliver frontier grade results at a fraction of closed API expense.
ServiceNow: Has positioned its platform as an agnostic AI orchestration layer using its Generative AI Controller and AI Control Tower framework. By dynamically routing tasks between its native domain models (Now LLM) and external foundation models (such as Azure OpenAI, Anthropic Claude, or Google Gemini), ServiceNow protects its gross margins while offering enterprise customers predictable, high ROI AI consumption.
This approach extends far beyond these examples into the broader enterprise software stack. Salesforce (Agentforce) and Shopify are scaling similar model agnostic architectures, positioning their platforms as the context, guardrail, and execution layers that connect enterprise data to any underlying model. By insulating themselves from model layer compute risks, these software platforms aim to capture the economic surplus of agentic workflows while maintaining their own software gross margins.
Unlocking Agentic AI: How Lower Costs Expand Enterprise ROI and Accelerate Demand
By lowering the cost of "agentic work units," these platforms can roll out AI driven products with compelling, undeniable ROI for their enterprise customers, driving adoption without compromising their own software gross margins.
When software applications deploy multi turn, autonomous agentic workflows, where a single user request triggers dozens of background reasoning, data retrieval, code execution, and verification loops, token consumption grows by orders of magnitude.
The structural collapse in token costs, driven by high performance open weight models and intelligent dynamic routing, fundamentally alters the adoption curve for enterprise automation:
Unlocking Previously Uneconomic Workflows: As the execution cost of an automated "agentic work unit" drops to a fraction of what it previously cost, vast categories of routine, multi step enterprise tasks suddenly clear the hurdle for compelling, positive ROI.
Massive Volume Expansion: Lower friction accelerates enterprise deployment from small, cautious pilots into site wide rollouts across entire workforces, driving exponential growth in the aggregate volume of agentic work units performed daily.
Accrual to Mission-Critical SaaS: Because these agentic units operate inside mission critical enterprise platforms where proprietary data, governance, and core business workflows already reside, incumbent SaaS leaders are uniquely positioned to capture this demand surge. With the shift in commercial models away from seat based to hybrid or unit based, they are able to monetise powerful agentic capabilities to drive top line ARR expansion with highly favourable unit economics.
-
Efficient Architectural Orchestration: Enterprise platforms maximise this ROI by deploying intelligent routing layers, directing high volume, routine steps to ultra cheap open models, mid tier tasks to open Mixture-of-Experts (MoE) architectures, and reserving closed frontier APIs for specialised, high complexity reasoning.
Accelerated Feature Velocity and R&D Efficiency: Lower model execution costs and high capability open weights don't just benefit end users, they enable SaaS platforms to build, test, and ship new complex features internally at significantly higher speeds and lower R&D overhead.
Corporate Earnings Validation: This transition is no longer theoretical; it is visible in recent enterprise earnings results. Platform leaders like ServiceNow have reported a 9x increase in customers running agentic AI in production over a nine-month span, while Workday has seen agentic AI new ACV grow over 200% year-over-year across more than 4,000 active customer deployments. By staying model agnostic and routing work to lower cost open models, these platforms are keeping software execution costs low while scaling high ROI automation for their users, capturing substantial market share in enterprise automation.
Lower token costs act as the primary growth engine for enterprise AI, expanding the universe of viable workflows, driving immense demand for agentic work units, and accelerating top line value creation across the application layer.
Portfolio Implications: Mapping the Beneficiaries
Lowering the unit cost of intelligence at the foundation layer redistributes long term value creation across three interconnected tiers within our portfolio:
Application layer (SaaS)
As covered above, SaaS platforms are the direct beneficiaries of falling agentic execution costs. Core portfolio holdings like ServiceNow, WiseTech Global and TechnologyOne sit squarely in this position. As established systems of record with proprietary workflow data and governed distribution, they are ideally placed to convert cheaper tokens into higher volumes of automated work units and ARR expansion with highly favourable unit economics. Falling token costs are the primary growth engine for application software, not a headwind.
Cloud hyperscalers (Microsoft Azure, AWS, Google Cloud)
Lower token prices expand the total workload pool hyperscalers can serve, driving demand across both closed APIs and open weight models. Enterprises are not going to download raw open weight models and run them independently. Instead, they rely on cloud providers to run governed, compliant and monitored instances with production ready enterprise tooling. Hyperscalers capture this expanding pool through three primary levers:
Governed hosting. Cloud platforms sit directly between open weight models and the enterprise, providing essential compliance, monitoring, security and identity management.
High margin adjacent services. The deployment of these workloads pulls through high margin enterprise data, retrieval, security and observability services, which often exceed the raw inference compute bill.
Pool expansion over substitution. Every dollar spent on model intelligence that moves away from closed lab APIs converts into hyperscaler inference compute plus adjacent service revenue. Rather than substituting spend, lower prices at the model layer simply expand the total volume of enterprise workloads running on public cloud infrastructure.
Hardware infrastructure (Nvidia and Micron)
The open weight expansion accelerates the ongoing structural shift from training dominated compute toward inference dominated compute. Because agentic inference workloads run continuously in the background as enterprise software usage scales, global demand for high performance GPUs, CPUs and high bandwidth memory (HBM) expands.
GPUs at the centre: Every additional agentic loop, every workflow that clears its new ROI hurdle, ultimately lands as GPU hours. That is why NVIDIA, despite the noise around cheaper open models, is a direct beneficiary of the shift rather than a casualty of it.
The memory bottleneck: Modern inference is bandwidth bound as much as compute bound. Long context windows, retrieval augmented generation, and agentic loops all lean heavily on High Bandwidth Memory. Every incremental data centre GPU ships with 8 to 12 stacks of HBM, which is why Micron is one of the cleanest secondary plays on inference volume growth.
Fund Strategy
While broader market narrative continues to debate whether proprietary model labs can sustain defensible competitive moats, our first principles framework views the commoditisation of foundation models as a powerful growth engine for the broader ecosystem.
By shifting toward open weight architectures and intelligent routing, the industry is collapsing the cost of reasoning. This cost collapse is not shrinking the market; it is unlocking vast categories of previously uneconomic workflows, driving exponential volume in agentic work units, and catalysing real world enterprise adoption.
Additional validation comes from the incumbents themselves. In July, Jensen Huang (NVIDIA) and Satya Nadella (Microsoft) both amplified an open letter urging Washington to protect open weight models. While framed around enhancing American competitiveness, the commercial motivation underlying this is clear: a thriving open source ecosystem accelerates inference volume on NVIDIA silicon and drives enterprise hosting onto Azure. When they actively champion open source models, it confirms where they see economic value accruing.
Within our fund strategy, we maintain high conviction positioning across the core beneficiaries of this shift:
Hardware & Memory Leaders (NVIDIA, Micron): Capitalising on the structural surge in global inference volume and the memory bandwidth demands of open architectures.
Cloud Hyperscalers: Capturing persistent compute hosting and orchestration revenue as enterprise workloads migrate onto private cloud infrastructure.
Software and Platform Leaders: Leveraging open architectures to deliver high ROI agentic capabilities, accelerating top line revenue growth while protecting software gross margin economics.
Disclosures: Nvidia (NVDA), Micron (MU), Uber (UBER), Block (XYZ), ServiceNow (NOW), WiseTech Global (WTC) and TechnologyOne (TNE) are current holdings across the ELM Australian and Global funds.
3 topics
6 stocks mentioned