The great open source AI acceleration

Open source models and dynamic routing are pushing token costs lower while driving demand for hardware, hyperscalers, and SaaS agentic AI.
Jai Mirchandani

ELM Responsible Investments

Earlier this year we argued, from first principles, that foundation models would be commoditised over time, and that value would accrue to the layers around them.

At the time, this was largely a theoretical framework. The market still expected a handful of closed frontier labs to retain most of the pricing power. Yet over recent quarters, the capability and adoption of high performance open source (open weight) models has accelerated faster than many expected.

To understand what this means for us as investors, we need to reframe how we look at open source AI. This is not about declaring winners and losers in a zero sum battle, nor is it a prediction that proprietary pioneers like OpenAI or Anthropic will simply disappear. Those businesses may well still capture enormous inference workloads and margin dollars, especially in premium enterprise use cases. The more useful question is not who loses at the frontier, but where the margin pool moves as the cost of a unit of intelligence falls and demand inflects.

Our answer, in plain terms: open source and cheaper models do not shrink demand for hardware or infrastructure. They shift the increasing pool of margin dollars away from the frontier model layer, where gross margins have been estimated at around 80 to 95%, and down into hardware, infrastructure, and the SaaS platforms that turn those cheaper tokens into workflows.

Cheaper Intelligence Drives More Demand

A common misconception in markets, vividly illustrated during the "DeepSeek moment" when low cost reasoning models triggered a temporary sell off across tech stocks, is that cheaper, open source models will reduce the overall need for AI infrastructure.

At a physical level, generating a token on an open weight model requires the same underlying compute as generating a token on a closed model of comparable parameter scale and architecture. Compute demand does not vanish simply because the model weights are free.

What changes fundamentally is the end user pricing and the elasticity effect:

  • Where the margin goes: Closed frontier APIs historically carried estimated gross margins of 80% to 95% at the model software layer. As open source architectures offer comparable reasoning at near zero software markup, those gross margin dollars are effectively redistributed down the stack to the infrastructure layer (hardware accelerators and cloud hyperscalers) and up to the application layer (SaaS).
  • Jevons Paradox (lower cost -> higher demand): When tokens become dramatically cheaper, people don't simply spend less, they use far more of them. Lower costs transform experimental chatbots into continuous agentic workflows where a single enterprise task can consume millions of background tokens.
OpenRouter acts as the routing layer, allowing applications to dynamically route tasks to the lowest cost model available. As OpenRouter token prices fell,  adoption of dynamic routing caused OpenRouter token volume to expand dramatically, demonstrating that falling unit costs have unlocked higher aggregate token demand.
OpenRouter acts as the routing layer, allowing applications to dynamically route tasks to the lowest cost model available. As OpenRouter token prices fell, adoption of dynamic routing caused OpenRouter token volume to expand dramatically, demonstrating that falling unit costs have unlocked higher aggregate token demand.

Dynamic Routing in Enterprise Software and Platforms

We are already seeing enterprise software platforms adapt to this reality by becoming model agnostic. Rather than locking into a single proprietary vendor, platforms are deploying intelligent routers that direct queries to the lowest cost model capable of completing the task.

Recent commentary from enterprise platforms highlights this transition:

Block: Highlighted on recent earnings calls that its platform is built to dynamically route workloads to the most cost effective intelligence available. This architecture ensures that as underlying model costs fall, the margin benefits flow directly to Block rather than being captured by a third party model vendor.

Uber: Has detailed how it pairs open source models with domain specific retrieval augmented generation (RAG) and internal fine tuning. For high volume workloads like dispatch support, eater recommendations, and code assistance, open models deliver results at a fraction of closed API expense.

ServiceNow: Has positioned its platform as an agnostic AI orchestration layer using its Generative AI Controller and AI Control Tower framework. By dynamically routing tasks between its native domain models (Now LLM) and external foundation models (such as Azure OpenAI, Anthropic Claude, or Google Gemini), ServiceNow protects its gross margins while offering enterprise customers predictable AI consumption.

This approach extends beyond these examples into the broader enterprise software stack. Salesforce (Agentforce) and Shopify are scaling similar model agnostic architectures, positioning their platforms as the context, guardrail, and execution layers that connect enterprise data to any underlying model. By insulating themselves from model layer compute risks, these software platforms aim to capture part of the value created by agentic workflows while maintaining their own software gross margins.

How cheaper tokens unlock agentic AI

By lowering the cost of agentic tasks, these platforms can roll out AI driven products with compelling ROI for their enterprise customers, driving adoption without compromising their own software gross margins.

When software applications deploy multi turn, autonomous agentic workflows, where a single user request triggers dozens of background reasoning, data retrieval, code execution, and verification loops, token consumption grows by orders of magnitude.

Massive Structural Demand Shift: Goldman Sachs Research projects global token consumption to surge 24x by 2030. The expansion is driven by autonomous enterprise and consumer agents.
Massive Structural Demand Shift: Goldman Sachs Research projects global token consumption to surge 24x by 2030. The expansion is driven by autonomous enterprise and consumer agents.

The structural collapse in token costs, driven by high performance open weight models and  dynamic routing, fundamentally alters the adoption curve for enterprise automation:

  • Unlocking Previously Uneconomic Workflows: As the execution cost of an automated agentic task drops to a fraction of what it previously cost, categories of routine, multi step enterprise tasks suddenly clear the hurdle for positive ROI.
  • Massive Volume Expansion: Lower friction accelerates enterprise deployment from cautious pilots into site wide rollouts across workforces, driving growth in the volume of agentic tasks performed daily.
  • Accrual to Mission-Critical SaaS: Because these agentic tasks operate inside mission critical enterprise platforms where proprietary data, governance, and core business workflows already reside, incumbent SaaS leaders are positioned to capture this demand surge. With the shift in commercial models away from seat based to hybrid or unit based, they are able to monetise agentic capabilities to drive top line ARR expansion with favourable unit economics.
  • Efficient Architectural Orchestration: Enterprise platforms maximise this ROI by deploying intelligent routing layers, directing high volume routine steps to cheap open models, mid tier tasks to open Mixture-of-Experts (MoE) architectures, and reserving closed frontier APIs for high complexity reasoning.
  • Accelerated Feature Velocity and R&D Efficiency: Lower model execution costs and high capability open weights don't just benefit end users, they enable SaaS platforms to build, test, and ship new features internally at higher speeds and lower R&D overhead.
  • What earnings reports are showing: This transition is no longer theoretical; it is visible in recent enterprise earnings results. Platform leaders like ServiceNow have reported a 9x increase in customers running agentic AI in production over a nine-month span, while Workday has seen agentic AI new ACV grow over 200% year-over-year across more than 4,000 active customer deployments. By staying model agnostic and routing work to lower cost open models, these platforms are keeping software execution costs low while scaling automation for their users, capturing market share in enterprise automation.

Lower token costs act as the primary growth engine for enterprise AI, expanding the universe of viable workflows, driving immense demand for agentic work units, and accelerating top line value creation across the application layer.

OpenAI's 30 July 2026 price cut on GPT-5.6 Luna provides a great example of this demand curve. By lowering input costs by ~90%, daily token consumption surged from tens of billions to nearly 800 billion tokens in under two weeks. For context, Luna’s share of total token usage on OpenRouter jumped from a fraction of a percent to over 7% in the same window. 
OpenAI's 30 July 2026 price cut on GPT-5.6 Luna provides a great example of this demand curve. By lowering input costs by ~90%, daily token consumption surged from tens of billions to nearly 800 billion tokens in under two weeks. For context, Luna’s share of total token usage on OpenRouter jumped from a fraction of a percent to over 7% in the same window. 

Portfolio Implications

Lowering the unit cost of intelligence at the foundation layer redistributes long term value creation across three tiers within our portfolio:

Application layer (SaaS)

As covered above, SaaS platforms are the direct beneficiaries of falling agentic execution costs. Core portfolio holdings like ServiceNow, WiseTech Global and TechnologyOne sit squarely in this position. As established systems of record with proprietary workflow data and governed distribution, they are placed to convert cheaper tokens into higher volumes of agentic tasks and ARR expansion with favourable unit economics. Falling token costs are the primary growth engine for application software, not a headwind.

Cloud hyperscalers (Microsoft Azure, AWS, Google Cloud)

Lower token prices expand the total workload pool hyperscalers can serve, driving demand across both closed APIs and open weight models. Enterprises are not going to download raw open weight models and run them independently. Instead, they rely on cloud providers to run governed, compliant and monitored instances with production ready enterprise tooling. Hyperscalers capture this expanding pool through three levers:

Governed hosting. Cloud platforms sit directly between open weight models and the enterprise, providing compliance, monitoring, security and identity management.

High margin adjacent services. The deployment of these workloads pulls through high margin enterprise data, retrieval, security and observability services, which often exceed the raw inference compute bill.

Pool expansion, not substitution. Every dollar spent on model intelligence that moves away from closed lab APIs converts into hyperscaler inference compute plus adjacent service revenue. Rather than substituting spend, lower prices at the model layer expand the total volume of enterprise workloads running on public cloud infrastructure.

Hardware infrastructure (Nvidia and Micron)

The open weight expansion accelerates the structural shift from training dominated compute toward inference dominated compute. Because agentic inference workloads run continuously in the background as enterprise software usage scales, global demand for high performance GPUs, CPUs and high bandwidth memory (HBM) expands.

GPUs at the centre: Every additional agentic loop, every workflow that clears its new ROI hurdle, lands as GPU hours. That is why NVIDIA, despite the noise around cheaper open models, is a beneficiary of the shift rather than a casualty of it.

The memory bottleneck: Modern inference is bandwidth bound as much as compute bound. Long context windows, retrieval augmented generation, and agentic loops all lean on High Bandwidth Memory. Every incremental data centre GPU ships with 8 to 12 stacks of HBM, which is why Micron is one of the cleanest secondary plays on inference volume growth.

Fund Strategy

While broader market narrative continues to debate whether proprietary model labs can sustain defensible competitive moats, our first principles framework views the commoditisation of foundation models as a growth engine for the layers around it.

By shifting toward open weight architectures and intelligent routing, the industry is collapsing the cost of reasoning. This cost collapse is not shrinking the market. It is unlocking categories of workflows that previously did not clear the ROI hurdle, and driving volume in agentic tasks.

In July, Jensen Huang (NVIDIA) and Satya Nadella (Microsoft) both amplified an open letter urging Washington to protect open weight models. While framed around enhancing American competitiveness, the commercial motivation is clear: a thriving open source ecosystem accelerates inference volume on NVIDIA silicon and drives enterprise hosting onto Azure. When they publicly champion open source, they are telling you where they see the value.

Within our fund strategy, we hold high conviction positions across the beneficiaries of this shift:

Hardware and Memory Leaders (NVIDIA, Micron): Direct beneficiaries of rising global inference volume and the memory bandwidth requirements of open architectures.

Cloud Hyperscalers: Capturing hosting and orchestration revenue as enterprise workloads land on Azure, AWS and Google Cloud.

Software and Platform Leaders: Using open architectures to deliver agentic capabilities, driving top line growth while protecting software gross margins.


Disclosures: Nvidia (NVDA), Micron (MU), Uber (UBER), Block (XYZ), ServiceNow (NOW), WiseTech Global (WTC) and TechnologyOne (TNE) are current holdings across the ELM Australian and Global funds.

........
This note has been prepared by ELM Responsible Investments (‘ELMRI’) ABN 70 607 177 711 AFSL 520428, for Australian wholesale clients for the purposes of section 761G of the Corporations Act 2001 (Cth). The information is not intended for general distribution or publication and must be retained in a confidential manner. Information contained herein consists of confidential proprietary information constituting the sole property of ELMRI and its investment activities; its use is restricted accordingly. This note is for general informational purposes only and does not purport to be comprehensive or to give advice. The views expressed are the views of the writer at the time of preparation and presenting and all forecasts, assumptions, opinions, data and other information are not warranted as to accuracy or completeness and are subject to change without notice. This is not an offer document and does not constitute an offer or invitation of investment recommendation to distribute or purchase securities, shares, units or other interests to enter into an investment agreement. No person should rely on the content and/or act on the basis of any material contained in this note. Any potential investor should consider their own circumstances and seek professional advice. ELMRI funds, its directors, employees, representatives and associates may have an interest in the named securities. Past performance is for illustrative purposes only and is not indicative of future performance.

Jai Mirchandani
Founder and CIO
ELM Responsible Investments

Jai has 18+ years of investment and financial markets experience. He currently manages the domestic and global portfolios at ELMRI, backing growth companies that are providing solutions to global challenges. Prior to ELMRI, Jai spent nine years at...

I would like to

Only to be used for sending genuine email enquiries to the Contributor. Livewire Markets Pty Ltd reserves its right to take any legal or other appropriate action in relation to misuse of this service.

Personal Information Collection Statement
Your personal information will be passed to the Contributor and/or its authorised service provider to assist the Contributor to contact you about your investment enquiry. They are required not to use your information for any other purpose. Our privacy policy explains how we store personal information and how you may access, correct or complain about the handling of personal information.

Comments

Sign In or Join Free to comment
The 10th annual Livewire Live 2026

One room. One day. The minds that move markets.

22 September 2026 Art Gallery of NSW, Sydney

Register Now