Ollama CEO Jeffrey Morgan: 80% of Enterprise Tokens Will Move to Open Source
Ollama co-founder and CEO Jeffrey Morgan stated on the Y Combinator podcast that the vast majority of token usage within enterprises will shift towards open-source weight models, estimating this to be around 80% to 90%. However, this does not imply that the budget will be allocated in the same proportion, as the open-source side may only account for 10% to 20% of costs.
His observation window is based on Ollama's own traffic: the platform claims to cover approximately 9 million developers, with about 178,000 GitHub stars and an 85% penetration rate in the Fortune 500. Since early 2026, token usage on Ollama Cloud has grown about 150 times, with key turning points being the programming agents introduced at the beginning of the year and the expansion of autonomous agents into finance, customer service, marketing, and sales around April with OpenClaw and Hermes.
The weekly token usage per developer was about 15 million before the emergence of "colleague-type agents," and has since multiplied several times; the context window for open-source models has expanded from about 128,000 to over a million, allowing long tasks to run continuously either locally or in the cloud. He cited a report from The Information stating that AT&T has shifted about 40% of its tokens to open source and is evaluating Chinese models.
Ollama was only finalized in mid-2023, having previously lingered for two years in the container desktop direction during the YC Winter 2021 batch; the founders, including Michael Chiang, previously worked on Kitematic, which was integrated into Docker, and developed Docker Desktop. The company completed a $65 million Series B round led by Theory Ventures in July this year, with total funding reaching approximately $88 million and a team of only about 14 people.
He predicts that the future will be a combination of local and cloud: data-sensitive loads will return to computers and intranets, while large models will move to the cloud; there is no need for a single "God model," but rather routing, customization, and operating systems above the model layer. U.S. labs still lag behind some Chinese open-source weights in providing ultra-large open-source models, and he views NVIDIA's Neotron Ultra as a signal of catching up.
In terms of market mechanisms, buyers are enterprises and programming agents looking to reduce unit intelligence costs, while sellers are closed-source cutting-edge labs charging by API. This is a cost-driven diversion: funds are shifting from high-priced closed-source APIs to self-hosted, fine-tunable open-source weights. Beneficiaries include inference gateways, hosting clouds, and open-source weight issuers, while those under pressure are closed-source platforms that can only rely on flagship models to charge high token fees but cannot prevent the overflow of long-tail workloads.
Source: Public Information
ABAB AI Insight
Morgan is not creating another chatbot, but rather a secondary application of Docker logic: simplifying complex installations into a single command. After Kitematic was acquired by Docker, he and his partner spent two years in YC until the emergence of Llama's open-source weights provided a distribution entry point. Whoever controls "how models are initiated by developers" will see the true usage, rather than just leaderboard scores.
The capital path is to run locally for free to build an installation base, then use cloud inference to convert installation numbers into token revenue. Benchmark's Peter Fenton bets that the same group that spread Docker Desktop to millions of developers will do the same here. Money is being carved out from enterprise IT API bills, flowing towards inference layers priced by GPU time rather than capped at millions of tokens; the cheaper the open-source weights, the more tokens consumed by upper-layer agents, making the entry point more valuable.
The analogy is not OpenAI's arms race against Anthropic, but rather Linux against Unix, Kubernetes against self-built clusters: expensive systems remain on critical paths, while cheaper systems capture the bulk. Products like Cursor have already proven that fine-tuning open-source weights can precede external services. The industry position is shifting from "catching up on evaluations" to "workload diversion"—closed-source maintains complex inference, while open-source absorbs coding and internal agents.
Structural changes belong to technological substitution. What is being substituted is not all intelligence, but rather the unit token cost. The mechanism is: programming agents can amplify an individual's weekly usage from millions to several times, increasing price elasticity; whoever can achieve the same capability at a lower order of magnitude will capture 80% of the usage, leaving 20% of the budget for the necessary cutting-edge models.
ABAB News · Cognitive Laws
- Technologies that account for 80% of usage often only account for 20% of the budget.
- Developer entry points are closer to real demand than model evaluations.
- Once agents appear, cheaper intelligence will consume more expensive intelligence.