Insights ·

Open-weight models now carry most of the tokens. The money still goes to closed ones

New data from Vercel, OpenRouter, Ramp and ETR shows US companies moving AI volume to open-weight models while spend stays with frontier labs. What it means for European organisations.

WEREALeggi in italianoLire en françaisAuf Deutsch lesenLeer en español

Between the beginning of September and the end of the month, four separate sets of data pointed in the same direction. On the platforms that route AI traffic for thousands of companies, open-weight models, the ones whose weights anyone can download and run, now process more tokens than proprietary models. A year ago they were a niche.

The shift is real, but it is easy to misread. Looking at where the money goes tells a different story, and the gap between the two is the most useful part.

What the data shows

  • Vercel AI Gateway, 17 September. Open-weight models processed 56% of all tokens on the gateway in August 2026, up from 13% in April and fewer than one in ten in December 2025. It is the first month they hold a majority.
  • OpenRouter, as reported by CNBC on 26 September. Chinese models, almost all of them open-weight (DeepSeek, Qwen, GLM, Kimi), accounted for between 57% and 67% of tokens in the week of 14 September, against 6% to 13% in February. In July, OpenRouter data reported by Bloomberg already put their share of tokens used by US firms on the platform at a record of about 58%.
  • Ramp AI Index, 9 September. Among US businesses, the share of tokens going to frontier models fell from a peak of 53% in August to 45%, as companies set cheaper models as defaults. The average price paid per million tokens fell 41% from its March peak, to $0.68.
  • Enterprise Technology Research survey, September. Open-weight models account for 34% of enterprise token usage, up from 23% a year earlier, and respondents expect 41% within twelve months. 42% of companies run at least one open-weight model in production. Cost is the main reason (69%), followed by control.

Tokens are not dollars

The same sources show the limits of the trend. On Vercel, the 56% of tokens handled by open-weight models corresponded to only 14% of spend. In the Ramp data, only 6.4% of US businesses that pay for AI pay for open-source models, while 43.8% pay Anthropic and 39.8% pay OpenAI. Surveys of large-company CIOs earlier this year still showed a preference for closed models, citing quality, security and the lack of in-house AI skills.

There are also blind spots. Routing platforms over-represent developers and fast-moving startups. Ramp sees card and invoice payments, so it cannot see an open model running on a company's own servers, which is exactly where many open-weight deployments live. The trend did not start this month either: it has been building since early 2026 and crossed the halfway mark during the summer.

The fair reading is this: companies are moving their volume, not their most critical work. High-volume, repetitive tasks (classification, extraction, summaries, the sub-steps of agent workflows, coding loops) go to open-weight models that cost a fraction per token. Hard reasoning and tasks where an error is expensive stay on frontier proprietary models.

Deciding which work can move to open models, and where they should run? Talk to the TrustTwin OS team.

Why it matters more in Europe

For US companies the driver is mostly cost. For European organisations there is a second reason that may weigh more: an open-weight model can run wherever you decide. On your own servers, on a node in a specific country, on an EU-only infrastructure, or on an employee's laptop. The data does not have to travel to a provider's API to be processed.

That freedom comes with its own work.

  • Open weights are not open data. You can inspect and run the model, not the data it was trained on.
  • Licences differ. Apache 2.0, Llama community licences and Gemma terms do not allow the same things.
  • Provenance gets reviewed. In the ETR survey, 62% of companies named security and compliance reviews as the main obstacle, and several of the most used models come from Chinese labs.
  • Someone has to run them. Hosting costs, updates and hardware sizing do not disappear because the licence is free.

From picking a model to routing work

When every company uses several models, the real decision is no longer which model to buy. It is which model handles which data, on which machine. That is a routing and policy problem, and it is the one TrustTwin OS was built to solve.

  • Local models, easily managed. With Agents Node, open-weight models run on the company's own machines and we keep them current. During inference, nothing leaves the computer.
  • Private agents with scoped knowledge. An agent knows only what the context it works in allows, whichever model is behind it.
  • Policy before placement. Data sensitivity and residency are checked before a workload runs. Work that must stay local stays local; if a residency rule cannot be guaranteed, the workload does not run.
  • A record of every decision. Where each job ran is logged, so the choice of model and machine can be audited later.
Agents Node local models screen listing open-weight models with download size, memory needs and licence
Agents Node: open-weight models such as gpt-oss, Qwen, Gemma and Llama, recommended for the machine they run on, with their licence and memory requirements.

The same approach runs in SweetHive, where small companies can choose local models for their agents and keep files and prompts on their own hardware.

Three questions for the next budget

  1. Which of our AI workloads are high-volume and routine, and what would they cost on an open-weight model?
  2. Which data must never leave our machines, our country or the EU, and do our current tools respect that by design?
  3. If we use several models tomorrow, who decides which one sees which data, and where is that decision recorded?

The move to open-weight models is not a vote against frontier models. It is the start of a mixed setup, and in a mixed setup the advantage goes to whoever governs the routing.


Sources

  • Vercel, AI Gateway Production Index, 17 September 2026: vercel.com/blog
  • CNBC, on Chinese models' share of OpenRouter traffic, 26 September 2026: cnbc.com
  • AI Weekly, summarising Bloomberg on OpenRouter data for US firms, 22 July 2026: aiweekly.co
  • Ramp, AI Index, 9 September 2026: ramp.com/data/ai-index-sept-2026
  • Techstrong.ai on the Enterprise Technology Research survey, September 2026: techstrong.ai
  • a16z, enterprise CIO survey, 2026: a16z.com