An Tran Solutions
An Tran Solutions
Back to Blog

A Third of the AI US Businesses Run Now Comes From China. Here's the Lesson for Your Costs

July 13, 20265 min readby An Tran
On this page

The story everyone tells about AI is a two-horse race: OpenAI versus Anthropic, Google squeezing in between, and the whole world holding its breath to see who ships the strongest model. You've been taught that backing the right "champion" is the strategic decision.

That story is out of date. While the press counts benchmark points for the top labs, a third or more of the AI tokens that US developers and businesses consume each week has quietly moved to Chinese models. Not out of reverse patriotism. Because they're up to 90% cheaper.

This isn't rumor. On July 7, 2026, CNBC published an investigation based on real routing data from OpenRouter, the intermediary platform tens of thousands of developers use to call APIs for every AI model. The numbers in it should make anyone setting a technology budget sit up straight.

The numbers that startled Silicon Valley

According to the OpenRouter data CNBC cites: as of mid-2026, models of Chinese origin accounted for 46.4% of all tokens routed in a peak week, ahead of US-origin models, which were down to 35.7%.

Let that sink in for a moment. The largest US provider on the platform, Anthropic, holds just 14.8%. Meanwhile DeepSeek alone carries 17.6%, equivalent to 5.13 trillion tokens a week, making it the single largest provider on the entire platform. Alibaba's Qwen is second at 13.9% (2.77 trillion tokens a week).

What's striking isn't the snapshot, it's the speed. The same dataset shows:

  • In the first half of 2025, Chinese models accounted for just 4.5% of traffic.
  • The average over the previous 12 months was 11%.
  • But since February 8, 2026, they have topped 30% every single week.
  • And in the week of February 9–15, 2026, Chinese models overtook US models in total tokens for the first time.

From 4.5% to 46% in just over a year. This isn't a trend in formation. It's a migration that has already happened; most people just haven't looked at the right chart.

"90% cheaper" isn't a promotion. It's a business model

Why did the flow reverse so fast? Justin Summerville of OpenRouter told CNBC plainly: "Chinese open-source models are consistently 60% to 90% cheaper than Anthropic's and OpenAI's flagship offerings."

One look at the price list explains it. Here's the cost per million output tokens, from byteiota's analysis of June–July 2026 pricing:

ModelPrice per 1M output tokens
DeepSeek V4 Flash$0.28
Qwen 3.6 Max$1.20
GLM-5.2 (MIT license)$4.40
Claude Sonnet 5$10.00*
GPT-5.5$30.00

*Claude Sonnet 5's price is a launch promotion running through the end of August 2026, after which it rises to $15 per million tokens. On input, the gap is even wider: DeepSeek V4 Flash charges $0.14 per million tokens, while GPT-5.5 charges $5.00, a difference of more than 35x.

A gap of nearly 100x between DeepSeek and GPT-5.5 isn't a rounding error. It's a fundamentally different business model.

And this is where the "more expensive means better" argument starts to crumble. Z.ai's GLM-5.2, released on June 13, 2026 with open weights under the MIT license, scores 62.1% on SWE-bench Pro (a benchmark of real-world coding ability), ahead of GPT-5.5 at 58.6%, at roughly one sixth of the cost. When a model six times cheaper does better on the test, the "quality reason" for paying ten times more becomes very hard to justify.

The cost matters even more for agentic tasks, where an AI loops through 50–200 calls to complete a single job. At that volume, the price gap doesn't add up linearly; it multiplies. An automated workflow running on GPT-5.5 can burn through dozens of times the budget of the same workflow on DeepSeek, for output that end users can barely tell apart.

Why this matters for your business

You might be thinking: "I don't build AI models, what does this have to do with me?" Everything. Because this signals a shift every business owner needs to understand: the AI model layer is becoming a commodity.

Two years ago, "the strongest AI" was a scarce, expensive asset, and whoever had better access had an edge. Today, capability that's good enough for 99% of business work can be bought for close to nothing. Customer support assistants, email triage, content drafts, document summaries, coding help: these jobs don't need the most expensive model on the planet. They need a model that's good enough, stable and cheap.

That has three blunt consequences for your budget:

One: if someone quotes you an "AI solution" on the strength of using the most premium model, push back. For most operational tasks, that's money you're paying for peace of mind, not for results. The difference in output quality keeps getting thinner, while the cost gap is dozens of times.

Two: your AI costs will go down, not up. Don't lock yourself into long-term contracts that assume AI stays expensive. The whole curve is heading down. What costs 100 today could cost 10 next year.

Three: don't bet the business on a single model. This is probably the most valuable lesson in the whole story.

But cheap doesn't mean risk-free

It would be dishonest of me to stop here and tell you to "move everything to Chinese models to save money". It isn't that simple, and anyone doing this work properly has to spell out the downside.

When you call a Chinese provider's API directly, your data passes through their infrastructure. That raises real questions:

  • Data sovereignty. China's National Security Law can compel domestic companies to cooperate in handing over data. For businesses with European customers, sending personal data this way can also run into Article 46 of the GDPR.
  • Geopolitical risk. The US Congress has opened investigations into some companies that integrate Chinese models. In the other direction, China is also considering tightening access to its models from abroad. You could end up in a split market where your technical dependency gets cut off by either government.

So the principle isn't "cheapest wins", it's routing by sensitivity. Public data and high-volume, low-risk tasks: cheap models. Customer data, sensitive information, output that speaks for your brand: pick models and providers you trust on data governance, and consider routing through a gateway hosted in the US or EU instead of calling direct. Cheap is one variable, not the only goal.

What actually creates an advantage (hint: not the model)

This is the conclusion I want you to take away. If the model layer is becoming a commodity anyone can rent for next to nothing, then the model cannot be your competitive advantage. Your competitors can rent the very same AI, at the very same price, on the same afternoon.

The advantage lies in things you can't rent off the shelf:

  • Your own data: knowledge of your customers, transaction history, industry context no public model has.
  • Your process: how you wire AI into operations so it actually produces results, rather than a feature bolted on for show.
  • Your distribution and brand: customer trust, which no API throws in for free.

AI is just a raw ingredient, cheaper and more available every month. What decides the outcome is what you cook with it. The smart business in 2026 doesn't ask "which model is strongest?" but "how do we put AI in the right place in our process to create real value, at a sustainable cost, without tying ourselves to one provider?"

That's exactly how I approach AI integration for websites and workflows: choose the model by the problem, not the brand; design so it's easy to swap when market prices move; and keep your data and customer experience at the core, because that's the part nobody can copy.

The AI race is still on. It's just that this year's quiet winner isn't the most expensive model. It's whoever understands the model was never the point.

Sources

Related articles

An Tran Solutions
September 30, 20268 min read

Spec-Driven Development Isn't Documentation. It's a Governance System for AI-Written Code

From vibe coding to spec-driven development: why longer specs won't save you, and a four-layer governance framework (constitution, risk-tiered specs, executable acceptance, control gates) that keeps AI-written code under your control. With data from Veracode, Thoughtworks and Martin Fowler.