An Tran Solutions
An Tran Solutions
Back to Blog

The Truth About Cheap Chinese AI: US Models Fell From 70% to 30% in a Year, and Almost Nobody Noticed

July 7, 20265 min readby An Tran
On this page

Everyone assumes "serious AI" means three American names: ChatGPT, Claude, Gemini. Serious businesses build on them. Everything else is a cheap knockoff.

The developer market just voted the other way — and almost nobody noticed.

On July 7, 2026, Xiaomi's MiMo-V2-Pro became the most-used model on OpenRouter by weekly tokens, according to Buildfast's news roundup. A phone maker. Ahead of OpenAI. On the largest AI routing platform for people who build products.

That isn't a fluke. It's the end point of a year-long trend that the "race for the biggest model" headlines missed entirely.

The number that shocks people

Look at the data, not the vibe.

According to a Bloomberg chart built from OpenRouter and Exponential View data, compiled by officechai, US models' share of tokens on OpenRouter (OpenAI, Google and Anthropic combined) fell from roughly 70% (June 2025) to roughly 30% (June 2026). In exactly 12 months.

On the other side, Chinese models now process the bulk of tokens. Depending on how you count, the figure varies: about 44–46% in April–June 2026 data, and as high as ~61% in a May 2026 count from Data Gravity. The difference comes down to methodology — counting only attributed tokens versus all of them — but the direction is not in dispute.

The more striking detail: DeepSeek alone accounts for about 16% of all tokens on OpenRouter, and OpenAI has slipped to fourth place by per-company token volume. Google — per Data Gravity's analysis — collapsed from 37% to 13%. Meta's Llama dropped off the leaderboard altogether.

This isn't revenue share. It's share of actual work — the tokens developers and real products burn every day. And by that measure, the "cheap knockoff" has become the workhorse of the AI economy.

Why? Not because they're smarter

This is where it's easy to get it wrong. Chinese models are winning not because they beat Claude or GPT on raw intelligence. They're winning on a calculation any business owner can do: price per unit of performance.

The specifics:

  • An hour-long coding session with Claude costs about $10; equivalent work with DeepSeek: under $0.50, according to Rest of World.
  • Data Gravity estimates DeepSeek-V4-Pro is ~12x cheaper than GPT-5.5 at comparable benchmark levels.
  • On OpenRouter in April 2026, Xiaomi's MiMo V2 Pro was priced at $1 / $3 per million input/output tokens, with a 1-million-token context window. Alibaba's Qwen 3.6 Plus was even free in preview, also with 1M context, according to digitalapplied's Q2 2026 report.

Flo Crivello, founder of Lindy — a company that switched from Anthropic to DeepSeek and saved millions of dollars — summed up the logic in Rest of World's reporting, roughly: you don't need God to write your emails, and if you can buy that lower tier of intelligence for a tenth of the price, it would be foolish not to.

That line is an entire strategy. And it ties straight into the second reason: open weights.

Unlike closed US APIs, most of the leading Chinese models are open-weight — you can download them, run them yourself, avoid vendor lock-in, and push cost per token down another notch. At the same time, the mix of work on OpenRouter shifted hard toward coding: from 11% to 50% of total usage in 12 months (Data Gravity). And code is "high-volume, repetitive, price-sensitive" work — exactly the slice cheap models eat whole.

Release speed is a weapon too: Moonshot shipped five major Kimi releases in under a year, faster than most companies' procurement cycles.

The real lesson for businesses: stop paying for the brand

If you're integrating AI into a product or a workflow, this is the most valuable part of this piece.

The most common mistake I see: picking one expensive, famous model for everything — classifying emails, summarizing tickets, generating product descriptions. That's buying a supercar to do the grocery run. You're paying a "brand fee" the developer market stopped paying a long time ago.

The right mental model is intelligence tiering — matching the tier of model to the value of the work:

  • Repetitive, high-volume, low-stakes work (classification, extraction, summarization, FAQ answers): use cheap open models. That's 80% of the real volume for most businesses.
  • Work that needs deep reasoning, creativity, or carries high risk (complex advice, architecture-level code, brand-facing content): that's where the expensive frontier model earns its keep.

Do it right and your AI bill can drop by an order of magnitude while the output quality users perceive doesn't change. That's not cost-cutting — it's engineering.

And here's the key architectural point: don't lock yourself into one vendor. The reason OpenRouter exists — and the reason it's where this shift is playing out — is that it lets you route between dozens of models behind a single API. A product designed to swap models easily can adopt next week's 12x-cheaper model by changing one line of config. A product welded to a single proprietary API pays full price, forever.

But "cheaper" doesn't automatically mean "right"

I won't sell you a one-sided story. There are two genuine reasons not to rush into moving everything to Chinese models.

First — quality at the top still matters. Anthropic CEO Dario Amodei argues that Chinese models are "optimized for benchmarks and distilled from US labs," and that in the long run raw capability, not price, decides what gets chosen (officechai). It's a contested argument — the market-share data is pushing back against him — but he isn't wrong at the top end: for genuinely hard work, the quality gap is still there and still worth paying for.

Second — data governance risk is real. Rest of World reports that Airbnb and Anysphere drew scrutiny from the US Congress for using Chinese open models (Qwen, Kimi). For any business handling customer data, the questions "where does the data flow, who controls the weights, which servers is it stored on" are not side issues. The open-weight advantage — self-hosting — is also the best answer to that concern: run the model on infrastructure you control, and the data never leaves the house.

So the conclusion is not "drop Claude, use DeepSeek." The conclusion is: stop defaulting to the highest price, and start designing on purpose.

What this signals

The 70% → 30% shift points to a bigger truth about the industry: the model layer is being commoditized. The advantage no longer lies in "who has the smartest model this week" but in what you build on top of the model — workflows, proprietary data, experience, and the flexibility to switch models.

For a business owner, the message fits in one sentence: your AI bill is a strategic choice, not a fixed cost. Whoever understands that will run at the same quality for a tenth of what competitors pay. Whoever doesn't will keep paying a brand fee to write emails.

When I integrate AI into a product, I design to exactly this principle: tier models by the value of the work and keep the ability to switch vendors — so you're never locked into a bill. If you want to know where your workflows are overpaying, talk to me.

The developer market has voted. The only remaining question is whether you read the results.

Sources

Related articles