TL;DR: If you're still picking models from last quarter's leaderboard screenshots, you're routing blind. OpenRouter through July 25, 2026 doesn't measure IQ — it measures where developers actually send paid traffic. Xiaomi Mimo V2.5 now leads at 1.4 trillion tokens/day, Chinese labs collectively hold ~46% share (up from under 2% a year ago), and closed-source flagships still own pricing power on hard tasks. This breakdown covers model/vendor/app rankings, the volume-vs-quality dumbbell market, the hidden roleplay economy, five August forecasts, and a six-step tiered routing guide. Background: June OpenRouter rankings analysis, OpenRouter API setup guide.
01

July OpenRouter Leaderboard: Xiaomi Tops Daily Volume, Chinese Labs Hit 46%

OpenRouter is the largest neutral model-routing platform, steering real paid requests from hundreds of thousands of developers daily. Its rankings don't answer "who's smartest" — they answer where production traffic actually flows. That gap between capability and popularity is the lens for everything below.

As of July 25, 2026, daily token volume Top 12:

RankModelVendorDaily Tokens30-Day Total
1Mimo V2.5Xiaomi1.4T31.2T
2DeepSeek V4 FlashDeepSeek943.9B23.6T
3Hy3Tencent590B23.4T
4Nemotron 3 Ultra 550B (free)NVIDIA428.6B9T
5DeepSeek V4 ProDeepSeek413.7B11.6T
6GLM 5.2Z.ai316.7B13.3T
7MiniMax M3MiniMax262.5B15.1T
8Step 3.7 FlashStepFun204.8B5.9T
9Kimi K3Moonshot157.6B1.6T (new entry)
10Ling 3.0 FlashAnt InclusionAI128.3B417.3B
11Gemini 3 Flash PreviewGoogle106.3B4T
12Claude Sonnet 5Anthropic99.5B3.6T

Seven of the top ten models are Chinese-lab releases. Zoom out to vendor share: Chinese vendors (DeepSeek, Xiaomi, Tencent, Z.ai, MiniMax, Moonshot, Alibaba) collectively hold ~46% — up from under 2% a year ago. The U.S. big three (OpenAI, Anthropic, Google) dropped from ~70% mid-year to 30%–36% combined.

VendorToken Share (approx.)Notes
DeepSeek16%–18%Most stable #1 single vendor
Xiaomi8%–18%Mimo V2.5 surge; highest volatility
Anthropic10%–15%Opus 5 shipped July 24
Tencent8%–13%Hy3 momentum
Google8%–13%Gemini 3 Flash

Price drives the shift: DeepSeek V4 Flash inputs run ~$0.05–0.14/M vs OpenAI GPT-5.5 at ~$5/M — a 35× spread. Monthly volume champions keep rotating: MiniMax M2.5 in February, MiMo-V2-Pro in April, Mimo V2.5 in July. Don't treat a single-day snapshot as a permanent pick — track who stays on the board.

01

Still picking by MMLU: Lab scores and production wallet votes often invert — month-end bills won't match your benchmark spreadsheet.

02

Ignoring rank volatility: Order shifts day to day (Mimo V2.5 moved between July 24–25). Always stamp a data cutoff date.

03

Confusing volume with quality: Cheap open models dominate total volume — that doesn't mean they belong on complex reasoning or high-risk agent decisions.

04

Single-model religion: Hard-coding one provider becomes tech debt the moment the monthly champion rotates.

05

API online, host offline: Close your laptop and the agent pipeline dies — accurate rankings can't fix runtime uptime.

02

High Volume Doesn't Mean High Quality: Closed Flagships Still Own Hard Tasks

OpenRouter ranks token throughput, not model quality. Fast, cheap models behind high-traffic apps climb the leaderboard easily — even when complex reasoning scores are mediocre.

Slice by dollar spend by task type (not token count) and the picture flips:

Task TypeSpend Share
General chat35.7%
Agent workflows30.4%
Code26.5%
Data processing7.5%

On classification and complex reasoning workloads, the leaderboard inverts: Claude Sonnet 4.6 and Claude Opus 4.7 tie at 13.5% spend share each; GPT-5.5 holds third at 11.6%. The cheap open models dominating total volume barely register.

The market is splitting naturally: affordable Chinese open models absorb high-volume, low-stakes, fault-tolerant workloads; closed flagships keep pricing power on high-value, low-tolerance hard tasks.

Anthropic's Claude Opus 5 (released July 24) is the clearest example: 43.3% on FrontierBench v0.1 (beating GPT-5.6 Sol's 37.5%), priced at $5/$25 per M — half Fable 5's input rate. See our Opus 5 and Kimi K3 distillation breakdown.

ModelInput $/MOutput $/MPositioning
DeepSeek V4 Flash~$0.05–0.14~$0.24–0.28Best value; agentic coding default
Nemotron 3 Ultra$0.42 (free tier available)$2.61U.S. full open weights; NVIDIA stack
GLM 5.2$0.45$3.31Closest open "Opus-class" planning
Kimi K3~$3~$15Largest open weights (1.4TB)
Claude Opus 5$5 (fast tier $10)$25 (fast tier $50)Closed flagship; top July benchmarks
03

The App Layer Secret: Coding Agents Rule, Roleplay Is Half the Iceberg

Model rankings show which brain gets picked. App rankings (openrouter.ai/apps) show what that brain actually does in production:

RankAppTypeShare (approx.)
1Hermes AgentPersonal agent / CLI agent~45%
2Kilo CodeCoding agent~13%
3OpenClawGeneral agent~9%
4Claude CodeCoding agent~6%
5DescriptContent production~4.5%
6piAgent~3.3%
7LemonadeCompanion / gaming~2.1% (new entry)
8ISEKAI ZERORoleplay~2.0% (new entry)
9Janitor AIRoleplay~1.8% (new entry)
10ClineCoding agent (IDE plugin)~1.7%

Hermes Agent (Nous Research's self-iterating open agent) holds ~45% of app-layer tokens alone. Coding agents occupy most of the rest. A telling fork lineage: Cline → Roo Code → Kilo Code share the same code ancestry — the "grandchild" Kilo Code now outruns grandfather Cline. First-mover advantage in open coding agents is fragile.

Nearly invisible in enterprise AI coverage: roleplay and companion apps — Janitor AI, ISEKAI ZERO, SillyTavern, HammerAI — consume a significant slice of open-model traffic. The OpenRouter × a16z State of AI report finds creative roleplay accounts for over half of all open-model volume. If you only read enterprise AI news, you miss half the market.

04

August Forecasts and a Six-Step Tiered Routing Playbook

Five August trend calls:

01

Chinese open-model share likely keeps climbing toward 50% this year — unless U.S. vendors make major pricing concessions.

02

Monthly volume champions will keep rotating — Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot aren't slowing the price war.

03

Anthropic may ship a cheaper tier in August (Haiku-class) to reclaim volume share — Opus 5 is already the fourth flagship in two months.

04

Community quantized Kimi K3 builds should land in 2–4 weeks — that's when smaller teams can actually run the 1.4TB weights.

05

Security and compliance become selection variables — ongoing OpenAI sandbox-escape headlines will raise vendor safety reputation weight in enterprise procurement.

Six-step tiered routing playbook:

01

Route by task tier: Chat, creative, and roleplay → cheap open models (DeepSeek V4 Flash, Mimo V2.5). Classification, complex reasoning, high-risk agent decisions → closed flagships (Claude Opus 5, GPT-5.5).

02

Adopt OpenRouter as a unified gateway: One key across hundreds of models; track weekly shifts at openrouter.ai/rankings. See our OpenRouter API guide.

03

For coding, benchmark DeepSeek V4 Flash and GLM 5.2 first: Best value vs best open planning quality; reserve Claude Opus 5 for steps that actually stall.

04

Set billing circuit breakers and daily caps: Price per M × daily call volume = threshold; default agent batches to cheap routes, escalate to flagship only on hard refactors.

05

Add vendor safety records to your scorecard: Autonomous agent deployments need tighter permission scopes and vendors with cleaner safety track records.

06

Provision 24/7 agent hosts: Move Hermes Agent, Kilo Code, and OpenClaw off laptops onto dedicated cloud Macs — launchd persistence, Keychain for multi-provider API keys. Compare pricing and help center specs.

05

Three Hard Numbers and Role-Based Takeaways

A

46% / 2%: Chinese vendors now hold ~46% of OpenRouter token share — under 2% a year ago. One of the steepest share migrations in AI in the past 12 months.

B

35× / 13.5%: DeepSeek V4 Flash vs GPT-5.5 price gap is ~35×; on complex-reasoning spend, Claude Sonnet 4.6 and Opus 4.7 tie at 13.5% each.

C

45% / 50%+: Hermes Agent holds ~45% of app-layer volume; roleplay drives over half of all open-model traffic — invisible in most enterprise coverage.

The July takeaway in one line: capability and popularity are diverging. Chinese open models win on price and volume; U.S. closed flagships keep hard-task pricing power and safety reputation moats. OpenRouter is fine for technical sandboxing, but average China-access latency runs ~180–250ms with no domestic invoicing — evaluate compliant relay options for production.

API routing alone can't replace agent hosting: laptops sleep, multi-key management gets messy, and local open-weight deployment needs 96GB+ unified memory — each path has hidden costs. For 24/7 multi-model agent pipelines, KVMNODE dedicated Mac Mini cloud rental is usually the better host: native Apple Silicon toolchain, flexible daily/weekly/monthly terms. See the pricing page or order directly.

Data as of: July 25, 2026 · Source: OpenRouter official rankings and public mirrors · Live numbers at openrouter.ai/rankings may differ