July OpenRouter Leaderboard: Xiaomi Tops Daily Volume, Chinese Labs Hit 46%
OpenRouter is the largest neutral model-routing platform, steering real paid requests from hundreds of thousands of developers daily. Its rankings don't answer "who's smartest" — they answer where production traffic actually flows. That gap between capability and popularity is the lens for everything below.
As of July 25, 2026, daily token volume Top 12:
| Rank | Model | Vendor | Daily Tokens | 30-Day Total |
|---|---|---|---|---|
| 1 | Mimo V2.5 | Xiaomi | 1.4T | 31.2T |
| 2 | DeepSeek V4 Flash | DeepSeek | 943.9B | 23.6T |
| 3 | Hy3 | Tencent | 590B | 23.4T |
| 4 | Nemotron 3 Ultra 550B (free) | NVIDIA | 428.6B | 9T |
| 5 | DeepSeek V4 Pro | DeepSeek | 413.7B | 11.6T |
| 6 | GLM 5.2 | Z.ai | 316.7B | 13.3T |
| 7 | MiniMax M3 | MiniMax | 262.5B | 15.1T |
| 8 | Step 3.7 Flash | StepFun | 204.8B | 5.9T |
| 9 | Kimi K3 | Moonshot | 157.6B | 1.6T (new entry) |
| 10 | Ling 3.0 Flash | Ant InclusionAI | 128.3B | 417.3B |
| 11 | Gemini 3 Flash Preview | 106.3B | 4T | |
| 12 | Claude Sonnet 5 | Anthropic | 99.5B | 3.6T |
Seven of the top ten models are Chinese-lab releases. Zoom out to vendor share: Chinese vendors (DeepSeek, Xiaomi, Tencent, Z.ai, MiniMax, Moonshot, Alibaba) collectively hold ~46% — up from under 2% a year ago. The U.S. big three (OpenAI, Anthropic, Google) dropped from ~70% mid-year to 30%–36% combined.
| Vendor | Token Share (approx.) | Notes |
|---|---|---|
| DeepSeek | 16%–18% | Most stable #1 single vendor |
| Xiaomi | 8%–18% | Mimo V2.5 surge; highest volatility |
| Anthropic | 10%–15% | Opus 5 shipped July 24 |
| Tencent | 8%–13% | Hy3 momentum |
| 8%–13% | Gemini 3 Flash |
Price drives the shift: DeepSeek V4 Flash inputs run ~$0.05–0.14/M vs OpenAI GPT-5.5 at ~$5/M — a 35× spread. Monthly volume champions keep rotating: MiniMax M2.5 in February, MiMo-V2-Pro in April, Mimo V2.5 in July. Don't treat a single-day snapshot as a permanent pick — track who stays on the board.
Still picking by MMLU: Lab scores and production wallet votes often invert — month-end bills won't match your benchmark spreadsheet.
Ignoring rank volatility: Order shifts day to day (Mimo V2.5 moved between July 24–25). Always stamp a data cutoff date.
Confusing volume with quality: Cheap open models dominate total volume — that doesn't mean they belong on complex reasoning or high-risk agent decisions.
Single-model religion: Hard-coding one provider becomes tech debt the moment the monthly champion rotates.
API online, host offline: Close your laptop and the agent pipeline dies — accurate rankings can't fix runtime uptime.
High Volume Doesn't Mean High Quality: Closed Flagships Still Own Hard Tasks
OpenRouter ranks token throughput, not model quality. Fast, cheap models behind high-traffic apps climb the leaderboard easily — even when complex reasoning scores are mediocre.
Slice by dollar spend by task type (not token count) and the picture flips:
| Task Type | Spend Share |
|---|---|
| General chat | 35.7% |
| Agent workflows | 30.4% |
| Code | 26.5% |
| Data processing | 7.5% |
On classification and complex reasoning workloads, the leaderboard inverts: Claude Sonnet 4.6 and Claude Opus 4.7 tie at 13.5% spend share each; GPT-5.5 holds third at 11.6%. The cheap open models dominating total volume barely register.
The market is splitting naturally: affordable Chinese open models absorb high-volume, low-stakes, fault-tolerant workloads; closed flagships keep pricing power on high-value, low-tolerance hard tasks.
Anthropic's Claude Opus 5 (released July 24) is the clearest example: 43.3% on FrontierBench v0.1 (beating GPT-5.6 Sol's 37.5%), priced at $5/$25 per M — half Fable 5's input rate. See our Opus 5 and Kimi K3 distillation breakdown.
| Model | Input $/M | Output $/M | Positioning |
|---|---|---|---|
| DeepSeek V4 Flash | ~$0.05–0.14 | ~$0.24–0.28 | Best value; agentic coding default |
| Nemotron 3 Ultra | $0.42 (free tier available) | $2.61 | U.S. full open weights; NVIDIA stack |
| GLM 5.2 | $0.45 | $3.31 | Closest open "Opus-class" planning |
| Kimi K3 | ~$3 | ~$15 | Largest open weights (1.4TB) |
| Claude Opus 5 | $5 (fast tier $10) | $25 (fast tier $50) | Closed flagship; top July benchmarks |
The App Layer Secret: Coding Agents Rule, Roleplay Is Half the Iceberg
Model rankings show which brain gets picked. App rankings (openrouter.ai/apps) show what that brain actually does in production:
| Rank | App | Type | Share (approx.) |
|---|---|---|---|
| 1 | Hermes Agent | Personal agent / CLI agent | ~45% |
| 2 | Kilo Code | Coding agent | ~13% |
| 3 | OpenClaw | General agent | ~9% |
| 4 | Claude Code | Coding agent | ~6% |
| 5 | Descript | Content production | ~4.5% |
| 6 | pi | Agent | ~3.3% |
| 7 | Lemonade | Companion / gaming | ~2.1% (new entry) |
| 8 | ISEKAI ZERO | Roleplay | ~2.0% (new entry) |
| 9 | Janitor AI | Roleplay | ~1.8% (new entry) |
| 10 | Cline | Coding agent (IDE plugin) | ~1.7% |
Hermes Agent (Nous Research's self-iterating open agent) holds ~45% of app-layer tokens alone. Coding agents occupy most of the rest. A telling fork lineage: Cline → Roo Code → Kilo Code share the same code ancestry — the "grandchild" Kilo Code now outruns grandfather Cline. First-mover advantage in open coding agents is fragile.
Nearly invisible in enterprise AI coverage: roleplay and companion apps — Janitor AI, ISEKAI ZERO, SillyTavern, HammerAI — consume a significant slice of open-model traffic. The OpenRouter × a16z State of AI report finds creative roleplay accounts for over half of all open-model volume. If you only read enterprise AI news, you miss half the market.
August Forecasts and a Six-Step Tiered Routing Playbook
Five August trend calls:
Chinese open-model share likely keeps climbing toward 50% this year — unless U.S. vendors make major pricing concessions.
Monthly volume champions will keep rotating — Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot aren't slowing the price war.
Anthropic may ship a cheaper tier in August (Haiku-class) to reclaim volume share — Opus 5 is already the fourth flagship in two months.
Community quantized Kimi K3 builds should land in 2–4 weeks — that's when smaller teams can actually run the 1.4TB weights.
Security and compliance become selection variables — ongoing OpenAI sandbox-escape headlines will raise vendor safety reputation weight in enterprise procurement.
Six-step tiered routing playbook:
Route by task tier: Chat, creative, and roleplay → cheap open models (DeepSeek V4 Flash, Mimo V2.5). Classification, complex reasoning, high-risk agent decisions → closed flagships (Claude Opus 5, GPT-5.5).
Adopt OpenRouter as a unified gateway: One key across hundreds of models; track weekly shifts at openrouter.ai/rankings. See our OpenRouter API guide.
For coding, benchmark DeepSeek V4 Flash and GLM 5.2 first: Best value vs best open planning quality; reserve Claude Opus 5 for steps that actually stall.
Set billing circuit breakers and daily caps: Price per M × daily call volume = threshold; default agent batches to cheap routes, escalate to flagship only on hard refactors.
Add vendor safety records to your scorecard: Autonomous agent deployments need tighter permission scopes and vendors with cleaner safety track records.
Provision 24/7 agent hosts: Move Hermes Agent, Kilo Code, and OpenClaw off laptops onto dedicated cloud Macs — launchd persistence, Keychain for multi-provider API keys. Compare pricing and help center specs.
Three Hard Numbers and Role-Based Takeaways
46% / 2%: Chinese vendors now hold ~46% of OpenRouter token share — under 2% a year ago. One of the steepest share migrations in AI in the past 12 months.
35× / 13.5%: DeepSeek V4 Flash vs GPT-5.5 price gap is ~35×; on complex-reasoning spend, Claude Sonnet 4.6 and Opus 4.7 tie at 13.5% each.
45% / 50%+: Hermes Agent holds ~45% of app-layer volume; roleplay drives over half of all open-model traffic — invisible in most enterprise coverage.
The July takeaway in one line: capability and popularity are diverging. Chinese open models win on price and volume; U.S. closed flagships keep hard-task pricing power and safety reputation moats. OpenRouter is fine for technical sandboxing, but average China-access latency runs ~180–250ms with no domestic invoicing — evaluate compliant relay options for production.
API routing alone can't replace agent hosting: laptops sleep, multi-key management gets messy, and local open-weight deployment needs 96GB+ unified memory — each path has hidden costs. For 24/7 multi-model agent pipelines, KVMNODE dedicated Mac Mini cloud rental is usually the better host: native Apple Silicon toolchain, flexible daily/weekly/monthly terms. See the pricing page or order directly.
Data as of: July 25, 2026 · Source: OpenRouter official rankings and public mirrors · Live numbers at openrouter.ai/rankings may differ