July 9 launch to July 30 price cut: what happened in OpenAI's 21-day sprint
On July 30, 2026, OpenAI adjusted GPT-5.6 pricing: Luna dropped from $1/$6 to $0.20/$1.20 per million tokens (-80%), Terra from $2.50/$15 to $2.00/$12 (-20%), Sol standard held at $5/$30, and a new Sol Fast SKU launched at $10/$60 with roughly 2.5× inference speed, replacing the old Priority Processing tier. The cuts landed days after Kimi K3's full weight release and OpenAI's own Sol GPU self-optimization blog — widely read as a competitive response to open-weight pressure.
| Date | Milestone | Key detail |
|---|---|---|
| Jul 9, 2026 | GPT-5.6 family ships | Sol, Terra, and Luna tiers go live; Sol positioned as flagship agentic coding model |
| Jul 16, 2026 | Kimi K3 hosted API | Moonshot AI launches Kimi K3 as a managed API, resetting coding-model price expectations |
| Jul 27, 2026 | Kimi K3 weights open | 2.8T-parameter MoE fully released; Artificial Analysis puts per-task cost at ~$0.94 |
| Jul 29, 2026 | Sol GPU optimization blog | OpenAI discloses Triton/Gluon kernel rewrites, speculative decoding, and KV cache work |
| Jul 30, 2026 | Official price cut | Luna -80%, Terra -20%, Sol standard unchanged; Sol Fast $10/$60 new SKU live |
The nuance: Luna and Terra got list-price cuts; Sol's "price hike" narrative comes from the new Fast tier — standard Sol didn't move, but latency-sensitive users now face 2× per-token pricing if they opt in.
Treating Luna's cut as a family-wide discount: Sol standard stays at $5/$30; only the new Fast mode adds a premium SKU — flagship total cost didn't drop.
Ignoring Kimi K3's pricing anchor: K3's $3/$15 headline matches Claude Sonnet 5, but vs K2.6 ($0.60/$2.50) it's roughly 6× more expensive — open-weight tiers are re-tiering too.
Booking savings from Sol's self-reported optimization: OpenAI claims 20% inference cost reduction and 15% throughput gain from GPU work — vendor-tested, not yet independently verified.
Assuming Sol Fast is just a Priority Processing rename: 2.5× speed at 2× price doesn't automatically lower cost per wall-clock hour — model it per SLA.
Picking sides in polarized Reddit threads: Community reaction split between celebrating "penny Luna" and questioning METR reward-hacking and self-reported benchmarks — enterprise picks need workload proof, not hot takes.
Before and after: GPT-5.6 tiers vs Kimi K3, DeepSeek, Claude, and more
| Model / SKU | Before (in/out per 1M) | After | Change | Role |
|---|---|---|---|---|
| GPT-5.6 Luna | $1.00 / $6.00 | $0.20 / $1.20 | -80% | Light / high-volume |
| GPT-5.6 Terra | $2.50 / $15.00 | $2.00 / $12.00 | -20% | Balanced / daily driver |
| GPT-5.6 Sol (standard) | $5.00 / $30.00 | $5.00 / $30.00 | Unchanged | Flagship agentic |
| GPT-5.6 Sol Fast (new) | — (Priority Processing) | $10.00 / $60.00 | New SKU | 2.5× speed, 2× price |
| Competitor | Input ($/M) | Output ($/M) | AA per-task cost* | Notes |
|---|---|---|---|---|
| Kimi K3 | $3.00 (cache $0.30) | $15.00 | ~$0.94 | ~6× vs K2.6 |
| DeepSeek V4 Pro | $0.435 (cache $0.0036) | $0.87 | — | Permanent 75% cut since May 2026 |
| DeepSeek V4 Flash | $0.14 | $0.28 | — | Lightweight tier |
| Claude Sonnet 5 | $3.00 (promo $2.00 through Aug 31) | $15.00 (promo $10.00) | — | Matches Kimi K3 standard rate |
| Gemini 3.5 Flash-Lite | ~$2.80 combined per 1M tokens | — | Google lightweight tier | |
| MAI-Code-1-Flash | $0.75 | $4.50 | — | GitHub Copilot only, no standalone API |
| GPT-5.6 Sol (standard) | $5.00 | $30.00 | ~$1.04 | AA per-task cost slightly above K3 |
*Artificial Analysis "per completed task" cost is a third-party methodology, not OpenAI official pricing. Competitor token rates from public July 2026 price sheets.
Artificial Analysis puts Kimi K3 at ~$0.94 per completed task and GPT-5.6 Sol at ~$1.04 — Sol still costs slightly more at the capability tier, but Luna's cut dramatically compresses unit cost for lightweight workloads.
Why Sol rewrote GPU kernels but OpenAI cut Luna instead of Sol
Bottom line: the July 29 engineering blog and July 30 tiered pricing are two sides of one strategy — use internal cost savings to fund lightweight SKU cuts, monetize latency via Fast mode, and keep standard Sol at flagship margin.
Sol's team says rewriting GPU kernels in Triton/Gluon, adding speculative decoding, and optimizing KV cache layout delivered roughly 20% inference cost reduction and 15% throughput improvement in internal tests, with FpSan numerical verification guarding precision. Those gains could fund Luna/Terra cuts — but standard Sol pricing didn't move, signaling OpenAI is routing cost wins toward volume tiers to counter DeepSeek V4 Flash and Luna-class competitors while preserving flagship margin.
Sol Fast pricing is straightforward: $10/$60 is 2× standard Sol for ~2.5× inference speed, replacing Priority Processing. Latency-sensitive production paths (real-time agent tool calls, interactive IDE completion) pay the premium; batch offline jobs should stay on standard Sol or discounted Terra. That contrasts with Kimi K3's "match Sonnet list price, differentiate on 1M context" play — OpenAI is tiering SKUs rather than cutting across the board.
After Kimi K3's July 27 weight release, Artificial Analysis pegged K3 per-task cost at $0.94 vs Sol at $1.04 — a 9.4% gap — while Reddit and X debates over METR reward hacking and vendor-reported benchmarks remain unresolved, with limited independent replication so far.
Six steps developers should take after the July 30 price cut
Export 30-day token detail: Split input, output, and cache hits by SKU (Luna / Terra / Sol / Sol Fast) to baseline pre-cut billing.
Remap task tiers: Route chat, classification, and simple completion to Luna ($0.20/$1.20); keep medium reasoning on Terra; reserve agentic / SWE work for standard Sol.
Model Sol Fast ROI: Benchmark latency-sensitive paths — 2.5× speed isn't worth 2× price for every workload; don't default Fast on batch jobs.
A/B Kimi K3 and DeepSeek V4 in parallel: Compare total cost on your own codebase using Artificial Analysis-style per-task methodology, not list $/M alone.
Update OpenRouter / LiteLLM routing: Write Luna's new rates into fallback chains; set daily spend caps so downgrade logic doesn't still point at old model IDs.
Hold a 7-day re-test window: Watch for further Sol or Terra moves; track Kimi K3 cache hit rates and K2.6 migration cost for existing users.
Controversies, three hard numbers, and a narrowing model-selection window
OpenAI's cuts split opinion: supporters call Luna's $0.20/M input a new era for high-volume apps; skeptics note Sol's 20% cost and 15% throughput figures are self-reported, METR reward-hacking discourse is still active, and Reddit benchmark-trust debates haven't settled. Meanwhile Kimi K3 beats Sol on AA per-task cost but is ~6× pricier than K2.6 — July 2026's "price cut season" isn't one-directional; vendors are re-tiering across speed premiums, SKU ladders, and open-weight anchors.
80% / Luna cut: Input $1→$0.20, output $6→$1.20 per OpenAI's July 30 announcement — the largest list-price move in the GPT-5.6 family.
$0.94 vs $1.04 / AA per-task cost: Artificial Analysis estimates Kimi K3 at ~$0.94 per completed task vs GPT-5.6 Sol at ~$1.04 — Sol still 9.4% higher at the capability tier, with Luna widening the gap on light workloads post-cut.
6× / K3 vs K2.6: Kimi K3 standard input $3/M vs K2.6 $0.60/M — the open-weight camp isn't uniformly cheaper; flagship open SKUs are re-anchoring too.
Note: Sol GPU optimization figures, METR controversies, and AA per-task costs are vendor or third-party estimates — validate on your own workloads; Sol Fast at 2× price doesn't automatically deliver 2× value.
Running multi-model A/B and agent load tests on a personal Mac hits sleep interruptions and network drops; relying entirely on cloud APIs during July's pricing churn means constant routing rewrites; waiting on enterprise GPU cluster approval rarely keeps pace with biweekly model and price iteration. For production iOS CI/CD, local inference, and 24/7 AI agent automation, KVMNODE dedicated Mac Mini M4 cloud rental is usually the better fit: Apple Silicon unified memory, full sudo access, flexible daily/weekly/monthly billing. See the pricing page, help center, or order directly.
Data as of July 31, 2026 · Sources: OpenAI July 30, 2026 pricing announcement, OpenAI July 29 Sol GPU optimization blog, Moonshot AI Kimi K3 release and weight open, Artificial Analysis per-task cost estimates, DeepSeek V4 official price sheet, Reddit / X community discussion, public METR-related coverage