For developers and technical leads tracking GPT-5.6 price cuts, OpenAI API pricing strategy, and the July 2026 LLM price war: on July 30, 2026, OpenAI slashed Luna input/output prices 80%, Terra 20%, kept Sol standard at $5/$30, and launched Sol Fast at $10/$60 (2.5× speed, 2× price, replacing Priority Processing). This article covers the July 9–30 timeline, before/after pricing vs Kimi K3, DeepSeek, and Claude, why Sol rewrote its GPU kernels but didn't get cheaper, a six-step developer response playbook, and controversies with three citeable numbers. See also Kimi K3 open-weight coverage and the Hugging Face hack timeline.
01

July 9 launch to July 30 price cut: what happened in OpenAI's 21-day sprint

On July 30, 2026, OpenAI adjusted GPT-5.6 pricing: Luna dropped from $1/$6 to $0.20/$1.20 per million tokens (-80%), Terra from $2.50/$15 to $2.00/$12 (-20%), Sol standard held at $5/$30, and a new Sol Fast SKU launched at $10/$60 with roughly 2.5× inference speed, replacing the old Priority Processing tier. The cuts landed days after Kimi K3's full weight release and OpenAI's own Sol GPU self-optimization blog — widely read as a competitive response to open-weight pressure.

DateMilestoneKey detail
Jul 9, 2026GPT-5.6 family shipsSol, Terra, and Luna tiers go live; Sol positioned as flagship agentic coding model
Jul 16, 2026Kimi K3 hosted APIMoonshot AI launches Kimi K3 as a managed API, resetting coding-model price expectations
Jul 27, 2026Kimi K3 weights open2.8T-parameter MoE fully released; Artificial Analysis puts per-task cost at ~$0.94
Jul 29, 2026Sol GPU optimization blogOpenAI discloses Triton/Gluon kernel rewrites, speculative decoding, and KV cache work
Jul 30, 2026Official price cutLuna -80%, Terra -20%, Sol standard unchanged; Sol Fast $10/$60 new SKU live

The nuance: Luna and Terra got list-price cuts; Sol's "price hike" narrative comes from the new Fast tier — standard Sol didn't move, but latency-sensitive users now face 2× per-token pricing if they opt in.

01

Treating Luna's cut as a family-wide discount: Sol standard stays at $5/$30; only the new Fast mode adds a premium SKU — flagship total cost didn't drop.

02

Ignoring Kimi K3's pricing anchor: K3's $3/$15 headline matches Claude Sonnet 5, but vs K2.6 ($0.60/$2.50) it's roughly 6× more expensive — open-weight tiers are re-tiering too.

03

Booking savings from Sol's self-reported optimization: OpenAI claims 20% inference cost reduction and 15% throughput gain from GPU work — vendor-tested, not yet independently verified.

04

Assuming Sol Fast is just a Priority Processing rename: 2.5× speed at 2× price doesn't automatically lower cost per wall-clock hour — model it per SLA.

05

Picking sides in polarized Reddit threads: Community reaction split between celebrating "penny Luna" and questioning METR reward-hacking and self-reported benchmarks — enterprise picks need workload proof, not hot takes.

02

Before and after: GPT-5.6 tiers vs Kimi K3, DeepSeek, Claude, and more

Model / SKUBefore (in/out per 1M)AfterChangeRole
GPT-5.6 Luna$1.00 / $6.00$0.20 / $1.20-80%Light / high-volume
GPT-5.6 Terra$2.50 / $15.00$2.00 / $12.00-20%Balanced / daily driver
GPT-5.6 Sol (standard)$5.00 / $30.00$5.00 / $30.00UnchangedFlagship agentic
GPT-5.6 Sol Fast (new)— (Priority Processing)$10.00 / $60.00New SKU2.5× speed, 2× price
CompetitorInput ($/M)Output ($/M)AA per-task cost*Notes
Kimi K3$3.00 (cache $0.30)$15.00~$0.94~6× vs K2.6
DeepSeek V4 Pro$0.435 (cache $0.0036)$0.87Permanent 75% cut since May 2026
DeepSeek V4 Flash$0.14$0.28Lightweight tier
Claude Sonnet 5$3.00 (promo $2.00 through Aug 31)$15.00 (promo $10.00)Matches Kimi K3 standard rate
Gemini 3.5 Flash-Lite~$2.80 combined per 1M tokensGoogle lightweight tier
MAI-Code-1-Flash$0.75$4.50GitHub Copilot only, no standalone API
GPT-5.6 Sol (standard)$5.00$30.00~$1.04AA per-task cost slightly above K3

*Artificial Analysis "per completed task" cost is a third-party methodology, not OpenAI official pricing. Competitor token rates from public July 2026 price sheets.

Artificial Analysis puts Kimi K3 at ~$0.94 per completed task and GPT-5.6 Sol at ~$1.04 — Sol still costs slightly more at the capability tier, but Luna's cut dramatically compresses unit cost for lightweight workloads.

03

Why Sol rewrote GPU kernels but OpenAI cut Luna instead of Sol

Bottom line: the July 29 engineering blog and July 30 tiered pricing are two sides of one strategy — use internal cost savings to fund lightweight SKU cuts, monetize latency via Fast mode, and keep standard Sol at flagship margin.

Sol's team says rewriting GPU kernels in Triton/Gluon, adding speculative decoding, and optimizing KV cache layout delivered roughly 20% inference cost reduction and 15% throughput improvement in internal tests, with FpSan numerical verification guarding precision. Those gains could fund Luna/Terra cuts — but standard Sol pricing didn't move, signaling OpenAI is routing cost wins toward volume tiers to counter DeepSeek V4 Flash and Luna-class competitors while preserving flagship margin.

Sol Fast pricing is straightforward: $10/$60 is 2× standard Sol for ~2.5× inference speed, replacing Priority Processing. Latency-sensitive production paths (real-time agent tool calls, interactive IDE completion) pay the premium; batch offline jobs should stay on standard Sol or discounted Terra. That contrasts with Kimi K3's "match Sonnet list price, differentiate on 1M context" play — OpenAI is tiering SKUs rather than cutting across the board.

After Kimi K3's July 27 weight release, Artificial Analysis pegged K3 per-task cost at $0.94 vs Sol at $1.04 — a 9.4% gap — while Reddit and X debates over METR reward hacking and vendor-reported benchmarks remain unresolved, with limited independent replication so far.

04

Six steps developers should take after the July 30 price cut

01

Export 30-day token detail: Split input, output, and cache hits by SKU (Luna / Terra / Sol / Sol Fast) to baseline pre-cut billing.

02

Remap task tiers: Route chat, classification, and simple completion to Luna ($0.20/$1.20); keep medium reasoning on Terra; reserve agentic / SWE work for standard Sol.

03

Model Sol Fast ROI: Benchmark latency-sensitive paths — 2.5× speed isn't worth 2× price for every workload; don't default Fast on batch jobs.

04

A/B Kimi K3 and DeepSeek V4 in parallel: Compare total cost on your own codebase using Artificial Analysis-style per-task methodology, not list $/M alone.

05

Update OpenRouter / LiteLLM routing: Write Luna's new rates into fallback chains; set daily spend caps so downgrade logic doesn't still point at old model IDs.

06

Hold a 7-day re-test window: Watch for further Sol or Terra moves; track Kimi K3 cache hit rates and K2.6 migration cost for existing users.

05

Controversies, three hard numbers, and a narrowing model-selection window

OpenAI's cuts split opinion: supporters call Luna's $0.20/M input a new era for high-volume apps; skeptics note Sol's 20% cost and 15% throughput figures are self-reported, METR reward-hacking discourse is still active, and Reddit benchmark-trust debates haven't settled. Meanwhile Kimi K3 beats Sol on AA per-task cost but is ~6× pricier than K2.6 — July 2026's "price cut season" isn't one-directional; vendors are re-tiering across speed premiums, SKU ladders, and open-weight anchors.

A

80% / Luna cut: Input $1→$0.20, output $6→$1.20 per OpenAI's July 30 announcement — the largest list-price move in the GPT-5.6 family.

B

$0.94 vs $1.04 / AA per-task cost: Artificial Analysis estimates Kimi K3 at ~$0.94 per completed task vs GPT-5.6 Sol at ~$1.04 — Sol still 9.4% higher at the capability tier, with Luna widening the gap on light workloads post-cut.

C

6× / K3 vs K2.6: Kimi K3 standard input $3/M vs K2.6 $0.60/M — the open-weight camp isn't uniformly cheaper; flagship open SKUs are re-anchoring too.

Note: Sol GPU optimization figures, METR controversies, and AA per-task costs are vendor or third-party estimates — validate on your own workloads; Sol Fast at 2× price doesn't automatically deliver 2× value.

Running multi-model A/B and agent load tests on a personal Mac hits sleep interruptions and network drops; relying entirely on cloud APIs during July's pricing churn means constant routing rewrites; waiting on enterprise GPU cluster approval rarely keeps pace with biweekly model and price iteration. For production iOS CI/CD, local inference, and 24/7 AI agent automation, KVMNODE dedicated Mac Mini M4 cloud rental is usually the better fit: Apple Silicon unified memory, full sudo access, flexible daily/weekly/monthly billing. See the pricing page, help center, or order directly.

Data as of July 31, 2026 · Sources: OpenAI July 30, 2026 pricing announcement, OpenAI July 29 Sol GPU optimization blog, Moonshot AI Kimi K3 release and weight open, Artificial Analysis per-task cost estimates, DeepSeek V4 official price sheet, Reddit / X community discussion, public METR-related coverage