What's actually confirmed vs. what's just Musk's word
Only one thing is on the record here: a tweet. Everything else — the "1.5 trillion parameters," the "significantly improved SFT & RL," the follow-up Grok 4.7 at 2.1 trillion parameters — comes from that single X post, with no accompanying model card, benchmark suite, or pricing page the way xAI published for Grok 4.5. That distinction matters for anyone deciding whether to build around this release.
| Date | Milestone | Key detail |
|---|---|---|
| Jul 8, 2026 | Grok 4.5 ships | Co-trained with Cursor on real developer sessions; 500K-token context; $2/$6 per million input/output tokens; published model card with 15 benchmark scores |
| Jul 16–26, 2026 | Kimi K3 goes open-weight | Moonshot AI releases 2.8T-parameter model with 1M-token context; tops Hugging Face trending; Musk praised it in benchmark comment threads |
| Jul 28, 2026 | Musk posts roadmap | Grok 4.6/4.7 timeline in reply to Rauch; same day 1,200+ employees publish "Pacing the Frontier" letter |
| ~Aug 7, 2026 | Grok 4.6 (target) | 1.5T parameters; positioned as SFT/RL upgrade, not raw scale-up |
| Late Aug–early Sep | Grok 4.7 (estimated) | 2.1T parameters; Musk says better than 4.6 except slightly slower to serve, with better token efficiency |
"Musk time" has a track record — xAI, Tesla, and SpaceX timelines from Musk have historically slipped by days to weeks. Treat "around August 7" as a target, not a guarantee.
Treating a tweet as an official announcement: No xAI blog post, model card, or product page confirms Grok 4.6 specs or date yet.
Assuming 1.5T params equals a capability leap: Musk explicitly framed the upgrade around SFT & RL, not parameter count alone.
Ignoring Kimi K3 competitive pressure: Grok 4.6's target lands ~10 days after K3's full open-weight release rattled the industry.
Conflating Grok 4.6 and 4.7 positioning: 4.6 is the post-training refinement SKU; 4.7 trades latency for scale and token efficiency.
Overlooking xAI safety controversies: In July 2026 xAI sued a user over CSAM generated by bypassing Grok safeguards — relevant vendor-risk context for enterprises.
The numbers so far: Grok vs the field
| Model | Date | Parameters | Context | Pricing (in/out per 1M) | Status |
|---|---|---|---|---|---|
| Grok 4.5 | Jul 8, 2026 | Undisclosed | 500K | $2 / $6 | Shipped, benchmarked |
| Grok 4.6 (announced) | ~Aug 7, 2026 | 1.5T | Undisclosed | Undisclosed | Tweet only, unshipped |
| Grok 4.7 (announced) | ~late Aug–Sep | 2.1T | Undisclosed | Undisclosed | Tweet only, unshipped |
| Kimi K3 | Jul 26, 2026 | 2.8T (MoE) | 1M | $0.30–3 in / $15 out | Shipped, verified scores |
| Claude Fable 5.1 (rumored) | Rumored Aug | Undisclosed | Undisclosed | Rumored $10/$50 | Unconfirmed by Anthropic |
| GPT-5.6 Sol | Jun 2026 | Undisclosed | Undisclosed | Undisclosed | Shipped |
All Grok 4.6/4.7 figures are unverified vendor claims from a single social media post — treat them as directional, not confirmed specs.
Kimi K3 topped the Frontend Code Arena leaderboard at 1,679 points — the first open-weight model to beat every closed model on that board — and ranked third on Artificial Analysis's Intelligence Index.
Why xAI is emphasizing post-training, not just scale
Bottom line: Grok 4.6's story is post-training refinement (SFT + RL), not a pure parameter bump.
Supervised fine-tuning (SFT) trains a model on curated example outputs to shape its behavior; reinforcement learning (RL) uses reward signals to teach a model which action sequences actually work — critical for multi-step agentic tasks. Musk's choice of words — "significantly improved SFT & RL" — signals xAI is doubling down on the same playbook that made Grok 4.5 competitive on agentic benchmarks: Grok 4.5 used roughly 15,954 output tokens per SWE-Bench Pro task versus Opus 4.8's 67,020, a 4.2x efficiency gap, largely credited to post-training on real Cursor developer sessions rather than sheer model size.
Grok 4.6's jump to 1.5T parameters is a real scale increase, but Musk's framing of Grok 4.7 — bigger at 2.1T, "better in every way except slightly slower to serve" — suggests xAI is deliberately building two SKUs with different trade-offs rather than one model for everything.
The same day Musk announced Grok 4.6/4.7, over 1,200 employees at OpenAI, Anthropic, Google DeepMind, and Meta published "Pacing the Frontier," asking the US government to help deliberately slow automated AI development — and both OpenAI and Anthropic endorsed it as companies. xAI is notably absent from that list.
Six steps to track Grok 4.6 before it ships
Watch official xAI channels: Follow x.ai/blog, @xai, and @elonmusk — treat official announcements as ground truth, not secondhand coverage.
Build in "Musk time" buffer: Historical slips of days to weeks are common; don't lock project timelines to a single date.
Baseline against Grok 4.5: Record current API pricing ($2/$6 per million tokens) and published benchmarks like SWE-Bench Pro 64.7% as your comparison yardstick.
Evaluate Kimi K3 and Claude Fable 5.1 in parallel: August may be a release pile-up month — avoid long-term contracts before multiple vendors ship.
Test token efficiency on real tasks: Run your own codebase through SWE/agent workloads after launch to compare total cost, not just per-token list price.
Include safety compliance review: xAI's July CSAM lawsuit is relevant vendor-risk context before enterprise integration.
Controversies, three hard numbers, and August's release bottleneck
If Musk's timeline holds, Grok 4.6 and Grok 4.7 will land in the same month as a rumored Claude Fable 5.1 and just weeks after Kimi K3's open-weight shock. For teams evaluating models, that compresses the useful shelf life of any single flagship to a matter of weeks — which makes token efficiency and real-world task cost, not leaderboard rank alone, the more durable basis for a model choice.
1.5T / Grok 4.6 parameters: Musk's July 28 X post — no third-party verification or official model card yet.
4.2x / token efficiency gap: Grok 4.5 used ~15,954 output tokens per SWE-Bench Pro task vs Opus 4.8's ~67,020 — a verified post-training advantage from 4.5.
1,679 pts / Kimi K3 Frontend Code Arena: First open-weight model to top that coding leaderboard; third on Artificial Analysis Intelligence Index.
Note: Grok 4.6 benchmarks and pricing are blank; xAI sued a user in July over CSAM generated by bypassing safeguards — factor vendor risk into enterprise decisions.
Running agent evals and multi-model A/B tests on a personal Mac hits sleep interruptions and network drops; relying entirely on cloud APIs during an August release pile-up means constant migration overhead; waiting on enterprise GPU cluster approval rarely keeps pace with frontier-model cadence. For production iOS CI/CD, local inference, and 24/7 AI agent automation, KVMNODE dedicated Mac Mini M4 cloud rental is usually the better fit: Apple Silicon unified memory, full sudo access, flexible daily/weekly/monthly billing. See the pricing page, help center, or order directly.
Data as of July 30, 2026 · Sources: xAI Grok 4.5 announcement, Musk's July 28 X post, Moonshot AI Kimi K3 release, Chain Tech Daily / Tron Weekly coverage, The Verge / Tech Times on "Pacing the Frontier," Ars Technica / TechCrunch on xAI safety controversies