Is Kimi K3 open source, or just open weight?
Short answer: no, not by the strict definition — and Moonshot AI itself does not claim otherwise. On July 27, 2026, Moonshot released the full weights and technical report for Kimi K3, a 2.8-trillion-parameter Mixture-of-Experts model with native vision understanding. The 1.56TB download hit #1 on Hugging Face's trending list within thirty minutes — making K3 the largest open-weight model ever shipped in full.
| Date | Milestone | Key detail |
|---|---|---|
| July 16 | API launch | Kimi App / Work / Code / API live; weights not yet public |
| July 17 | WAIC 2026 | State media reports call it the world's largest open model by parameter count |
| July 22–23 | Distillation dispute | White House adviser Kratsios accuses Moonshot of industrial distillation against Anthropic's Fable model |
| July 27 | Full weight release | Complete weights + tech report + MoonEP / AgentEnv (FlashKDA already open) |
| July 28 | MOFCOM response | China's Ministry of Commerce accuses Washington of "AI hegemonism" |
Conflating open weight with open source: Moonshot never uses "open source" — training data and full training code are not public.
Underestimating license gates: K3 adds a $20M MaaS revenue threshold and 100M MAU attribution requirement absent from the K2 family.
Assuming consumer hardware can run it: 2.8T parameters across 896 experts — Moonshot recommends 64+ accelerator supernodes.
Chasing parameters over cost: K3 is the open-weight capability ceiling, but ~$0.95/task vs GLM-5.2's ~$0.47.
Ignoring geopolitical compliance: The US-China distillation dispute layered on WAIC 2026 timing adds supply-chain risk for enterprise buyers.
Kimi K3 specs and architecture: KDA, AttnRes, and Per-Head Muon
| Spec | Value |
|---|---|
| Total parameters | 2.8 trillion |
| Active parameters | ~104 billion |
| Expert config | 896 routed experts, 16 activated per token (~1.8% sparsity) |
| Attention | Kimi Delta Attention (KDA) + Gated Multi-Head Latent Attention (MLA) |
| Context window | 1,000,000 tokens |
| Multimodality | Native vision (ViT-V2, 27 layers) |
| Weight format | MXFP4 weights, MXFP8 activations (quantization-aware from SFT) |
| Download size | ~1.56TB (Hugging Face) |
| License | Custom — Moonshot calls it "open weight," not "open source" |
Kimi Delta Attention (KDA) replaces a single scalar forgetting gate with channel-wise gating: every feature dimension gets its own decay rate, implemented as chunkwise DPLR for linear-time recurrence. K3 interleaves KDA with Gated MLA layers — the hybrid design behind the 1M-token window and minimal KV cache footprint.
Attention Residuals (AttnRes) replace uniform residual accumulation with selective, input-dependent aggregation across all preceding layers — roughly 25% higher training efficiency for under 2% added parameters.
Per-Head Muon optimizes each attention head independently. Stable LatentMoE uses Quantile Balancing with MoonEP's mathematical proof of redundant-expert upper bounds per compute rank.
| Component | Problem solved | Key gain |
|---|---|---|
| KDA | KV cache explosion on long sequences | Linear-time recurrence, minimal VRAM |
| AttnRes | Early-layer signal dilution at depth | ~25% training efficiency boost |
| Per-Head Muon | Uniform optimizer across heads | Frontier-class results at lower active-param count |
| Stable LatentMoE | Expert load imbalance at scale | 896 routed, 16 active — 1.8% sparsity |
Three infrastructure releases and benchmarks vs GPT-5.6, Claude, GLM-5.2
| Technology | Role | Key numbers |
|---|---|---|
| MoonEP | MoE communication at extreme scale | Duplicates overloaded experts; proves redundant-expert upper bound per rank |
| FlashKDA | CUTLASS KDA kernel | 1.72×–2.22× faster prefill on H20 vs flash-linear-attention baseline |
| AgentEnv | Firecracker microVM agent sandbox | Checkpoint ~133ms, resume ~49ms, up to 6.5× memory overcommit |
SWE-bench Verified (independent evaluation, Vals AI, July 2026):
| Model | Score | Release date |
|---|---|---|
| Claude Opus 5 | 97% | 2026-07-24 |
| GPT-5.6 Sol | 96.2% | 2026-07-09 |
| Claude Fable 5 | 95% | 2026-06-09 |
| Kimi K3 | 93.4% | 2026-07-16 |
| GLM-5.2 (max) | 51 (AA Index) | Previous open-weight leader |
On the Artificial Analysis Intelligence Index (max reasoning), Kimi K3 scores ~57 — #3 overall, #1 among open-weight models. It is the capability ceiling of the open-weight tier, but at roughly $0.95/task it costs more than GLM-5.2 (~$0.47). K3 currently ranks #1 on Arena.ai's Frontend Code Arena leaderboard.
License commercial gates and six-step deployment checklist
Moonshot consistently calls this an "open-weight" release, never "open source." The license is a fully custom document with two commercial thresholds that did not exist in the K2 family:
| Gate | Trigger | Requirement |
|---|---|---|
| MaaS revenue gate | Model-as-a-Service revenue exceeds $20M over any trailing 12 months | Separate commercial agreement with Moonshot AI |
| Attribution gate | 100M+ monthly active users or $20M+ monthly revenue | Prominently display "Kimi K3" in the product UI |
| Token type | Price per million tokens |
|---|---|
| Input (cache hit) | $0.30 |
| Input (cache miss) | $3.00 |
| Output (incl. reasoning trace) | $15.00 |
Moonshot's disaggregated Mooncake serving architecture reportedly holds cache hit rates above 90% on typical coding workloads — most real-world usage lands closer to the $0.30 rate.
Register API access: Create a key at platform.kimi.ai, set base_url to https://api.moonshot.ai/v1, model ID kimi-k3.
Review the LICENSE: Before downloading Hugging Face weights, legal should confirm whether MaaS revenue or MAU gates apply to your business.
OpenAI SDK-compatible integration: Most existing projects only need a base URL and API key change.
Cache prefix optimization: Reuse system prompts and tool-definition prefixes — coding agents can hit 90%+ cache rates, effective input cost drops to $0.30/M.
Self-hosting hardware check: Confirm 64+ accelerator supernode capacity; for most teams, OpenRouter or similar hosts (7 providers already live) is the practical path.
Pick an inference stack: Deploy via vLLM (KDA + prefix caching) or SGLang; FlashKDA drops in as a direct chunk_kda backend replacement.
from openai import OpenAI
client = OpenAI(
api_key="your_moonshot_api_key",
base_url="https://api.moonshot.ai/v1"
)
response = client.chat.completions.create(
model="kimi-k3",
messages=[{"role": "user", "content": "Analyze the complexity of this code..."}]
)US-China AI dispute context and three hard numbers
Kimi K3's open-weight release landed during WAIC 2026 and days after the US-China "model distillation" dispute escalated: White House science adviser Michael Kratsios accused Moonshot of running a large-scale distillation campaign against Anthropic's Fable model, with Treasury Secretary Scott Bessent floating sanctions. China's Ministry of Commerce pushed back on July 28, accusing Washington of "AI hegemonism." Three days after K3 shipped, Alibaba unveiled Qwen3.8-Max-Preview — widely read as a direct response, pushing the open-weight race into the "3T club."
2.8T / #1 open-weight: The largest open-weight model ever shipped in full — 1.56TB on Hugging Face, #1 trending within 30 minutes.
93.4% / #3 globally: Independent SWE-bench Verified score of 93.4%; AA Index ~57 — third overall, first among open-weight models.
11 days / three infra releases: From API launch to full weight release in 11 days, with MoonEP, FlashKDA, and AgentEnv open-sourced alongside.
The practical alternatives all carry hidden costs: running Kimi Code agents on a personal Mac breaks on sleep and network drops; self-hosting full K3 needs 64+ accelerators most teams cannot provision; relying on a single closed API sacrifices K3's 1M-token flat pricing advantage. For production iOS CI/CD pipelines and 24/7 AI agent automation, KVMNODE dedicated Mac Mini cloud rental is usually the better fit: native Apple Silicon, full sudo access, flexible daily/weekly/monthly billing. See the pricing page, help center, or order directly.
Data as of July 28, 2026 · Sources: Moonshot AI official blog, Hugging Face, GitHub moonshotai/FlashKDA, Vals AI SWE-bench Verified, Artificial Analysis Intelligence Index