Architecture
MoE + KDA Hybrid Attention
Sparse MoE built on Kimi Delta Attention (KDA), a hybrid linear-attention mechanism with attention residuals
Parameters
2.8T Total / 896 Experts
2.8 trillion total parameters across 896 experts, with 16 experts activated per token
Context Length
1M Tokens
Up to 1M tokens for Allegretto tier and above; Moderato tier gets 256K
Modalities
Text + Vision
Native visual understanding built in from the ground up
Reasoning
Thinking Effort: max
Ships with max thinking effort (reasoning_effort: max); low and high tiers roll out later
Code Capability
Flagship Coding
Moonshot's strongest model for coding, game/3D and knowledge tasks — coding scores surpass Claude Fable 5
API Price (per 1M tokens)
$3 In / $15 Out
Cached input just $0.30/1M — Mooncake serving keeps coding cache rates above 90%, cutting real input cost ~4×
Open Source
Open Weights · Modified MIT
Full model weights land by July 27, 2026 under a Modified MIT license — the first open 3T-class model
* All specifications are sourced from Moonshot AI's official documentation and launch announcements (July 2026).