Kimi K3 is an open-weight large language model launched by Moonshot AI that features a 1-million token Context window built for reasoning tasks across very long inputs. For example, a developer could paste an entire multi-file code repository into one prompt and have the model reason across all of it at once. Founded in 2023, Moonshot followed three earlier Kimi generations, K2, K2.5, and K2.6, before shipping K3 in July 2026. It runs on a 2.8-trillion-parameter mixture-of-experts architecture, meaning it activates only a subset of its parameters per token, and Moonshot credits most of its roughly 2.5 times efficiency gain over K2 to two changes it calls Kimi Delta Attention and Attention Residuals.
Moonshot priced Kimi K3's inference cost at USD 3 per million input tokens and USD 15 per million output tokens, undercutting the USD 10 and USD 50 Anthropic charges for Claude Fable 5, a gap that pressures how closed-model providers justify higher prices for similar coding work. Reporting on independent benchmark placements found K3 scoring second on AA-Briefcase, a private test of long-horizon agentic knowledge work, behind Fable 5 and ahead of GPT-5.6 Sol, suggesting the price gap does not come with a proportional capability gap.
Agent Swarm, a feature built into Kimi K3, lets one orchestrator agent direct up to 300 sub-agents across as many as 4,000 parallel workflow steps, fanning out to search, download, categorize, and summarize hundreds of articles into thematic folders without a person assigning each step. That scale means a deployer must set permissions and cost limits for hundreds of sub-agents at once rather than reviewing one model's output. Semgrep, a company that builds code-scanning tools, tested K3 on security-relevant coding tasks and found its precision fell off on large, enterprise-scale codebases, meaning it grew worse at telling a real vulnerability from a false flag as the codebase got larger. This gap lands hardest with write access to production code, where a sub-agent can commit on a bad finding before a person reviews it.



