Moonshot AI
Kimi K3
Near-S intelligence held in A tier by real-world cost and quota friction.
Canonical model record
Current identity, limits, and pricing
- Status
- Current
- API model ID
kimi-k3- Context
- 1.05M
- Max output
- 131K
- API price / 1M tokens
- $3 input · $15 output
Suggested for this guide
Kimi
Use the same Kimi family tested on this page. Check the current plan and model access before subscribing.
Best for: Claude Code setups and cost-aware model work
Check Kimi plansPartner link. It supports Superbash Learn at no extra cost to you.Visual prompt runs
Benchmark runs
Open each generated scene, or compare the same prompt across models.

Helm's Deep
Fortress siege scene testing scale, lighting, architecture, and cinematic atmosphere.

Hogwarts Broom Flight Simulator
Broom-flight scene testing depth, motion cues, castle scale, and fantasy mood.

Jabberwock
Dark fantasy encounter testing creature design, forest mood, and narrative staging.

Low Poly World
Stylized island build testing composition, color, and low-poly worldbuilding.

Mechanical Watch Simulator
Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Office Life
Workplace vignette testing everyday scene logic, objects, and believable office detail.

Petri Dish
Microscopic ecosystem testing organic forms, scientific clarity, and cellular detail.

Universe Simulator
Cosmic system testing orbital structure, glowing bodies, scale, and simulation readability.

Vice City
Neon coastal city testing vehicles, architecture, atmosphere, and dense urban layout.

Yingzao Fashi Assembly
Timber assembly scene testing structure, joinery, construction order, and material clarity.
Official benchmark profile
How Kimi K3 scores beyond our visual tests.
Moonshot reports Kimi K3 at maximum reasoning effort across coding, research, tool use, document work, and vision. The official table mixes provider harnesses on some agent tasks, so each row keeps its evaluation setting.
GPQA Diamond
93.5%DeepSWE
67.5%Terminal-Bench 2.1
88.3%Full official benchmark table12 rows with source settings and peer charts
| Benchmark | Area | Score | Setting / comparison |
|---|---|---|---|
| GPQA Diamond | Academic reasoning | 93.5% | Maximum reasoning effort, temperature 1.0, top-p 0.95. |
| DeepSWE | Coding | 67.5% | Kimi Code harness on DeepSWE v1.1 tasks.The official leaderboard score under mini-SWE-agent is 67.3%. |
| Terminal-Bench 2.1 | Agentic coding | 88.3% | Kimi Code harness; peers use each provider's cited best harness. |
| SWE-Marathon | Long-horizon coding | 42.0% | Claude Code for Kimi and Claude; Codex for GPT-5.6 on the July 9 H20-calibrated task branch. |
| BrowseComp | Agentic research | 91.2% | Context compaction at 300K tokens; the full 1M-context run without context management scored 90.4%. |
| DeepSearchQA (F1) | Deep research | 95.0% | Maximum reasoning effort, temperature 1.0, top-p 1.0. |
| MCPMark-Verified | Tool use | 94.5% | Maximum reasoning effort, temperature 1.0, top-p 1.0. |
| AutomationBench | Workflow automation | 30.8% | 600-task public subset with the official benchmark setup. |
| SpreadsheetBench 2 | Spreadsheet work | 34.8% | Claude Code for Kimi and Claude; Codex for GPT models. |
| OmniDocBench | Document understanding | 91.1% | Maximum reasoning effort; average over three multimodal runs. |
| MMMU-Pro | Multimodal reasoning | 81.6% / 83.4% | Without tools / with Python; original input order; average over three runs. |
| CharXiv (RQ) | Chart reasoning | 84.8% / 91.3% | Without tools / with Python; average over three runs. |

Superbash commentary
Our take
Kimi K3 sits between A and S on capability, but lands in A because it is no longer the cheap Chinese-model bargain the team expected. It integrates broadly and remains highly capable, yet the team exhausted both short-window and total monthly plan allowances by mid-August.
Best for
Watch out
Why it is ranked here
Evidence and commentary
Superbash editorial model ranking
Takeaway: Kimi K3 is currently placed in Tier A.
The August 2026 editorial roster places Kimi K3 at rank 4.
Open source →Kimi K3: Open Frontier Intelligence
Takeaway: Moonshot’s release table reports competitive coding, agentic, and vision results, including 88.3% on Terminal-Bench 2.1 and 81.2% on FrontierSWE.
These are vendor-reported results at maximum reasoning effort and vary by agent harness, so they establish K3’s frontier capability but are not a like-for-like overall score.
Open source →Superbash visual benchmark suite: Kimi K3
Takeaway: K3 completed all nine shared visual benchmark briefs in the Superbash suite.
The outputs provide a practical check of visual implementation quality alongside published benchmark results.
Open source →Keep learning