Moonshot AI

Tier A · Strong with clear operating rules.

Kimi K3

Kimi K3 has near-S capability, but cost and quota limits make it an A-tier operating choice.

API pricing

USD per million tokens · Kimi API
Input
$3
Output
$15
Cached input
$0.3
Context window
1.05M
Pricing details and specifications

Suggested for this guide

Kimi

Use the same Kimi family tested on this page. Check the current plan and model access before subscribing.

Best for: Claude Code setups and cost-aware model work

Check Kimi plansPartner link. It supports Superbash Learn at no extra cost to you.

Try the visual runs

Open the generated scene to test it, or compare that prompt with another model.

10 runs

Mechanical Watch Simulator

Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Editorial verdict

Where this model fits

Back to all benchmarks →

Kimi K3 handles long coding, tool use, and visual iteration well. The team exhausted short-window and monthly plan allowances during regular use, so this placement accounts for access and cost as well as raw capability.

Best for

  • Long-horizon coding in Kimi Code
  • Large-repository and tool-using tasks
  • Visual iteration against screenshots
  • Knowledge work with long context

Watch out

  • Monthly and short-window plan limits affected the team’s use
  • Premium inference pricing
  • Use a K3-compatible harness and preserve thinking history
  • Avoid switching to K3 mid-session

Why it is ranked here

  1. Moonshot reports 67.5 on DeepSWE, 88.3 on Terminal-Bench 2.1, 81.2 on FrontierSWE, and 77.8 on Program Bench at maximum effort. Harness differences prevent a universal ranking.
  2. K3 completed nine Superbash visual prompts, providing inspectable output beyond vendor scores.
  3. The A-tier placement reflects exhausted subscription quotas and cost, not a downgrade in raw intelligence.

Evidence

2026-09-07

Superbash editorial model ranking

Takeaway: Kimi K3 is currently placed in Tier A.

The September 2026 editorial roster places Kimi K3 at rank 4.

Open source →
2026-07-16

Kimi K3: Open Frontier Intelligence

Takeaway: Moonshot’s release table reports competitive coding, agentic, and vision results, including 88.3% on Terminal-Bench 2.1 and 81.2% on FrontierSWE.

These are vendor-reported results at maximum reasoning effort and vary by agent harness, so they establish K3’s frontier capability but are not a like-for-like overall score.

Open source →
2026-07-16

Superbash visual benchmark suite: Kimi K3

Takeaway: K3 completed all nine shared visual benchmark briefs in the Superbash suite.

The outputs provide a practical check of visual implementation quality alongside published benchmark results.

Open source →

How Kimi K3 compares

Three relevant peers, side by side.

  • Kimi K3 This model
  • Claude Fable 5
  • GPT-5.6 Sol
  • Claude Opus 4.8

6 benchmarks · 4 models

GPQA Diamond

Academic reasoning · Higher is better

Source
Kimi K393.5%
Claude Fable 592.6%
GPT-5.6 Sol94.1%
Claude Opus 4.8Not reported

DeepSWE

Coding · Higher is better

Source
Kimi K367.5%
Claude Fable 570.0%
GPT-5.6 Sol73.0%
Claude Opus 4.859.0%

Terminal-Bench 2.1

Agentic coding · Higher is better

Source
Kimi K388.3%
Claude Fable 588.0%
GPT-5.6 Sol88.8%
Claude Opus 4.884.6%

SWE-Marathon

Long-horizon coding · Higher is better

Source
Kimi K342.0%
Claude Fable 535.0%
GPT-5.6 Sol39.0%
Claude Opus 4.840.0%

BrowseComp

Agentic research · Higher is better

Source
Kimi K391.2%
Claude Fable 588.0%
GPT-5.6 Sol90.4%
Claude Opus 4.8Not reported

DeepSearchQA (F1)

Deep research · Higher is better

Source
Kimi K395.0%
Claude Fable 594.2%
GPT-5.6 SolNot reported
Claude Opus 4.893.1%

Official benchmark profile

Published scores outside our visual tests

Moonshot AI sourceAugust 2026Source report →
Academic reasoning

GPQA Diamond

93.5%
Coding

DeepSWE

67.5%
Agentic coding

Terminal-Bench 2.1

88.3%
Full official benchmark table12 scores and peer comparisons
BenchmarkAreaScoreComparison
GPQA DiamondAcademic reasoning93.5%
Kimi K393.5%
GPT-5.6 Sol94.1%
Claude Fable 592.6%
GLM 5.291.2%
DeepSWECoding67.5%
Kimi K367.5%
GPT-5.567.0%
Claude Fable 570.0%
GPT-5.6 Sol73.0%
Terminal-Bench 2.1Agentic coding88.3%
Kimi K388.3%
Claude Fable 588.0%
GPT-5.6 Sol88.8%
Claude Opus 4.884.6%
SWE-MarathonLong-horizon coding42.0%
Kimi K342.0%
Claude Opus 4.840.0%
GPT-5.6 Sol39.0%
Claude Fable 535.0%
BrowseCompAgentic research91.2%
Kimi K391.2%
GPT-5.6 Sol90.4%
Claude Fable 588.0%
GPT-5.584.4%
DeepSearchQA (F1)Deep research95.0%
Kimi K395.0%
Claude Fable 594.2%
Claude Opus 4.893.1%
MCPMark-VerifiedTool use94.5%
Kimi K394.5%
GPT-5.6 Sol92.9%
GPT-5.592.9%
Claude Fable 587.4%
AutomationBenchWorkflow automation30.8%
Kimi K330.8%
GPT-5.6 Sol29.7%
Claude Fable 529.1%
Claude Opus 4.827.2%
SpreadsheetBench 2Spreadsheet work34.8%
Kimi K334.8%
Claude Fable 534.7%
GPT-5.6 Sol32.4%
Claude Opus 4.831.6%
OmniDocBenchDocument understanding91.1%
Kimi K391.1%
Claude Fable 589.8%
GPT-5.589.4%
Claude Opus 4.887.9%
MMMU-ProMultimodal reasoning81.6% / 83.4%
Kimi K381.6% / 83.4%
Claude Fable 581.2% / 86.5%
GPT-5.581.2% / 83.2%
GPT-5.6 Sol83.0% / 84.6%
CharXiv (RQ)Chart reasoning84.8% / 91.3%
Kimi K384.8% / 91.3%
GPT-5.6 Sol84.6% / 89.1%
GPT-5.584.1% / 89.0%
Claude Fable 588.9% / 93.5%
Sources, pricing details and methodology

Verified model facts

Model specifications

Pricing source · checked 2026-09-08 ↗
Status
Current
API model ID
kimi-k3
Context
1.05M
Max output
131K
License
Not publicly verified
Access route
Not publicly verified
Modalities
Not publicly verified
Hardware
Not publicly verified
Kimi API · USD / 1M tokens
$3 input · $0.3 cached input · $15 output

Prices exclude applicable taxes and separately billed tools.

Benchmark sources

Moonshot reports Kimi K3 at maximum reasoning effort across coding, research, tool use, document work, and vision. The official table mixes provider harnesses on some agent tasks, so each row keeps its evaluation setting.

Kimi K3 Technical Report

Evaluation settings

  • Maximum reasoning effort, temperature 1.0, top-p 0.95.
  • Kimi Code harness on DeepSWE v1.1 tasks.
  • Kimi Code harness; peers use each provider's cited best harness.
  • Claude Code for Kimi and Claude; Codex for GPT-5.6 on the July 9 H20-calibrated task branch.
  • Context compaction at 300K tokens; the full 1M-context run without context management scored 90.4%.
  • Maximum reasoning effort, temperature 1.0, top-p 1.0.
  • 600-task public subset with the official benchmark setup.
  • Claude Code for Kimi and Claude; Codex for GPT models.
  • Maximum reasoning effort; average over three multimodal runs.
  • Without tools / with Python; original input order; average over three runs.
  • Without tools / with Python; average over three runs.
DeepSWE
The official leaderboard score under mini-SWE-agent is 67.3%.