Thinking Machines Lab

Inkling

Version 1.0 benchmark runs across the shared Superbash visual prompts.

API pricing

USD per million tokens · Tinker serverless inference beta (256K)
Input
$1
Output
$4.05
Cached input
$0.17
Context window
1M
Pricing details and specifications

Try the active visual tests

Open the generated scene to test it, or compare that prompt with another model.

3 runs

Yingzao Fashi Assembly

Run date: 2026-07-16

Specialist test · evaluated separately

Timber assembly scene testing structure, joinery, construction order, and material clarity.

Retired tests · 6 archived runs

These tests are no longer in the active suite. Their original results remain available for inspection.

How Inkling compares

Three relevant peers, side by side.

  • Inkling This model
  • Claude Fable 5
  • GPT-5.6 Sol
  • Gemini 3.1 Pro

6 benchmarks · 4 models

AIME 2026

Math · Higher is better

Source
Inkling97.1%
Claude Fable 599.9%
GPT-5.6 Sol99.9%
Gemini 3.1 ProNot reported

GPQA Diamond

Academic reasoning · Higher is better

Source
Inkling87.2%
Claude Fable 592.6%
GPT-5.6 Sol94.1%
Gemini 3.1 Pro94.1%

SWE-bench Verified

Agentic coding · Higher is better

Source
Inkling77.6%
Claude Fable 595.0%
GPT-5.6 Sol82.2%
Gemini 3.1 ProNot reported

SWE-bench Pro Public

Agentic coding · Higher is better

Source
Inkling54.3%
Claude Fable 580.0%
GPT-5.6 Sol64.6%
Gemini 3.1 ProNot reported

Terminal-Bench 2.1 Best Harness

Terminal agents · Higher is better

Source
Inkling63.8%
Claude Fable 584.6%
GPT-5.6 Sol89.5%
Gemini 3.1 ProNot reported

MCP Atlas

Tool use · Higher is better

Source
Inkling76.0%
Claude Fable 583.3%
GPT-5.6 Sol81.8%
Gemini 3.1 ProNot reported

Official benchmark profile

Published scores outside our visual tests

Thinking Machines sourceJuly 2026Source report →
Math

AIME 2026

97.1%
Academic reasoning

GPQA Diamond

87.2%
Agentic coding

SWE-bench Verified

77.6%
Full official benchmark table12 scores and peer comparisons
BenchmarkAreaScoreComparison
AIME 2026Math97.1%
Inkling97.1%
GLM 5.299.2%
Claude Fable 599.9%
GPT-5.6 Sol99.9%
GPQA DiamondAcademic reasoning87.2%
Inkling87.2%
Kimi K2.691.1%
Claude Fable 592.6%
GPT-5.6 Sol94.1%
SWE-bench VerifiedAgentic coding77.6%
Inkling77.6%
GLM 5.280.0%
DeepSeek V4 Pro80.6%
GPT-5.6 Sol82.2%
SWE-bench Pro PublicAgentic coding54.3%
Inkling54.3%
GLM 5.262.1%
GPT-5.6 Sol64.6%
Claude Fable 580.0%
Terminal-Bench 2.1 Best HarnessTerminal agents63.8%
Inkling63.8%
GLM 5.282.7%
Claude Fable 584.6%
GPT-5.6 Sol89.5%
MCP AtlasTool use76.0%
Inkling76.0%
GLM 5.277.8%
GPT-5.6 Sol81.8%
Claude Fable 583.3%
BrowseComp with context managementAgentic research77.1%
Inkling77.1%
DeepSeek V4 Pro83.4%
Gemini 3.1 Pro85.9%
Claude Fable 588.0%
IFBenchInstruction following79.8%
Inkling79.8%
Nemotron 3 Ultra81.4%
Gemini 3.1 Pro77.1%
DeepSeek V4 Pro76.5%
Global-MMLU-LiteMultilingual knowledge88.7%
Inkling88.7%
GLM 5.289.2%
GPT-5.6 Sol91.8%
Gemini 3.1 Pro92.7%
MMMU-Pro Standard 10Multimodal reasoning73.5%
Inkling73.5%
Gemini 3.1 Pro82.0%
GPT-5.6 Sol83.0%
Claude Fable 584.2%
VoiceBenchAudio91.4%
Inkling91.4%
Gemini 3.1 Pro94.3%
FORTRESS AdversarialSafety78.0%
Inkling78.0%
Nemotron 3 Ultra77.6%
GPT-5.6 Sol82.4%
Claude Fable 596.0%
Sources, pricing details and methodology

Verified model facts

Model specifications

Pricing source · checked 2026-09-08 ↗
Status
Current
API model ID
Not publicly verified
Context
1M
Max output
Not publicly verified
License
Not publicly verified
Access route
Not publicly verified
Modalities
Not publicly verified
Hardware
Not publicly verified
Tinker serverless inference beta (256K) · USD / 1M tokens
$1 input · $0.17 cached input · $4.05 output

Applies to thinkingmachines/Inkling:peft:262144:sampling-nvfp4, with a 256K context limit on this route. Training, other Tinker configurations, self-hosting, and third-party providers have different costs.

Benchmark sources

Thinking Machines presents Inkling as a customizable open-weights multimodal model rather than the strongest model in every category. These rows use the current website model card at effort 0.99 and temperature 1.0.

Inkling Model Card

Evaluation settings

  • Effort 0.99 and temperature 1.0.
  • Externally reported score reproduced in the Inkling model card.
  • Effort 0.99, temperature 1.0, and a 256K-token coding trajectory cap.
  • Effort 0.99 and temperature 1.0; web-search-contaminated rollouts score zero.
  • Current Thinking Machines website model-card value.
  • Model-card run with context management enabled.
  • Externally reported score; long image edge resized under the model-card protocol.
  • Adversarial refusal evaluation from the model card.
MCP Atlas
The older Hugging Face README reports 74.1%; the website model card is used here.