Thinking Machines Lab
Inkling
Version 1.0 benchmark runs across the shared Superbash visual prompts.
API pricing
USD per million tokens · Tinker serverless inference beta (256K)- Input
- $1
- Output
- $4.05
- Cached input
- $0.17
- Context window
- 1M
Try the active visual tests
Open the generated scene to test it, or compare that prompt with another model.

Helm's Deep
Run date: 2026-07-16
Fortress siege scene testing scale, lighting, architecture, and cinematic atmosphere.

Hogwarts Broom Flight Simulator
Run date: 2026-07-16
Broom-flight scene testing depth, motion cues, castle scale, and fantasy mood.

Yingzao Fashi Assembly
Run date: 2026-07-16
Specialist test · evaluated separately
Timber assembly scene testing structure, joinery, construction order, and material clarity.
Retired tests · 6 archived runs
These tests are no longer in the active suite. Their original results remain available for inspection.
How Inkling compares
Three relevant peers, side by side.
- Inkling This model
- Claude Fable 5
- GPT-5.6 Sol
- Gemini 3.1 Pro
6 benchmarks · 4 models
Official benchmark profile
Published scores outside our visual tests
AIME 2026
97.1%GPQA Diamond
87.2%SWE-bench Verified
77.6%Full official benchmark table12 scores and peer comparisons
| Benchmark | Area | Score | Comparison |
|---|---|---|---|
| AIME 2026 | Math | 97.1% | |
| GPQA Diamond | Academic reasoning | 87.2% | |
| SWE-bench Verified | Agentic coding | 77.6% | |
| SWE-bench Pro Public | Agentic coding | 54.3% | |
| Terminal-Bench 2.1 Best Harness | Terminal agents | 63.8% | |
| MCP Atlas | Tool use | 76.0% | |
| BrowseComp with context management | Agentic research | 77.1% | |
| IFBench | Instruction following | 79.8% | |
| Global-MMLU-Lite | Multilingual knowledge | 88.7% | |
| MMMU-Pro Standard 10 | Multimodal reasoning | 73.5% | |
| VoiceBench | Audio | 91.4% | |
| FORTRESS Adversarial | Safety | 78.0% |
Sources, pricing details and methodology
Verified model facts
Model specifications
- Status
- Current
- API model ID
- Not publicly verified
- Context
- 1M
- Max output
- Not publicly verified
- License
- Not publicly verified
- Access route
- Not publicly verified
- Modalities
- Not publicly verified
- Hardware
- Not publicly verified
- Tinker serverless inference beta (256K) · USD / 1M tokens
- $1 input · $0.17 cached input · $4.05 output
Applies to thinkingmachines/Inkling:peft:262144:sampling-nvfp4, with a 256K context limit on this route. Training, other Tinker configurations, self-hosting, and third-party providers have different costs.
Benchmark sources
Thinking Machines presents Inkling as a customizable open-weights multimodal model rather than the strongest model in every category. These rows use the current website model card at effort 0.99 and temperature 1.0.
Inkling Model CardEvaluation settings
- Effort 0.99 and temperature 1.0.
- Externally reported score reproduced in the Inkling model card.
- Effort 0.99, temperature 1.0, and a 256K-token coding trajectory cap.
- Effort 0.99 and temperature 1.0; web-search-contaminated rollouts score zero.
- Current Thinking Machines website model-card value.
- Model-card run with context management enabled.
- Externally reported score; long image edge resized under the model-card protocol.
- Adversarial refusal evaluation from the model card.
- MCP Atlas
- The older Hugging Face README reports 74.1%; the website model card is used here.