MiniMax
MiniMax M3
A former daily burner that has fallen behind the August field.
API pricing
USD per million tokens · MiniMax API Standard (≤512K input tokens)- Input
- $0.3
- Output
- $1.2
- Cached input
- $0.06
- Context window
- 1M
Suggested for this guide
MiniMax
Use the same MiniMax family tested on this page. Check the current plan and model access before subscribing.
Best for: Long-context work and cost-efficient execution
Check MiniMax plansPartner link. It supports Superbash Learn at no extra cost to you.Try the visual runs
Open the generated scene to test it, or compare that prompt with another model.

Helm's Deep
Fortress siege scene testing scale, lighting, architecture, and cinematic atmosphere.

Hogwarts Broom Flight Simulator
Broom-flight scene testing depth, motion cues, castle scale, and fantasy mood.

Jabberwock
Dark fantasy encounter testing creature design, forest mood, and narrative staging.

Low Poly World
Stylized island build testing composition, color, and low-poly worldbuilding.

Mechanical Watch Simulator
Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Office Life
Workplace vignette testing everyday scene logic, objects, and believable office detail.

Petri Dish
Microscopic ecosystem testing organic forms, scientific clarity, and cellular detail.

Stormwind Trebuchet Simulator
Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Universe Simulator
Cosmic system testing orbital structure, glowing bodies, scale, and simulation readability.

Vice City
Neon coastal city testing vehicles, architecture, atmosphere, and dense urban layout.

Yingzao Fashi Assembly
Timber assembly scene testing structure, joinery, construction order, and material clarity.
How MiniMax M3 compares
Three relevant peers, side by side.
- MiniMax M3 This model
- Claude Opus 4.7
- Gemini 3.1 Pro
- GPT-5.5
6 benchmarks · 4 models
Official benchmark profile
Published scores outside our visual tests
SWE-bench Pro
59.0%Terminal-Bench 2.1
66.0%VIBE V2
50.1%Full official benchmark table11 scores and peer comparisons
| Benchmark | Area | Score | Comparison |
|---|---|---|---|
| SWE-bench Pro | Agentic coding | 59.0% | |
| Terminal-Bench 2.1 | Terminal agents | 66.0% | |
| VIBE V2 | Full-stack coding | 50.1% | |
| SVG-Bench | Visual coding | 63.7% | |
| KernelBench Hard | GPU kernels | 28.8% | |
| BrowseComp | Web research | 83.5% | |
| GDPval rubrics | Professional work | 74.7% | |
| BankerToolBench | Finance tools | 76.1% | |
| MCP Atlas | Tool use | 74.2% | |
| OSWorld-Verified | Computer use | 75.2% | |
| SWE-fficiency | Coding efficiency | 34.8% |
Sources, pricing details and methodology
Verified model facts
Model specifications
- Status
- Current
- API model ID
MiniMax-M3- Context
- 1M
- Max output
- Not publicly verified
- License
- Not publicly verified
- Access route
- Not publicly verified
- Modalities
- Not publicly verified
- Hardware
- Not publicly verified
- MiniMax API Standard (≤512K input tokens) · USD / 1M tokens
- $0.3 input · $0.06 cached input · $1.2 output
Published permanent 50% discount. Above 512K input tokens: $0.60 input, $0.12 cached input, and $2.40 output per 1M tokens. Priority pricing differs.
Benchmark sources
MiniMax publishes M3 results across coding, browser work, tool use, spreadsheets, and computer control. Its strongest relative result in the official comparison chart is SVG-Bench; terminal and OS-control scores remain below the largest closed models.
MiniMax M3: Frontier Coding, 1M Context, Native MultimodalityEvaluation settings
- MiniMax infrastructure with Claude Code scaffolding and official-aligned evaluation logic.
- 8C16G sandbox, two-hour timeout, 128K output cap, Terminus 2.
- Internal build-from-scratch benchmark with Claude Code and a three-run average.
- Internal text/image build and edit tasks with VLM render verification; three-run average.
- Claude Code on NVIDIA Blackwell sm_120; average submitted TFLOPs over theoretical peak across nine questions.
- WebExplorer agent framework with history discarded above 64K tokens.
- Public GDPval cases and rubrics with pointwise scoring aligned to GDPval-AA.
- Public dataset; Claude Code except GPT uses Codex; MiniMax M2.7 judge.
- Official public set and codebase with Gemini 2.5 Pro as judge.
- 361 samples, official codebase, 1920x1080, relative coordinates, and a 200-step cap.
- MiniMax open-source dataset and workflow with a two-hour timeout in Claude Code.

Editorial verdict
Where this model fits
MiniMax M3 now sits in C tier. It was one of the team’s favorite cheap daily burners only a few months ago, but newer models widened the gap: M3 now needs more bug checking, can confidently claim work is fixed when it is not, and is neither as cheap as DeepSeek nor as capable as Kimi.
Best for
Watch out
Why it is ranked here
Evidence
Superbash editorial model ranking
Takeaway: MiniMax M3 is currently placed in Tier C.
The September 2026 editorial roster places MiniMax M3 at rank 15.
Open source →NEW AI Model Tier List for Vibe Coding!
Takeaway: MiniMax M3 needs a strong harness and supervision; the August roster no longer treats it as a current workhorse.
The older source records why the team adopted M3, while the new recording explains why newer peers have overtaken it.
Open source →Related guides