MiniMax

Tier C · Situational tools.

MiniMax M3

A former daily burner that has fallen behind the August field.

API pricing

USD per million tokens · MiniMax API Standard (≤512K input tokens)
Input
$0.3
Output
$1.2
Cached input
$0.06
Context window
1M
Pricing details and specifications

Suggested for this guide

MiniMax

Use the same MiniMax family tested on this page. Check the current plan and model access before subscribing.

Best for: Long-context work and cost-efficient execution

Check MiniMax plansPartner link. It supports Superbash Learn at no extra cost to you.

Try the visual runs

Open the generated scene to test it, or compare that prompt with another model.

11 runs

Mechanical Watch Simulator

Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Stormwind Trebuchet Simulator

Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Editorial verdict

Where this model fits

Back to all benchmarks →

MiniMax M3 now sits in C tier. It was one of the team’s favorite cheap daily burners only a few months ago, but newer models widened the gap: M3 now needs more bug checking, can confidently claim work is fixed when it is not, and is neither as cheap as DeepSeek nor as capable as Kimi.

Best for

  • Low-stakes experiments
  • Existing MiniMax workflows with strong review
  • Cheap continuation work when newer options are unavailable

Watch out

  • Frequent bugs and false completion claims in team use
  • Has not received the same recent step-change as peers
  • Use a stronger reviewer for important work
  • Do not use it for consequential life decisions

Why it is ranked here

  1. The team used MiniMax M3 heavily for two to three months because it was cheap enough for speculative projects.
  2. By August, the team was spending too much time checking bugs and confident but inaccurate completion claims.
  3. DeepSeek now wins the price role and Kimi wins the capability role, leaving M3 as a situational C-tier option.

Evidence

2026-09-07

Superbash editorial model ranking

Takeaway: MiniMax M3 is currently placed in Tier C.

The September 2026 editorial roster places MiniMax M3 at rank 15.

Open source →
2026-06-17

NEW AI Model Tier List for Vibe Coding!

Takeaway: MiniMax M3 needs a strong harness and supervision; the August roster no longer treats it as a current workhorse.

The older source records why the team adopted M3, while the new recording explains why newer peers have overtaken it.

Open source →

How MiniMax M3 compares

Three relevant peers, side by side.

  • MiniMax M3 This model
  • Claude Opus 4.7
  • Gemini 3.1 Pro
  • GPT-5.5

6 benchmarks · 4 models

SWE-bench Pro

Agentic coding · Higher is better

Source
MiniMax M359.0%
Claude Opus 4.764.3%
Gemini 3.1 Pro54.2%
GPT-5.558.6%

Terminal-Bench 2.1

Terminal agents · Higher is better

Source
MiniMax M366.0%
Claude Opus 4.766.1%
Gemini 3.1 Pro70.0%
GPT-5.578.2%

VIBE V2

Full-stack coding · Higher is better

Source
MiniMax M350.1%
Claude Opus 4.755.8%
Gemini 3.1 Pro28.0%
GPT-5.550.5%

SVG-Bench

Visual coding · Higher is better

Source
MiniMax M363.7%
Claude Opus 4.762.3%
Gemini 3.1 Pro59.2%
GPT-5.558.2%

KernelBench Hard

GPU kernels · Higher is better

Source
MiniMax M328.8%
Claude Opus 4.730.7%
Gemini 3.1 Pro18.6%
GPT-5.520.9%

BrowseComp

Web research · Higher is better

Source
MiniMax M383.5%
Claude Opus 4.779.3%
Gemini 3.1 Pro85.9%
GPT-5.584.4%

Official benchmark profile

Published scores outside our visual tests

MiniMax sourceJune 2026Source report →
Agentic coding

SWE-bench Pro

59.0%
Terminal agents

Terminal-Bench 2.1

66.0%
Full-stack coding

VIBE V2

50.1%
Full official benchmark table11 scores and peer comparisons
BenchmarkAreaScoreComparison
SWE-bench ProAgentic coding59.0%
MiniMax M359.0%
GPT-5.558.6%
Gemini 3.1 Pro54.2%
Claude Opus 4.764.3%
Terminal-Bench 2.1Terminal agents66.0%
MiniMax M366.0%
Claude Opus 4.766.1%
Gemini 3.1 Pro70.0%
GPT-5.578.2%
VIBE V2Full-stack coding50.1%
MiniMax M350.1%
GPT-5.550.5%
Claude Opus 4.755.8%
Gemini 3.1 Pro28.0%
SVG-BenchVisual coding63.7%
MiniMax M363.7%
Claude Opus 4.762.3%
Gemini 3.1 Pro59.2%
GPT-5.558.2%
KernelBench HardGPU kernels28.8%
MiniMax M328.8%
Claude Opus 4.730.7%
GPT-5.520.9%
Gemini 3.1 Pro18.6%
BrowseCompWeb research83.5%
MiniMax M383.5%
GPT-5.584.4%
Gemini 3.1 Pro85.9%
Claude Opus 4.779.3%
GDPval rubricsProfessional work74.7%
MiniMax M374.7%
Claude Opus 4.779.8%
GPT-5.580.6%
Gemini 3.1 Pro57.8%
BankerToolBenchFinance tools76.1%
MiniMax M376.1%
Claude Opus 4.781.3%
GPT-5.570.0%
Gemini 3.1 Pro67.0%
MCP AtlasTool use74.2%
MiniMax M374.2%
GPT-5.575.3%
Claude Opus 4.777.0%
Gemini 3.1 Pro69.2%
OSWorld-VerifiedComputer use75.2%
MiniMax M375.2%
Gemini 3.1 Pro76.2%
GPT-5.578.7%
Claude Opus 4.782.8%
SWE-fficiencyCoding efficiency34.8%
Sources, pricing details and methodology

Verified model facts

Model specifications

Pricing source · checked 2026-09-08 ↗
Status
Current
API model ID
MiniMax-M3
Context
1M
Max output
Not publicly verified
License
Not publicly verified
Access route
Not publicly verified
Modalities
Not publicly verified
Hardware
Not publicly verified
MiniMax API Standard (≤512K input tokens) · USD / 1M tokens
$0.3 input · $0.06 cached input · $1.2 output

Published permanent 50% discount. Above 512K input tokens: $0.60 input, $0.12 cached input, and $2.40 output per 1M tokens. Priority pricing differs.

Benchmark sources

MiniMax publishes M3 results across coding, browser work, tool use, spreadsheets, and computer control. Its strongest relative result in the official comparison chart is SVG-Bench; terminal and OS-control scores remain below the largest closed models.

MiniMax M3: Frontier Coding, 1M Context, Native Multimodality

Evaluation settings

  • MiniMax infrastructure with Claude Code scaffolding and official-aligned evaluation logic.
  • 8C16G sandbox, two-hour timeout, 128K output cap, Terminus 2.
  • Internal build-from-scratch benchmark with Claude Code and a three-run average.
  • Internal text/image build and edit tasks with VLM render verification; three-run average.
  • Claude Code on NVIDIA Blackwell sm_120; average submitted TFLOPs over theoretical peak across nine questions.
  • WebExplorer agent framework with history discarded above 64K tokens.
  • Public GDPval cases and rubrics with pointwise scoring aligned to GDPval-AA.
  • Public dataset; Claude Code except GPT uses Codex; MiniMax M2.7 judge.
  • Official public set and codebase with Gemini 2.5 Pro as judge.
  • 361 samples, official codebase, 1920x1080, relative coordinates, and a 200-step cap.
  • MiniMax open-source dataset and workflow with a two-hour timeout in Claude Code.