Z.ai

Tier A · Strong with clear operating rules.

GLM 5.3 Flash

An A-tier efficiency build: the full 16-prompt visual set delivered at a fraction of frontier cost.

API pricing

USD per million tokens · Z.ai API promotional rate
Input
$0.075
Output
$0.25
Cached input
$0.015
Context window
1M
Pricing details and specifications

Try the visual runs

Open the generated scene to test it, or compare that prompt with another model.

16 runs

City Scroll Journey

Cinematic scroll journey testing AI-generated scene continuity, scroll-scrubbed camera motion, and art-directed landing craft.

Mechanical Watch Simulator

Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Stormwind Trebuchet Simulator

Counterweight siege simulation testing coupled mechanics, trajectory prediction, projectile cameras, and interactive tuning.

Editorial verdict

Where this model fits

Back to all benchmarks →

GLM 5.3 Flash is Z.ai's efficiency variant of GLM 5.3: a 320B-parameter open-weight model (18B activated) with native text, image, and video input at roughly a tenth of frontier list price. Independent builder agents completed all sixteen Superbash visual prompts in one series, every artifact validating and running, and Z.ai reports coding results approaching Claude Opus 4.8.

Best for

  • Cheap full-length coding and build work
  • Agent loops with heavy token budgets
  • GLM Coding Plan workflows
  • Cost-efficient iteration before a strong review

Watch out

  • The final prompts were verified with light checks rather than full browser QA
  • Editorial review of the run outputs is still pending
  • Vendor results are launch-reported and harness-dependent

Why it is ranked here

  1. Z.ai documents GLM-5.3-Flash with a one-million-token context, 128K-class output, and DeepSWE v1.1 at 63.4% versus GLM-5.2 at 46.2%.
  2. The completed 16-run independent series covers the full prompt range with honest manifest warnings where work was simplified.
  3. A tier reflects the proven completion record and extreme cost efficiency; the placement stays measured because the outputs have not yet been editorially reviewed.

Evidence

2026-09-07

Superbash editorial model ranking

Takeaway: GLM 5.3 Flash is currently placed in Tier A.

The September 2026 editorial roster places GLM 5.3 Flash at rank 8.

Open source →
2026-08-26

GLM-5.3-Flash

Takeaway: Use GLM 5.3 Flash for cost-conscious coding work, then review the output.

The provider page documents the hybrid sparse-plus-linear architecture, multimodal input, and Coding Plan availability with three times the GLM-5.3 quota.

Open source →

How GLM 5.3 Flash compares

Three relevant peers, side by side.

  • GLM 5.3 Flash This model
  • GLM 5.2
  • GPT-5.6 Sol
  • Kimi K3

1 benchmarks · 4 models

GPQA Diamond

Academic reasoning · Higher is better

Source
GLM 5.3 Flash87.5%
GLM 5.291.2%
GPT-5.6 Sol94.1%
Kimi K393.5%

Official benchmark profile

GLM-5.3-Flash efficiency-tier coding and agent results.

Z.ai sourceAugust 2026Source report →
Coding

DeepSWE v1.1

63.4%
Agentic work

AutomationBench

48.8%
Coding

Z.ai Code Bench v1.0

29.0%
Full official benchmark table6 scores and peer comparisons
BenchmarkAreaScoreComparison
DeepSWE v1.1Coding63.4%
AutomationBenchAgentic work48.8%
Z.ai Code Bench v1.0Coding29.0%
Artificial Analysis Intelligence Index v4.1.1Composite intelligence57
GPQA DiamondAcademic reasoning87.5%
GLM 5.3 Flash87.5%
GLM 5.291.2%
Kimi K393.5%
GPT-5.6 Sol94.1%
TAU-BenchTool use80.7%
Sources, pricing details and methodology

Verified model facts

Model specifications

Pricing source · checked 2026-09-08 ↗
Status
Current
API model ID
glm-5.3-flash
Context
1M
Max output
131K
License
Not publicly verified
Access route
Not publicly verified
Modalities
Not publicly verified
Hardware
Not publicly verified
Z.ai API promotional rate · USD / 1M tokens
$0.075 input · $0.015 cached input · $0.25 output

Official 50% promotion through September 9, 2026 at 24:00 UTC+8. List rates are $0.15 input, $0.03 cached input, and $0.50 output per 1M tokens. Cached-input storage is currently free for a limited time.

Benchmark sources

Z.ai positions GLM-5.3-Flash as the efficiency variant of GLM-5.3: a 320B-parameter open-weight model (18B activated) with hybrid sparse-plus-linear attention, native text, image, and video input, and roughly one-tenth the list price of frontier APIs. Z.ai reports it approaches Claude Opus 4.8 across six coding and agentic benchmarks.

GLM-5.3-Flash

GLM-5.3-Flash is fully available on the GLM Coding Plan with three times the GLM-5.3 quota, and is listed by multiple third-party inference providers.

Evaluation settings

  • Z.ai launch result; GLM-5.2 scored 46.2% in the same comparison.
  • Z.ai launch result; GLM-5.2 scored 26.2% in the same comparison.
  • Claude Code 2.1.207 harness at maximum reasoning effort; Claude Opus 4.8 scored 29.5% on the same harness.
  • Z.ai-reported score at a discounted $0.045 per task, which Z.ai says pushes the cost-quality Pareto frontier.
  • Highest provider-reported OpenRouter result (GMICloud); other providers reported 84.3–84.5%.
  • Highest provider-reported OpenRouter result (NovitaAI); the Z.ai direct endpoint reported 77.3%.