← All open-weight models

Chinese AI labs · Vision & multimodal

GLM-5.3-Flash

The cheapest capable multimodal model in this entire report and plain MIT weights — but the posted rate is a 50%-off promo running to 9 Sep 2026, so model your budget on double it.

Open weights · Permissive (Apache, MIT, OpenMDW) · Data checked September 6, 2026 · Not a hands-on evaluation

Model ID glm-5.3-flash

Parameters
320B total · 18B active320B total / 18B active MoE (sparse + linear hybrid attention)
Architecture
Mixture of experts
Context window
1,000,000 tokensMaximum input
Max output
131,072 tokensPer response
Input price
$0.075 per 1M tokens
Output price
$0.25 per 1M tokens
Cached input
$0.015 per 1M tokens
Blended price
$0.1188 per 1M tokens3:1 input to output
Knowledge cutoff
Not recorded
Released
2026-08-25
Status
ga
Size band
120B and above

License and openness

GLM-5.3-Flash is released under the MIT. This is a permissive license. It generally allows use, modification and redistribution, including commercial use, subject to attribution and notice requirements. License terms can change between versions; confirm the text published with the weights you download.

Capabilities

Input modalitiestext, image, video, file
Output modalitiestext

What it is for

At a three-to-one input-to-output ratio, GLM-5.3-Flash costs $0.12 per million tokens blended. A workload of one million input and 330,000 output tokens per day would run about $4.72 per month at list price, before caching or batch discounts.

Cost at list price

Input$0.075 per 1M tokens
Output$0.25 per 1M tokens
Cached input$0.015 per 1M tokens
Blended, 3:1$0.1188 per 1M tokens
Estimate a workload across all models ↗

Sources

Sources checked September 6, 2026. Specifications and prices change without notice; confirm against the provider before you commit.

Other open-weight models from Chinese AI labs

See all 17 from Chinese AI labs ↗

Browse all open-weight models ↗