Seminal AI
§1

GLM-5.3-Flash

The cheapest capable multimodal model in this entire report and plain MIT weights — but the posted rate is a 50%-off promo running to 9 Sep 2026, so model your budget on double it.

Data checked 2026-09-06

Chinese AI labs ga open weights glm-5.3-flash

Context window
1M tokens in
Max output
131.1K tokens out
Input
$0.07 per 1M tokens
Output
$0.25 per 1M tokens
Cached input
$0.01 per 1M tokens
Blended
$0.12 3:1 in:out
Released
2026-08-25
Parameters
320B total / 18B active MoE (sparse + linear hybrid attention)
Licence
MIT

A What it is for

At a three-to-one input-to-output ratio, GLM-5.3-Flash costs $0.12 per million tokens blended. A workload of one million input and 330,000 output tokens per day would run about $4.72 per month at list price, before caching or batch discounts.

B Capabilities

native multimodal reasoning function calling context caching
Input modalitiestext, image, video, file
Output modalitiestext

C Others from Chinese AI labs

Model Context Max out In $/M Out $/M Blended Status
Kimi K3 kimi-k3 1M $3.00 $15.00 $6.00 ga open
Kimi K2.7 Code kimi-k2.7-code 262.1K $0.95 $4.00 $1.71 ga open
Kimi K2.7 Code Highspeed kimi-k2.7-code-highspeed 262.1K $1.90 $8.00 $3.42 ga open
Kimi K2.6 kimi-k2.6 262.1K $0.95 $4.00 $1.71 ga open
GLM-5.3 glm-5.3 1M 131.1K $1.40 $4.40 $2.15 ga open
GLM-5.2 glm-5.2 1M 131.1K $1.40 $4.40 $2.15 ga open
GLM-5 glm-5 200K 131.1K $1.00 $3.20 $1.55 ga open
GLM-4.7 glm-4.7 200K 131.1K $0.60 $2.20 $1.00 ga open
GLM-4.7-FlashX glm-4.7-flashx 200K 131.1K $0.07 $0.40 $0.15 ga
GLM-4.7-Flash glm-4.7-flash 200K 131.1K $0.00 $0.00 $0.00 ga
GLM-4.6V glm-4.6v 128K $0.30 $0.90 $0.45 ga
GLM-4.5-Air glm-4.5-air $0.20 $1.10 $0.43 ga open

D Verify before you commit

Model pricing changes without notice and this page is a snapshot. Confirm against the vendor's own page before you build a budget on it.

Source: docs.z.ai · checked 2026-09-06 · Chinese AI labs pricing · API docs