ERNIE-4.5-21B-A3B-Thinking
Small Apache-2.0 reasoning model that runs on a single modern GPU; not a line item in Qianfan's price table, so treat it as self-host-only.
Data checked 2026-09-06Chinese AI labs ga open weights
- Context window
- 131.1K tokens in
- Max output
- —
- Input
- — per 1M tokens
- Output
- — per 1M tokens
- Parameters
- 21B total / 3B active MoE (64 text experts, 6 active + 2 shared, 28 layers)
- Licence
- Apache-2.0
A What it is for
B Capabilities
reasoning long context
| Input modalities | text |
|---|---|
| Output modalities | text |
C Others from Chinese AI labs
| Model | Context | Max out | In $/M | Out $/M | Blended | Status |
|---|---|---|---|---|---|---|
Kimi K3
kimi-k3
|
1M | — | $3.00 | $15.00 | $6.00 | ga open |
Kimi K2.7 Code
kimi-k2.7-code
|
262.1K | — | $0.95 | $4.00 | $1.71 | ga open |
Kimi K2.7 Code Highspeed
kimi-k2.7-code-highspeed
|
262.1K | — | $1.90 | $8.00 | $3.42 | ga open |
Kimi K2.6
kimi-k2.6
|
262.1K | — | $0.95 | $4.00 | $1.71 | ga open |
GLM-5.3
glm-5.3
|
1M | 131.1K | $1.40 | $4.40 | $2.15 | ga open |
GLM-5.3-Flash
glm-5.3-flash
|
1M | 131.1K | $0.07 | $0.25 | $0.12 | ga open |
GLM-5.2
glm-5.2
|
1M | 131.1K | $1.40 | $4.40 | $2.15 | ga open |
GLM-5
glm-5
|
200K | 131.1K | $1.00 | $3.20 | $1.55 | ga open |
GLM-4.7
glm-4.7
|
200K | 131.1K | $0.60 | $2.20 | $1.00 | ga open |
GLM-4.7-FlashX
glm-4.7-flashx
|
200K | 131.1K | $0.07 | $0.40 | $0.15 | ga |
GLM-4.7-Flash
glm-4.7-flash
|
200K | 131.1K | $0.00 | $0.00 | $0.00 | ga |
GLM-4.6V
glm-4.6v
|
128K | — | $0.30 | $0.90 | $0.45 | ga |
D Verify before you commit
Model pricing changes without notice and this page is a snapshot. Confirm against the vendor's own page before you build a budget on it.
Source: huggingface.co · checked 2026-09-06 · Chinese AI labs pricing · API docs