Qwen3.8-2.4T-A95B
The published base weights behind Qwen3.8-Max and the largest open-weight model in existence — text-only, thinking-mode-only, 262K native context (YaRN-extensible to ~1.01M).
Data checked 2026-09-06
Alibaba Qwen
ga
open weights
qwen3.8-2.4t-a95b
- Context window
- 1M tokens in
- Max output
- 131.1K tokens out
- Input
- $2.00 per 1M tokens
- Output
- $6.00 per 1M tokens
- Blended
- $3.00 3:1 in:out
- Released
- 2026-08
- Parameters
- 2.4T-A95B MoE (512 experts, 11 active/token)
- Licence
- Qwen3.8-Max License (custom, non-Apache)
A What it is for
Note it ships under a custom licence, NOT Apache 2.0; the hosted qwen3.8-max adds vision, non-thinking mode and 1M context by default.
At a three-to-one input-to-output ratio, Qwen3.8-2.4T-A95B costs $3.00 per million tokens blended. A workload of one million input and 330,000 output tokens per day would run about $119.40 per month at list price, before caching or batch discounts.
B Capabilities
| Input modalities | text |
|---|---|
| Output modalities | text |
| Batch discount | 50% off asynchronous requests |
C Others from Alibaba Qwen
| Model | Context | Max out | In $/M | Out $/M | Blended | Status |
|---|---|---|---|---|---|---|
Qwen3.8-Max
qwen3.8-max
|
1M | 131.1K | $2.00 | $6.00 | $3.00 | ga |
Qwen3.7-Max
qwen3.7-max
|
1M | 131.1K | $2.50 | $7.50 | $3.75 | ga |
Qwen3-Max
qwen3-max
|
262.1K | 65.5K | $1.20 | $6.00 | $2.40 | ga |
Qwen-Max (legacy)
qwen-max
|
32.8K | 8.2K | $1.60 | $6.40 | $2.80 | deprecated |
Qwen3.7-Plus
qwen3.7-plus
|
1M | 131.1K | $0.40 | $1.60 | $0.70 | ga |
Qwen3.5-Plus
qwen3.5-plus
|
1M | 131.1K | $0.40 | $2.40 | $0.90 | ga |
Qwen-Plus (rolling)
qwen-plus
|
1M | 32.8K | $0.40 | $1.20 | $0.60 | ga |
Qwen3.8-Flash
qwen3.8-flash
|
1M | 131.1K | $0.15 | $0.47 | $0.23 | ga |
Qwen3.7-Flash
qwen3.7-flash
|
1M | 65.5K | $0.03 | $0.13 | $0.06 | ga |
Qwen3.5-Flash
qwen3.5-flash
|
1M | 65.5K | $0.10 | $0.40 | $0.18 | ga |
Qwen-Flash (rolling)
qwen-flash
|
1M | 32.8K | $0.05 | $0.40 | $0.14 | ga |
Qwen-Turbo
qwen-turbo
|
1M | 16.4K | $0.05 | $0.20 | $0.09 | deprecated |
How this model is billed. Real and served. Model ID is lowercase `qwen3.8-2.4t-a95b`, listed under the billing page section 'Text generation - Qwen (open source) > Qwen3.8'. Singapore/International table row reads: Mode 'Non-Thinking and Thinking modes', Input tokens per request '0<Token<=1M', Input $2 per 1M tokens, Output $6 per 1M tokens (output price covers chain of thought + answer), free quota 1 million tokens valid 90 days. A context-caching discount applies. China (Beijing) region prices differ: $1.65 in / $4.951 out. Context window inferred from the single published request tier (up to 1M input tokens); the docs' model tables do not restate it. Note this model is NOT listed on the /model-studio/models overview or /text-generation pages, only in the billing tables.
D Verify before you commit
Model pricing changes without notice and this page is a snapshot. Confirm against the vendor's own page before you build a budget on it.
Source: alibabacloud.com · checked 2026-09-06 · Alibaba Qwen pricing · API docs