← Explore Seminal AI · AI reference · Snapshot captured September 7, 2026 · Understand model costs
Seminal AI
§1

Qwen3.8-2.4T-A95B

The published base weights behind Qwen3.8-Max and the largest open-weight model in existence — text-only, thinking-mode-only, 262K native context (YaRN-extensible to ~1.01M).

Data checked 2026-09-06

Alibaba Qwen ga open weights qwen3.8-2.4t-a95b

Context window
1M tokens in
Max output
131.1K tokens out
Input
$2.00 per 1M tokens
Output
$6.00 per 1M tokens
Blended
$3.00 3:1 in:out
Released
2026-08
Parameters
2.4T-A95B MoE (512 experts, 11 active/token)
Licence
Qwen3.8-Max License (custom, non-Apache)

A What it is for

Note it ships under a custom licence, NOT Apache 2.0; the hosted qwen3.8-max adds vision, non-thinking mode and 1M context by default.

At a three-to-one input-to-output ratio, Qwen3.8-2.4T-A95B costs $3.00 per million tokens blended. A workload of one million input and 330,000 output tokens per day would run about $119.40 per month at list price, before caching or batch discounts.

B Capabilities

tools reasoning batch streaming structured-output
Input modalitiestext
Output modalitiestext
Batch discount50% off asynchronous requests

C Others from Alibaba Qwen

Model Context Max out In $/M Out $/M Blended Status
Qwen3.8-Max qwen3.8-max 1M 131.1K $2.00 $6.00 $3.00 ga
Qwen3.7-Max qwen3.7-max 1M 131.1K $2.50 $7.50 $3.75 ga
Qwen3-Max qwen3-max 262.1K 65.5K $1.20 $6.00 $2.40 ga
Qwen-Max (legacy) qwen-max 32.8K 8.2K $1.60 $6.40 $2.80 deprecated
Qwen3.7-Plus qwen3.7-plus 1M 131.1K $0.40 $1.60 $0.70 ga
Qwen3.5-Plus qwen3.5-plus 1M 131.1K $0.40 $2.40 $0.90 ga
Qwen-Plus (rolling) qwen-plus 1M 32.8K $0.40 $1.20 $0.60 ga
Qwen3.8-Flash qwen3.8-flash 1M 131.1K $0.15 $0.47 $0.23 ga
Qwen3.7-Flash qwen3.7-flash 1M 65.5K $0.03 $0.13 $0.06 ga
Qwen3.5-Flash qwen3.5-flash 1M 65.5K $0.10 $0.40 $0.18 ga
Qwen-Flash (rolling) qwen-flash 1M 32.8K $0.05 $0.40 $0.14 ga
Qwen-Turbo qwen-turbo 1M 16.4K $0.05 $0.20 $0.09 deprecated

How this model is billed. Real and served. Model ID is lowercase `qwen3.8-2.4t-a95b`, listed under the billing page section 'Text generation - Qwen (open source) > Qwen3.8'. Singapore/International table row reads: Mode 'Non-Thinking and Thinking modes', Input tokens per request '0<Token<=1M', Input $2 per 1M tokens, Output $6 per 1M tokens (output price covers chain of thought + answer), free quota 1 million tokens valid 90 days. A context-caching discount applies. China (Beijing) region prices differ: $1.65 in / $4.951 out. Context window inferred from the single published request tier (up to 1M input tokens); the docs' model tables do not restate it. Note this model is NOT listed on the /model-studio/models overview or /text-generation pages, only in the billing tables.

D Verify before you commit

Model pricing changes without notice and this page is a snapshot. Confirm against the vendor's own page before you build a budget on it.

Source: alibabacloud.com · checked 2026-09-06 · Alibaba Qwen pricing · API docs