Seminal AI
§1

Qwen3-Next-80B-A3B-Thinking

Very cheap open reasoning at $0.15 input — hybrid-attention architecture with 3B active params makes long-context inference unusually fast per dollar.

Data checked 2026-09-06

Alibaba Qwen ga open weights qwen3-next-80b-a3b-thinking

Context window
262.1K tokens in
Max output
32.8K tokens out
Input
$0.15 per 1M tokens
Output
$1.20 per 1M tokens
Blended
$0.41 3:1 in:out
Released
2025-09
Parameters
80B-A3B MoE
Licence
Apache-2.0

A What it is for

At a three-to-one input-to-output ratio, Qwen3-Next-80B-A3B-Thinking costs $0.41 per million tokens blended. A workload of one million input and 330,000 output tokens per day would run about $16.38 per month at list price, before caching or batch discounts.

B Capabilities

tools reasoning batch streaming structured-output fine-tuning
Input modalitiestext
Output modalitiestext
Batch discount50% off asynchronous requests

C Others from Alibaba Qwen

Model Context Max out In $/M Out $/M Blended Status
Qwen3.8-Max qwen3.8-max 1M 131.1K $2.00 $6.00 $3.00 ga
Qwen3.7-Max qwen3.7-max 1M 131.1K $2.50 $7.50 $3.75 ga
Qwen3-Max qwen3-max 262.1K 65.5K $1.20 $6.00 $2.40 ga
Qwen-Max (legacy) qwen-max 32.8K 8.2K $1.60 $6.40 $2.80 deprecated
Qwen3.7-Plus qwen3.7-plus 1M 131.1K $0.40 $1.60 $0.70 ga
Qwen3.5-Plus qwen3.5-plus 1M 131.1K $0.40 $2.40 $0.90 ga
Qwen-Plus (rolling) qwen-plus 1M 32.8K $0.40 $1.20 $0.60 ga
Qwen3.8-Flash qwen3.8-flash 1M 131.1K $0.15 $0.47 $0.23 ga
Qwen3.7-Flash qwen3.7-flash 1M 65.5K $0.03 $0.13 $0.06 ga
Qwen3.5-Flash qwen3.5-flash 1M 65.5K $0.10 $0.40 $0.18 ga
Qwen-Flash (rolling) qwen-flash 1M 32.8K $0.05 $0.40 $0.14 ga
Qwen-Turbo qwen-turbo 1M 16.4K $0.05 $0.20 $0.09 deprecated

D Verify before you commit

Model pricing changes without notice and this page is a snapshot. Confirm against the vendor's own page before you build a budget on it.

Source: alibabacloud.com · checked 2026-09-06 · Alibaba Qwen pricing · API docs