NVIDIA Nemotron 3 Super 120B-A12B
The practical sweet spot of the Nemotron line: 12B active params keeps throughput high for high-volume ticket automation and RAG while retaining a 256K default window.
Data checked 2026-09-06
Other notable labs
ga
open weights
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
- Context window
- 1M tokens in
- Max output
- —
- Input
- — per 1M tokens
- Output
- — per 1M tokens
- Knowledge cutoff
- 2025-06 (pre-training), 2026-02 (post-training)
- Released
- 2026-03-11
- Parameters
- 120B total / 12B active (LatentMoE, Mamba-2 + MoE)
- Licence
- NVIDIA Nemotron Open Model License
A What it is for
B Capabilities
reasoning agentic long-context tool-use multilingual
| Input modalities | text |
|---|---|
| Output modalities | text |
C Others from Other notable labs
| Model | Context | Max out | In $/M | Out $/M | Blended | Status |
|---|---|---|---|---|---|---|
Jamba Large 1.7
jamba-large-1.7
|
256K | 4.1K | $2.00 | $8.00 | $3.50 | ga open |
Jamba Mini 1.7
jamba-mini-1.7
|
256K | 4.1K | $0.20 | $0.40 | $0.25 | ga open |
Jamba2 Mini
ai21labs/AI21-Jamba2-Mini
|
256K | — | — | — | — | ga open |
Jamba2 3B
ai21labs/AI21-Jamba2-3B
|
256K | — | — | — | — | ga open |
Jamba Reasoning 3B
ai21labs/AI21-Jamba-Reasoning-3B
|
— | — | — | — | — | ga open |
Reka Core
reka-core
|
— | — | $2.00 | $6.00 | $3.00 | ga |
Reka Flash
reka-flash
|
— | — | $0.80 | $2.00 | $1.10 | ga |
Reka Edge (reka-edge-2603)
reka-edge
|
— | — | $0.10 | $0.10 | $0.10 | ga open |
Reka Flash 3.1
RekaAI/reka-flash-3.1
|
— | — | — | — | — | ga open |
Olmo 3.1 32B Think
allenai/Olmo-3.1-32B-Think
|
65.5K | 32.8K | — | — | — | ga open |
Olmo 3.1 32B Instruct
allenai/Olmo-3.1-32B-Instruct
|
65.5K | — | — | — | — | ga open |
Olmo 3 32B Base
allenai/Olmo-3-32B
|
65.5K | — | — | — | — | ga open |
D Verify before you commit
Model pricing changes without notice and this page is a snapshot. Confirm against the vendor's own page before you build a budget on it.
Source: huggingface.co · checked 2026-09-06 · Other notable labs pricing · API docs