← All open-weight models

Other notable labs · General language

NVIDIA Nemotron 3 Ultra 550B-A55B

Frontier-scale open weights for on-prem agentic workloads — but it needs 8x GB200/B200 or 16x H100 minimum, so it is a datacentre commitment, not a download.

Open weights · Permissive (Apache, MIT, OpenMDW) · Data checked September 6, 2026 · Not a hands-on evaluation

Model ID nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16

Parameters
550B total · 55B active550B total / 55B active (LatentMoE, Mamba-2 + MoE)
Architecture
Mixture of experts
Context window
1,000,000 tokensMaximum input
Max output
Not recordedPer response
Input price
Not published
Output price
Not published
Cached input
Not published
Blended price
Not available3:1 input to output
Knowledge cutoff
2025-09 (pre-training), 2026-05 (post-training)
Released
2026-06-04
Status
gaFlagship in its lineup
Size band
120B and above

License and openness

NVIDIA Nemotron 3 Ultra 550B-A55B is released under the OpenMDW-1.1. This is a permissive license. It generally allows use, modification and redistribution, including commercial use, subject to attribution and notice requirements. License terms can change between versions; confirm the text published with the weights you download.

Capabilities

Input modalitiestext
Output modalitiestext

Cost at list price

InputNot published
OutputNot published
Cached inputNot published
Blended, 3:1Not available

The provider publishes no per-token price for this model. Any price you see elsewhere belongs to a third-party host, and running the weights yourself has hardware costs instead.

Estimate a workload across all models ↗

Sources

Sources checked September 6, 2026. Specifications and prices change without notice; confirm against the provider before you commit.

Other open-weight models from Other notable labs

See all 26 from Other notable labs ↗

Browse all open-weight models ↗