Seminal AI
§1

NVIDIA Nemotron 3.5 Lightning 30B-A3B

Newest and most deployable Nemotron: 256K context on a single H100 with only 3B active params, and reasoning can be switched off per-request to cut token spend.

Data checked 2026-09-06

Other notable labs ga open weights nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Context window
1M tokens in
Max output
Input
per 1M tokens
Output
per 1M tokens
Knowledge cutoff
2025-09 (pre-training), 2026-05 (post-training)
Released
2026-08-11
Parameters
30B total / 3B active (hybrid Mamba-2 + MoE)
Licence
OpenMDW-1.1

A What it is for

B Capabilities

reasoning toggleable-thinking long-context code tool-use
Input modalitiestext
Output modalitiestext

C Others from Other notable labs

Model Context Max out In $/M Out $/M Blended Status
Jamba Large 1.7 jamba-large-1.7 256K 4.1K $2.00 $8.00 $3.50 ga open
Jamba Mini 1.7 jamba-mini-1.7 256K 4.1K $0.20 $0.40 $0.25 ga open
Jamba2 Mini ai21labs/AI21-Jamba2-Mini 256K ga open
Jamba2 3B ai21labs/AI21-Jamba2-3B 256K ga open
Jamba Reasoning 3B ai21labs/AI21-Jamba-Reasoning-3B ga open
Reka Core reka-core $2.00 $6.00 $3.00 ga
Reka Flash reka-flash $0.80 $2.00 $1.10 ga
Reka Edge (reka-edge-2603) reka-edge $0.10 $0.10 $0.10 ga open
Reka Flash 3.1 RekaAI/reka-flash-3.1 ga open
Olmo 3.1 32B Think allenai/Olmo-3.1-32B-Think 65.5K 32.8K ga open
Olmo 3.1 32B Instruct allenai/Olmo-3.1-32B-Instruct 65.5K ga open
Olmo 3 32B Base allenai/Olmo-3-32B 65.5K ga open

D Verify before you commit

Model pricing changes without notice and this page is a snapshot. Confirm against the vendor's own page before you build a budget on it.

Source: huggingface.co · checked 2026-09-06 · Other notable labs pricing · API docs