OpenAI · General language
gpt-oss-20b
On-device or low-latency local inference with tool calling; runs in ~16GB. No OpenAI-hosted price is published on the pricing page.
Open weights · Permissive (Apache, MIT, OpenMDW) · Data checked September 6, 2026 · Not a hands-on evaluation
Model ID gpt-oss-20b
- Parameters
- 21B total · 3.6B active21B total / 3.6B active MoE
- Architecture
- Mixture of experts
- Context window
- 131,072 tokensMaximum input
- Max output
- 131,072 tokensPer response
- Input price
- Not published
- Output price
- Not published
- Cached input
- Not published
- Blended price
- Not available3:1 input to output
- Knowledge cutoff
- 2024-06-01
- Released
- 2025-08-05
- Status
- ga
- Size band
- 15B to 40B
License and openness
gpt-oss-20b is released under the Apache-2.0. This is a permissive license. It generally allows use, modification and redistribution, including commercial use, subject to attribution and notice requirements. License terms can change between versions; confirm the text published with the weights you download.
Capabilities
| Input modalities | text |
|---|---|
| Output modalities | text |
What it is for
No OpenAI-hosted price is published on the pricing page.
Cost at list price
| Input | Not published |
|---|---|
| Output | Not published |
| Cached input | Not published |
| Blended, 3:1 | Not available |
The provider publishes no per-token price for this model. Any price you see elsewhere belongs to a third-party host, and running the weights yourself has hardware costs instead.
Estimate a workload across all models ↗Sources
- Source: huggingface.co ↗
- OpenAI pricing ↗
- API docs ↗
- Full reference record for gpt-oss-20b ↗
- OpenAI in the reference ↗
Sources checked September 6, 2026. Specifications and prices change without notice; confirm against the provider before you commit.
Other open-weight models from OpenAI
- gpt-oss-120b · 117B total · 5.1B active