Microsoft · Reasoning
Phi-4-mini-flash-reasoning
Hybrid SambaY/Gated-Memory-Unit architecture giving up to ~10x higher decoding throughput than Phi-4-mini-reasoning on long generations — self-host it (or use Foundry managed compute) when tokens/sec on a single GPU is the binding constraint; there is no per-token Azure meter for it.
Open weights · Permissive (Apache, MIT, OpenMDW) · Data checked September 6, 2026 · Not a hands-on evaluation
Model ID Phi-4-mini-flash-reasoning
- Parameters
- 3.8B
- Architecture
- Not recorded
- Context window
- 65,536 tokensMaximum input
- Max output
- Not recordedPer response
- Input price
- Not published
- Output price
- Not published
- Cached input
- Not published
- Blended price
- Not available3:1 input to output
- Knowledge cutoff
- 2025-02
- Released
- 2025-06
- Status
- ga
- Size band
- Up to 4B
License and openness
Phi-4-mini-flash-reasoning is released under the MIT. This is a permissive license. It generally allows use, modification and redistribution, including commercial use, subject to attribution and notice requirements. License terms can change between versions; confirm the text published with the weights you download.
Capabilities
| Input modalities | text |
|---|---|
| Output modalities | text |
Cost at list price
| Input | Not published |
|---|---|
| Output | Not published |
| Cached input | Not published |
| Blended, 3:1 | Not available |
The provider publishes no per-token price for this model. Any price you see elsewhere belongs to a third-party host, and running the weights yourself has hardware costs instead.
Estimate a workload across all models ↗Sources
- Source: huggingface.co ↗
- Microsoft pricing ↗
- API docs ↗
- Full reference record for Phi-4-mini-flash-reasoning ↗
- Microsoft in the reference ↗
Sources checked September 6, 2026. Specifications and prices change without notice; confirm against the provider before you commit.
Other open-weight models from Microsoft
- MAI-DS-R1 · 671B total · 37B active
- Phi-4 · 14.7B
- Phi-4-mini-instruct · 3.8B
- Phi-4-multimodal-instruct · 5.6B
- Phi-4-reasoning · 14.7B
- Phi-4-reasoning-plus · 14.7B
- Phi-4-mini-reasoning · 3.8B
- Phi-4-Reasoning-Vision-15B · 15B
- Phi-mini-MoE-instruct · 7.6B total · 2.4B active
- Phi-tiny-MoE-instruct · 3.8B total · 1.1B active
- Phi-Ground-Any · Not recorded
- Phi-3.5-mini-instruct · 3.8B