Phi-4-mini-flash-reasoning
Hybrid SambaY/Gated-Memory-Unit architecture giving up to ~10x higher decoding throughput than Phi-4-mini-reasoning on long generations — self-host it (or use Foundry managed compute) when tokens/sec on a single GPU is the binding constraint; there is no per-token Azure meter for it.
Data checked 2026-09-06
Microsoft
ga
open weights
Phi-4-mini-flash-reasoning
- Context window
- 65.5K tokens in
- Max output
- —
- Input
- — per 1M tokens
- Output
- — per 1M tokens
- Knowledge cutoff
- 2025-02
- Released
- 2025-06
- Parameters
- 3.8B
- Licence
- MIT
A What it is for
B Capabilities
reasoning
| Input modalities | text |
|---|---|
| Output modalities | text |
C Others from Microsoft
| Model | Context | Max out | In $/M | Out $/M | Blended | Status |
|---|---|---|---|---|---|---|
MAI-Thinking-1
MAI-Thinking-1
|
256K | 64K | $2.00 | $8.00 | $3.50 | preview |
MAI-Cyber-1-Flash
MAI-Cyber-1-Flash
|
— | — | $0.60 | $3.50 | $1.32 | preview |
MAI-Image-2.6
MAI-Image-2.6
|
32K | — | — | — | — | preview |
MAI-Image-2.6-Flash
MAI-Image-2.6-Flash
|
32K | — | — | — | — | preview |
MAI-Image-2.5-Pro
MAI-Image-2.5-Pro
|
32K | — | $5.00 | $106.00 | $30.25 | preview |
MAI-Image-2.5
MAI-Image-2.5
|
32K | — | $5.00 | $47.00 | $15.50 | preview |
MAI-Image-2.5-Flash
MAI-Image-2.5-Flash
|
32K | — | $1.75 | $19.50 | $6.19 | preview |
MAI-Voice-2
MAI-Voice-2
|
— | — | — | — | — | preview |
MAI-Voice-2-Flash
MAI-Voice-2-Flash
|
— | — | — | — | — | preview |
MAI-Voice-1
MAI-Voice-1
|
— | — | — | — | — | deprecated |
MAI-Transcribe-2
MAI-Transcribe-2
|
— | — | — | — | — | preview |
MAI-Transcribe-1.5
MAI-Transcribe-1.5
|
— | — | — | — | — | preview |
D Verify before you commit
Model pricing changes without notice and this page is a snapshot. Confirm against the vendor's own page before you build a budget on it.
Source: huggingface.co · checked 2026-09-06 · Microsoft pricing · API docs