Seminal AI
§1

Phi-mini-MoE-instruct

SlimMoE compression of Phi-3.5-MoE down to 7.6B total / 2.4B active — near-Phi-3.5-MoE quality at a third of the memory, but a hard 4K context kills it for RAG.

Data checked 2026-09-06

Microsoft ga open weights microsoft/Phi-mini-MoE-instruct

Context window
4.1K tokens in
Max output
Input
per 1M tokens
Output
per 1M tokens
Knowledge cutoff
2023-10
Released
2025-06-23
Parameters
7.6B-A2.4B MoE
Licence
MIT

A What it is for

Weights only, no hosted endpoint.

B Capabilities

Not documented.
Input modalitiestext
Output modalitiestext

C Others from Microsoft

Model Context Max out In $/M Out $/M Blended Status
MAI-Thinking-1 MAI-Thinking-1 256K 64K $2.00 $8.00 $3.50 preview
MAI-Cyber-1-Flash MAI-Cyber-1-Flash $0.60 $3.50 $1.32 preview
MAI-Image-2.6 MAI-Image-2.6 32K preview
MAI-Image-2.6-Flash MAI-Image-2.6-Flash 32K preview
MAI-Image-2.5-Pro MAI-Image-2.5-Pro 32K $5.00 $106.00 $30.25 preview
MAI-Image-2.5 MAI-Image-2.5 32K $5.00 $47.00 $15.50 preview
MAI-Image-2.5-Flash MAI-Image-2.5-Flash 32K $1.75 $19.50 $6.19 preview
MAI-Voice-2 MAI-Voice-2 preview
MAI-Voice-2-Flash MAI-Voice-2-Flash preview
MAI-Voice-1 MAI-Voice-1 deprecated
MAI-Transcribe-2 MAI-Transcribe-2 preview
MAI-Transcribe-1.5 MAI-Transcribe-1.5 preview

D Verify before you commit

Model pricing changes without notice and this page is a snapshot. Confirm against the vendor's own page before you build a budget on it.

Source: huggingface.co · checked 2026-09-06 · Microsoft pricing · API docs