← All open-weight models

Microsoft · Vision & multimodal

Phi-4-multimodal-instruct

The only Phi that takes audio as well as images — good for cheap on-device-class speech understanding, OCR and chart reading in one 5.6B model. Watch the meter: audio input is billed separately at $4.00 per 1M audio tokens, 50x the text input rate.

Open weights · Permissive (Apache, MIT, OpenMDW) · Data checked September 6, 2026 · Not a hands-on evaluation

Model ID Phi-4-multimodal-instruct

Parameters
5.6B
Architecture
Not recorded
Context window
131,072 tokensMaximum input
Max output
4,096 tokensPer response
Input price
$0.08 per 1M tokens
Output price
$0.32 per 1M tokens
Cached input
Not published
Blended price
$0.14 per 1M tokens3:1 input to output
Knowledge cutoff
2024-06
Released
2025-02
Status
ga
Size band
4B to 15B

License and openness

Phi-4-multimodal-instruct is released under the MIT. This is a permissive license. It generally allows use, modification and redistribution, including commercial use, subject to attribution and notice requirements. License terms can change between versions; confirm the text published with the weights you download.

Capabilities

Input modalitiestext, image, audio
Output modalitiestext

What it is for

Watch the meter: audio input is billed separately at $4.00 per 1M audio tokens, 50x the text input rate.

At a three-to-one input-to-output ratio, Phi-4-multimodal-instruct costs $0.14 per million tokens blended. A workload of one million input and 330,000 output tokens per day would run about $5.57 per month at list price, before caching or batch discounts.

Cost at list price

Input$0.08 per 1M tokens
Output$0.32 per 1M tokens
Cached inputNot published
Blended, 3:1$0.14 per 1M tokens
Estimate a workload across all models ↗

Sources

Sources checked September 6, 2026. Specifications and prices change without notice; confirm against the provider before you commit.

Other open-weight models from Microsoft

See all 23 from Microsoft ↗

Browse all open-weight models ↗