Seminal AI
§1

MAI-Voice-2

Highest-fidelity MAI TTS for audiobooks, podcasts and long-form narration — ~46 prebuilt voices across 15 languages/18 locales with SSML style+styledegree control and gated instant voice cloning.

Data checked 2026-09-06

Microsoft preview MAI-Voice-2

Context window
Max output
Input
per 1M tokens
Output
per 1M tokens
Released
2026-06-02
Licence
proprietary

A What it is for

Not token-billed: it runs on the Azure Speech Neural HD Text to Speech meter at $22.00 per 1M characters (Personal Voice cloning is $24.00 per 1M characters).

B Capabilities

streaming
Input modalitiestext
Output modalitiesaudio

C Others from Microsoft

Model Context Max out In $/M Out $/M Blended Status
MAI-Thinking-1 MAI-Thinking-1 256K 64K $2.00 $8.00 $3.50 preview
MAI-Cyber-1-Flash MAI-Cyber-1-Flash $0.60 $3.50 $1.32 preview
MAI-Image-2.6 MAI-Image-2.6 32K preview
MAI-Image-2.6-Flash MAI-Image-2.6-Flash 32K preview
MAI-Image-2.5-Pro MAI-Image-2.5-Pro 32K $5.00 $106.00 $30.25 preview
MAI-Image-2.5 MAI-Image-2.5 32K $5.00 $47.00 $15.50 preview
MAI-Image-2.5-Flash MAI-Image-2.5-Flash 32K $1.75 $19.50 $6.19 preview
MAI-Voice-2-Flash MAI-Voice-2-Flash preview
MAI-Voice-1 MAI-Voice-1 deprecated
MAI-Transcribe-2 MAI-Transcribe-2 preview
MAI-Transcribe-1.5 MAI-Transcribe-1.5 preview
MAI-Transcribe-1 MAI-Transcribe-1 deprecated

D Verify before you commit

Model pricing changes without notice and this page is a snapshot. Confirm against the vendor's own page before you build a budget on it.

Source: learn.microsoft.com · checked 2026-09-06 · Microsoft pricing · API docs

Some figures on this page are recorded at medium confidence — the vendor does not publish them in a single authoritative place. Treat them as indicative.