Seminal AI
§1

Muse Voice Transcribe 1.0

Billed per audio-hour, not per token: $0.18 per hour of audio processed, streaming and file transcription priced the same, with zero-data-retention at price parity and no discounted training-eligible tier — both token prices are -1 because Meta publishes none.

Data checked 2026-09-06

Meta ga muse-voice-transcribe-1.0

Context window
Max output
Input
per 1M tokens
Output
per 1M tokens

A What it is for

Good for voice agents and call/meeting intelligence, but rule it out if you need word-level timestamps, sound-event or emotion detection, or speech synthesis; it does none of those, and the 8-concurrent-stream cap constrains fan-out.

B Capabilities

speech-to-text realtime-streaming speaker-diarization voice-activity-detection endpointing keyword-biasing turn-level-timestamps multilingual
Input modalitiesaudio
Output modalitiestext

C Others from Meta

Model Context Max out In $/M Out $/M Blended Status
Muse Spark 1.3 muse-spark-1.3 1M $1.25 $4.25 $2.00 ga
Muse Spark 1.3 (Contributor tier) muse-spark-1.3-contributor 1M $0.10 $0.20 $0.13 ga
Muse Spark 1.2 muse-spark-1.2 1M $1.25 $4.25 $2.00 ga
Muse Spark 1.2 (Contributor tier) muse-spark-1.2-contributor 1M $0.10 $0.20 $0.13 ga
Muse Spark 1.1 muse-spark-1.1 1M $1.25 $4.25 $2.00 ga
Muse Image 1.0 muse-image-1.0 ga
Muse Glimmer 30B meta-models/Muse-Glimmer-30B 131.1K ga open
Llama 4 Scout meta-llama/Llama-4-Scout-17B-16E-Instruct 10M ga open
Llama 4 Maverick meta-llama/Llama-4-Maverick-17B-128E-Instruct 1M ga open
Llama 3.3 70B Instruct meta-llama/Llama-3.3-70B-Instruct 128K ga open
Llama 3.2 90B Vision Instruct meta-llama/Llama-3.2-90B-Vision-Instruct 128K ga open
Llama 3.2 11B Vision Instruct meta-llama/Llama-3.2-11B-Vision-Instruct 128K ga open

D Verify before you commit

Model pricing changes without notice and this page is a snapshot. Confirm against the vendor's own page before you build a budget on it.

Source: dev.meta.ai · checked 2026-09-06 · Meta pricing · API docs