Seminal AI
§1

Llama 3.2 90B Vision Instruct

A cross-attention vision adapter bolted onto Llama 3.1 70B — capable at charts, documents and captioning, but image+text is English-only and it is superseded by Llama 4's native early fusion; pick it only when you need a 3.x-lineage vision model.

Data checked 2026-09-06

Meta ga open weights meta-llama/Llama-3.2-90B-Vision-Instruct

Context window
128K tokens in
Max output
Input
per 1M tokens
Output
per 1M tokens
Knowledge cutoff
2023-12
Released
2024-09-25
Parameters
90B (88.8B)
Licence
Llama 3.2 Community License Agreement

A What it is for

Prices -1: weights only.

B Capabilities

chat image-understanding document-understanding chart-reasoning fine-tuning
Input modalitiestext, image
Output modalitiestext

C Others from Meta

Model Context Max out In $/M Out $/M Blended Status
Muse Spark 1.3 muse-spark-1.3 1M $1.25 $4.25 $2.00 ga
Muse Spark 1.3 (Contributor tier) muse-spark-1.3-contributor 1M $0.10 $0.20 $0.13 ga
Muse Spark 1.2 muse-spark-1.2 1M $1.25 $4.25 $2.00 ga
Muse Spark 1.2 (Contributor tier) muse-spark-1.2-contributor 1M $0.10 $0.20 $0.13 ga
Muse Spark 1.1 muse-spark-1.1 1M $1.25 $4.25 $2.00 ga
Muse Image 1.0 muse-image-1.0 ga
Muse Voice Transcribe 1.0 muse-voice-transcribe-1.0 ga
Muse Glimmer 30B meta-models/Muse-Glimmer-30B 131.1K ga open
Llama 4 Scout meta-llama/Llama-4-Scout-17B-16E-Instruct 10M ga open
Llama 4 Maverick meta-llama/Llama-4-Maverick-17B-128E-Instruct 1M ga open
Llama 3.3 70B Instruct meta-llama/Llama-3.3-70B-Instruct 128K ga open
Llama 3.2 11B Vision Instruct meta-llama/Llama-3.2-11B-Vision-Instruct 128K ga open

D Verify before you commit

Model pricing changes without notice and this page is a snapshot. Confirm against the vendor's own page before you build a budget on it.

Source: github.com · checked 2026-09-06 · Meta pricing · API docs