← All open-weight models

Meta · Vision & multimodal

Llama 3.2 90B Vision Instruct

A cross-attention vision adapter bolted onto Llama 3.1 70B — capable at charts, documents and captioning, but image+text is English-only and it is superseded by Llama 4's native early fusion; pick it only when you need a 3.x-lineage vision model. Prices -1: weights only.

Open weights · Community or custom license · Data checked September 6, 2026 · Not a hands-on evaluation

Model ID meta-llama/Llama-3.2-90B-Vision-Instruct

Parameters
90B90B (88.8B)
Architecture
Not recorded
Context window
128,000 tokensMaximum input
Max output
Not recordedPer response
Input price
Not published
Output price
Not published
Cached input
Not published
Blended price
Not available3:1 input to output
Knowledge cutoff
2023-12
Released
2024-09-25
Status
ga
Size band
40B to 120B

License and openness

Llama 3.2 90B Vision Instruct is released under the Llama 3.2 Community License Agreement. This is a community or custom license with its own conditions, which can include acceptable-use rules, user thresholds or naming requirements. Read it before commercial use. License terms can change between versions; confirm the text published with the weights you download.

Capabilities

Input modalitiestext, image
Output modalitiestext

What it is for

Prices -1: weights only.

Cost at list price

InputNot published
OutputNot published
Cached inputNot published
Blended, 3:1Not available

The provider publishes no per-token price for this model. Any price you see elsewhere belongs to a third-party host, and running the weights yourself has hardware costs instead.

Estimate a workload across all models ↗

Sources

Sources checked September 6, 2026. Specifications and prices change without notice; confirm against the provider before you commit.

Other open-weight models from Meta

Browse all open-weight models ↗