Seminal AI
§1

Llama 3.1 8B Instruct

Still the default single-GPU fine-tuning baseline across the open-source ecosystem because tooling support is universal; choose Llama 3.2 3B instead if you need something smaller and Muse Glimmer 30B if licence permissiveness matters.

Data checked 2026-09-06

Meta ga open weights meta-llama/Llama-3.1-8B-Instruct

Context window
128K tokens in
Max output
Input
per 1M tokens
Output
per 1M tokens
Knowledge cutoff
2023-12
Released
2024-07-23
Parameters
8B
Licence
Llama 3.1 Community License Agreement

A What it is for

Prices -1 — weights only.

B Capabilities

chat multilingual tool-calling fine-tuning
Input modalitiestext
Output modalitiestext

C Others from Meta

Model Context Max out In $/M Out $/M Blended Status
Muse Spark 1.3 muse-spark-1.3 1M $1.25 $4.25 $2.00 ga
Muse Spark 1.3 (Contributor tier) muse-spark-1.3-contributor 1M $0.10 $0.20 $0.13 ga
Muse Spark 1.2 muse-spark-1.2 1M $1.25 $4.25 $2.00 ga
Muse Spark 1.2 (Contributor tier) muse-spark-1.2-contributor 1M $0.10 $0.20 $0.13 ga
Muse Spark 1.1 muse-spark-1.1 1M $1.25 $4.25 $2.00 ga
Muse Image 1.0 muse-image-1.0 ga
Muse Voice Transcribe 1.0 muse-voice-transcribe-1.0 ga
Muse Glimmer 30B meta-models/Muse-Glimmer-30B 131.1K ga open
Llama 4 Scout meta-llama/Llama-4-Scout-17B-16E-Instruct 10M ga open
Llama 4 Maverick meta-llama/Llama-4-Maverick-17B-128E-Instruct 1M ga open
Llama 3.3 70B Instruct meta-llama/Llama-3.3-70B-Instruct 128K ga open
Llama 3.2 90B Vision Instruct meta-llama/Llama-3.2-90B-Vision-Instruct 128K ga open

D Verify before you commit

Model pricing changes without notice and this page is a snapshot. Confirm against the vendor's own page before you build a budget on it.

Source: github.com · checked 2026-09-06 · Meta pricing · API docs