Search by name, size, license or capability. Filters work without an account and can be shared through the page URL.
149 open-weight models · Data checked September 6, 2026 · Not hands-on evaluations
OpenAI · General language Self-hosting a frontier-ish reasoning model on a single H100 with full chain-of-thought visibility and no per-token bill. OpenAI publishes no hosted…
- Parameters
- 117B total · 5.1B active
- Context
- 131K tokens
- Input / output
- Not published
- License
- Apache-2.0
OpenAI · General language On-device or low-latency local inference with tool calling; runs in ~16GB. No OpenAI-hosted price is published on the pricing page.
- Parameters
- 21B total · 3.6B active
- Context
- 131K tokens
- Input / output
- Not published
- License
- Apache-2.0
Google DeepMind · Vision & multimodal Largest dense open-weight Gemma for self-hosting on server-grade GPUs with a 256K window; hosted access via the Gemini API is free-tier only (the…
- Parameters
- 31B dense
- Context
- 262K tokens
- Input / output
- $0.00 / $0.00
- License
- Gemma 4 license (Gemma Terms of Use)
Google DeepMind · Vision & multimodal Best open-weight throughput-per-dollar in the family — MoE activating only 4B params per token, so it serves near-4B-dense speed at 26B-class…
- Parameters
- 26B total · 4B active
- Context
- 262K tokens
- Input / output
- $0.00 / $0.00
- License
- Gemma 4 license (Gemma Terms of Use)
Google DeepMind · Vision & multimodal Encoder-free multimodal mid-size open model with a 256K window — the sweet spot for single-GPU self-hosting; download-only (not exposed on the…
- Parameters
- 12B
- Context
- 262K tokens
- Input / output
- $0.00 / $0.00
- License
- Gemma 4 license (Gemma Terms of Use)
Google DeepMind · Vision & multimodal On-device/edge deployment with a 128K window; download-only, with QAT quantized builds for phone and laptop inference engines.
- Parameters
- 4B
- Context
- 128K tokens
- Input / output
- $0.00 / $0.00
- License
- Gemma 4 license (Gemma Terms of Use)
Google DeepMind · Vision & multimodal Smallest Gemma 4 — ultra-mobile and embedded inference at 128K context; download-only, no hosted endpoint.
- Parameters
- 2B
- Context
- 128K tokens
- Input / output
- $0.00 / $0.00
- License
- Gemma 4 license (Gemma Terms of Use)
Meta · Vision & multimodal The one Meta open-weight model with a genuinely permissive licence (Apache 2.0, no Community-License user-count or naming clauses) — a 30B dense…
- Parameters
- 30B dense
- Context
- 131K tokens
- Input / output
- Not published
- License
- Apache License 2.0
Meta · Vision & multimodal The long-context option in Meta's open-weight line — a 10M-token window and single-H100 deployability at int4 make it the pick for whole-repository…
- Parameters
- 109B total · 17B active
- Context
- 10M tokens
- Input / output
- Not published
- License
- Llama 4 Community License Agreement
Meta · Vision & multimodal Higher-quality sibling to Scout with the same 17B activated cost per token but 400B total weights — better reasoning and image understanding, at the…
- Parameters
- 400B total · 17B active
- Context
- 1M tokens
- Input / output
- Not published
- License
- Llama 4 Community License Agreement
The best quality-per-GPU text-only Llama: 405B-class instruction following at 70B serving cost, which is why it remains the most widely hosted Llama…
- Parameters
- 70B
- Context
- 128K tokens
- Input / output
- Not published
- License
- Llama 3.3 Community License Agreement
Meta · Vision & multimodal A cross-attention vision adapter bolted onto Llama 3.1 70B — capable at charts, documents and captioning, but image+text is English-only and it is…
- Parameters
- 90B
- Context
- 128K tokens
- Input / output
- Not published
- License
- Llama 3.2 Community License Agreement
Meta · Vision & multimodal Small enough to fine-tune for a single narrow vision task (receipt or form extraction, screenshot classification) on one commodity GPU; not a…
- Parameters
- 11B
- Context
- 128K tokens
- Input / output
- Not published
- License
- Llama 3.2 Community License Agreement
On-device and edge summarisation, rewriting and query expansion where latency and privacy beat raw capability; note Meta's own quantized builds drop…
- Parameters
- 3B
- Context
- 128K tokens
- Input / output
- Not published
- License
- Llama 3.2 Community License Agreement
The smallest Llama, for phone- and microcontroller-class deployment or as a speculative-decoding draft model; expect to fine-tune it for one task…
- Parameters
- 1B
- Context
- 128K tokens
- Input / output
- Not published
- License
- Llama 3.2 Community License Agreement
Now mainly a teacher model: its licence explicitly permits using outputs to train other models, which is the remaining reason to run something this…
- Parameters
- 405B
- Context
- 128K tokens
- Input / output
- Not published
- License
- Llama 3.1 Community License Agreement
Superseded at the same size and cost by Llama 3.3 70B — keep it only for pinned reproducibility of existing 3.1 evaluations or fine-tunes. Weights…
- Parameters
- 70B
- Context
- 128K tokens
- Input / output
- Not published
- License
- Llama 3.1 Community License Agreement
Still the default single-GPU fine-tuning baseline across the open-source ecosystem because tooling support is universal; choose Llama 3.2 3B instead…
- Parameters
- 8B
- Context
- 128K tokens
- Input / output
- Not published
- License
- Llama 3.1 Community License Agreement
DeepSeek · General language The default DeepSeek pick: near-Pro reasoning at a third of the price with the same 1M context, ideal for high-volume agentic coding and…
- Parameters
- 304B total · 13B active
- Context
- 1M tokens
- Input / output
- $0.44 / $1.32
- License
- MIT
DeepSeek · General language DeepSeek's frontier tier — worth the 3x premium over Flash only for hard agentic/production coding, deep reasoning and tool-heavy workflows where…
- Parameters
- 1.7T total · 49B active
- Context
- 1M tokens
- Input / output
- $1.32 / $3.96
- License
- MIT
DeepSeek · Vision & multimodal The only DeepSeek model that takes images — use it for screenshot-driven and chart/document-reading agents at identical Flash pricing; it is…
- Parameters
- 305B
- Context
- 1M tokens
- Input / output
- $0.44 / $1.32
- License
- MIT
Alibaba Qwen · General language The published base weights behind Qwen3.8-Max and the largest open-weight model in existence — text-only, thinking-mode-only, 262K native context…
- Parameters
- 2.4T total · 95B active
- Context
- 1M tokens
- Input / output
- $2.00 / $6.00
- License
- Qwen3.8-Max License (custom, non-Apache)
Alibaba Qwen · Vision & multimodal The standout self-host pick — a truly Apache-2.0, 27B dense multimodal model with 262K context and adjustable reasoning effort (xhigh/medium/low)…
- Parameters
- 27B dense
- Context
- 1M tokens
- Input / output
- $0.50 / $3.00
- License
- Apache-2.0
Alibaba Qwen · Vision & multimodal Only 3B active parameters, so it runs fast on modest hardware while punching well above its weight on coding benchmarks — the sweet spot for local…
- Parameters
- 35B total · 3B active
- Context
- 262K tokens
- Input / output
- $0.375 / $2.25
- License
- Apache-2.0
Alibaba Qwen · Vision & multimodal Dense 27B for workloads where MoE routing hurts quality consistency; now priced above the newer Qwen3.8-27B ($0.60/$3.60 vs $0.50/$3) — prefer…
- Parameters
- 27B dense
- Context
- 262K tokens
- Input / output
- $0.60 / $3.60
- License
- Apache-2.0
Alibaba Qwen · General language The largest fully Apache-2.0 Qwen model — the one to self-host when licence purity matters more than raw scale and the custom-licensed Qwen3.8-2.4T…
- Parameters
- 397B total · 17B active
- Context
- 262K tokens
- Input / output
- $0.60 / $3.60
- License
- Apache-2.0
Alibaba Qwen · General language Mid-size Apache-2.0 MoE that fits a single 8-GPU node in FP8 — a reasonable step down from the 397B when memory is the constraint.
- Parameters
- 122B total · 10B active
- Context
- 262K tokens
- Input / output
- $0.40 / $3.20
- License
- Apache-2.0
Alibaba Qwen · General language Cheapest hosted Qwen3.5 open model and an easy local run at 3B active params; superseded on quality by Qwen3.6-35B-A3B for a small price premium.
- Parameters
- 35B total · 3B active
- Context
- 262K tokens
- Input / output
- $0.25 / $2.00
- License
- Apache-2.0
Alibaba Qwen · General language Dense 27B at the lowest price in the 27B line ($0.30/$2.40) — good value if you don't need the multimodal input that Qwen3.6-27B and Qwen3.8-27B add.
- Parameters
- 27B dense
- Context
- 262K tokens
- Input / output
- $0.30 / $2.40
- License
- Apache-2.0
Best coding value in the catalogue — 5x cheaper than qwen3-coder-plus at the base tier and open-weight. Tiers: 0-32K $0.30/$1.50, 32-128K…
- Parameters
- Not recorded
- Context
- 262K tokens
- Input / output
- $0.30 / $1.50
- License
- Apache-2.0
The heavyweight open coding model for hard multi-file agentic tasks; expensive and tier-sensitive (0-32K $1.50/$7.50, 32-128K $2.70/$13.50, 128-200K…
- Parameters
- 480B total · 35B active
- Context
- 200K tokens
- Input / output
- $1.50 / $7.50
- License
- Apache-2.0
Laptop-to-workstation coding model with 3B active params; the practical choice for local IDE autocomplete and small refactors. Tiers: 0-32K…
- Parameters
- 30B total · 3B active
- Context
- 200K tokens
- Input / output
- $0.45 / $2.25
- License
- Apache-2.0
Very cheap open reasoning at $0.15 input — hybrid-attention architecture with 3B active params makes long-context inference unusually fast per dollar.
- Parameters
- 80B total · 3B active
- Context
- 262K tokens
- Input / output
- $0.15 / $1.20
- License
- Apache-2.0
Alibaba Qwen · General language Non-thinking sibling for latency-sensitive work where you don't want reasoning tokens billed — same price as the thinking variant, so choose on…
- Parameters
- 80B total · 3B active
- Context
- 262K tokens
- Input / output
- $0.15 / $1.20
- License
- Apache-2.0
The 2025 open reasoning workhorse, still one of the cheapest large thinking models at $0.23 input; Qwen3.5-397B beats it but needs far more memory…
- Parameters
- 235B total · 22B active
- Context
- 262K tokens
- Input / output
- $0.23 / $2.30
- License
- Apache-2.0
Alibaba Qwen · General language Excellent bulk-throughput value at $0.23/$0.92 — 2.5x cheaper output than its thinking twin, ideal for summarisation and extraction at scale.
- Parameters
- 235B total · 22B active
- Context
- 262K tokens
- Input / output
- $0.23 / $0.92
- License
- Apache-2.0
Small MoE reasoner that self-hosts on a single 48GB card in FP8; output pricing is high relative to size, so prefer local deployment over the API…
- Parameters
- 30B total · 3B active
- Context
- 262K tokens
- Input / output
- $0.20 / $2.40
- License
- Apache-2.0
Alibaba Qwen · General language Fast, cheap non-thinking small MoE — a solid drop-in for classification, routing and tool-dispatch layers.
- Parameters
- 30B total · 3B active
- Context
- 262K tokens
- Input / output
- $0.20 / $0.80
- License
- Apache-2.0
Alibaba Qwen · General language Original hybrid-thinking Qwen3 flagship; output is $2.80 non-thinking / $8.40 thinking. The -2507 refresh is better and 3x cheaper — migrate.
- Parameters
- 235B total · 22B active
- Context
- 131K tokens
- Input / output
- $0.70 / $2.80
- License
- Apache-2.0
Alibaba Qwen · General language The most widely deployed open Qwen3 dense model and a well-supported fine-tuning base; output $0.64 non-thinking. Newer 27B models beat it, but…
- Parameters
- 32B dense
- Context
- 131K tokens
- Input / output
- $0.16 / $0.64
- License
- Apache-2.0
Alibaba Qwen · General language Original hybrid-mode 30B MoE; output $0.80 non-thinking / $2.40 thinking. Superseded by the -2507 split variants at the same price.
- Parameters
- 30B total · 3B active
- Context
- 131K tokens
- Input / output
- $0.20 / $0.80
- License
- Apache-2.0
Alibaba Qwen · General language Oddly priced above Qwen3-32B on the API ($0.35 vs $0.16) — only worth it as a self-hosted fine-tune target on a single 24-32GB GPU. Output $1.40…
- Parameters
- 14B dense
- Context
- 131K tokens
- Input / output
- $0.35 / $1.40
- License
- Apache-2.0
Alibaba Qwen · General language Edge/consumer-GPU tier — runs quantised on 8-12GB VRAM. Output $0.70 non-thinking / $2.10 thinking. Smaller 4B/1.7B/0.6B siblings exist on Hugging…
- Parameters
- 8B dense
- Context
- 131K tokens
- Input / output
- $0.18 / $0.70
- License
- Apache-2.0
Largest open vision-language reasoner — use for chart/diagram reasoning and GUI-agent work where a small VL model fails; output is 2.5x the instruct…
- Parameters
- 235B total · 22B active
- Context
- 262K tokens
- Input / output
- $0.40 / $4.00
- License
- Apache-2.0
Alibaba Qwen · Vision & multimodal Top open VL model for straightforward description/extraction at $0.40/$1.60 — no reasoning-token overhead.
- Parameters
- 235B total · 22B active
- Context
- 262K tokens
- Input / output
- $0.40 / $1.60
- License
- Apache-2.0
Cheapest open VL reasoner on the API ($0.16/$0.64) and dense enough to self-host on two 40GB cards — strong default for visual QA pipelines.
- Parameters
- 32B dense
- Context
- 262K tokens
- Input / output
- $0.16 / $0.64
- License
- Apache-2.0
Alibaba Qwen · Vision & multimodal Same price as the thinking variant with lower latency — pick this for OCR-adjacent and captioning work that needs no deliberation.
- Parameters
- 32B dense
- Context
- 262K tokens
- Input / output
- $0.16 / $0.64
- License
- Apache-2.0
Fast MoE visual reasoner (3B active) — good local throughput, but on the API the 32B-thinking model is cheaper on output ($0.64 vs $2.40).
- Parameters
- 30B total · 3B active
- Context
- 262K tokens
- Input / output
- $0.20 / $2.40
- License
- Apache-2.0
Alibaba Qwen · Vision & multimodal Best latency-per-dollar open VL model for high-volume image pipelines you host yourself.
- Parameters
- 30B total · 3B active
- Context
- 262K tokens
- Input / output
- $0.20 / $0.80
- License
- Apache-2.0
Smallest open VL reasoner — the one to fine-tune for a narrow visual domain on a single consumer GPU.
- Parameters
- 8B dense
- Context
- 262K tokens
- Input / output
- $0.18 / $2.10
- License
- Apache-2.0
Alibaba Qwen · Vision & multimodal Edge-deployable multimodal model for on-device captioning and screenshot understanding.
- Parameters
- 8B dense
- Context
- 262K tokens
- Input / output
- $0.18 / $0.70
- License
- Apache-2.0
Alibaba Qwen · Vision & multimodal The small open any-to-any model people actually self-host for offline voice assistants; output pricing varies by modality and is not stated as a…
- Parameters
- 7B dense
- Context
- 33K tokens
- Input / output
- Not published
- License
- Apache-2.0
Mistral AI · Vision & multimodal Mistral's current flagship for long-horizon agentic work and agentic coding — but note it costs 3x the input and 5x the output of Mistral Large 3,…
- Parameters
- 128B dense
- Context
- 256K tokens
- Input / output
- $1.50 / $7.50
- License
- Modified MIT
Mistral AI · Vision & multimodal The value pick of the whole lineup: frontier-size open-weight MoE at $0.50/$1.50, cheaper than Medium 3.5 and most rivals' mid-tier models — default…
- Parameters
- 675B total · 41B active
- Context
- 256K tokens
- Input / output
- $0.50 / $1.50
- License
- Apache-2.0
Mistral AI · Vision & multimodal Best price/performance workhorse — a hybrid instruct+reasoning+coding model with only 6.5B active params, so it is fast and cheap while still…
- Parameters
- 119B total · 6.5B active
- Context
- 256K tokens
- Input / output
- $0.15 / $0.60
- License
- Apache-2.0
Mistral AI · Vision & multimodal Largest edge-class Ministral — flat $0.20 in and out makes it the cheapest option when your workload is output-heavy; separate Base, Instruct and…
- Parameters
- 14B dense
- Context
- 256K tokens
- Input / output
- $0.20 / $0.20
- License
- Apache-2.0
Mistral AI · Vision & multimodal The sweet spot for high-volume classification, extraction and routing where you still want vision and 256K context; also the recommended local model…
- Parameters
- 8B dense
- Context
- 256K tokens
- Input / output
- $0.15 / $0.15
- License
- Apache-2.0
Mistral AI · Vision & multimodal Cheapest model Mistral serves and the one to run on-device — pick it for latency-critical edge inference or trivially structured tasks, not for…
- Parameters
- 3B dense
- Context
- 256K tokens
- Input / output
- $0.10 / $0.10
- License
- Apache-2.0
Mistral AI · General language Use when you need a 1M-token window on Mistral's EU-hosted infrastructure — it is a third-party Z.ai model served without Mistral modifications,…
- Parameters
- Not recorded
- Context
- 1M tokens
- Input / output
- $1.40 / $4.40
- License
- Open (third-party, Z.ai — licence set by…
Free Labs endpoint for Lean 4 formal proof engineering, autoformalization and automated theorem proving — genuinely $0 while Mistral gathers…
- Parameters
- 119B total · 6.5B active
- Context
- 256K tokens
- Input / output
- $0.00 / $0.00
- License
- Apache-2.0
Mistral AI · Audio understanding The one Voxtral that does audio *understanding* rather than plain transcription — chat and Q&A over speech via /v1/chat/completions. Text tokens are…
- Parameters
- 24B dense
- Context
- 32K tokens
- Input / output
- $0.10 / $0.40
- License
- Apache-2.0
Mistral AI · Speech (transcription or voice) Live/streaming transcription at $0.006 per audio minute — double the batch model's rate, so only use it when you actually need low-latency partial…
- Parameters
- 4B dense
- Context
- Not recorded
- Input / output
- Not published
- License
- Apache-2.0
Mistral AI · Speech (transcription or voice) Text-to-speech with zero-shot voice cloning (no transcript needed for the voice prompt), 9 languages and ~90ms time-to-first-audio. Billed per…
- Parameters
- 4B dense
- Context
- Not recorded
- Input / output
- Not published
- License
- CC-BY-NC-4.0
Mistral AI · Safety & moderation Compact self-hosted safety classifier: you write policy questions in natural language and it returns yes/no for prompt moderation, response…
- Parameters
- 3.8B dense
- Context
- 32K tokens
- Input / output
- Not published
- License
- Apache-2.0
Microsoft's safety- and unblocking-post-trained variant of DeepSeek-R1 (671B MoE, 37B active, MIT). It has dropped out of the current Foundry…
- Parameters
- 671B total · 37B active
- Context
- 164K tokens
- Input / output
- Not published
- License
- MIT
Microsoft · General language The dense 14B workhorse of the family: strongest general Phi quality per dollar for math, code and reasoning when a 16K context is enough — if you…
- Parameters
- 14.7B
- Context
- 16K tokens
- Input / output
- $0.125 / $0.50
- License
- MIT
Microsoft · General language Cheapest hosted Phi and the default for high-volume classification, extraction and routing across 23 languages with a real 128K window; no native…
- Parameters
- 3.8B
- Context
- 131K tokens
- Input / output
- $0.075 / $0.30
- License
- MIT
Microsoft · Vision & multimodal The only Phi that takes audio as well as images — good for cheap on-device-class speech understanding, OCR and chart reading in one 5.6B model.…
- Parameters
- 5.6B
- Context
- 131K tokens
- Input / output
- $0.08 / $0.32
- License
- MIT
14B SFT-distilled reasoner at Phi-4 prices — the value pick for math and STEM chains of thought when you cannot justify MAI-Thinking-1 or a frontier…
- Parameters
- 14.7B
- Context
- 33K tokens
- Input / output
- $0.125 / $0.50
- License
- MIT
RL-tuned sibling of Phi-4-reasoning at identical token prices — higher accuracy but noticeably longer chains of thought, so it costs more per answer…
- Parameters
- 14.7B
- Context
- 33K tokens
- Input / output
- $0.125 / $0.50
- License
- MIT
The cheapest reasoning model Microsoft sells: 3.8B with a 128K window in and out, aimed at edge/embedded math tutoring and step-by-step solvers…
- Parameters
- 3.8B
- Context
- 128K tokens
- Input / output
- $0.075 / $0.30
- License
- MIT
Hybrid SambaY/Gated-Memory-Unit architecture giving up to ~10x higher decoding throughput than Phi-4-mini-reasoning on long generations — self-host…
- Parameters
- 3.8B
- Context
- 66K tokens
- Input / output
- Not published
- License
- MIT
Newest Phi release — a 15B multimodal reasoner with explicit <think> traces for chart/diagram/document math and GUI element localization…
- Parameters
- 15B
- Context
- 16K tokens
- Input / output
- Not published
- License
- MIT
Microsoft · General language SlimMoE compression of Phi-3.5-MoE down to 7.6B total / 2.4B active — near-Phi-3.5-MoE quality at a third of the memory, but a hard 4K context kills…
- Parameters
- 7.6B total · 2.4B active
- Context
- 4K tokens
- Input / output
- Not published
- License
- MIT
Microsoft · General language Smallest SlimMoE variant (1.1B active) for CPU and NPU inference where every GB counts; 4K context and an Oct-2023 cutoff make it a component model,…
- Parameters
- 3.8B total · 1.1B active
- Context
- 4K tokens
- Input / output
- Not published
- License
- MIT
Microsoft · Vision & multimodal Research GUI-grounding model (Phi-3-V based) that maps a natural-language instruction to on-screen coordinates — a building block for computer-use…
- Parameters
- Not recorded
- Context
- Not recorded
- Input / output
- Not published
- License
- MIT
Microsoft · General language Still deployable and still metered on Azure, but Microsoft has dropped it from the current Microsoft-models doc table and Phi-4-mini-instruct is…
- Parameters
- 3.8B
- Context
- 131K tokens
- Input / output
- $0.13 / $0.52
- License
- MIT
Microsoft · General language The largest Phi ever shipped (16x3.8B MoE, 6.6B active) — historically interesting and still self-hostable under MIT, but superseded on…
- Parameters
- 41.9B total · 6.6B active
- Context
- 131K tokens
- Input / output
- $0.16 / $0.64
- License
- MIT
Microsoft · Vision & multimodal Multi-image and video-frame reasoning at 4.2B; use Phi-4-multimodal-instruct instead unless you specifically need this checkpoint's multi-frame…
- Parameters
- 4.2B
- Context
- 131K tokens
- Input / output
- $0.13 / $0.52
- License
- MIT
Microsoft · General language Legacy 14B long-context Phi-3; still metered and catalogued but strictly worse and pricier than Phi-4 — kept only for pinned reproducibility.
- Parameters
- 14B
- Context
- 131K tokens
- Input / output
- $0.17 / $0.68
- License
- MIT
Microsoft · General language Short-context variant of Phi-3-medium; no reason to choose it today over Phi-4 at a lower price with a larger window.
- Parameters
- 14B
- Context
- 4K tokens
- Input / output
- $0.17 / $0.68
- License
- MIT
Microsoft · General language Legacy 7B long-context model; superseded on every axis by Phi-4-mini-instruct at half the price.
- Parameters
- 7B
- Context
- 131K tokens
- Input / output
- $0.15 / $0.60
- License
- MIT
Microsoft · General language Short-context legacy 7B; migrate to Phi-4-mini-instruct. (The Foundry catalog record misreports its window as 131072; the model card and name are…
- Parameters
- 7B
- Context
- 8K tokens
- Input / output
- $0.15 / $0.60
- License
- MIT
Microsoft · General language The model that launched the SLM category; still one of the most downloaded Phi checkpoints for offline/edge use, but on Azure Phi-4-mini-instruct is…
- Parameters
- 3.8B
- Context
- 131K tokens
- Input / output
- $0.13 / $0.52
- License
- MIT
Microsoft · General language The canonical tiny Phi for phones and NPUs (also shipped as GGUF and ONNX DirectML/CUDA/CPU builds); pick it only for offline deployment where the…
- Parameters
- 3.8B
- Context
- 4K tokens
- Input / output
- $0.13 / $0.52
- License
- MIT
Microsoft · Vision & multimodal Original Phi vision model, managed-compute/self-host only (no serverless per-token meter); superseded by Phi-3.5-vision-instruct and then…
- Parameters
- 4.2B
- Context
- 131K tokens
- Input / output
- Not published
- License
- MIT
Microsoft · General language Base (non-instruct) 2.7B research model, still the most-downloaded Microsoft checkpoint on Hugging Face and a common fine-tuning starting point —…
- Parameters
- 2.7B
- Context
- 2K tokens
- Input / output
- Not published
- License
- MIT
Cohere · Vision & multimodal Cohere's flagship: the only model that folds vision, reasoning, agentic tool use and 48-language coverage into one set of weights, and it runs on 1x…
- Parameters
- 218B total · 25B active
- Context
- 128K tokens
- Input / output
- $0.00 / $0.00
- License
- Apache-2.0
Cohere · General language The long-context (256K) dense workhorse for RAG, tool use and 23-language agents, and the only Command A variant with a real 500 req/min production…
- Parameters
- 111B
- Context
- 256K tokens
- Input / output
- Not published
- License
- CC-BY-NC-4.0
Pick when you need an explicit thinking budget on nuanced multi-step or agentic problems in 23 languages and can host it (4x H100 for production);…
- Parameters
- 111B
- Context
- 256K tokens
- Input / output
- Not published
- License
- CC-BY-NC-4.0
Cohere · General language A dedicated 23-language machine-translation model for regulated shops that must translate sensitive documents inside their own perimeter; the 8K…
- Parameters
- 111B
- Context
- 8K tokens
- Input / output
- Not published
- License
- CC-BY-NC-4.0
Cohere · Vision & multimodal Document/chart/OCR understanding with up to 20 images per request — but it does NOT support tool use, so for agentic multimodal work go to Command…
- Parameters
- 111B
- Context
- 128K tokens
- Input / output
- Not published
- License
- CC-BY-NC-4.0
Cohere · General language Still Live with a 500 req/min production limit, but at $2.50/$10.00 it is priced under the 'existing customers' FAQ and comprehensively beaten by…
- Parameters
- 104B
- Context
- 128K tokens
- Input / output
- $2.50 / $10.00
- License
- CC-BY-NC-4.0
Cohere · General language The cheapest Cohere model with a published price, a real 500 req/min production limit and full tool use — the practical default for high-volume RAG…
- Parameters
- 32B
- Context
- 128K tokens
- Input / output
- $0.15 / $0.60
- License
- CC-BY-NC-4.0
Cohere · General language Cohere's cheapest served model by a wide margin ($0.0375/1M in) with a 128K window — right for latency-sensitive chatbots, classification and…
- Parameters
- 8B
- Context
- 128K tokens
- Input / output
- $0.0375 / $0.15
- License
- CC-BY-NC-4.0
Cohere · General language Deprecated 2025-09-15 and 3.3x the price of command-r-08-2024 for worse results; the alias `command-r` also points here. No reason to choose it —…
- Parameters
- 35B
- Context
- 128K tokens
- Input / output
- $0.50 / $1.50
- License
- CC-BY-NC-4.0
Cohere · General language Deprecated 2025-09-15 and the most expensive model Cohere still serves; the alias `command-r-plus` resolves here. Move to command-r-plus-08-2024 for…
- Parameters
- 104B
- Context
- 128K tokens
- Input / output
- $3.00 / $15.00
- License
- CC-BY-NC-4.0
Cohere's first agentic coding model: Apache-2.0, free API key plus free weights, and a 3B active footprint that runs locally — aimed at repo-level…
- Parameters
- 30B total · 3B active
- Context
- Not recorded
- Input / output
- $0.00 / $0.00
- License
- Apache-2.0
Cohere · Speech (transcription or voice) Open (Apache-2.0) 2B Conformer ASR covering 14 languages with a real-time factor up to 3x faster than similar-size models; no timestamps, no…
- Parameters
- 2B
- Context
- Not recorded
- Input / output
- Not published
- License
- Apache-2.0
Cohere · Speech (transcription or voice) Arabic-specialised fine-tune of Cohere Transcribe — use it over the base model for any Arabic audio; same 25MB file cap, and it is not yet on…
- Parameters
- 2B
- Context
- Not recorded
- Input / output
- Not published
- License
- Apache-2.0
Cohere · General language Research-grade 23-language model with a 128K window at a flat $0.50/$1.50 — the cheapest way to test Cohere-family multilingual quality, but…
- Parameters
- 32B
- Context
- 128K tokens
- Input / output
- $0.50 / $1.50
- License
- CC-BY-NC-4.0
Cohere · Vision & multimodal Open multilingual vision-language research model across 23 languages; the pricing FAQ covers only Aya Expanse, so its API price is unstated — for…
- Parameters
- 32B
- Context
- 16K tokens
- Input / output
- Not published
- License
- CC-BY-NC-4.0
Cohere · General language The balanced 3.35B/70-language Tiny Aya variant — the one to start with when you want broad low-resource language coverage on small hardware (GGUF…
- Parameters
- 3.35B
- Context
- 8K tokens
- Input / output
- Not published
- License
- CC-BY-NC-4.0
Cohere · General language Region-specialised Tiny Aya tuned for West Asian and African languages; choose it over Tiny Aya Global only when your traffic is concentrated in…
- Parameters
- 3.35B
- Context
- 8K tokens
- Input / output
- Not published
- License
- CC-BY-NC-4.0
Cohere · General language Region-specialised Tiny Aya tuned for South Asian languages; worth the swap from Tiny Aya Global for Indic-heavy workloads. No public API price.
- Parameters
- 3.35B
- Context
- 8K tokens
- Input / output
- Not published
- License
- CC-BY-NC-4.0
Cohere · General language Region-specialised Tiny Aya tuned for European and Asia-Pacific languages; the pick for EU/APAC-focused small-model deployments. No public API price.
- Parameters
- 3.35B
- Context
- 8K tokens
- Input / output
- Not published
- License
- CC-BY-NC-4.0
Chinese AI labs · Vision & multimodal Frontier long-context agent work where you can hit cache: the 10x cache-hit discount ($0.30 vs $3.00) matters far more than the headline rate, and…
- Parameters
- 2.8T total · 104B active
- Context
- 1M tokens
- Input / output
- $2.98 / $14.90
- License
- Kimi K3 License (custom, license_name…
The value pick for autonomous coding agents — roughly a third of K3's input cost and a quarter of its output cost, with 256k context and weights you…
- Parameters
- 1T total · 32B active
- Context
- 262K tokens
- Input / output
- $0.95 / $4.00
- License
- Modified MIT
Same weights as kimi-k2.7-code at exactly 2x the price — only worth it when interactive latency, not cost, is the constraint.
- Parameters
- 1T total · 32B active
- Context
- 262K tokens
- Input / output
- $1.90 / $8.00
- License
- Modified MIT
Chinese AI labs · Vision & multimodal Cheapest multimodal option in the Kimi line — pick it over K2.7 Code when you need image/video input and over K3 when 256k context is enough.
- Parameters
- 1T total · 32B active
- Context
- 262K tokens
- Input / output
- $0.97 / $4.02
- License
- Modified MIT
Chinese AI labs · General language Best frontier-class price/context ratio here — 1M context at $1.40 in / $4.40 out, but note reasoning cannot be disabled, so budget for thinking…
- Parameters
- 753B
- Context
- 1M tokens
- Input / output
- $1.40 / $4.40
- License
- GLM-5.3 License (custom; license_name…
Chinese AI labs · Vision & multimodal The cheapest capable multimodal model in this entire report and plain MIT weights — but the posted rate is a 50%-off promo running to 9 Sep 2026, so…
- Parameters
- 320B total · 18B active
- Context
- 1M tokens
- Input / output
- $0.075 / $0.25
- License
- MIT
Chinese AI labs · General language Same base model and same price as GLM-5.3 but under plain MIT and with thinking optionally off — the one to self-host, or to use when you need…
- Parameters
- 753B
- Context
- 1M tokens
- Input / output
- $1.40 / $4.40
- License
- MIT
Chinese AI labs · General language Cheaper than GLM-5.2/5.3 if 200k context is enough; otherwise the newer siblings are worth the extra 40 cents per million input.
- Parameters
- 754B
- Context
- 200K tokens
- Input / output
- $1.00 / $3.20
- License
- MIT
Chinese AI labs · General language Solid mid-tier coding/agent workhorse at under half GLM-5.2's price; step down to it when you don't need million-token context.
- Parameters
- Not recorded
- Context
- 200K tokens
- Input / output
- $0.60 / $2.20
- License
- MIT
Chinese AI labs · General language Legacy small model kept alive for existing integrations — new builds should start on GLM-4.7-FlashX or GLM-5.3-Flash instead.
- Parameters
- Not recorded
- Context
- Not recorded
- Input / output
- $0.20 / $1.10
- License
- MIT
Chinese AI labs · Vision & multimodal Outstanding cost-per-context: 1M window at $0.30/$1.20 with open weights — just watch the tier break, since anything over 512k input bills at double.
- Parameters
- 427B total · 23B active
- Context
- 1M tokens
- Input / output
- $0.30 / $1.20
- License
- MiniMax Model License (license_name…
Chinese AI labs · General language Same price as M3 with a fifth of the context and no vision — only choose it if you have already validated against these exact weights.
- Parameters
- 229B
- Context
- 205K tokens
- Input / output
- $0.30 / $1.20
- License
- MiniMax Model License (custom; see…
Chinese AI labs · General language A 2x latency surcharge on identical weights — justify it with a measured p95 requirement, not a hunch.
- Parameters
- 229B
- Context
- 205K tokens
- Input / output
- $0.60 / $2.40
- License
- MiniMax Model License (custom; see…
Chinese AI labs · General language Absurdly cheap 224k-context reasoning (¥0.5/¥2 per 1M = $0.07/$0.30) — the budget choice for bulk long-document work if you can live with a…
- Parameters
- 80B total · 13B active
- Context
- 224K tokens
- Input / output
- $0.074 / $0.297
- License
- Tencent Hunyuan A13B Community License
Chinese AI labs · General language Open weights only — Tencent sells no first-party per-token endpoint for it, so budget for GPUs (770B params) or a third-party host; the Apache-2.0…
- Parameters
- 770B total · 49B active
- Context
- 1M tokens
- Input / output
- Not published
- License
- Apache-2.0
Chinese AI labs · General language Apache-2.0, 21B active — the most self-hostable strong model in this report, and the practical Tencent choice when Hy4-preview's 770B is too big for…
- Parameters
- 295B total · 21B active
- Context
- 256K tokens
- Input / output
- Not published
- License
- Apache-2.0
Chinese AI labs · Reasoning Small Apache-2.0 reasoning model that runs on a single modern GPU; not a line item in Qianfan's price table, so treat it as self-host-only.
- Parameters
- 21B total · 3B active
- Context
- 131K tokens
- Input / output
- Not published
- License
- Apache-2.0
Other notable labs · General language Long-document grounded QA where you need a 256K window and citations-faithful answers on a budget; the hybrid SSM design keeps long-context cost far…
- Parameters
- 398B total · 94B active
- Context
- 262K tokens
- Input / output
- $2.00 / $8.00
- License
- Jamba Open Model License
Other notable labs · General language High-volume RAG and summarisation over long inputs; one of the cheapest 256K-context commercial endpoints available.
- Parameters
- 52B total · 12B active
- Context
- 256K tokens
- Input / output
- $0.20 / $0.40
- License
- Jamba Open Model License
Other notable labs · General language Self-hosted enterprise QA where you want 256K context and Apache-2.0 freedom; answers without the token overhead of a reasoning model. Not on AI21's…
- Parameters
- 52B total · 12B active
- Context
- 256K tokens
- Input / output
- Not published
- License
- Apache 2.0
Other notable labs · General language On-device RAG on iOS/Android/desktop when you need an unusually large 256K window from a 3B model.
- Parameters
- 3B
- Context
- 256K tokens
- Input / output
- Not published
- License
- Apache 2.0
Other notable labs · Reasoning Local reasoning at 3B scale; evaluate against Qwen and LFM2.5 thinking models before committing, since AI21 publishes no hosted endpoint for it.
- Parameters
- 3B
- Context
- Not recorded
- Input / output
- Not published
- License
- Apache 2.0
Other notable labs · Vision & multimodal Cheapest way to bolt image/video understanding onto a high-volume pipeline, and the weights are downloadable if you'd rather run it yourself; check…
- Parameters
- 7B
- Context
- Not recorded
- Input / output
- $0.10 / $0.10
- License
- reka-edge-2603-license
Other notable labs · Reasoning The strongest fully-reproducible reasoning model: choose it when you must audit or re-derive the training pipeline, not when you need best-in-class…
- Parameters
- 32B
- Context
- 66K tokens
- Input / output
- Not published
- License
- Apache 2.0
Other notable labs · General language General chat variant of the fully-open 32B; Ai2 flags it as intended for research and educational use, so read the Responsible Use Guidelines before…
- Parameters
- 32B
- Context
- 66K tokens
- Input / output
- Not published
- License
- Apache 2.0
Other notable labs · General language The base checkpoint to fine-tune when you need documented provenance for every training token (Dolma 3, ~9.3T tokens) plus intermediate checkpoints.
- Parameters
- 32B
- Context
- 66K tokens
- Input / output
- Not published
- License
- Apache 2.0
Other notable labs · General language Small fully-open chat model for academic baselines and ablation studies where a licence-clean, data-transparent 7B is the requirement.
- Parameters
- 7B
- Context
- 66K tokens
- Input / output
- Not published
- License
- Apache 2.0
Other notable labs · Reasoning Cheap local long-chain-of-thought experiments; pair with the RL-Zero checkpoints if you are studying RL recipes rather than deploying.
- Parameters
- 7B
- Context
- 66K tokens
- Input / output
- Not published
- License
- Apache 2.0
Other notable labs · Vision & multimodal Best open choice when you need pointing/grounding output — coordinates and timestamps rather than prose — on short video and multi-image inputs;…
- Parameters
- 9B
- Context
- 16K tokens
- Input / output
- Not published
- License
- Apache 2.0 (some training data is…
Other notable labs · Vision & multimodal The efficiency-tuned Molmo 2 for single-GPU video grounding when the 8B is too heavy.
- Parameters
- 4B
- Context
- 16K tokens
- Input / output
- Not published
- License
- Apache 2.0 (some training data is…
Other notable labs · Vision & multimodal The only end-to-end open VLM here whose language backbone is also fully open (Olmo, not Qwen) — pick it when backbone provenance is the point.
- Parameters
- 7B
- Context
- 16K tokens
- Input / output
- Not published
- License
- Apache 2.0 (some training data is…
Other notable labs · General language Frontier-scale open weights for on-prem agentic workloads — but it needs 8x GB200/B200 or 16x H100 minimum, so it is a datacentre commitment, not a…
- Parameters
- 550B total · 55B active
- Context
- 1M tokens
- Input / output
- Not published
- License
- OpenMDW-1.1
Other notable labs · General language The practical sweet spot of the Nemotron line: 12B active params keeps throughput high for high-volume ticket automation and RAG while retaining a…
- Parameters
- 120B total · 12B active
- Context
- 1M tokens
- Input / output
- Not published
- License
- NVIDIA Nemotron Open Model License
Other notable labs · General language Newest and most deployable Nemotron: 256K context on a single H100 with only 3B active params, and reasoning can be switched off per-request to cut…
- Parameters
- 30B total · 3B active
- Context
- 1M tokens
- Input / output
- Not published
- License
- OpenMDW-1.1
Other notable labs · Vision & multimodal Document intelligence — invoices, forms, charts — at up to 4 images of 2048x1536 per request; a strong self-hosted alternative to paid OCR/VLM APIs.
- Parameters
- 12.6B
- Context
- 128K tokens
- Input / output
- Not published
- License
- NVIDIA Open Model License Agreement
Other notable labs · General language Apache-2.0 enterprise workhorse with 128K native context (512K extensible) and three effort levels, so you can dial reasoning cost per request;…
- Parameters
- 30B dense
- Context
- 128K tokens
- Input / output
- Not published
- License
- Apache 2.0
Other notable labs · General language The best value in the Granite line for self-hosted agents: 128K context and reliable tool calling on a single mid-range GPU, 12 languages, no…
- Parameters
- 8B dense
- Context
- 128K tokens
- Input / output
- Not published
- License
- Apache 2.0
Other notable labs · General language Smallest Granite 4.2 for CPU or edge deployment where you still want a 128K window; GGUF, nvfp4 and mxfp4 builds ship alongside.
- Parameters
- 3B dense
- Context
- 128K tokens
- Input / output
- Not published
- License
- Apache 2.0
Other notable labs · General language Largest LFM: strong agentic tool use at 1.5B active params and 18.5K tok/s at high concurrency — but Liquid explicitly says it is weak on heavy…
- Parameters
- 8.3B total · 1.5B active
- Context
- 128K tokens
- Input / output
- Not published
- License
- lfm1.0
Other notable labs · General language The default on-device model here: 131K context, 16 languages, and competitive with models 4x its size on tool use — just note the bespoke lfm1.0…
- Parameters
- 2.69B
- Context
- 131K tokens
- Input / output
- Not published
- License
- lfm1.0
Other notable labs · Vision & multimodal Vision on a laptop or NPU: ~3.3GB memory and 228 tok/s on an M5 Max, at the cost of a comparatively short 32K window.
- Parameters
- 3.1B
- Context
- 33K tokens
- Input / output
- Not published
- License
- lfm1.0
Other notable labs · Reasoning MIT-licensed multimodal reasoner sized for a single GPU; ServiceNow claims ~30% fewer reasoning tokens than Apriel 1.5, which matters more than raw…
- Parameters
- 15B
- Context
- 131K tokens
- Input / output
- Not published
- License
- MIT
Other notable labs · General language Narrow but useful: a small model RL-trained specifically for multi-turn MCP tool calling. Also ships at 4B and 14B. Treat it as a research artifact,…
- Parameters
- 8B
- Context
- Not recorded
- Input / output
- Not published
- License
- Apache 2.0