Google DeepMind
Google DeepMind is Google's consolidated AI research division (London/Mountain View), formed in April 2023 by merging DeepMind (founded London 2010, acquired by Google 2014) with Google Brain. It is wholly owned by Alphabet and is not independently funded.
Data checked 2026-09-06- Models tracked
- 46
- OpenAI-compatible API
- yes
- API base URL
- https://generativelanguage.googleapis.com/v1beta
It builds the Gemini frontier model family plus the open-weight Gemma models, Veo (video), Lyria (music), Imagen, and the Gemini Robotics ER embodied-reasoning line. Models ship through two commercial surfaces: the Gemini Developer API / Google AI Studio (ai.google.dev, self-serve API keys) and the enterprise Vertex AI / Gemini Enterprise Agent Platform on Google Cloud; the positioning is long-context (1M-token), natively multimodal, tool-integrated models at aggressive price-performance versus OpenAI and Anthropic.
A Access and limits
API keys are self-serve from Google AI Studio (aistudio.google.com) with no credit card required for the free tier. FREE TIER: most Gemini text models (3.8/3.7/3.6/3.5 Flash, 3.5 & 3.1 Flash-Lite, 3 Flash Preview, 2.5 Pro/Flash/Flash-Lite, TTS, Live, Transcribe, embeddings, robotics, Gemma 4) are free of charge at limited rates; free-tier content IS used to improve Google's products, while paid-tier content is not. Notably NOT on the free tier: Gemini 3.1 Pro Preview, all Nano Banana image models, Gemini Omni Flash, 2.5 Pro TTS, 2.5 Computer Use, Veo and Lyria.
Google no longer publishes free-tier RPM/TPM/RPD in the docs — per-model live limits are shown at aistudio.google.com/rate-limit. PAID USAGE TIERS: Free (active project) -> Tier 1 (link an active billing account, $250 billing cap) -> Tier 2 (paid $100 + 3 days, $2,000 cap) -> Tier 3 (paid $1,000 + 30 days, $20,000-$100,000+ cap); a rolling 10-minute spend cap also applies ($10 / $50 / $200 for Tiers 1/2/3). CONSUMPTION MODES: Standard; Batch API and Flex inference at 50% of standard; Priority inference at 180% of standard (1.8x) with 0.3x the standard rate limit.
CONTEXT-LENGTH TIERS: Gemini 3.1 Pro Preview, 2.5 Pro and 2.5 Computer Use are cheaper for prompts <=200k tokens and step up above 200k — the recorded input/output prices are the <=200k base tier. INTRODUCTORY PRICING: Gemini 3.8/3.7/3.6 Flash are at half price ($0.75/$3.75) only through December 31, 2026; standard $1.50/$7.50 begins January 1, 2027. TOOLS ARE BILLED SEPARATELY: Google Search grounding on Gemini 3.x models is 5,000 free requests/month shared across all Gemini 3.x models, then $14 per 1,000 requests (Gemini 2.5 models: 1,500 RPD free, then $35 per 1,000 grounded prompts); Google Maps grounding $25 per 1,000 grounded prompts after free quota; File Search embeddings $0.15/1M tokens; code execution and URL context are billed as ordinary model tokens.
Context-cache storage is billed by the hour on top of cached-read tokens ($1.00/1M tokens/hour for most Flash models, $4.50/1M/hour for Pro-class and 3.8/3.7/3.6 Flash after the intro period, $0.50/1M/hour during the intro period). VERTEX AI: the same models are sold on Vertex AI / Gemini Enterprise Agent Platform at identical global-endpoint prices, but non-global (regional) endpoints carry roughly a 10% premium (e.g. Gemini 3.8 Flash $0.825 input non-global vs $0.75 global); Vertex adds provisioned throughput, FSP committed-use discounts, VPC-SC and data-residency controls.
Batch API is capped by enqueued tokens per model per tier (e.g. Tier 1: 3M tokens for 3.8 Flash, 10M for 3.5 Flash-Lite; Tier 2: 400M-500M), with 100 concurrent batch jobs, 2GB input file limit and 20GB file storage.
C Calling the API
Google DeepMind accepts requests in the OpenAI Chat Completions format, so any
OpenAI-compatible client works by changing the base URL. Example uses
gemini-embedding-001.
curl https://generativelanguage.googleapis.com/v1beta/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-embedding-001",
"messages": [{"role": "user", "content": "Hello"}]
}'
from openai import OpenAI
client = OpenAI(
base_url="https://generativelanguage.googleapis.com/v1beta",
api_key=os.environ["API_KEY"],
)
response = client.chat.completions.create(
model="gemini-embedding-001",
messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)
A compatibility layer is not always the provider's full API — caching, strict tool schemas and structured output are commonly unsupported there. Check the official docs before relying on an advanced feature.
B Models
| Model | Context | Max out | In $/M | Out $/M | Blended | Status |
|---|---|---|---|---|---|---|
Gemma 4 31B Instruct
gemma-4-31b-it
|
262.1K | — | $0.00 | $0.00 | $0.00 | ga open |
Gemma 4 26B A4B Instruct
gemma-4-26b-a4b-it
|
262.1K | — | $0.00 | $0.00 | $0.00 | ga open |
| Gemma 4 12B | 262.1K | — | $0.00 | $0.00 | $0.00 | ga open |
| Gemma 4 E4B | 128K | — | $0.00 | $0.00 | $0.00 | ga open |
| Gemma 4 E2B | 128K | — | $0.00 | $0.00 | $0.00 | ga open |
Gemini Embedding 001
gemini-embedding-001
|
2K | — | $0.15 | $0.00 | $0.11 | ga |
Gemini Embedding 2
gemini-embedding-2
|
8.2K | — | $0.20 | $0.00 | $0.15 | ga |
Gemini 2.5 Flash-Lite
gemini-2.5-flash-lite
|
1M | 65.5K | $0.10 | $0.40 | $0.18 | ga |
Gemini 3.1 Flash-Lite
gemini-3.1-flash-lite
|
1M | 65.5K | $0.25 | $1.50 | $0.56 | ga |
Gemini 3.5 Flash-Lite
gemini-3.5-flash-lite
|
1M | 65.5K | $0.30 | $2.50 | $0.85 | ga |
Gemini 2.5 Flash
gemini-2.5-flash
|
1M | 65.5K | $0.30 | $2.50 | $0.85 | ga |
Gemini 2.5 Flash Native Audio (Live API)
gemini-2.5-flash-native-audio-preview-12-2025
|
131.1K | 8.2K | $0.50 | $2.00 | $0.88 | preview |
Gemini 3 Flash Preview
gemini-3-flash-preview
|
1M | 65.5K | $0.50 | $3.00 | $1.13 | preview |
Gemini 3.8 Flash
gemini-3.8-flash
|
1M | 65.5K | $0.75 | $3.75 | $1.50 | ga |
Gemini 3.7 Flash
gemini-3.7-flash
|
1M | 65.5K | $0.75 | $3.75 | $1.50 | ga |
Gemini 3.6 Flash
gemini-3.6-flash
|
1M | 65.5K | $0.75 | $3.75 | $1.50 | ga |
Gemini 3.1 Flash Live Preview
gemini-3.1-flash-live-preview
|
131.1K | 65.5K | $0.75 | $4.50 | $1.69 | preview |
Gemini Robotics ER 1.6 Preview
gemini-robotics-er-1.6-preview
|
131.1K | 65.5K | $1.00 | $5.00 | $2.00 | preview |
Gemini 2.5 Flash Preview TTS
gemini-2.5-flash-preview-tts
|
8.2K | 16.4K | $0.50 | $10.00 | $2.88 | preview |
Gemini 3.5 Flash
gemini-3.5-flash
|
1M | 65.5K | $1.50 | $9.00 | $3.38 | ga |
Gemini 2.5 Pro
gemini-2.5-pro
|
1M | 65.5K | $1.25 | $10.00 | $3.44 | ga |
Gemini 2.5 Computer Use Preview
gemini-2.5-computer-use-preview-10-2025
|
128K | 64K | $1.25 | $10.00 | $3.44 | preview |
Gemini Robotics ER 2 Preview
gemini-robotics-er-2-preview
|
131.1K | 65.5K | $2.00 | $10.00 | $4.00 | preview |
Gemini Robotics ER 2 Streaming Preview
gemini-robotics-er-2-streaming-preview
|
131.1K | 65.5K | $2.00 | $10.00 | $4.00 | preview |
Gemini 3.1 Pro Preview
gemini-3.1-pro-preview
|
1M | 65.5K | $2.00 | $12.00 | $4.50 | preview |
Gemini 3.5 Transcribe
gemini-3.5-transcribe
|
— | — | $2.00 | $12.00 | $4.50 | ga |
Gemini Omni Flash
gemini-omni-1.1-flash
|
1M | — | $1.50 | $17.50 | $5.50 | ga |
Gemini Omni Flash Preview
gemini-omni-flash-preview
|
1M | — | $1.50 | $17.50 | $5.50 | preview |
Gemini 3.1 Flash TTS Preview
gemini-3.1-flash-tts-preview
|
8.2K | 16.4K | $1.00 | $20.00 | $5.75 | preview |
Gemini 2.5 Pro Preview TTS
gemini-2.5-pro-preview-tts
|
8.2K | 16.4K | $1.00 | $20.00 | $5.75 | preview |
Gemini 3.1 Flash Lite Image (Nano Banana 2 Lite)
gemini-3.1-flash-lite-image
|
65.5K | 4.1K | $0.25 | $30.00 | $7.69 | ga |
Gemini 3.5 Transcribe Live
gemini-3.5-transcribe-live
|
— | — | $3.50 | $21.00 | $7.88 | ga |
Gemini 3.5 Live Translate Preview
gemini-3.5-live-translate-preview
|
131.1K | 65.5K | $3.50 | $21.00 | $7.88 | preview |
Gemini 3.1 Flash Image (Nano Banana 2)
gemini-3.1-flash-image
|
131.1K | 32.8K | $0.50 | $60.00 | $15.38 | ga |
Gemini 3 Pro Image (Nano Banana Pro)
gemini-3-pro-image
|
65.5K | 32.8K | $2.00 | $120.00 | $31.50 | ga |
Gemini 2.5 Flash Image (Nano Banana)
gemini-2.5-flash-image
|
65.5K | 32.8K | $0.30 | — | — | ga |
Veo 3.1
veo-3.1-generate-preview
|
— | — | — | — | — | preview |
Veo 3.1 Fast
veo-3.1-fast-generate-preview
|
— | — | — | — | — | preview |
Veo 3.1 Lite
veo-3.1-lite-generate-preview
|
— | — | — | — | — | preview |
Lyria 3.5
lyria-3.5
|
— | — | — | — | — | ga |
Lyria 3 Clip Preview
lyria-3-clip-preview
|
— | — | — | — | — | preview |
Lyria 3 Pro Preview
lyria-3-pro-preview
|
— | — | — | — | — | preview |
Lyria RealTime
lyria-realtime-exp
|
— | — | — | — | — | preview |
Gemini Deep Research
deep-research-preview-04-2026
|
1M | 65.5K | — | — | — | preview |
Gemini Deep Research Max
deep-research-max-preview-04-2026
|
1M | 65.5K | — | — | — | preview |
Antigravity Agent
antigravity-preview-05-2026
|
1M | 65.5K | — | — | — | preview |