Seminal AI
§2

Google DeepMind

Google DeepMind is Google's consolidated AI research division (London/Mountain View), formed in April 2023 by merging DeepMind (founded London 2010, acquired by Google 2014) with Google Brain. It is wholly owned by Alphabet and is not independently funded.

Data checked 2026-09-06
Models tracked
46
OpenAI-compatible API
yes
API base URL
https://generativelanguage.googleapis.com/v1beta

It builds the Gemini frontier model family plus the open-weight Gemma models, Veo (video), Lyria (music), Imagen, and the Gemini Robotics ER embodied-reasoning line. Models ship through two commercial surfaces: the Gemini Developer API / Google AI Studio (ai.google.dev, self-serve API keys) and the enterprise Vertex AI / Gemini Enterprise Agent Platform on Google Cloud; the positioning is long-context (1M-token), natively multimodal, tool-integrated models at aggressive price-performance versus OpenAI and Anthropic.

A Access and limits

API keys are self-serve from Google AI Studio (aistudio.google.com) with no credit card required for the free tier. FREE TIER: most Gemini text models (3.8/3.7/3.6/3.5 Flash, 3.5 & 3.1 Flash-Lite, 3 Flash Preview, 2.5 Pro/Flash/Flash-Lite, TTS, Live, Transcribe, embeddings, robotics, Gemma 4) are free of charge at limited rates; free-tier content IS used to improve Google's products, while paid-tier content is not. Notably NOT on the free tier: Gemini 3.1 Pro Preview, all Nano Banana image models, Gemini Omni Flash, 2.5 Pro TTS, 2.5 Computer Use, Veo and Lyria.

Google no longer publishes free-tier RPM/TPM/RPD in the docs — per-model live limits are shown at aistudio.google.com/rate-limit. PAID USAGE TIERS: Free (active project) -> Tier 1 (link an active billing account, $250 billing cap) -> Tier 2 (paid $100 + 3 days, $2,000 cap) -> Tier 3 (paid $1,000 + 30 days, $20,000-$100,000+ cap); a rolling 10-minute spend cap also applies ($10 / $50 / $200 for Tiers 1/2/3). CONSUMPTION MODES: Standard; Batch API and Flex inference at 50% of standard; Priority inference at 180% of standard (1.8x) with 0.3x the standard rate limit.

CONTEXT-LENGTH TIERS: Gemini 3.1 Pro Preview, 2.5 Pro and 2.5 Computer Use are cheaper for prompts <=200k tokens and step up above 200k — the recorded input/output prices are the <=200k base tier. INTRODUCTORY PRICING: Gemini 3.8/3.7/3.6 Flash are at half price ($0.75/$3.75) only through December 31, 2026; standard $1.50/$7.50 begins January 1, 2027. TOOLS ARE BILLED SEPARATELY: Google Search grounding on Gemini 3.x models is 5,000 free requests/month shared across all Gemini 3.x models, then $14 per 1,000 requests (Gemini 2.5 models: 1,500 RPD free, then $35 per 1,000 grounded prompts); Google Maps grounding $25 per 1,000 grounded prompts after free quota; File Search embeddings $0.15/1M tokens; code execution and URL context are billed as ordinary model tokens.

Context-cache storage is billed by the hour on top of cached-read tokens ($1.00/1M tokens/hour for most Flash models, $4.50/1M/hour for Pro-class and 3.8/3.7/3.6 Flash after the intro period, $0.50/1M/hour during the intro period). VERTEX AI: the same models are sold on Vertex AI / Gemini Enterprise Agent Platform at identical global-endpoint prices, but non-global (regional) endpoints carry roughly a 10% premium (e.g. Gemini 3.8 Flash $0.825 input non-global vs $0.75 global); Vertex adds provisioned throughput, FSP committed-use discounts, VPC-SC and data-residency controls.

Batch API is capped by enqueued tokens per model per tier (e.g. Tier 1: 3M tokens for 3.8 Flash, 10M for 3.5 Flash-Lite; Tier 2: 400M-500M), with 100 concurrent batch jobs, 2GB input file limit and 20GB file storage.

C Calling the API

Google DeepMind accepts requests in the OpenAI Chat Completions format, so any OpenAI-compatible client works by changing the base URL. Example uses gemini-embedding-001.

curl https://generativelanguage.googleapis.com/v1beta/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-embedding-001",
    "messages": [{"role": "user", "content": "Hello"}]
  }'
from openai import OpenAI

client = OpenAI(
    base_url="https://generativelanguage.googleapis.com/v1beta",
    api_key=os.environ["API_KEY"],
)

response = client.chat.completions.create(
    model="gemini-embedding-001",
    messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)

A compatibility layer is not always the provider's full API — caching, strict tool schemas and structured output are commonly unsupported there. Check the official docs before relying on an advanced feature.

B Models

Model Context Max out In $/M Out $/M Blended Status
Gemma 4 31B Instruct gemma-4-31b-it 262.1K $0.00 $0.00 $0.00 ga open
Gemma 4 26B A4B Instruct gemma-4-26b-a4b-it 262.1K $0.00 $0.00 $0.00 ga open
Gemma 4 12B 262.1K $0.00 $0.00 $0.00 ga open
Gemma 4 E4B 128K $0.00 $0.00 $0.00 ga open
Gemma 4 E2B 128K $0.00 $0.00 $0.00 ga open
Gemini Embedding 001 gemini-embedding-001 2K $0.15 $0.00 $0.11 ga
Gemini Embedding 2 gemini-embedding-2 8.2K $0.20 $0.00 $0.15 ga
Gemini 2.5 Flash-Lite gemini-2.5-flash-lite 1M 65.5K $0.10 $0.40 $0.18 ga
Gemini 3.1 Flash-Lite gemini-3.1-flash-lite 1M 65.5K $0.25 $1.50 $0.56 ga
Gemini 3.5 Flash-Lite gemini-3.5-flash-lite 1M 65.5K $0.30 $2.50 $0.85 ga
Gemini 2.5 Flash gemini-2.5-flash 1M 65.5K $0.30 $2.50 $0.85 ga
Gemini 2.5 Flash Native Audio (Live API) gemini-2.5-flash-native-audio-preview-12-2025 131.1K 8.2K $0.50 $2.00 $0.88 preview
Gemini 3 Flash Preview gemini-3-flash-preview 1M 65.5K $0.50 $3.00 $1.13 preview
Gemini 3.8 Flash gemini-3.8-flash 1M 65.5K $0.75 $3.75 $1.50 ga
Gemini 3.7 Flash gemini-3.7-flash 1M 65.5K $0.75 $3.75 $1.50 ga
Gemini 3.6 Flash gemini-3.6-flash 1M 65.5K $0.75 $3.75 $1.50 ga
Gemini 3.1 Flash Live Preview gemini-3.1-flash-live-preview 131.1K 65.5K $0.75 $4.50 $1.69 preview
Gemini Robotics ER 1.6 Preview gemini-robotics-er-1.6-preview 131.1K 65.5K $1.00 $5.00 $2.00 preview
Gemini 2.5 Flash Preview TTS gemini-2.5-flash-preview-tts 8.2K 16.4K $0.50 $10.00 $2.88 preview
Gemini 3.5 Flash gemini-3.5-flash 1M 65.5K $1.50 $9.00 $3.38 ga
Gemini 2.5 Pro gemini-2.5-pro 1M 65.5K $1.25 $10.00 $3.44 ga
Gemini 2.5 Computer Use Preview gemini-2.5-computer-use-preview-10-2025 128K 64K $1.25 $10.00 $3.44 preview
Gemini Robotics ER 2 Preview gemini-robotics-er-2-preview 131.1K 65.5K $2.00 $10.00 $4.00 preview
Gemini Robotics ER 2 Streaming Preview gemini-robotics-er-2-streaming-preview 131.1K 65.5K $2.00 $10.00 $4.00 preview
Gemini 3.1 Pro Preview gemini-3.1-pro-preview 1M 65.5K $2.00 $12.00 $4.50 preview
Gemini 3.5 Transcribe gemini-3.5-transcribe $2.00 $12.00 $4.50 ga
Gemini Omni Flash gemini-omni-1.1-flash 1M $1.50 $17.50 $5.50 ga
Gemini Omni Flash Preview gemini-omni-flash-preview 1M $1.50 $17.50 $5.50 preview
Gemini 3.1 Flash TTS Preview gemini-3.1-flash-tts-preview 8.2K 16.4K $1.00 $20.00 $5.75 preview
Gemini 2.5 Pro Preview TTS gemini-2.5-pro-preview-tts 8.2K 16.4K $1.00 $20.00 $5.75 preview
Gemini 3.1 Flash Lite Image (Nano Banana 2 Lite) gemini-3.1-flash-lite-image 65.5K 4.1K $0.25 $30.00 $7.69 ga
Gemini 3.5 Transcribe Live gemini-3.5-transcribe-live $3.50 $21.00 $7.88 ga
Gemini 3.5 Live Translate Preview gemini-3.5-live-translate-preview 131.1K 65.5K $3.50 $21.00 $7.88 preview
Gemini 3.1 Flash Image (Nano Banana 2) gemini-3.1-flash-image 131.1K 32.8K $0.50 $60.00 $15.38 ga
Gemini 3 Pro Image (Nano Banana Pro) gemini-3-pro-image 65.5K 32.8K $2.00 $120.00 $31.50 ga
Gemini 2.5 Flash Image (Nano Banana) gemini-2.5-flash-image 65.5K 32.8K $0.30 ga
Veo 3.1 veo-3.1-generate-preview preview
Veo 3.1 Fast veo-3.1-fast-generate-preview preview
Veo 3.1 Lite veo-3.1-lite-generate-preview preview
Lyria 3.5 lyria-3.5 ga
Lyria 3 Clip Preview lyria-3-clip-preview preview
Lyria 3 Pro Preview lyria-3-pro-preview preview
Lyria RealTime lyria-realtime-exp preview
Gemini Deep Research deep-research-preview-04-2026 1M 65.5K preview
Gemini Deep Research Max deep-research-max-preview-04-2026 1M 65.5K preview
Antigravity Agent antigravity-preview-05-2026 1M 65.5K preview

D Official references