| Amazon Bedrock |
~100+ models, ~18 providers: Claude, Nova, Llama, Mistral, Cohere, and more | 30+ AWS commercial regions plus GovCloud; geographic (US/EU/APAC/JP/AU) and Global cross-Region inference profiles | Data at rest stays in source region; Global CRIS processes cross-region. Prompts not used for training. ZDR granted per account/model via account team | Yes — hourly model units (1-month/6-month terms) plus a Reserved tier priced per 1K tokens-per-minute for 1 or 3 months | Native AWS IAM, SCPs, KMS CMK, PrivateLink, CloudTrail, Bedrock Guardrails |
On-demand from $0.035/1M input tokens (Amazon Nova Micro) up to ~$10/1M input for the top Claude tier; Nova Pro ~$0.80 in / $3.20 out per 1M. Batch 50% off, Flex 50% off, Priority +75%. Provisioned Throughput model units run roughly $21-$50/hour depending on model and term. |
| Gemini Enterprise Agent Platform (formerly Vertex AI) |
Gemini 3.x family (Pro, Flash, Flash Lite, image models) plus Model Garden third-party and open weights including Claude, Llama, Mistral | ~40 Google Cloud regions and multi-regions, plus a global endpoint with no regional isolation | ML processing occurs in the requested region/multi-region; data at rest stays put. Global endpoint waives residency. 24h in-memory cache is on by default and can be disabled for ZDR | Yes — generative AI scale units (GSUs), 1-month or 1-year commitments | Cloud IAM, VPC Service Controls, CMEK, Org Policy, Cloud Audit Logs |
Gemini 3 Pro $2.00/1M input and $12.00/1M output at ≤200K context, rising to $4.00/$18.00 above 200K. Provisioned Throughput roughly $42-$158 per GSU-hour depending on model and region (third-party figure), with 20-45% off for 1-month or 1-year commitments. |
| Microsoft Foundry (formerly Azure AI Foundry) |
10,000+ in the Foundry Models catalog: OpenAI GPT-5.x/GPT-6, Anthropic, Meta, DeepSeek, Mistral and open weights | Global, Data Zone (US, EU, and new APAC), and Regional deployments across up to 27 Azure regions | Data Zone and Regional deployment types constrain processing geography; Global does not. Abuse-monitoring exemption available on request for zero retention | Yes — PTUs billed hourly, 15 PTU minimum (Global/Data Zone), with 1-month and 1-year Azure Reservations | Microsoft Entra ID, Azure RBAC, Private Link, CMK, Azure Policy, unified Foundry control plane |
GPT-5.6 Sol $5/$30, Terra $2/$12, Luna $0.20/$1.20 per 1M input/output at Global Standard. Batch is 50% off Global Standard. PTUs bill hourly per deployed PTU with a 15 PTU minimum for Global and Data Zone; from 1 Sept 2026 EU Data Zone is +9% over Global, non-US Regional +7-16%, and the new APAC Data Zone +20%. |
| Cloudflare Workers AI |
50+ open-weight models (Llama, Mistral, DeepSeek V4, Kimi, GLM, embeddings, image, ASR); no closed frontier models | Runs across Cloudflare's global network; no user-selectable inference region | Weak — inference and AI Gateway logs cannot be pinned to the EU; Data Localization Suite is an Enterprise add-on that does not cover hosted inference | No — pay-per-use only, no reserved capacity or capacity SLA | Cloudflare account roles and scoped API tokens plus Worker bindings; no federation with cloud IAM systems |
10,000 neurons/day free; $0.011 per 1,000 neurons beyond that on Workers Paid. Example: Llama 3.2 1B costs 2,457 neurons per 1M input tokens and 18,252 per 1M output; DeepSeek V4 Pro 120,000 in / 360,000 out per 1M. |
| IBM watsonx.ai |
IBM Granite 4.1/4.2 (Apache 2.0) plus Llama, Ministral and other open weights; no closed frontier models | Foundation-model inference concentrated in Dallas (us-south) and Frankfurt (eu-de); on-prem via Cloud Pak for Data and on IBM Z/LinuxONE with Spyre | Region-pinned by IBM Cloud region; self-managed and Z/LinuxONE deployments give full control. Prompts not used for training | Partial — dedicated model hosting / on-demand GPU deployment on the Standard plan rather than a token-rate reservation | IBM Cloud IAM, resource groups, Key Protect/Hyper Protect Crypto, watsonx.governance integration |
Free playground: 300,000 tokens/month, 20 CUH, 100 text-extraction documents. Essentials from $0/month pay-as-you-go with ML at $0.55/Capacity Unit-Hour and text extraction at $0.0403/page. Standard from $1,110/month with ML at $0.45/CUH and $0.0318/page. Inference billed in Resource Units of 1,000 tokens. |
| Oracle OCI Generative AI |
Cohere Command A family, Llama 4 Maverick/Scout, Llama 3.3 70B, gpt-oss-120b/20b, Gemini 2.5 Pro/Flash, Grok 4.3 and 4.20, Embed 4, Rerank 4; no Claude | Selected OCI commercial regions per model, plus sovereign and US Classified Cloud regions | Region-scoped; dedicated AI clusters are single-tenant and reachable only from your tenancy, which is the strongest isolation story here | Yes — dedicated AI clusters for hosting and fine-tuning, 744 unit-hour minimum per hosting cluster | OCI IAM policies and compartments, Vault/KMS, private endpoints, audit service |
— |
| Snowflake Cortex AI |
Anthropic, OpenAI (GPT 5.6 preview), Google Gemini 3.1 Pro, Meta Llama, Mistral, DeepSeek-V4-Flash, GLM-5.3, plus Snowflake Arctic embeddings | Snowflake account regions across AWS, Azure and GCP; many models require cross-region inference to reach | Home-region pinning costs $2.20/credit vs $2.00 for global routing; data stays inside the Snowflake governance boundary and is not used for training | No — consumption-only; no reserved token capacity | Snowflake RBAC, network policies, Tri-Secret Secure; administrators allowlist which models and providers users may call |
AI Credits at $2.00 each with cross-region routing (ANY_REGION/AWS_GLOBAL) or $2.20 pinned to the home region. Per-model AI Function rates run roughly $0.12 per 1M tokens on small open models to about $5.10 per 1M on frontier models. Cortex Search adds serving compute per GB-month plus per-token embedding cost. |
| Databricks Mosaic AI / Agent Bricks |
Very broad: GPT-6 Astra and GPT-5.x, Claude Opus 5/Sonnet 5/Fable 5.1, Gemini 3.x, Llama 4, Qwen 3.5, GLM 5.3, Kimi K3, DeepSeek V4, Grok 4.6, plus embeddings | Databricks workspace regions on AWS, Azure and GCP; model availability varies by workspace region | Inference runs within the workspace's cloud region for Databricks-hosted models; external model routing leaves the boundary. Data governed by Unity Catalog | Yes — provisioned throughput endpoints are the recommended production mode, billed in DBU/hr | Unity Catalog permissions, workspace SSO/SCIM, cloud IAM passthrough, AI Gateway policy and rate limiting |
Mosaic AI serving from $0.07/DBU; Foundation Model APIs priced in DBUs per 1M tokens (pay-per-token); GPU Model Serving roughly 10.48-628 DBU/hr depending on instance class. Cloud infrastructure is billed separately by AWS/Azure/GCP. External model routing incurs gateway DBUs on top of the upstream provider's own token fees. |
| Amazon SageMaker AI |
JumpStart catalog of open weights plus any model you containerise; no Claude, GPT or Gemini | Nearly all AWS commercial regions plus GovCloud | Strong — endpoints run in your chosen region and VPC, nothing leaves unless you route it out; no third-party model provider involvement | N/A by name — you provision instances directly; serverless inference bills per request | AWS IAM, VPC isolation, KMS CMK, PrivateLink, CloudTrail, SageMaker Role Manager |
— |
| Alibaba Cloud Model Studio |
Qwen3.x Max / Plus / Flash text models, plus Qwen vision, audio, embedding and image models; Qwen-centric catalog | International endpoint (Singapore) plus Beijing, Tokyo, Frankfurt and Virginia deployments | Endpoint choice determines jurisdiction; Beijing endpoint is subject to PRC data law. Verify retention terms per region before committing | Yes — dedicated/provisioned instance options for enterprise contracts, in addition to pay-per-token | Alibaba Cloud RAM roles and API keys; weaker federation with non-Alibaba identity systems |
International (Singapore) endpoint: qwen3.8-max $2.00 in / $6.00 out per 1M; qwen3.7-max $2.50/$7.50; qwen3.8-flash $0.15/$0.47; qwen3.7-plus $0.40-$1.20 in / $1.60-$4.80 out (tiered by context). Mainland/Beijing endpoint runs roughly 60-70% cheaper. |
| DigitalOcean Gradient AI Platform |
Open models (Llama, Ministral and similar) plus resold OpenAI, Anthropic and Meta endpoints; small catalog | DigitalOcean datacenter regions; no per-request region pinning for serverless inference | Limited — no published residency guarantee for serverless inference; resold models inherit the upstream provider's terms | No — usage-based only; dedicated capacity means renting GPU Droplets instead | DigitalOcean model access keys and team roles; no enterprise identity federation |
Open-model serverless inference roughly $0.18-$0.99 per 1M tokens (Ministral 3 14B ~$0.20/M, Llama 3.3 70B ~$0.65/M). Web search $10 per 1,000 requests, web fetch $3 per 1,000 (not charged with Anthropic models). BYOM weights $5/month. Prepaid balance required. |
| NVIDIA NIM / NVIDIA AI Enterprise |
Open-weight models packaged as NIM containers (Llama, Mistral, Nemotron, embedding, speech, vision); no closed frontier models | Wherever you run it — on-prem, any cloud, DGX Cloud; the hosted catalog is US-centric | Strongest in this list for self-hosting: weights and prompts never leave your infrastructure | N/A — you provision GPUs; capacity is whatever you own or rent | None built in — inherits your Kubernetes RBAC, ingress and secrets management; NGC keys for image pull |
NVIDIA AI Enterprise from $4,500 per GPU per year, or about $1 per GPU-hour in the cloud plus the CSP instance cost. build.nvidia.com hosted endpoints are free for prototyping at roughly 40 RPM via the NVIDIA Developer Program; downloadable NIM on up to 16 GPUs for dev/test. |
| OVHcloud AI Endpoints |
40+ open-weight models: Llama 3.3 70B, Mixtral, Mistral, Qwen, code/reasoning models, plus image, embedding and speech | Gravelines, France (EU) — single-country hosting | Strong for EU: French datacentre, EU jurisdiction, customer inputs not used for training; Fast API tier advertises enhanced privacy handling | Partial — Fast API tier with a minimum monthly commitment for guaranteed throughput | OVHcloud Public Cloud IAM and API tokens; limited enterprise identity federation |
Per-token billing from about $0.04 per 1M input tokens on the smallest models to about $0.91 per 1M on the largest. Three modes: BaseAPI (standard per-token), Batch API (discounted, off-peak), Fast API (per-token with a minimum monthly commitment for guaranteed throughput). |
| Scaleway Generative APIs |
Open weights only: GLM 5.2, DeepSeek-V4-Flash, Qwen3.6-35b-a3b, Mistral family, embeddings; context windows 22K-256K | Paris, France (EU); also reachable via Hugging Face Inference Providers | Strong for EU: French company, French datacentres, European data sovereignty positioning | No — serverless pay-per-token only on Generative APIs | Scaleway IAM (projects, applications, API keys); limited enterprise federation |
Per-token billing from around €0.20 per 1M tokens; third-party trackers put the range at roughly $0.12-$2.08 per 1M input tokens depending on model and context window. |
| IONOS AI Model Hub |
Open weights only: Llama 3.1 8B-405B, Mistral variants, gpt-oss-120b, Qwen3 Coder, FLUX.1/FLUX.2 image models, multilingual embeddings | IONOS datacentres in Germany | Processing confined to Germany; customer input excluded from training. Strong single-jurisdiction guarantee | No — token-per-use billing only | IONOS Cloud contract users and API tokens; no enterprise identity federation |
— |
| Red Hat OpenShift AI |
None bundled — you deploy open weights (vLLM-compatible) or your own models; optimised model repo published on Hugging Face | Wherever you run OpenShift: on-prem, edge, AWS/Azure/GCP/IBM Cloud | Complete control — nothing leaves your cluster; the strongest residency story alongside self-hosted NIM | N/A — capacity equals the GPUs you allocate to the cluster | OpenShift RBAC and OAuth integration with enterprise IdPs (LDAP, OIDC, SAML) |
— |
| SAP AI Core / Generative AI Hub |
Curated third-party foundation models (OpenAI, Anthropic, Google and open weights) via the generative AI hub; catalog lags upstream launches | SAP BTP regions across AWS, Azure, GCP and Alibaba Cloud; model availability varies sharply by region | Follows the BTP subaccount region and SAP's data processing agreement; check per-model terms since models are hosted by upstream providers | No published token-rate reservation; capacity governed by the service plan and contract | SAP BTP Identity Authentication / Identity Provisioning, XSUAA roles, subaccount entitlements |
— |