Seminal AI
§6

Cloud AI platforms

Cloud AI platforms are the managed layers that hyperscalers and data platforms put in front of foundation models: one API endpoint, one billing relationship, and the cloud's existing identity, networking and audit controls. They matter less for model access — most frontier models are reachable directly from their labs — and more for the surrounding plumbing: IAM instead of bearer tokens, VPC/PrivateLink instead of the public internet, regional pinning for GDPR and sector regulators, and reserved capacity so a production workload does not sit behind a shared rate limit.

Data checked 2026-09-06

The category churned hard in 2026: Vertex AI was rebranded to the Gemini Enterprise Agent Platform at Cloud Next in April 2026, Azure AI Foundry became Microsoft Foundry, and Databricks folded Mosaic AI into Agent Bricks — the APIs largely survived, the documentation and naming did not. Pricing has also fragmented beyond per-token rates into service tiers (Bedrock Flex/Priority/Reserved), abstract units (Snowflake AI Credits, Databricks DBUs, SAP capacity units, Cloudflare neurons), and reservation instruments (Azure PTU reservations, Google GSUs) that make cross-platform comparison genuinely hard.

A How to choose

Start from where your data and your identity system already live, not from the model catalog — the catalogs have converged, and in 2026 Bedrock, Foundry, Google's Agent Platform and Databricks all serve the same Claude, GPT-5.6/GPT-6, Gemini 3.x and DeepSeek/GLM/Qwen families within weeks of each other. The second axis is residency: if a regulator will ask where inference physically happened, ignore the marketing and check whether the platform offers a regional endpoint, because Bedrock's Global cross-Region inference, Google's global endpoint and Snowflake's cross-region routing all trade residency for capacity, and Snowflake charges a 10% premium ($2.20 vs $2.00 per AI Credit) for pinning to your home region. The third axis is capacity economics: on-demand is right until you have several months of stable volume, and only then do reservations pay — the rough break-evens are ~150-200M tokens/month for Azure PTUs and ~$50/day of single-model spend for Google GSUs, and Bedrock's Reserved tier still requires an account-team conversation rather than a console click.

Fourth, be honest about markup: warehouse-resident platforms (Snowflake Cortex, Databricks) are convenient because the data never leaves the governance boundary, but you pay a platform unit on top of the model, and if your workload is a standalone service rather than SQL over your own tables, that convenience is pure cost. Do not reach for the popular default reflexively: Bedrock is the wrong choice if you need a model AWS has not onboarded or you want fine-tuning of a closed model; Google's Agent Platform is the wrong choice if you need feature parity on the global endpoint, which still lacks tuning, batch prediction and context caching; Microsoft Foundry is the wrong choice for a small EU-only workload now that EU Data Zone deployments carry a 9% uplift and non-US regional deployments 7-16%. For EU sovereignty with real teeth, OVHcloud, Scaleway and IONOS give you single-jurisdiction hosting that no hyperscaler data zone matches — at the cost of open-weight models only.

And if your requirement is "run these weights on our own GPUs with a support contract," this whole category is the wrong shelf: NVIDIA AI Enterprise or Red Hat AI Inference Server is the answer, and you should budget for GPU fleet operations, not tokens.

B At a glance

Name Models availableRegionsData residency / ZDRProvisioned throughputIAM integration Pricing
Amazon Bedrock ~100+ models, ~18 providers: Claude, Nova, Llama, Mistral, Cohere, and more30+ AWS commercial regions plus GovCloud; geographic (US/EU/APAC/JP/AU) and Global cross-Region inference profilesData at rest stays in source region; Global CRIS processes cross-region. Prompts not used for training. ZDR granted per account/model via account teamYes — hourly model units (1-month/6-month terms) plus a Reserved tier priced per 1K tokens-per-minute for 1 or 3 monthsNative AWS IAM, SCPs, KMS CMK, PrivateLink, CloudTrail, Bedrock Guardrails On-demand from $0.035/1M input tokens (Amazon Nova Micro) up to ~$10/1M input for the top Claude tier; Nova Pro ~$0.80 in / $3.20 out per 1M. Batch 50% off, Flex 50% off, Priority +75%. Provisioned Throughput model units run roughly $21-$50/hour depending on model and term.
Gemini Enterprise Agent Platform (formerly Vertex AI) Gemini 3.x family (Pro, Flash, Flash Lite, image models) plus Model Garden third-party and open weights including Claude, Llama, Mistral~40 Google Cloud regions and multi-regions, plus a global endpoint with no regional isolationML processing occurs in the requested region/multi-region; data at rest stays put. Global endpoint waives residency. 24h in-memory cache is on by default and can be disabled for ZDRYes — generative AI scale units (GSUs), 1-month or 1-year commitmentsCloud IAM, VPC Service Controls, CMEK, Org Policy, Cloud Audit Logs Gemini 3 Pro $2.00/1M input and $12.00/1M output at ≤200K context, rising to $4.00/$18.00 above 200K. Provisioned Throughput roughly $42-$158 per GSU-hour depending on model and region (third-party figure), with 20-45% off for 1-month or 1-year commitments.
Microsoft Foundry (formerly Azure AI Foundry) 10,000+ in the Foundry Models catalog: OpenAI GPT-5.x/GPT-6, Anthropic, Meta, DeepSeek, Mistral and open weightsGlobal, Data Zone (US, EU, and new APAC), and Regional deployments across up to 27 Azure regionsData Zone and Regional deployment types constrain processing geography; Global does not. Abuse-monitoring exemption available on request for zero retentionYes — PTUs billed hourly, 15 PTU minimum (Global/Data Zone), with 1-month and 1-year Azure ReservationsMicrosoft Entra ID, Azure RBAC, Private Link, CMK, Azure Policy, unified Foundry control plane GPT-5.6 Sol $5/$30, Terra $2/$12, Luna $0.20/$1.20 per 1M input/output at Global Standard. Batch is 50% off Global Standard. PTUs bill hourly per deployed PTU with a 15 PTU minimum for Global and Data Zone; from 1 Sept 2026 EU Data Zone is +9% over Global, non-US Regional +7-16%, and the new APAC Data Zone +20%.
Cloudflare Workers AI 50+ open-weight models (Llama, Mistral, DeepSeek V4, Kimi, GLM, embeddings, image, ASR); no closed frontier modelsRuns across Cloudflare's global network; no user-selectable inference regionWeak — inference and AI Gateway logs cannot be pinned to the EU; Data Localization Suite is an Enterprise add-on that does not cover hosted inferenceNo — pay-per-use only, no reserved capacity or capacity SLACloudflare account roles and scoped API tokens plus Worker bindings; no federation with cloud IAM systems 10,000 neurons/day free; $0.011 per 1,000 neurons beyond that on Workers Paid. Example: Llama 3.2 1B costs 2,457 neurons per 1M input tokens and 18,252 per 1M output; DeepSeek V4 Pro 120,000 in / 360,000 out per 1M.
IBM watsonx.ai IBM Granite 4.1/4.2 (Apache 2.0) plus Llama, Ministral and other open weights; no closed frontier modelsFoundation-model inference concentrated in Dallas (us-south) and Frankfurt (eu-de); on-prem via Cloud Pak for Data and on IBM Z/LinuxONE with SpyreRegion-pinned by IBM Cloud region; self-managed and Z/LinuxONE deployments give full control. Prompts not used for trainingPartial — dedicated model hosting / on-demand GPU deployment on the Standard plan rather than a token-rate reservationIBM Cloud IAM, resource groups, Key Protect/Hyper Protect Crypto, watsonx.governance integration Free playground: 300,000 tokens/month, 20 CUH, 100 text-extraction documents. Essentials from $0/month pay-as-you-go with ML at $0.55/Capacity Unit-Hour and text extraction at $0.0403/page. Standard from $1,110/month with ML at $0.45/CUH and $0.0318/page. Inference billed in Resource Units of 1,000 tokens.
Oracle OCI Generative AI Cohere Command A family, Llama 4 Maverick/Scout, Llama 3.3 70B, gpt-oss-120b/20b, Gemini 2.5 Pro/Flash, Grok 4.3 and 4.20, Embed 4, Rerank 4; no ClaudeSelected OCI commercial regions per model, plus sovereign and US Classified Cloud regionsRegion-scoped; dedicated AI clusters are single-tenant and reachable only from your tenancy, which is the strongest isolation story hereYes — dedicated AI clusters for hosting and fine-tuning, 744 unit-hour minimum per hosting clusterOCI IAM policies and compartments, Vault/KMS, private endpoints, audit service
Snowflake Cortex AI Anthropic, OpenAI (GPT 5.6 preview), Google Gemini 3.1 Pro, Meta Llama, Mistral, DeepSeek-V4-Flash, GLM-5.3, plus Snowflake Arctic embeddingsSnowflake account regions across AWS, Azure and GCP; many models require cross-region inference to reachHome-region pinning costs $2.20/credit vs $2.00 for global routing; data stays inside the Snowflake governance boundary and is not used for trainingNo — consumption-only; no reserved token capacitySnowflake RBAC, network policies, Tri-Secret Secure; administrators allowlist which models and providers users may call AI Credits at $2.00 each with cross-region routing (ANY_REGION/AWS_GLOBAL) or $2.20 pinned to the home region. Per-model AI Function rates run roughly $0.12 per 1M tokens on small open models to about $5.10 per 1M on frontier models. Cortex Search adds serving compute per GB-month plus per-token embedding cost.
Databricks Mosaic AI / Agent Bricks Very broad: GPT-6 Astra and GPT-5.x, Claude Opus 5/Sonnet 5/Fable 5.1, Gemini 3.x, Llama 4, Qwen 3.5, GLM 5.3, Kimi K3, DeepSeek V4, Grok 4.6, plus embeddingsDatabricks workspace regions on AWS, Azure and GCP; model availability varies by workspace regionInference runs within the workspace's cloud region for Databricks-hosted models; external model routing leaves the boundary. Data governed by Unity CatalogYes — provisioned throughput endpoints are the recommended production mode, billed in DBU/hrUnity Catalog permissions, workspace SSO/SCIM, cloud IAM passthrough, AI Gateway policy and rate limiting Mosaic AI serving from $0.07/DBU; Foundation Model APIs priced in DBUs per 1M tokens (pay-per-token); GPU Model Serving roughly 10.48-628 DBU/hr depending on instance class. Cloud infrastructure is billed separately by AWS/Azure/GCP. External model routing incurs gateway DBUs on top of the upstream provider's own token fees.
Amazon SageMaker AI JumpStart catalog of open weights plus any model you containerise; no Claude, GPT or GeminiNearly all AWS commercial regions plus GovCloudStrong — endpoints run in your chosen region and VPC, nothing leaves unless you route it out; no third-party model provider involvementN/A by name — you provision instances directly; serverless inference bills per requestAWS IAM, VPC isolation, KMS CMK, PrivateLink, CloudTrail, SageMaker Role Manager
Alibaba Cloud Model Studio Qwen3.x Max / Plus / Flash text models, plus Qwen vision, audio, embedding and image models; Qwen-centric catalogInternational endpoint (Singapore) plus Beijing, Tokyo, Frankfurt and Virginia deploymentsEndpoint choice determines jurisdiction; Beijing endpoint is subject to PRC data law. Verify retention terms per region before committingYes — dedicated/provisioned instance options for enterprise contracts, in addition to pay-per-tokenAlibaba Cloud RAM roles and API keys; weaker federation with non-Alibaba identity systems International (Singapore) endpoint: qwen3.8-max $2.00 in / $6.00 out per 1M; qwen3.7-max $2.50/$7.50; qwen3.8-flash $0.15/$0.47; qwen3.7-plus $0.40-$1.20 in / $1.60-$4.80 out (tiered by context). Mainland/Beijing endpoint runs roughly 60-70% cheaper.
DigitalOcean Gradient AI Platform Open models (Llama, Ministral and similar) plus resold OpenAI, Anthropic and Meta endpoints; small catalogDigitalOcean datacenter regions; no per-request region pinning for serverless inferenceLimited — no published residency guarantee for serverless inference; resold models inherit the upstream provider's termsNo — usage-based only; dedicated capacity means renting GPU Droplets insteadDigitalOcean model access keys and team roles; no enterprise identity federation Open-model serverless inference roughly $0.18-$0.99 per 1M tokens (Ministral 3 14B ~$0.20/M, Llama 3.3 70B ~$0.65/M). Web search $10 per 1,000 requests, web fetch $3 per 1,000 (not charged with Anthropic models). BYOM weights $5/month. Prepaid balance required.
NVIDIA NIM / NVIDIA AI Enterprise Open-weight models packaged as NIM containers (Llama, Mistral, Nemotron, embedding, speech, vision); no closed frontier modelsWherever you run it — on-prem, any cloud, DGX Cloud; the hosted catalog is US-centricStrongest in this list for self-hosting: weights and prompts never leave your infrastructureN/A — you provision GPUs; capacity is whatever you own or rentNone built in — inherits your Kubernetes RBAC, ingress and secrets management; NGC keys for image pull NVIDIA AI Enterprise from $4,500 per GPU per year, or about $1 per GPU-hour in the cloud plus the CSP instance cost. build.nvidia.com hosted endpoints are free for prototyping at roughly 40 RPM via the NVIDIA Developer Program; downloadable NIM on up to 16 GPUs for dev/test.
OVHcloud AI Endpoints 40+ open-weight models: Llama 3.3 70B, Mixtral, Mistral, Qwen, code/reasoning models, plus image, embedding and speechGravelines, France (EU) — single-country hostingStrong for EU: French datacentre, EU jurisdiction, customer inputs not used for training; Fast API tier advertises enhanced privacy handlingPartial — Fast API tier with a minimum monthly commitment for guaranteed throughputOVHcloud Public Cloud IAM and API tokens; limited enterprise identity federation Per-token billing from about $0.04 per 1M input tokens on the smallest models to about $0.91 per 1M on the largest. Three modes: BaseAPI (standard per-token), Batch API (discounted, off-peak), Fast API (per-token with a minimum monthly commitment for guaranteed throughput).
Scaleway Generative APIs Open weights only: GLM 5.2, DeepSeek-V4-Flash, Qwen3.6-35b-a3b, Mistral family, embeddings; context windows 22K-256KParis, France (EU); also reachable via Hugging Face Inference ProvidersStrong for EU: French company, French datacentres, European data sovereignty positioningNo — serverless pay-per-token only on Generative APIsScaleway IAM (projects, applications, API keys); limited enterprise federation Per-token billing from around €0.20 per 1M tokens; third-party trackers put the range at roughly $0.12-$2.08 per 1M input tokens depending on model and context window.
IONOS AI Model Hub Open weights only: Llama 3.1 8B-405B, Mistral variants, gpt-oss-120b, Qwen3 Coder, FLUX.1/FLUX.2 image models, multilingual embeddingsIONOS datacentres in GermanyProcessing confined to Germany; customer input excluded from training. Strong single-jurisdiction guaranteeNo — token-per-use billing onlyIONOS Cloud contract users and API tokens; no enterprise identity federation
Red Hat OpenShift AI None bundled — you deploy open weights (vLLM-compatible) or your own models; optimised model repo published on Hugging FaceWherever you run OpenShift: on-prem, edge, AWS/Azure/GCP/IBM CloudComplete control — nothing leaves your cluster; the strongest residency story alongside self-hosted NIMN/A — capacity equals the GPUs you allocate to the clusterOpenShift RBAC and OAuth integration with enterprise IdPs (LDAP, OIDC, SAML)
SAP AI Core / Generative AI Hub Curated third-party foundation models (OpenAI, Anthropic, Google and open weights) via the generative AI hub; catalog lags upstream launchesSAP BTP regions across AWS, Azure, GCP and Alibaba Cloud; model availability varies sharply by regionFollows the BTP subaccount region and SAP's data processing agreement; check per-model terms since models are hosted by upstream providersNo published token-rate reservation; capacity governed by the service plan and contractSAP BTP Identity Authentication / Identity Provisioning, XSUAA roles, subaccount entitlements

C Entries

Amazon Bedrock

Bedrock exposes ~100+ models from roughly 18 providers (Anthropic Claude, Amazon Nova, Meta Llama, Mistral, Cohere and others) behind one IAM-authenticated API, with no infrastructure to run. Beyond plain on-demand it now sells four tiers: Batch at 50% off, Flex at a 50% discount for latency-tolerant traffic, Priority at a 75% premium for latency-sensitive traffic, and a Reserved tier (launched November 2025, extended to Claude Opus 4.5 and Haiku 4.5 in January 2026) where you buy input and output tokens-per-minute separately for 1 or 3 months and overflow to Standard. Provisioned Throughput sold in hourly model units still exists but is now mostly for custom and fine-tuned models. AgentCore adds a managed agent runtime, registry and orchestration layer on top, billed separately.

Models available~100+ models, ~18 providers: Claude, Nova, Llama, Mistral, Cohere, and more
Regions30+ AWS commercial regions plus GovCloud; geographic (US/EU/APAC/JP/AU) and Global cross-Region inference profiles
Data residency / ZDRData at rest stays in source region; Global CRIS processes cross-region. Prompts not used for training. ZDR granted per account/model via account team
Provisioned throughputYes — hourly model units (1-month/6-month terms) plus a Reserved tier priced per 1K tokens-per-minute for 1 or 3 months
IAM integrationNative AWS IAM, SCPs, KMS CMK, PrivateLink, CloudTrail, Bedrock Guardrails
Fine-tuningYes for Nova, Llama and select open models; not available for Claude

Watch out: Per-model, per-region access still has to be requested and quotas are granted per region, so a working dev account does not guarantee production capacity. Token prices for third-party models generally track the vendor's own list price, so you get no discount over calling Anthropic or Mistral directly — you are paying in convenience, not dollars. Global cross-Region inference gives you higher quotas but processes requests in whatever commercial region has capacity, which is a compliance conversation even though data at rest stays in your source region. Reserved tier access requires an AWS account team, not a console button, and zero-data-retention is granted case-by-case per account and per model rather than being a default. Closed models like Claude cannot be fine-tuned here.

On-demand from $0.035/1M input tokens (Amazon Nova Micro) up to ~$10/1M input for the top Claude tier; Nova Pro ~$0.80 in / $3.20 out per 1M. Batch 50% off, Flex 50% off, Priority +75%. Provisioned Throughput model units run roughly $21-$50/hour depending on model and term.

Gemini Enterprise Agent Platform (formerly Vertex AI)

Google renamed Vertex AI to the Gemini Enterprise Agent Platform at Cloud Next in April 2026, with migration completing in late May; API endpoints, SDKs and the aiplatform namespace were left intact, so existing integrations kept working while the documentation moved. It serves first-party Gemini 3.x models alongside a Model Garden of third-party and open weights (Claude, Llama, Mistral and others), plus training, tuning, batch prediction and endpoints as sub-features of an agent-first platform. Capacity is bought either on-demand or as Provisioned Throughput measured in generative AI scale units (GSUs) with 1-month or 1-year commitments. Regional endpoints keep ML processing in the chosen geography; the global endpoint improves availability and 429 rates but explicitly gives up residency.

Models availableGemini 3.x family (Pro, Flash, Flash Lite, image models) plus Model Garden third-party and open weights including Claude, Llama, Mistral
Regions~40 Google Cloud regions and multi-regions, plus a global endpoint with no regional isolation
Data residency / ZDRML processing occurs in the requested region/multi-region; data at rest stays put. Global endpoint waives residency. 24h in-memory cache is on by default and can be disabled for ZDR
Provisioned throughputYes — generative AI scale units (GSUs), 1-month or 1-year commitments
IAM integrationCloud IAM, VPC Service Controls, CMEK, Org Policy, Cloud Audit Logs
Fine-tuningYes — supervised tuning and distillation for Gemini and selected Model Garden models (regional endpoints only)

Watch out: The April/May 2026 rebrand left a genuinely confusing documentation estate — search results, blog posts and third-party guides still say 'Vertex AI', and some URLs redirect while others do not. The global endpoint is not feature-equivalent: tuning, batch prediction and context caching are regional-only, so the availability win costs you capabilities as well as residency. In-memory caching retains prompt data at project level for up to 24 hours by default, and you must explicitly disable it via API to get zero retention. Per-region, per-model quota is the usual operational bottleneck, and GSU pricing varies by model and region in a way that is hard to model before you commit. Model Garden third-party models bill separately and reach fewer regions than Gemini itself.

Gemini 3 Pro $2.00/1M input and $12.00/1M output at ≤200K context, rising to $4.00/$18.00 above 200K. Provisioned Throughput roughly $42-$158 per GSU-hour depending on model and region (third-party figure), with 20-45% off for 1-month or 1-year commitments.

Microsoft Foundry (formerly Azure AI Foundry)

Rebranded from Azure AI Studio to Azure AI Foundry and now to Microsoft Foundry, this is a single Azure resource provider that unifies models, agents and tools with one RBAC, networking and policy surface. It lists more than 10,000 models (Microsoft, OpenAI, Anthropic, Meta and others), splits deployments into Global, Data Zone and Regional types, and offers prompt agents (declarative) and hosted agents (bring your own container). Capacity is bought per token, or as provisioned throughput units billed hourly per PTU regardless of tokens consumed, with Azure Reservations for 1-month or 1-year terms. The Assistants API has been superseded by the Responses API and monthly api-version params by stable /openai/v1/ routes.

Models available10,000+ in the Foundry Models catalog: OpenAI GPT-5.x/GPT-6, Anthropic, Meta, DeepSeek, Mistral and open weights
RegionsGlobal, Data Zone (US, EU, and new APAC), and Regional deployments across up to 27 Azure regions
Data residency / ZDRData Zone and Regional deployment types constrain processing geography; Global does not. Abuse-monitoring exemption available on request for zero retention
Provisioned throughputYes — PTUs billed hourly, 15 PTU minimum (Global/Data Zone), with 1-month and 1-year Azure Reservations
IAM integrationMicrosoft Entra ID, Azure RBAC, Private Link, CMK, Azure Policy, unified Foundry control plane
Fine-tuningYes — supervised fine-tuning, DPO and distillation for selected OpenAI and open models

Watch out: Two years of renames left a split estate — Foundry (classic) portal with hub-based projects versus the new portal with Foundry projects — and migration guidance is now a required read rather than a nice-to-have. A reservation is a financial discount, not a capacity guarantee: you must deploy first and reserve second, scaling a deployment down releases capacity permanently with no promise you can get it back, and deployments cannot be paused, only deleted. PTU quota is fragmented by deployment type, so Global reservations do not cover Regional deployments. The September 2026 uplifts make non-US deployments materially more expensive than Global, which penalises exactly the residency-conscious workloads that need them.

GPT-5.6 Sol $5/$30, Terra $2/$12, Luna $0.20/$1.20 per 1M input/output at Global Standard. Batch is 50% off Global Standard. PTUs bill hourly per deployed PTU with a 15 PTU minimum for Global and Data Zone; from 1 Sept 2026 EU Data Zone is +9% over Global, non-US Regional +7-16%, and the new APAC Data Zone +20%.

Cloudflare Workers AI

Workers AI runs a catalog of open-weight models on serverless GPUs across Cloudflare's network, callable from a Worker binding or the REST API with no capacity planning. Billing uses 'neurons', an abstracted GPU-work unit: 10,000 per day free, then $0.011 per 1,000, with per-model neuron costs published per million tokens (Llama 3.2 1B is 2,457 neurons per 1M input tokens; larger frontier open models such as DeepSeek V4 Pro run 120,000 in / 360,000 out). It composes tightly with AI Gateway (caching, retries, fallback), Vectorize, D1, R2 and Durable Objects, which is the real reason to pick it. Some frontier open models require a Workers Paid plan or prepaid AI Gateway credits.

Models available50+ open-weight models (Llama, Mistral, DeepSeek V4, Kimi, GLM, embeddings, image, ASR); no closed frontier models
RegionsRuns across Cloudflare's global network; no user-selectable inference region
Data residency / ZDRWeak — inference and AI Gateway logs cannot be pinned to the EU; Data Localization Suite is an Enterprise add-on that does not cover hosted inference
Provisioned throughputNo — pay-per-use only, no reserved capacity or capacity SLA
IAM integrationCloudflare account roles and scoped API tokens plus Worker bindings; no federation with cloud IAM systems
Fine-tuningLimited — LoRA adapters on selected models only

Watch out: There is no data residency story for hosted-model inference: you cannot pin inference or AI Gateway logs to the EU, and the Data Localization Suite is an Enterprise add-on that does not cover the inference plane, so this is the wrong platform for a workload a European regulator will inspect. The catalog is open-weight only and much smaller than the hyperscalers — no Claude, GPT or Gemini served natively. There is no provisioned throughput or capacity SLA, so you get no protection against noisy-neighbour latency for production traffic. Neuron pricing is hard to forecast without benchmarking your actual prompt shapes, and Regional Services pinning applies only to the Worker's custom hostname, not to subrequests, Queues or Cron triggers.

10,000 neurons/day free; $0.011 per 1,000 neurons beyond that on Workers Paid. Example: Llama 3.2 1B costs 2,457 neurons per 1M input tokens and 18,252 per 1M output; DeepSeek V4 Pro 120,000 in / 360,000 out per 1M.

IBM watsonx.ai

watsonx.ai pairs IBM's own Apache-2.0-licensed Granite models (4.1 released April 2026; 4.2 in 3B/8B/30B released August 2026) with third-party open models such as Llama and Ministral, wrapped in Prompt Lab, guardrails, AutoAI, notebooks and document understanding. Billing runs two meters simultaneously: Resource Units for inference, where 1 RU equals 1,000 tokens counting input and output together, and Capacity Unit Hours for machine-learning compute. Version 2.4 (June 2026) expanded model choice, synthetic data generation and runtime support, and the platform now also runs on IBM Z and LinuxONE with Spyre accelerators, plus on-prem via Cloud Pak for Data. Its distinguishing pitch is governance integration with watsonx.governance rather than model breadth.

Models availableIBM Granite 4.1/4.2 (Apache 2.0) plus Llama, Ministral and other open weights; no closed frontier models
RegionsFoundation-model inference concentrated in Dallas (us-south) and Frankfurt (eu-de); on-prem via Cloud Pak for Data and on IBM Z/LinuxONE with Spyre
Data residency / ZDRRegion-pinned by IBM Cloud region; self-managed and Z/LinuxONE deployments give full control. Prompts not used for training
Provisioned throughputPartial — dedicated model hosting / on-demand GPU deployment on the Standard plan rather than a token-rate reservation
IAM integrationIBM Cloud IAM, resource groups, Key Protect/Hyper Protect Crypto, watsonx.governance integration
Fine-tuningYes — LoRA fine-tuning on the Standard plan (GPU-dependent pricing); prompt tuning on lower tiers

Watch out: Foundation-model inference and Prompt Lab are effectively limited to Dallas and Frankfurt, which is a hard stop if you need APAC or in-country processing. The catalog has no closed frontier models — no GPT-5.x, Claude or Gemini — so if your evaluation says you need one of those, watsonx is out before you start. The dual RU-plus-CUH metering is genuinely confusing to forecast, and RUs counting input and output at the same rate penalises output-heavy workloads relative to per-token pricing elsewhere. The Standard plan's $1,110/month floor makes it a poor fit for small or bursty production loads, and the surrounding portal (Studio, Machine Learning, Governance as separate provisioned services) carries more setup overhead than a single API key.

Free playground: 300,000 tokens/month, 20 CUH, 100 text-extraction documents. Essentials from $0/month pay-as-you-go with ML at $0.55/Capacity Unit-Hour and text extraction at $0.0403/page. Standard from $1,110/month with ML at $0.45/CUH and $0.0318/page. Inference billed in Resource Units of 1,000 tokens.

Oracle OCI Generative AI

OCI Generative AI serves a curated set of pretrained models — Cohere Command A / A Reasoning / A Vision, Meta Llama 4 Maverick and Scout, Llama 3.3 70B, OpenAI gpt-oss-120b/20b, Google Gemini 2.5 Pro/Flash/Flash-Lite, xAI Grok 4.3 and Grok 4.20, plus Cohere Embed 4 and Rerank 4 — through a REST API in your tenancy. It bills on-demand inference by character (one transaction equals one character) rather than by token, which is unusual and forces you to redo your cost model. Dedicated AI clusters give single-tenant hosting and fine-tuning capacity, with a minimum commitment of 744 unit-hours per hosting cluster (about one month) and 1 unit-hour per fine-tuning job; imported models are exempt from the 744-hour floor. It is also available in Oracle's US Classified Cloud regions, which very few competitors match.

Models availableCohere Command A family, Llama 4 Maverick/Scout, Llama 3.3 70B, gpt-oss-120b/20b, Gemini 2.5 Pro/Flash, Grok 4.3 and 4.20, Embed 4, Rerank 4; no Claude
RegionsSelected OCI commercial regions per model, plus sovereign and US Classified Cloud regions
Data residency / ZDRRegion-scoped; dedicated AI clusters are single-tenant and reachable only from your tenancy, which is the strongest isolation story here
Provisioned throughputYes — dedicated AI clusters for hosting and fine-tuning, 744 unit-hour minimum per hosting cluster
IAM integrationOCI IAM policies and compartments, Vault/KMS, private endpoints, audit service
Fine-tuningYes — on select models (e.g. Llama 3.3 70B, Cohere) using fine-tuning clusters, 1 unit-hour minimum

Watch out: Character-based on-demand billing is a real friction: every token-denominated benchmark, budget and vendor comparison has to be converted before it means anything, and the exact per-character rates live only on the OCI price list rather than the docs. The catalog churns aggressively — Cohere Command R and R+, the Embed 3 family, Rerank 3.5 and Llama 3.2 Vision are deprecated, and Command R/R+ 16K and Llama 3.1 70B are already retired — so expect forced migrations on a roughly annual cadence. There is no Anthropic Claude, which rules it out for a large share of 2026 agent workloads. The 744 unit-hour minimum on hosting clusters is a month-long commitment in all but name, and the surrounding tooling (agent frameworks, evaluation, observability) is thinner than the big three.

usage-based

Snowflake Cortex AI

Cortex AI puts models behind SQL and REST surfaces — AI_COMPLETE and the AISQL function family, Cortex Search, Cortex Agents, Cortex Code, AI Parse Doc — so inference runs inside Snowflake's governance boundary and your table data never leaves it. Since 1 April 2026 it bills in a separate AI Credit currency priced at $2.00 with global cross-region routing enabled, or $2.20 if you pin requests to your home region for residency; Standard and Enterprise editions now pay the same inference rate. The catalog spans Anthropic, OpenAI (GPT 5.6 in private preview), Google Gemini 3.1 Pro, Meta, Mistral and open models including DeepSeek-V4-Flash and GLM-5.3, and a Cortex AI Gateway can now route dynamically between models on quality/cost.

Models availableAnthropic, OpenAI (GPT 5.6 preview), Google Gemini 3.1 Pro, Meta Llama, Mistral, DeepSeek-V4-Flash, GLM-5.3, plus Snowflake Arctic embeddings
RegionsSnowflake account regions across AWS, Azure and GCP; many models require cross-region inference to reach
Data residency / ZDRHome-region pinning costs $2.20/credit vs $2.00 for global routing; data stays inside the Snowflake governance boundary and is not used for training
Provisioned throughputNo — consumption-only; no reserved token capacity
IAM integrationSnowflake RBAC, network policies, Tri-Secret Secure; administrators allowlist which models and providers users may call
Fine-tuningYes — Cortex Fine-tuning on selected open models, billed in credits

Watch out: This is only a good deal if the data is already in Snowflake; as a general-purpose model API you are paying a platform margin for models you could call directly. Most of the best models require enabling cross-region inference, which is the cheaper credit rate precisely because it abandons residency — and if your account is in Europe while the model runs in AWS US, you also pick up data transfer charges that do not appear in the credit price. The AI Credit abstraction makes true $/1M-token comparison against Bedrock or Foundry a spreadsheet exercise rather than a glance. There is no provisioned throughput or dedicated capacity, and model availability generally lags the upstream provider by weeks, with headline models often arriving in private or public preview first.

AI Credits at $2.00 each with cross-region routing (ANY_REGION/AWS_GLOBAL) or $2.20 pinned to the home region. Per-model AI Function rates run roughly $0.12 per 1M tokens on small open models to about $5.10 per 1M on frontier models. Cortex Search adds serving compute per GB-month plus per-token embedding cost.

Databricks Mosaic AI / Agent Bricks

Databricks' AI stack — Mosaic AI model serving, Foundation Model APIs, AI Gateway and the Agent Bricks agent-building surface introduced at Data + AI Summit — serves an unusually broad hosted catalog directly against Unity Catalog data: OpenAI GPT-6 Astra and the GPT-5.x line, Anthropic Claude Opus 5 / Sonnet 5 / Fable 5.1 / Haiku 4.5, Google Gemini 3.x, Meta Llama 4, Qwen 3.5, GLM 5.3, Kimi K3, DeepSeek V4, xAI Grok 4.6 and open embedding models. Consumption normalises into DBUs rather than tokens: Foundation Model APIs are priced in DBUs per 1M tokens for pay-per-token, GPU model serving runs roughly 10.48-628 DBU/hr, and provisioned throughput endpoints are the recommended production path. Governance, lineage and evaluation come from Unity Catalog and MLflow rather than a bolt-on.

Models availableVery broad: GPT-6 Astra and GPT-5.x, Claude Opus 5/Sonnet 5/Fable 5.1, Gemini 3.x, Llama 4, Qwen 3.5, GLM 5.3, Kimi K3, DeepSeek V4, Grok 4.6, plus embeddings
RegionsDatabricks workspace regions on AWS, Azure and GCP; model availability varies by workspace region
Data residency / ZDRInference runs within the workspace's cloud region for Databricks-hosted models; external model routing leaves the boundary. Data governed by Unity Catalog
Provisioned throughputYes — provisioned throughput endpoints are the recommended production mode, billed in DBU/hr
IAM integrationUnity Catalog permissions, workspace SSO/SCIM, cloud IAM passthrough, AI Gateway policy and rate limiting
Fine-tuningYes — fine-tuning and continued pretraining on open models, served through provisioned throughput

Watch out: DBU pricing is a genuine obstacle to cost comparison — you cannot read a $/1M-token number off the page, and you also pay the underlying cloud's compute bill on top. Brand churn is real: Mosaic AI has been folded into the Agent Bricks framing, so half the documentation and most blog posts you find use retired names. Models get retired on a published policy and deprecations land regularly (Gemini 2.5 Pro/Flash retire 2 October 2026; Claude Sonnet 4 is deprecated), so pinned model IDs need an owner. Routing to external providers means you pay their token fees directly plus Databricks gateway DBUs — a markup for the privilege of one control plane. If your data is not in the lakehouse, none of this is worth the platform tax.

Mosaic AI serving from $0.07/DBU; Foundation Model APIs priced in DBUs per 1M tokens (pay-per-token); GPU Model Serving roughly 10.48-628 DBU/hr depending on instance class. Cloud infrastructure is billed separately by AWS/Azure/GCP. External model routing incurs gateway DBUs on top of the upstream provider's own token fees.

Amazon SageMaker AI

SageMaker AI is the other half of AWS's AI story and the right answer whenever Bedrock's serverless abstraction is the problem rather than the solution. You pick the container, the instance type and the serving stack, deploy open weights from JumpStart or your own artifacts, and pay per instance-hour for real-time endpoints or per request for serverless inference. The catalog reachable through JumpStart is wider than Bedrock's for open models, but the closed frontier models Bedrock resells — Claude foremost — are not available here. It integrates with Bedrock in both directions: Bedrock agents can call SageMaker endpoints as tools, and the two now share a development environment.

Models availableJumpStart catalog of open weights plus any model you containerise; no Claude, GPT or Gemini
RegionsNearly all AWS commercial regions plus GovCloud
Data residency / ZDRStrong — endpoints run in your chosen region and VPC, nothing leaves unless you route it out; no third-party model provider involvement
Provisioned throughputN/A by name — you provision instances directly; serverless inference bills per request
IAM integrationAWS IAM, VPC isolation, KMS CMK, PrivateLink, CloudTrail, SageMaker Role Manager
Fine-tuningYes — full control including full fine-tuning, LoRA, continued pretraining and distributed training

Watch out: You own the operational burden: instance sizing, autoscaling, cold starts, container builds and upgrades are all yours, and an idle GPU endpoint bills 24/7 whether or not it serves a request. No Claude and no other closed frontier models, so a mixed strategy usually means running both this and Bedrock. Economics only favour it at scale — the common rule of thumb puts the crossover at roughly 200M tokens/day of steady traffic, below which Bedrock's per-token pricing usually wins. Serverless inference reduces the idle cost but brings cold-start latency and concurrency ceilings. Budgeting is also harder to state up front because the bill is instance-hours and storage rather than a published per-token rate, so no single price_summary figure is meaningful.

usage-based

Alibaba Cloud Model Studio

Model Studio (Bailian) is Alibaba's managed platform for the Qwen family plus assorted third-party and open models, exposed through an OpenAI-compatible API with separate Chinese Mainland and International endpoints. Current international pricing puts qwen3.8-max at $2.00 in / $6.00 out per 1M tokens, qwen3.8-flash at $0.15 / $0.47, and the plus tier on tiered context-length pricing from $0.40 in. New accounts get a 1M-token free quota per eligible model valid 90 days — but only on the Singapore endpoint. It is the cheapest credible route to frontier-adjacent quality at the flash tier, and the natural choice if you are serving users in Southeast Asia.

Models availableQwen3.x Max / Plus / Flash text models, plus Qwen vision, audio, embedding and image models; Qwen-centric catalog
RegionsInternational endpoint (Singapore) plus Beijing, Tokyo, Frankfurt and Virginia deployments
Data residency / ZDREndpoint choice determines jurisdiction; Beijing endpoint is subject to PRC data law. Verify retention terms per region before committing
Provisioned throughputYes — dedicated/provisioned instance options for enterprise contracts, in addition to pay-per-token
IAM integrationAlibaba Cloud RAM roles and API keys; weaker federation with non-Alibaba identity systems
Fine-tuningYes — SFT and LoRA on selected Qwen models

Watch out: Jurisdiction is the deciding factor for most Western buyers: even on the Singapore endpoint, procurement, legal and sector regulators frequently rule out a PRC-headquartered provider, and the Mainland endpoint's much lower prices come with PRC data law attached. The free quota exists only in Singapore, and the older free OAuth API tier was discontinued on 15 April 2026, so trial paths have narrowed. The catalog is Qwen-centric — you will not find Claude or GPT here — and enterprise controls (federated IAM, private connectivity, audit) are less mature than the big three. English documentation and console quality are noticeably behind the Chinese originals.

International (Singapore) endpoint: qwen3.8-max $2.00 in / $6.00 out per 1M; qwen3.7-max $2.50/$7.50; qwen3.8-flash $0.15/$0.47; qwen3.7-plus $0.40-$1.20 in / $1.60-$4.80 out (tiered by context). Mainland/Beijing endpoint runs roughly 60-70% cheaper.

DigitalOcean Gradient AI Platform

Gradient AI (formerly the GenAI Platform) gives DigitalOcean customers a model access key and an OpenAI-shaped endpoint covering open models plus resold OpenAI, Anthropic and Meta models, with no infrastructure to provision. Published open-model rates run roughly $0.18-$0.99 per 1M tokens — Ministral 3 14B around $0.20/M, Llama 3.3 70B around $0.65/M — and agent tooling is metered separately at $10 per 1,000 web-search requests and $3 per 1,000 web-fetch requests, with bring-your-own-model weights billed at $5/month. Serverless inference is prepaid only: a positive account balance is required before requests are served. GPU Droplets sit alongside for teams that want raw H100/H200 capacity instead.

Models availableOpen models (Llama, Ministral and similar) plus resold OpenAI, Anthropic and Meta endpoints; small catalog
RegionsDigitalOcean datacenter regions; no per-request region pinning for serverless inference
Data residency / ZDRLimited — no published residency guarantee for serverless inference; resold models inherit the upstream provider's terms
Provisioned throughputNo — usage-based only; dedicated capacity means renting GPU Droplets instead
IAM integrationDigitalOcean model access keys and team roles; no enterprise identity federation
Fine-tuningLimited — bring-your-own-model weights ($5/month) rather than a managed tuning service

Watch out: Prepaid-only billing means a lapsed balance is an outage, which is a poor fit for production without monitoring on the balance itself. The catalog is small and the resold closed models carry no price advantage over calling the provider directly. There is no provisioned throughput, no meaningful data residency commitment, and IAM is DigitalOcean tokens rather than a federated enterprise identity system — so regulated buyers will not clear it. Compliance attestations and regional coverage are well behind AWS, Azure and Google, and the platform has been renamed once already, which is worth noting when you pin documentation links.

Open-model serverless inference roughly $0.18-$0.99 per 1M tokens (Ministral 3 14B ~$0.20/M, Llama 3.3 70B ~$0.65/M). Web search $10 per 1,000 requests, web fetch $3 per 1,000 (not charged with Anthropic models). BYOM weights $5/month. Prepaid balance required.

NVIDIA NIM / NVIDIA AI Enterprise

NIM packages models as optimised inference containers with an OpenAI-compatible API, so the same artifact runs on a workstation, in your Kubernetes cluster, or on DGX Cloud. The hosted catalog at build.nvidia.com is free for prototyping through the NVIDIA Developer Program at roughly 40 requests per minute, and Developer Program members can also run downloadable NIMs on up to 16 GPUs for development and testing. Production use requires an NVIDIA AI Enterprise licence, listed from $4,500 per GPU per year or about $1 per GPU-hour in the cloud on top of the instance cost, with multi-year, perpetual and 75% education/Inception discounts documented as of June 2026. This is a serving layer with support, not a multi-tenant token API.

Models availableOpen-weight models packaged as NIM containers (Llama, Mistral, Nemotron, embedding, speech, vision); no closed frontier models
RegionsWherever you run it — on-prem, any cloud, DGX Cloud; the hosted catalog is US-centric
Data residency / ZDRStrongest in this list for self-hosting: weights and prompts never leave your infrastructure
Provisioned throughputN/A — you provision GPUs; capacity is whatever you own or rent
IAM integrationNone built in — inherits your Kubernetes RBAC, ingress and secrets management; NGC keys for image pull
Fine-tuningYes via NeMo/NeMo Customizer; NIM serves the resulting weights including LoRA adapters

Watch out: You still buy or rent every GPU, so the licence is additive to a much larger hardware bill, and per-GPU pricing punishes fleets that sit idle. There is no per-token billing and no elastic multi-tenant capacity, so this is not comparable to Bedrock or Foundry on cost-per-request without doing your own utilisation maths. No closed frontier models: GPT, Claude and Gemini are not available as NIMs. The honest alternative for many teams is self-hosted vLLM at zero licence cost, and the case for NIM rests on support, hardened builds and TensorRT-LLM optimisations rather than capability. The hosted catalog's rate limits make it a prototyping surface only.

NVIDIA AI Enterprise from $4,500 per GPU per year, or about $1 per GPU-hour in the cloud plus the CSP instance cost. build.nvidia.com hosted endpoints are free for prototyping at roughly 40 RPM via the NVIDIA Developer Program; downloadable NIM on up to 16 GPUs for dev/test.

OVHcloud AI Endpoints

AI Endpoints is OVHcloud's serverless, OpenAI-compatible API over 40+ open-weight models — Llama 3.3 70B, Mixtral, Mistral, Qwen, code and reasoning models, plus image, embedding and speech endpoints — served from the Gravelines datacentre in France under French/EU jurisdiction with no training on customer inputs. Billing is per token across three modes: BaseAPI for general traffic, a discounted Batch API for off-peak processing, and a Fast API tier with a minimum monthly commitment that buys guaranteed throughput and enhanced privacy handling. Published input rates start around $0.04 per 1M tokens and top out near $0.91 per 1M for the largest models, which is competitive with US serverless providers rather than a sovereignty premium.

Models available40+ open-weight models: Llama 3.3 70B, Mixtral, Mistral, Qwen, code/reasoning models, plus image, embedding and speech
RegionsGravelines, France (EU) — single-country hosting
Data residency / ZDRStrong for EU: French datacentre, EU jurisdiction, customer inputs not used for training; Fast API tier advertises enhanced privacy handling
Provisioned throughputPartial — Fast API tier with a minimum monthly commitment for guaranteed throughput
IAM integrationOVHcloud Public Cloud IAM and API tokens; limited enterprise identity federation
Fine-tuningNot on AI Endpoints — use OVHcloud AI Training / AI Deploy with your own GPUs

Watch out: Open-weight models only — there is no Claude, GPT or Gemini, so if your evaluation depends on a frontier closed model this platform cannot serve it at any price. Single-datacentre hosting means poor latency for users outside Europe and a narrower blast radius for availability than a multi-region hyperscaler. Throughput ceilings are lower and the Fast API tier's monthly commitment is the only guaranteed-capacity option. The surrounding ecosystem is thin: no managed agent runtime, no equivalent of Bedrock Guardrails or Unity Catalog governance, and IAM is OVHcloud's own rather than something you can federate deeply. Best treated as an inference endpoint, not a platform.

Per-token billing from about $0.04 per 1M input tokens on the smallest models to about $0.91 per 1M on the largest. Three modes: BaseAPI (standard per-token), Batch API (discounted, off-peak), Fast API (per-token with a minimum monthly commitment for guaranteed throughput).

Scaleway Generative APIs

Scaleway's Generative APIs serve open-weight models from Paris datacentres through an OpenAI-compatible endpoint, and Scaleway is also available as a Hugging Face Inference Provider. The 2026 catalog tracks the open frontier closely: GLM 5.2 arrived as one of the strongest open-weight agentic/coding models, Qwen3.6-35b-a3b landed in April 2026 as a small agentic model, and DeepSeek-V4-Flash in July 2026 as the cost-efficient long-horizon option. Pricing is per token, published from around €0.20 per 1M with third-party trackers putting the range at roughly $0.12-$2.08 per 1M input tokens across context windows from 22K to 256K. For a French company with French datacentres, the sovereignty claim is about as clean as it gets in this category.

Models availableOpen weights only: GLM 5.2, DeepSeek-V4-Flash, Qwen3.6-35b-a3b, Mistral family, embeddings; context windows 22K-256K
RegionsParis, France (EU); also reachable via Hugging Face Inference Providers
Data residency / ZDRStrong for EU: French company, French datacentres, European data sovereignty positioning
Provisioned throughputNo — serverless pay-per-token only on Generative APIs
IAM integrationScaleway IAM (projects, applications, API keys); limited enterprise federation
Fine-tuningNot offered on Generative APIs — deploy tuned weights on Scaleway GPU instances instead

Watch out: Model turnover is aggressive and disruptive: Devstral 2, Voxtral Small, Gemma 3, Pixtral and Qwen 3 Coder were deprecated and taken to end-of-life on 1 July 2026, so pinned model IDs need an owner and a migration plan on roughly a two-quarter cadence. Open-weight only — no Claude, GPT or Gemini. Capacity and rate limits are well below hyperscaler scale, and there is no provisioned-throughput product to buy your way past them. Region footprint is essentially Paris, so non-EU latency is poor. Enterprise controls (federated identity, private connectivity, audit depth) are limited compared with the big three, and euro-denominated pricing complicates like-for-like comparison with dollar-quoted competitors.

Per-token billing from around €0.20 per 1M tokens; third-party trackers put the range at roughly $0.12-$2.08 per 1M input tokens depending on model and context window.

IONOS AI Model Hub

IONOS AI Model Hub serves open-weight text, image and embedding models through an OpenAI-compatible REST API from IONOS datacentres in Germany, with processing confined to Germany and customer input explicitly excluded from training. The catalog covers Llama 3.1 from 8B to 405B, Mistral variants, gpt-oss-120b, Qwen3 Coder, FLUX.1 and FLUX.2 for image generation, and multilingual embedding models, priced per token (per image for FLUX). It is the narrowest offering in this chapter but also the most straightforward German-jurisdiction answer for teams whose compliance requirement is literally 'processed in Germany'.

Models availableOpen weights only: Llama 3.1 8B-405B, Mistral variants, gpt-oss-120b, Qwen3 Coder, FLUX.1/FLUX.2 image models, multilingual embeddings
RegionsIONOS datacentres in Germany
Data residency / ZDRProcessing confined to Germany; customer input excluded from training. Strong single-jurisdiction guarantee
Provisioned throughputNo — token-per-use billing only
IAM integrationIONOS Cloud contract users and API tokens; no enterprise identity federation
Fine-tuningNo managed fine-tuning service

Watch out: Pricing and the current model list could not be confirmed against an official pricing page during this review — third-party guides quote roughly $0.17 per 1M tokens for Llama 3.1 8B, but treat that as indicative only; confidence for this entry is deliberately low. The catalog is the smallest in this chapter and holds no closed frontier models. Germany-only hosting gives you jurisdiction but also a single point of latency and availability. There is no provisioned throughput, no managed agent or RAG tooling worth the name, and no managed fine-tuning, so you are buying an inference endpoint plus a legal posture. Documentation and support are oriented to the German market and thinner in English.

usage-based

Red Hat OpenShift AI

OpenShift AI is a platform for building, tuning and serving models on your own Kubernetes, not a model API. Its serving core is Red Hat AI Inference Server, a hardened supported distribution of vLLM shipped with LLM compression tooling and an optimised model repository on Hugging Face; it is included in OpenShift AI and RHEL AI and also sold standalone. Subscriptions follow Red Hat's usual shape — AI Inference is priced per accelerator, while OpenShift AI mirrors OpenShift units per vCPU or per bare-metal node and requires the underlying OpenShift entitlement. A 60-day self-supported subscription is available from the customer portal for evaluation. It is the mainstream choice when portability across on-prem, edge and multiple clouds is the requirement.

Models availableNone bundled — you deploy open weights (vLLM-compatible) or your own models; optimised model repo published on Hugging Face
RegionsWherever you run OpenShift: on-prem, edge, AWS/Azure/GCP/IBM Cloud
Data residency / ZDRComplete control — nothing leaves your cluster; the strongest residency story alongside self-hosted NIM
Provisioned throughputN/A — capacity equals the GPUs you allocate to the cluster
IAM integrationOpenShift RBAC and OAuth integration with enterprise IdPs (LDAP, OIDC, SAML)
Fine-tuningYes — training, tuning and LLM compression pipelines via OpenShift AI and RHEL AI

Watch out: No models are included — you source the weights, and no closed frontier model can be run here at all. Costs stack in a way that is easy to underestimate: OpenShift entitlement plus OpenShift AI plus AI Inference per accelerator plus the GPUs themselves, and Red Hat does not publish list prices, so budgeting requires a sales conversation. There is no per-token billing, so comparing it with Bedrock or Foundry means modelling your own utilisation. It demands genuine Kubernetes and GPU operations skill; a team that wanted an API key will find this a much larger commitment than expected. If you do not already run OpenShift, adopting it to get model serving is the tail wagging the dog.

subscription

SAP AI Core / Generative AI Hub

The generative AI hub in SAP AI Core is the sanctioned route to foundation models for SAP BTP applications — orchestration, grounding, prompt management and content filtering wired into SAP's own identity, data and extensibility model. Consumption is measured in 'GenAI tokens', a virtual unit whose conversion ratio from real input and output tokens varies per model, then converted again into BTP capacity units for billing; output tokens generally convert less favourably than input. The hub is available only under the extended service plan. Choose it because your application is an SAP extension, not because it is the best or cheapest way to reach a given model.

Models availableCurated third-party foundation models (OpenAI, Anthropic, Google and open weights) via the generative AI hub; catalog lags upstream launches
RegionsSAP BTP regions across AWS, Azure, GCP and Alibaba Cloud; model availability varies sharply by region
Data residency / ZDRFollows the BTP subaccount region and SAP's data processing agreement; check per-model terms since models are hosted by upstream providers
Provisioned throughputNo published token-rate reservation; capacity governed by the service plan and contract
IAM integrationSAP BTP Identity Authentication / Identity Provisioning, XSUAA roles, subaccount entitlements
Fine-tuningLimited — orchestration, grounding and prompt management are the focus rather than model tuning

Watch out: The double conversion — real tokens to GenAI tokens to BTP capacity units, with per-model ratios — makes forecasting and vendor comparison genuinely painful, and SAP does not publish plain per-1M-token dollar rates, so no honest price_summary figure exists. The generative AI hub sits behind the extended service plan, so it is gated by contract rather than a credit card. Model availability lags the upstream providers, sometimes by months, so you will not get a new frontier model here on launch day. Outside an SAP BTP context this platform has no reason to exist for you, and even inside one, calling a model provider directly is usually cheaper if you do not need SAP's grounding and governance integration.

enterprise