Seminal AI
§2

Chinese AI labs

Seven distinct companies grouped for convenience — they do not share infrastructure, licences, or pricing conventions. Moonshot AI (Kimi) has become the most Western-developer-friendly of the group: prices posted in USD, an api.moonshot.ai endpoint that speaks both OpenAI and Anthropic wire formats, and a documented deprecation list (the entire moonshot-v1 line, kimi-k2, kimi-k2.5 and kimi-latest are retired in favour of kimi-k3).

Data checked 2026-09-06
Models tracked
38
OpenAI-compatible API
yes
API base URL
https://api.moonshot.ai/v1 (Kimi, also /anthropic) · https://api.z.ai/api/paas/v4 (GLM) · https://api.minimax.cn/v1 (MiniMax, also /anthropic) · https://api.hunyuan.cloud.tencent.com/v1 (Tencent) · https://qianfan.baidubce.com/v2 (Baidu) · wss://spark-api.xf-yun.com/v4.0/chat (iFlytek) · https://ark.cn-beijing.volces.com/api/v3 (Volcengine Ark) / https://ai.byteplus.com (BytePlus ModelArk)

Kimi K3 is a 2.8T-parameter / 104B-active MoE with MXFP4 quantisation-aware training, released under a bespoke "Kimi K3 License"; the K2.x line stays under Modified MIT. Z.ai (Zhipu) is the price leader at the frontier: GLM-5.3 matches GLM-5.2 at $1.40/$4.40 with a 1M context, and GLM-5.3-Flash — a 320B-total/18B-active natively multimodal model — undercuts everything else in this report at $0.075/$0.25. GLM-5.2, GLM-5 and GLM-5.3-Flash are plain MIT on Hugging Face; only GLM-5.3 itself carries a custom licence. MiniMax sells one current model (M3, 428B/~23B active, 1M context) at $0.30/$1.20 and keeps M2.7 as a cheaper 204k-context option. Its "highspeed" SKUs are the same weights at 2x price for throughput, and a Priority tier bills at 1.5x standard. ByteDance ships Seed 2.1 (Pro and Turbo) through Volcengine Ark domestically and BytePlus ModelArk internationally. Seed 2.1 is closed-weight — ByteDance-Seed's Hugging Face org publishes research models (Cola-DLM, Seed1.5-VL, Stable-DiffCoder) but not the flagship. Pricing is not machine-readable from either doc site. Tencent has effectively bifurcated: the paid Hunyuan API is now a catalogue of small/specialist models (a13b, translation, role-play, vision) at very low CNY prices, while the serious frontier work — Hy3 (295B/21B, 256k) and Hy4-preview (770B/49B, 1M) — ships as Apache-2.0 open weights with no first-party per-token endpoint. That makes Tencent the most generous open-weight licensor in this group and the least useful as an API vendor. Baidu is the opposite: ERNIE 5.1 and 5.0 are closed, API-only, and priced in CNY per 1k tokens, while the open ERNIE 4.5 family (Apache-2.0) is a generation behind. ERNIE 5.1 undercut ERNIE 5.0 by ~33% on input.

iFlytek Spark is the weakest fit for this manual: real API, published model list and context windows, but no retrievable public rate card and a sales-contact-driven pricing model. Its top tier is "Spark 4.0 Ultra" (now running an X1.5 fast-thinking mode); Spark Max retires 10 Mar 2026.

A Access and limits

CURRENCY CONVERSION: Tencent Hunyuan and Baidu ERNIE publish prices only in CNY. All CNY figures below were converted at 1 USD = 6.7344 CNY (open.er-api.com daily rate, snapshot dated Sun 06 Sep 2026 00:02 UTC). Baidu quotes 元 per 1,000 tokens; those were multiplied by 1,000 before conversion (e.g. ERNIE 5.1 input 0.004 元/1k = 4 元/M = $0.594/M). Tencent already quotes 元 per 1M tokens. Moonshot/Kimi, Z.ai and MiniMax's international platform all publish USD directly — no conversion applied to those.

MiniMax has two storefronts with different currencies for the same models: platform.minimaxi.com (CN, CNY: MiniMax-M3 ¥2.10 in / ¥8.40 out per 1M) and platform.minimax.io (international, USD: $0.30 in / $1.20 out). The USD figures are recorded here as first-party. Both carry a "permanent 50% off" banner on M3, i.e. these ARE the posted rates.

BYTEDANCE GAP — READ THIS: Volcengine Ark (docs.volcengine.com/docs/82379/1544106) and BytePlus ModelArk (docs.byteplus.com/en/docs/ModelArk/1544106) are both client-rendered SPAs that return no pricing text to a fetcher, and the doc-centre JSON API refuses unauthenticated reads ("未授权访问"). ByteDance definitely sells Doubao/Seed per-token, but no per-token number could be read off an official page, so Seed 2.1 prices are -1 / low confidence rather than guessed. Do not backfill these from an aggregator.

IFLYTEK GAP: Spark has a genuine public API (WebSocket, wss://spark-api.xf-yun.com/...) and the model/domain table at xfyun.cn/doc/spark/Web.html is readable, but every xfyun.cn pricing/product page (/services/bm4, /services/cbm) returns HTTP 500 and xinghuo.xfyun.cn renders client-side. iFlytek's docs also direct buyers to a WeChat sales contact for "one-on-one quotes", so a public rate card may not exist. All Spark prices are -1.

TENCENT SPLIT: Tencent's published price list (cloud.tencent.com/document/product/1729/97731) covers only hunyuan-a13b, the translation/role models and the vision models. Its flagship chat model hunyuan-turbos-latest appears in the API docs (1729/111007) but NOT in the price table — the page states billing functions are "gradually migrating to TokenHub", and tokenhub.tencent.com does not resolve. hunyuan-turbos-latest is therefore priced -1. Tencent's newest open weights (Hy4-preview, Hy3) are Apache-2.0 on Hugging Face and are NOT in any first-party price list, so both prices are -1 per the open-weight rule.

OPEN-WEIGHT / API OVERLAP: Moonshot (all Kimi models), Z.ai (all GLM-5.x), MiniMax (M3, M2.7) and Tencent (hunyuan-a13b) publish weights AND sell a first-party API for the same checkpoint — those carry real prices plus open_weights true. Baidu's ERNIE 4.5 open weights (Apache-2.0) are a different lineage from the paid ERNIE 5.x endpoints; the open ERNIE-4.5-21B-A3B-Thinking is priced -1 because it is not a line item in Qianfan's price table.

CACHE PRICING: Kimi, Z.ai, MiniMax and Baidu all bill cache-hit input at a large discount (Kimi K3 $0.30 vs $3.00 cache-miss — a 10x spread, the single biggest cost lever in this group). Z.ai currently lists cached-input storage as "Limited-time Free". Z.ai's GLM-5.3-Flash prices are marked 50% off through 9 Sep 2026; the posted (discounted) numbers are recorded.

CONTEXT TIERING: MiniMax M3 doubles its rate above 512k input tokens ($0.60/$2.40). Baidu ERNIE 5.x has a 32k/128k tier break. Recorded prices are the base (cheapest) tier; the higher tier is called out in each model's best_for.

Where a context window could not be confirmed on an official page it is set to -1 rather than assumed.

C Calling the API

Chinese AI labs accepts requests in the OpenAI Chat Completions format, so any OpenAI-compatible client works by changing the base URL. Example uses glm-5.3-flash.

curl https://api.moonshot.ai/v1 (Kimi, also /anthropic) · https://api.z.ai/api/paas/v4 (GLM) · https://api.minimax.cn/v1 (MiniMax, also /anthropic) · https://api.hunyuan.cloud.tencent.com/v1 (Tencent) · https://qianfan.baidubce.com/v2 (Baidu) · wss://spark-api.xf-yun.com/v4.0/chat (iFlytek) · https://ark.cn-beijing.volces.com/api/v3 (Volcengine Ark) / https://ai.byteplus.com (BytePlus ModelArk)/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3-flash",
    "messages": [{"role": "user", "content": "Hello"}]
  }'
from openai import OpenAI

client = OpenAI(
    base_url="https://api.moonshot.ai/v1 (Kimi, also /anthropic) · https://api.z.ai/api/paas/v4 (GLM) · https://api.minimax.cn/v1 (MiniMax, also /anthropic) · https://api.hunyuan.cloud.tencent.com/v1 (Tencent) · https://qianfan.baidubce.com/v2 (Baidu) · wss://spark-api.xf-yun.com/v4.0/chat (iFlytek) · https://ark.cn-beijing.volces.com/api/v3 (Volcengine Ark) / https://ai.byteplus.com (BytePlus ModelArk)",
    api_key=os.environ["API_KEY"],
)

response = client.chat.completions.create(
    model="glm-5.3-flash",
    messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)

A compatibility layer is not always the provider's full API — caching, strict tool schemas and structured output are commonly unsupported there. Check the official docs before relying on an advanced feature.

B Models

Model Context Max out In $/M Out $/M Blended Status
GLM-4.7-Flash glm-4.7-flash 200K 131.1K $0.00 $0.00 $0.00 ga
GLM-5.3-Flash glm-5.3-flash 1M 131.1K $0.07 $0.25 $0.12 ga open
hunyuan-a13b hunyuan-a13b 224K 32.8K $0.07 $0.30 $0.13 ga open
GLM-4.7-FlashX glm-4.7-flashx 200K 131.1K $0.07 $0.40 $0.15 ga
ERNIE-4.5-Turbo-128K ernie-4.5-turbo-128k 128K $0.12 $0.47 $0.21 ga
hunyuan-translation-lite hunyuan-translation-lite 4K 4K $0.15 $0.45 $0.22 ga
hunyuan-translation hunyuan-translation 4K 4K $0.18 $0.54 $0.27 ga
GLM-4.5-Air glm-4.5-air $0.20 $1.10 $0.43 ga open
GLM-4.6V glm-4.6v 128K $0.30 $0.90 $0.45 ga
MiniMax-M3 MiniMax-M3 1M $0.30 $1.20 $0.52 ga open
MiniMax-M2.7 MiniMax-M2.7 204.8K $0.30 $1.20 $0.52 ga open
hunyuan-role-latest hunyuan-role-latest 28K 4.1K $0.36 $1.43 $0.62 ga
hunyuan-vision-1.5-instruct hunyuan-vision-1.5-instruct 24K 16K $0.45 $1.34 $0.67 ga
hunyuan-t1-vision hunyuan-t1-vision-20250916 28K 20K $0.45 $1.34 $0.67 ga
hunyuan-turbos-vision-video hunyuan-turbos-vision-video 24K 8K $0.45 $1.34 $0.67 ga
ERNIE-4.5-Turbo-VL-32K ernie-4.5-turbo-vl-32k 32K $0.45 $1.34 $0.67 ga
GLM-4.7 glm-4.7 200K 131.1K $0.60 $2.20 $1.00 ga open
MiniMax-M2.7-highspeed MiniMax-M2.7-highspeed 204.8K $0.60 $2.40 $1.05 ga open
ERNIE 5.1 ernie-5.1 128K $0.59 $2.67 $1.11 ga
GLM-5 glm-5 200K 131.1K $1.00 $3.20 $1.55 ga open
ERNIE 5.0 ernie-5.0 128K $0.89 $3.56 $1.56 ga
Kimi K2.7 Code kimi-k2.7-code 262.1K $0.95 $4.00 $1.71 ga open
Kimi K2.6 kimi-k2.6 262.1K $0.95 $4.00 $1.71 ga open
GLM-5.3 glm-5.3 1M 131.1K $1.40 $4.40 $2.15 ga open
GLM-5.2 glm-5.2 1M 131.1K $1.40 $4.40 $2.15 ga open
Kimi K2.7 Code Highspeed kimi-k2.7-code-highspeed 262.1K $1.90 $8.00 $3.42 ga open
Kimi K3 kimi-k3 1M $3.00 $15.00 $6.00 ga open
Seed 2.1 Pro ga
Seed 2.1 Turbo dola-seed-2-1-turbo ga
hunyuan-turbos-latest hunyuan-turbos-latest ga
Hunyuan Hy4-preview 1M preview open
Hunyuan Hy3 256K ga open
ERNIE-4.5-21B-A3B-Thinking 131.1K ga open
Spark 4.0 Ultra 4.0Ultra 32.8K 32.8K ga
Spark Pro-128K pro-128k 131.1K 131.1K ga
Spark Max-32K max-32k 32.8K 32.8K ga
Spark Lite lite 8.2K 4.1K ga
Spark Max generalv3.5 8.2K 8.2K deprecated

D Official references