Veo 3.1 generates video with a synchronized audio track (dialogue, ambience, SFX) in a single pass, available through the Gemini API, Vertex AI, and the Flow consumer app. Google publishes three tiers at very different prices: Lite at $0.05/s (720p) and $0.08/s (1080p), Fast at $0.10/s (720p), $0.12/s (1080p) and $0.30/s (4K), and Standard at $0.40/s (720p/1080p) and $0.60/s (4K). Unlike most competitors, Google charges only for successfully generated videos. The practical difference from Runway or Luma is that audio is included rather than a separate pipeline stage.
| Max clip length | ~8s per generation; longer via scene extension in Flow |
|---|
| Resolution / fps | 720p / 1080p / 4K, 24fps |
|---|
| Native audio | Yes — dialogue, SFX, ambience, on by default |
|---|
| Image-to-video | Yes, plus first/last frame conditioning |
|---|
| Price per second | $0.05 (Lite 720p) – $0.60 (Standard 4K) |
|---|
| API availability | Gemini API + Vertex AI; also resold via Runway, fal, Replicate |
|---|
Watch out: Per-call clip length is short (roughly 8 seconds before you resort to scene extension), so anything narrative requires stitching. Standard at $0.40/s is expensive enough that iterating on it is a budget mistake — the Fast tier is what you should actually be prompting against. Safety filters are among the strictest in the category and reject a lot of benign human-subject prompts; there is no way to tune this. 4K is a Fast/Standard-only option.
Veo 3.1 Standard $0.40/s (720p/1080p), $0.60/s (4K); Fast $0.10/s (720p), $0.12/s (1080p), $0.30/s (4K); Lite $0.05/s (720p), $0.08/s (1080p). Consumer access via Google AI Pro/Ultra subscriptions.
Sora 2 and Sora 2 Pro generate video with synchronized audio and were the models that made AI video mainstream in late 2025. OpenAI discontinued the Sora web and mobile apps on 26 April 2026 and has scheduled the API for shutdown on 24 September 2026, reportedly because the product cost roughly $1M/day to run against negligible revenue. Pricing at time of writing is $0.10/s for sora-2 at 720p and $0.30/$0.50/$0.70 per second for sora-2-pro at 720p/1024p/1080p, halved under the Batch API. OpenAI has announced no successor video model.
| Max clip length | Up to ~12s (4/8/12s options) |
|---|
| Resolution / fps | 720p / 1024p / 1080p, ~24–30fps |
|---|
| Native audio | Yes |
|---|
| Image-to-video | Yes, via input reference image |
|---|
| Price per second | $0.10 – $0.70 (50% off in batch) |
|---|
| API availability | OpenAI API — terminating 24 September 2026 |
|---|
Watch out: This is a dead product with under three weeks of API life left as of 5 September 2026. Do not start anything on it. Consumer apps are already shut down and account data is scheduled for deletion. Note that OpenAI's own pricing page still lists the models with no deprecation banner, so the pricing page is not a reliable signal of availability here — the shutdown is documented in the OpenAI help centre and widely reported. Sora Pro at $0.70/s was also the most expensive mainstream option, which is part of why it did not survive.
sora-2: $0.10/s (720p). sora-2-pro: $0.30/s (720p), $0.50/s (1024p), $0.70/s (1080p). Batch API: 50% of these rates. Access ends 24 Sep 2026.
· deprecated
Gen-4.5 launched December 2025 and topped the Artificial Analysis Video Arena; it remains Runway's flagship because Runway has not shipped a frontier successor since. It is priced at 12 credits per second on both the app and the developer API, where credits are $0.01 each, so $0.12/s or $0.60 for a 5-second clip. Its practical edge is physical plausibility, human motion and camera language rather than raw resolution, plus professional delivery formats (ProRes, PNG sequences, HDR) that no other vendor in this list offers at API level. Runway's more interesting 2026 move is becoming a router: the same API now serves Veo 3.1, Seedance 2.5, Wan 3.0, Hailuo3 and Grok Imagine 1.5 alongside its own models.
| Max clip length | ~10s per generation (extendable) |
|---|
| Resolution / fps | 1080p native; 2K/4K via upscaling, 24fps |
|---|
| Native audio | No — separate audio models required |
|---|
| Image-to-video | Yes, plus references and keyframes |
|---|
| Price per second | $0.12 (Gen-4.5); Gen-4 Turbo $0.05 |
|---|
| API availability | dev.runwayml.com, plus a cost/latency/quality model router |
|---|
Watch out: No native audio generation in Gen-4.5 — you pay separately for Runway's audio and voice models and then have to sync them. The model is nine months old and Runway has publicly declined to say when Gen-5 lands, so you are buying a maturing rather than improving asset. App credits and developer API credits are separate pools, which surprises teams that prototype in the app and then discover their subscription buys nothing at the API. The credit abstraction also makes cross-vendor cost comparison harder than a published per-second rate.
API: 12 credits/s = $0.12/s ($0.01/credit, $10 minimum top-up). App plans: Free (125 one-time credits), Standard $12/mo annual (625 credits/mo), Pro $28/mo (2,250), Max $76/mo (9,500). ProRes/HDR delivery adds 5–40 credits/s.
Aleph is Runway's in-context video editing model rather than a generator: you feed it real or previously generated footage plus an instruction and it modifies the shot — changing lighting, removing an object, altering style, generating a new camera angle on the same scene. Version 2.0 shipped in May 2026. It is billed at 28 credits per second ($0.28/s) with a 56-credit ($0.56) minimum per job. It is the cleanest answer to 'I have footage and need to change it', a job that text-to-video models handle badly because they regenerate everything.
| Max clip length | Short shots; degrades on long clips |
|---|
| Resolution / fps | 1080p, plus ProRes/HDR delivery options |
|---|
| Native audio | No |
|---|
| Image-to-video | N/A — this is video-to-video |
|---|
| Price per second | $0.28 (56-credit / $0.56 minimum) |
|---|
| API availability | Yes, Runway developer API |
|---|
Watch out: More than twice the price per second of Gen-4.5, and it is the wrong tool if you have no source footage. Edits are not frame-accurate or deterministic, so it is a look-development tool rather than a VFX finishing tool — you will not get a rotoscope-grade matte out of it. Results degrade on long or high-motion clips, and the 56-credit minimum makes short test jobs disproportionately expensive.
28 credits/s = $0.28/s via the Runway API, 56-credit ($0.56) minimum per generation. Available on Runway app plans from Standard $12/mo.
Kling 3.0 launched 4 February 2026 and is the strongest non-Western competitor to Veo on motion quality and character consistency. It generates 3–15 second clips with native audio in Chinese, English, Japanese, Korean and Spanish, and adds a multi-shot storyboard mode that produces up to six camera cuts inside a single generation — genuinely useful, and something Veo and Runway do not offer. Via fal, Kling v3 Pro is $0.112/s with audio off, $0.168/s with audio on, and $0.196/s with voice control. Kuaishou's own API sells prepaid resource packages rather than a simple per-second rate, listing 6–8 credits/s without audio and 9–12 credits/s with it.
| Max clip length | 15s (3–15s range), multi-shot up to 6 cuts |
|---|
| Resolution / fps | 720p/1080p standard; up to 4K at 60fps claimed |
|---|
| Native audio | Yes — multilingual, multi-speaker, optional |
|---|
| Image-to-video | Yes, plus start/end frame control |
|---|
| Price per second | $0.112 (no audio) – $0.196 (audio + voice control) |
|---|
| API availability | Yes — kling.ai dev platform, fal, Replicate, Runway, PiAPI |
|---|
Watch out: The 4K/60fps headline is a marketing ceiling, not what most generations return — plan on 1080p. Pricing is genuinely hard to pin down: Kuaishou's own developer platform sells opaque prepaid 'resource unit' packages rather than a public per-second rate, and the app and API billing systems are entirely separate, so most teams end up buying through fal, Replicate or Runway and paying a reseller margin. Chinese-jurisdiction data handling will fail some enterprise reviews. Non-Chinese/English audio is machine-translated to English rather than natively generated.
Via fal: $0.112/s (no audio), $0.168/s (audio), $0.196/s (audio + voice control) — a 5s clip is $0.56–$0.98. Consumer plans: Free (66 daily credits), Standard ~$10/mo, Pro ~$37/mo, Premier ~$92/mo, Ultra ~$180/mo. Official API sells prepaid packages (e.g. $9.80 trial, $700 for 5,000 units).
Seedance 2.0 is ByteDance's flagship generator, notable for spanning the widest resolution range in the category — 480p through genuine 4K — with a single model family and matched audio. It takes text, images, video and audio as input and supports first-frame and last-frame conditioning. Volcengine bills it in tokens (roughly ¥46/M for pure generation, ¥28/M when editing existing video), which works out to about $0.14/s on average, with published per-second equivalents from $0.067 at 480p to $0.778 at 4K. Faster and cheaper variants (Seedance 2.0 Fast, 2.0 Mini) exist for iteration.
| Max clip length | 15s (4–15s range) |
|---|
| Resolution / fps | 480p / 720p / 1080p / 4K |
|---|
| Native audio | Yes — matched audio track |
|---|
| Image-to-video | Yes, first-frame and last-frame control |
|---|
| Price per second | ~$0.067 (480p) – ~$0.78 (4K) |
|---|
| API availability | Yes — BytePlus/Volcengine, plus fal, Runway, OpenRouter |
|---|
Watch out: The 4K tier is expensive enough ($0.78/s direct, up to $1.50/s through Runway) that it is a finishing-only option — generating at 4K for iteration will destroy a budget fast. Token-based billing makes cost forecasting materially harder than a flat per-second rate, and quoted per-second equivalents vary by a factor of two between resellers. Access splits between Volcengine (China) and BytePlus (international) with different terms, and ByteDance provenance is a non-starter for some US government-adjacent buyers.
~$0.067/s at 480p rising to ~$0.778/s at 4K; averages about $0.14/s (¥1/s) per Volcengine. Token billing: ¥46/M tokens for generation, ¥28/M for video-input editing. On Runway: 36–150 credits/s ($0.36–$1.50/s) for 2.0, 29 credits/s for 2.0 Fast.
Released 7 August 2026, Seedance 2.5 trades resolution for duration and controllability: it produces up to 30 seconds of audio-video in a single generation with lip-synced dialogue lifted from quoted lines in the prompt, and accepts up to 50 reference assets (roughly 30 images, 10 video clips, 10 audio clips) to lock characters, products and voices across shots. That reference budget is the largest of any model here and is the reason to pick it over Kling or Veo for serial content. The live API currently exposes only 480p and 720p output tiers, unlike Seedance 2.0's 4K.
| Max clip length | 30s (4–30s range) in one pass |
|---|
| Resolution / fps | 480p / 720p only on the live API |
|---|
| Native audio | Yes — optional, with multilingual lip-synced dialogue |
|---|
| Image-to-video | Yes — plus up to 50 image/video/audio references |
|---|
| Price per second | ~$0.10 (480p) – ~$0.47 (720p); reseller quotes vary widely |
|---|
| API availability | Yes — BytePlus/Volcengine, fal, Runway, OpenRouter |
|---|
Watch out: No 1080p or 4K on the live API as of September 2026, so it is a storytelling model you will have to upscale for delivery — a real regression from Seedance 2.0. Published pricing is inconsistent across resellers by a factor of four or more (from ~$0.10/s to over $2/s depending on who is quoting and which resolution), and reference-video inputs carry extra charges, so model your actual costs before committing. Same ByteDance jurisdiction concerns as 2.0. Treat every price here as needing verification against your chosen provider.
From ~$0.1028/s at 480p; token-billed at ~$0.0214 per 1,000 tokens, working out to roughly $0.22/s at 480p and ~$0.47/s at 720p. On Runway: 20–68 credits/s ($0.20–$0.68/s) plus surcharges for input/reference video.
MiniMax shipped H3 on 29 July 2026 at $0.13/s, generating 5–15 second clips at up to 2K with a matched audio track and first/last-frame steering. On 2 September 2026 it followed with H3 Max, a faster variant at $0.05/s (480p) to $0.08/s (768p) that drops the resolution ceiling and, per the current model cards, the audio. MiniMax's historical strength is human motion and expression on a modest budget, and the H3 Max price point makes it one of the cheapest credible options for high-volume iteration. Some sources describe H3 as open-weights; that claim is not confirmed on MiniMax's own channels.
| Max clip length | 15s (5–15s range) |
|---|
| Resolution / fps | H3 up to 2K; H3 Max 480p/768p |
|---|
| Native audio | H3 yes; H3 Max not documented |
|---|
| Image-to-video | Yes — first-frame and last-frame keyframes |
|---|
| Price per second | H3 $0.13; H3 Max $0.05–$0.08 |
|---|
| API availability | Yes — MiniMax platform, fal, OpenRouter, Runway |
|---|
Watch out: H3 Max, the cheap and fast option, appears to have no audio and tops out at 768p — so the variant you want for cost is the one that is weakest on the two features that most define the 2026 category. Resolution ceilings across the family (2K on H3, 768p on Max) are below Veo, Kling and Seedance. Both models are days-to-weeks old as of this writing, so behaviour and pricing are still moving. Treat the widely repeated 'open-weights' description of H3 with suspicion until MiniMax publishes a licence and checkpoint.
H3: $0.13/s (2K) plus $0.04 per reference image after the first five. H3 Max: from $0.05/s (480p) to $0.08/s (768p). Via Runway: H3 Max 5–8 credits/s, Hailuo3 10–15 credits/s. Consumer Hailuo plans run roughly $10–$95/mo.
Luma's Ray3.2 is the only model in this list with a native 16-bit HDR pipeline and EXR export, which is why it shows up in colour-managed film and commercial pipelines where everything else outputs flat 8-bit H.264. API billing is per 5-second block: $0.15/$0.30/$1.20 for a 5s clip at 540p/720p/1080p, doubling for HDR and tripling for HDR+EXR. Video-to-video runs up to 20 seconds with 5-second forward and backward extensions, and there is a per-second reframe (aspect-ratio conversion) endpoint from $0.06/s.
| Max clip length | 10s generated; up to 20s via video-to-video extension |
|---|
| Resolution / fps | 540p / 720p / 1080p, 16-bit HDR + EXR option |
|---|
| Native audio | No |
|---|
| Image-to-video | Yes, plus keyframe and video-to-video modes |
|---|
| Price per second | $0.03 (540p) / $0.06 (720p) / $0.24 (1080p); HDR 2× |
|---|
| API availability | Yes — Luma API, pay-as-you-go, no minimum |
|---|
Watch out: 1080p is priced badly relative to the field: $1.20 for a 5-second clip is $0.24/s, roughly double Runway and triple Kling, and HDR doubles that again to $0.48/s. There is no native audio generation. Clips are 5 or 10 seconds, shorter than Kling, Seedance or Wan. Model naming is a mess in public — Ray3, Ray3.14 and Ray3.2 all appear in circulation, and only Ray3.2 is on the current API pricing page. Subscription credits do not roll over.
API per 5s block: 540p $0.15, 720p $0.30, 1080p $1.20; per 10s: $0.45 / $0.90 / $3.60. Video-to-video 5s: $0.72–$2.16. Reframe $0.06–$0.36/s. HDR 2×, HDR+EXR 3×. App plans: Plus $30/mo, Pro $90/mo, Ultra $300/mo.
Pika competes on manipulation primitives rather than raw generation quality. Its differentiators are named tools: Pikaffects (melt, explode, squish), Pikadditions (drop a character or object into real footage), Pikaswaps (replace an object in a scene), Pikaframes (interpolate between start and end keyframes), and Pikaformance (audio-driven lip-sync at 3 credits/second). Plans on annual billing are Free (80 credits), Standard $8/mo (700), Pro $28/mo (2,300, watermark-free with commercial rights) and Fancy $76/mo (6,000). API access runs through dev.pika.art and, for the 2.2 generation, through fal at roughly $0.20 per 5s clip at 720p and $0.45 at 1080p.
| Max clip length | ~5s per generation, extendable |
|---|
| Resolution / fps | 480p / 720p / 1080p by plan tier |
|---|
| Native audio | No generative soundtrack; audio-driven lip-sync via Pikaformance |
|---|
| Image-to-video | Yes, plus video-to-video and keyframe interpolation |
|---|
| Price per second | ~$0.04–$0.09 equivalent via fal (Pika 2.2); credit-based in app |
|---|
| API availability | Yes — dev.pika.art and fal-hosted endpoints |
|---|
Watch out: Base generation quality is clearly behind Veo 3.1, Kling 3.0 and Seedance; if you want a photoreal shot, this is the wrong tool. Clips are short (typically 5s, extendable) and the credit system makes per-second cost hard to compare against per-second-priced competitors. Commercial use and watermark removal require the $28/mo Pro tier. The API story is fragmented — Pika's own dev portal and a fal-hosted 2.2 endpoint sit alongside the newer 2.5 app model, and 2.5 API parity is not clearly documented.
Annual: Free $0 (80 credits), Standard $8/mo (700), Pro $28/mo (2,300), Fancy $76/mo (6,000); monthly billing ~20% more; add-on credits ~$26.67 per 1,000. Via fal (Pika 2.2): ~$0.20 per 5s at 720p, ~$0.45 at 1080p.
Wan 3.0 entered public beta on 6 August 2026 and is built around a single capability nobody else matches cleanly: a native 30-second generation with automatic scene splitting, including a document-to-video mode that turns a script or deck into a multi-shot sequence. It accepts up to 20 reference assets, does integrated sound design, and handles text rendering better than most. On Runway's API it is the cheapest listed video model at 5–20 credits/s ($0.05–$0.20/s for 480p–1080p). Direct access is via Alibaba's Model Studio and Qwen Cloud, currently behind an application process.
| Max clip length | 30s native, with automatic scene splitting |
|---|
| Resolution / fps | 480p – 1080p on published routes; fps not documented |
|---|
| Native audio | Yes — integrated sound design |
|---|
| Image-to-video | Yes — image-to-30s, plus up to 20 multimodal references |
|---|
| Price per second | $0.05 (480p) – $0.20 (1080p) via Runway |
|---|
| API availability | Application-gated on Model Studio / Qwen Cloud; open via Runway and fal |
|---|
Watch out: Alibaba pre-announced Wan 3.0 as an Apache 2.0 open-weights release and then did not ship weights — the open Wan line stops at 2.2, and the official GitHub repo contains only a README and a licence file. If your plan depended on self-hosting Wan 3.0, that plan is dead. It is still in beta with application-gated direct access, and Alibaba does not publish clear USD per-second pricing, so most Western teams reach it through Runway or fal at a margin. Resolution tops out at 1080p on the routes with published pricing.
Via Runway API: 5–20 credits/s = $0.05/s (480p) to $0.20/s (1080p) — the cheapest model on that platform. Direct Alibaba Model Studio / Qwen Cloud pricing is not clearly published in USD; confirm before committing.
· beta
Wan 2.2 is the last openly licensed generation of Alibaba's video line and remains the default choice for anyone who needs weights on their own hardware — Apache 2.0, widely supported in ComfyUI and diffusers, with a large ecosystem of LoRAs and community fine-tunes. Alibaba continued shipping Apache 2.0 components into August 2026 (Wan2.2-Animate-2-14B, base and distilled checkpoints), but the frontier moved to closed Wan 2.5 and 3.0. If you want hosted convenience instead, Wan 2.5 runs around $0.05/s on fal, still the cheapest hosted rate in the category.
| Max clip length | ~5s typical per generation; extendable with community tooling |
|---|
| Resolution / fps | 720p / 1080p, 24fps (hardware-dependent) |
|---|
| Native audio | No |
|---|
| Image-to-video | Yes, plus animate/reference variants |
|---|
| Price per second | $0 self-hosted (GPU cost); ~$0.05 hosted on fal (Wan 2.5) |
|---|
| API availability | Self-host, or hosted on fal, Replicate and Runway |
|---|
Watch out: Quality is a generation or two behind Veo 3.1, Kling 3.0 and Seedance 2.x, and the gap is visible in motion coherence and prompt adherence. No native audio at this generation. You own the ops: multi-GPU inference, VRAM management, queueing, safety filtering and content moderation are all yours. Self-hosting only beats API pricing above roughly $3–5k/month of hosted spend. And the strategic lesson of 2026 applies — Alibaba's open line stopped at 2.2, so do not assume future Wan releases will be open.
Weights: free under Apache 2.0 (you pay GPU cost). Hosted Wan 2.5 on fal: $0.05/s (~20 seconds of output per $1).
· open source
Grok Imagine Video 1.5 is xAI's text-and-image-to-video model, launched mid-2026 and priced aggressively — xAI's own docs list a $0.080/s base rate, with tiered pricing reported at $0.08/s (480p), $0.14/s (720p) and $0.25/s (1080p) plus $0.01 per input image. It generates synchronized audio, supports the batch API, is deployed in us-east-1 and us-west-2, and allows 10 requests per second with effectively unlimited token throughput. It briefly topped video leaderboards at a fraction of Sora's price, which is a reasonable summary of xAI's strategy here.
| Max clip length | Not documented by xAI |
|---|
| Resolution / fps | Reported 480p / 720p / 1080p; not stated in official docs |
|---|
| Native audio | Yes — synchronized audio |
|---|
| Image-to-video | Yes ($0.01 per input image) |
|---|
| Price per second | $0.08 base; reported up to $0.25 at 1080p |
|---|
| API availability | Yes — xAI API (preview), batch supported; also on fal and Runway |
|---|
Watch out: The model ID is literally `grok-imagine-video-1.5-preview` — it is a preview service and xAI's docs do not publish max duration, fps or resolution tiers, so the widely quoted resolution-based prices come from secondary sources rather than xAI. That is a poor basis for a production cost model. Content moderation is notably looser than Google's or OpenAI's, which is a liability as often as a feature for brand work. Only two US regions, no data residency options, and no track record of long-term API stability from xAI in this product line.
$0.080/s base per xAI docs; reported tiers $0.08/s (480p), $0.14/s (720p), $0.25/s (1080p), plus $0.01 per input image. Batch API supported. Rate limit 10 req/s.
· beta
LTX-2 (January 2026, updated to LTX-2.5 in August) is the only openly licensed model here that generates video and audio together: 14B video parameters plus 5B audio, up to 20 seconds at native 4K and 50fps with expressive lip sync. Lightricks released weights, inference code and training code, plus a distilled variant and NVFP8 quantization that cuts model size ~30% and roughly doubles throughput. The licence is free for academic use and for companies under $10M ARR, which makes it the default for startups and researchers who need audio-video without a per-second bill. Hosted endpoints are available on fal for teams that do not want to run GPUs.
| Max clip length | 20s with synchronized audio |
|---|
| Resolution / fps | Native 4K at 50fps |
|---|
| Native audio | Yes — 5B-parameter audio model with lip sync |
|---|
| Image-to-video | Yes |
|---|
| Price per second | $0 self-hosted under the ARR threshold; hosted rates vary |
|---|
| API availability | Self-host (weights + training code on GitHub/HF) or hosted on fal |
|---|
Watch out: The licence is source-available, not open source: cross $10M ARR and you need a commercial deal on terms Lightricks does not publish, which is a real procurement risk if you build a business on it. Quality trails Veo 3.1 and Kling 3.0 on prompt adherence and motion, particularly on complex human action. 'Runs on consumer GPUs' means the distilled/quantized variants at reduced settings — native 4K/50fps at 20 seconds needs serious hardware. As with any self-hosted model, safety filtering and moderation are entirely your problem.
Free under the LTX-2 licence for academic use and companies under $10M ARR; commercial licence required above that (terms not publicly listed). Hosted rates vary by provider — verify on fal or Replicate before budgeting.
· open source
ShengShu's Vidu line is built around reference-conditioned generation — feeding in images of specific characters, objects or settings and getting them back consistently across shots — and Q3 extends that with a dedicated reference-to-video endpoint and the higher-consistency Q3-mix variant. Credits are $0.005 each and pricing is explicit: Q3-pro runs 24/20/9 credits per second at 1080p/720p/540p ($0.12/$0.10/$0.045), Q3-turbo runs 13/11/7 ($0.065/$0.055/$0.035), and Q3-mix reference-to-video goes up to 29 credits/s. Clips run 1–16 seconds. ShengShu offers a 50% off-peak discount, which is unusual and genuinely useful for batch pipelines.
| Max clip length | 16s (1–16s range; reference mode 3–16s) |
|---|
| Resolution / fps | 540p / 720p / 1080p |
|---|
| Native audio | Not listed for Q3 (Q2 charged a separate audio add-on) |
|---|
| Image-to-video | Yes — plus reference-to-video and start/end frame modes |
|---|
| Price per second | $0.035 (turbo 540p) – $0.145 (Q3-mix reference) |
|---|
| API availability | Yes — platform.vidu.com, pay-as-you-go; also WaveSpeed, Atlas Cloud |
|---|
Watch out: Audio is the weak spot: Vidu Q2 charged a flat 15-credit add-on for audio and the Q3 pricing tables do not list native audio at all, so assume you are doing a separate audio pass. Output quality is a tier below Veo 3.1 and Kling 3.0 on photorealism — Vidu's reputation is stylized and anime work more than live-action realism. The credit abstraction plus off-peak/peak variance makes cost modelling fiddly. ShengShu is a smaller Chinese vendor with a shorter operating history than Kuaishou or ByteDance, which carries its own continuity risk.
Credits $0.005 each. Q3-pro: 1080p 24 cr/s ($0.12), 720p 20 cr/s ($0.10), 540p 9 cr/s ($0.045). Q3-turbo: 13/11/7 cr/s ($0.065/$0.055/$0.035). Q3-mix reference-to-video up to 29 cr/s ($0.145). 50% off-peak discount available.
PixVerse is the effects-template player: its strength is fast, stylized, anime-adjacent output driven by a large library of one-tap effect presets, which is why it dominates short-form social workflows in Asia. V5 added native audio generation with automatic soundtracks and SFX, and V5.6 improved audio-visual sync. Consumer pricing is Standard $8/mo (1,200 credits, 720p) and Pro $24/mo (6,000 credits, 1080p). Programmatic access is either PixVerse's own API plans, which start at $100/mo for 15,000 credits with no free tier, or pay-as-you-go through resellers from about $0.025/s.
| Max clip length | ~5–8s typical per generation |
|---|
| Resolution / fps | 540p / 720p / 1080p by tier |
|---|
| Native audio | Yes — automatic soundtrack and SFX since V5 |
|---|
| Image-to-video | Yes, plus effect-template modes |
|---|
| Price per second | ~$0.025/s via resellers; credit-based on official API plans |
|---|
| API availability | Yes, but $100/mo minimum on PixVerse; pay-as-you-go via Together AI, Atlas Cloud, Replicate |
|---|
Watch out: The API is committed-spend only — $100/month minimum with no free tier and no pay-as-you-go on PixVerse's own platform, which is a hard stop for prototyping and pushes small teams to resellers. Photorealism is not the strength; for live-action-looking output you want Veo, Kling or Seedance. The template-driven approach means outputs look recognisably PixVerse, which is a liability for brand work. Credit costs per generation vary by effect and resolution in ways the pricing page does not fully enumerate.
Consumer: Standard $8/mo (1,200 credits, 720p), Pro $24/mo (6,000 credits, 1080p). Official API plans: Essential $100/mo (15,000 credits, ~333 clips at 540p/5s), Scale $1,500/mo (239,230 credits), Business $6,000/mo (1,069,500 credits) — no free tier. Resellers: ~$0.025/s pay-as-you-go, or ~$0.30/video via Together AI.
The Firefly Video Model exists to be legally defensible: Adobe trains on licensed and public-domain content and indemnifies commercial output, which is the entire reason enterprises pick it over models that generate better video for less. Plans are Standard $9.99/mo (2,000 credits), Pro $19.99/mo (4,000), Pro Plus $49.99/mo (10,000, roughly 100 five-second videos), and Premium $199.99/mo (50,000 credits with unlimited Firefly Video Model generation). Credits are consumed only by premium features — video, translation, sound effects and partner models. Firefly also aggregates third-party models (Veo, Runway and others) inside the same interface.
| Max clip length | ~5s per generation |
|---|
| Resolution / fps | 1080p, 24fps |
|---|
| Native audio | Sound effects and translation available as separate credit-billed features |
|---|
| Image-to-video | Yes |
|---|
| Price per second | Not published per-second; ~100 five-second clips on the $49.99/mo tier |
|---|
| API availability | Firefly Services for enterprise; no public per-second consumer API |
|---|
Watch out: The model itself is behind the field — 5-second 1080p clips with weaker motion and prompt adherence than Veo 3.1, Kling 3.0 or Seedance, at a higher effective cost. Critically, the indemnity covers only Adobe's own Firefly models; the third-party partner models exposed in the same UI are not indemnified, and teams routinely miss this distinction and lose the one benefit they were paying for. Pricing is subscription-and-credit rather than per-second, so it does not slot into a cost model alongside the API-native vendors, and there is no straightforward public per-second developer API.
Standard $9.99/mo (2,000 credits), Pro $19.99/mo (4,000), Pro Plus $49.99/mo (10,000, ~100 5s videos), Premium $199.99/mo (50,000 credits + unlimited Firefly Video Model). No public per-second API rate.
Higgsfield is primarily a front end and credit wallet over other people's models — Kling 3.0, Seedance 2.0, Veo 3.1 and others — plus its own Soul V2 and Soul Cinema character/cinematic-look models and a large library of camera-motion presets. Plans run roughly Free ($0, no commercial use), Starter ~$19/mo annual (270 credits), Plus $47/mo annual or $59 monthly (1,200 credits), and Ultra $99/mo annual or $129 monthly (3,000 credits), with concurrency scaling from 1 to 8 parallel generations. Representative costs: Kling 3.0 at ~7–10 credits per 5s clip, Seedance 2.0 1080p at ~45 credits per 5s clip.
| Max clip length | Inherited from the underlying model (5–30s) |
|---|
| Resolution / fps | Up to 1080p+ depending on the model selected |
|---|
| Native audio | Depends on the underlying model (Kling/Veo/Seedance yes) |
|---|
| Image-to-video | Yes, plus camera-motion presets and character tools |
|---|
| Price per second | Credit-based, not per-second; ~45 credits per 5s Seedance 1080p clip |
|---|
| API availability | No clearly documented public developer API |
|---|
Watch out: You are paying a margin on models you could buy directly, and the credit abstraction actively obscures what each generation costs upstream — at ~45 credits per 5s Seedance 1080p clip, the $19 Starter tier buys you about six of them. It has no frontier first-party video model of its own, so its capability is entirely downstream of vendors who can reprice or pull access. There is no clearly documented developer API, making it a poor fit for anything programmatic. Advertised pricing has been reported to swing widely month to month, so verify against the live pricing page rather than any secondary source, including this one.
Free $0 (no commercial use); Starter ~$19/mo annual (270 credits); Plus $47/mo annual or $59/mo monthly (1,200 credits); Ultra $99/mo annual or $129/mo monthly (3,000 credits). Per-generation: Kling 3.0 ~7–10 credits per 5s, Seedance 2.0 1080p ~45 credits per 5s.