Seminal AI
§6

Image generation models

Text-to-image models are now bought like any other API primitive: metered per image or per megapixel, in the $0.01–$0.15 band, with editing and multi-reference composition as table stakes rather than separate products. The 2026 shape of the market is a top tier of reasoning-based multimodal models from OpenAI, Google, Microsoft and Meta that plan a layout before rendering and therefore get text and instruction-following right; a middle tier of specialist models (Ideogram for typography, Recraft for vector/brand systems, FLUX for open weights and self-hosting); and a shrinking tail of standalone image startups.

Data checked 2026-09-06

Consolidation has been brutal this year: Google retired Imagen entirely in favour of Nano Banana, Reve shut its API two weeks after taking OpenAI investment, Freepik rebranded to Magnific, and Stability shipped no new image model at all. Resolution has stopped being the differentiator — several models now do native 2K–4K — and the real axes are text fidelity, edit locality, licence terms and whether you can run the weights yourself. Anyone integrating today should assume the specific model ID they pick will be deprecated inside 12 months and build an abstraction layer accordingly.

A How to choose

Start by asking whether your prompt contains words that must appear correctly in the output. If it does — packaging, UI mockups, ads, infographics — you want a reasoning-based model (GPT Image 2, Nano Banana Pro, MAI-Image-2.6) or a typography specialist (Ideogram 4.0), and you should not use FLUX, Stable Diffusion or Midjourney, all of which still garble small text. If it does not, the market has a clearing price of roughly $0.03–$0.04 per production frame and a dozen models will do; pick on latency and licence, not Elo.

Second axis: editing. "Generate a new image" and "change this one thing and leave everything else alone" are different products. Nano Banana 2, GPT Image 2, Seedream 5.0 and MAI-2.6 do scoped, mask-free instruction edits well; FLUX Kontext and Recraft's raster ops are the alternatives if you need self-hosting or vector output.

Third axis: licence and provenance. If you need indemnification against training-data claims, Adobe Firefly is the only mainstream option that offers it contractually, and you pay for that in output quality. If you need to run weights on your own hardware — regulated data, air-gapped, or per-image economics at millions of frames — FLUX.2 Klein is the honest answer; Ideogram 4.0's weights are non-commercial and Stable Diffusion 3.5 is now two years stale.

Fourth axis: volume economics. At scale the spread is real: Meta Muse Image at $0.01, Nano Banana 2 at $0.067 (or half that in batch), Nano Banana Pro at $0.134, GPT Image 2 high at roughly $0.12. A 10x price gap that buys a 5% quality lift is a bad trade for thumbnails and a good one for hero assets, so tier your pipeline rather than standardising on one model.

On the popular defaults specifically: do not reach for Midjourney if you are building a product — it still has no official public API in September 2026, its ToS prohibits the unofficial wrappers people use, and that is a business risk, not an inconvenience; it remains excellent for humans making images by hand. Do not default to GPT Image 2 for high-volume background generation — the quality is top of the leaderboard but you are paying reasoning-model latency and price for frames nobody will scrutinise. And do not pick Reve or any single-product image startup for a new integration this year; Reve's API sunset on 14 August 2026 with about three weeks' notice is the cautionary tale, and the pattern of acqui-hires means the vendor list will look different by mid-2027.

B At a glance

Name Max resolutionPrice per imageText renderingEditing / inpaintingCommercial licence Pricing
OpenAI GPT Image 2 3840px long edge (up to ~8.3MP)~$0.01 / $0.03 / $0.12 (low/med/high at 1024²)Best in class; multi-scriptYes — mask-based and instruction editsPermitted; user owns outputs Metered on tokens, not flat per image: $5.00/M text input, $8.00/M image input, $30.00/M image output; batch API is 50% off. At 1024x1024 that works out to roughly $0.01 low, $0.03 medium, $0.12 high. gpt-image-1.5 ($32/M output) and gpt-image-1-mini ($8/M output) remain listed as cheaper siblings.
Google Nano Banana 2 (Gemini 3.1 Flash Image) 4096x4096 (4K)$0.067 at 1K; $0.151 at 4KGood; weaker on dense small typeYes — conversational multi-turn editsPermitted; SynthID watermark applied $0.045 (0.5K) / $0.067 (1K) / $0.101 (2K) / $0.151 (4K) per image; input $0.50/M tokens; batch pricing is 50% off across all tiers. Lite: $0.0336 at 1K.
Google Nano Banana Pro (Gemini 3 Pro Image) 4096x4096 (4K)$0.134 at 1K/2K; $0.24 at 4KVery good; better than Flash tierYes — instruction edits, multi-referencePermitted; SynthID watermark applied $0.134 per image at 1K or 2K; $0.24 at 4K; input $2.00/M tokens; batch 50% off.
Microsoft MAI-Image-2.6 ~1.5K, dynamic aspect ratios$38/M output tokens; no per-image rate cardStrong; #1 on AA editing benchmarkYes — up to 5 reference imagesPermitted under Azure/Foundry terms $5.00/M text input, $8.00/M image input, $38.00/M image output (as listed on aggregators; Microsoft's launch post cites price-per-Elo leadership but publishes no per-image rate card). Free to test in MAI Playground and on Arena.
Meta Muse Image ~1600px$0.01Inconsistent — varies run to runYes — edit and multi-image composePermitted; terms not detailed in docs $0.01 per image at production volumes on Meta Model API. Free for consumers in the Meta AI app, WhatsApp DMs and Instagram Stories.
Black Forest Labs FLUX.2 Pro ~4MP (≈2048x2048)$0.03 at 1MP; $0.045/MP for editsWeak — headlines onlyYes — FLUX Tools, Kontext-style editsPermitted on hosted API $0.03/MP text-to-image (≈$0.03 for a 1024² frame), $0.045/MP editing, $0.015/MP reference-image input, $0.015/MP for each MP beyond the first. Flex $0.05/MP, Max $0.07/MP. 1 credit = $0.01. Roughly $0.055/image on Replicate and $0.073 on fal if you route through a reseller.
Black Forest Labs FLUX.2 Max ~4MP (≈2048x2048)$0.07 at 1MP (generation and editing)Weak — same family limitationYes — same price as generationPermitted on hosted API $0.07/MP for text-to-image and $0.07/MP for editing (≈$0.07 for a 1024² frame). 1 credit = $0.01.
Black Forest Labs FLUX.2 Klein ~4MP (≈2048x2048)$0.014–$0.015 at 1MP hosted; GPU cost if self-hostedPoor — worst of the FLUX tiersYes — same rate as generationPaid BFL tier required for commercial self-hosting Hosted: from $0.014/MP (Klein 4B) and $0.015/MP (Klein 9B), same rate for text-to-image and editing. Self-hosted: weights are free to download; you pay only GPU compute. Commercial self-hosting requires a BFL licence tier (Builder through Enterprise, with 10K–100K images/month allowances).
Midjourney V8.2 Native 2KNo per-image rate — $10–$120/mo GPU-hour plansImproved in V8 but behind leadersYes — V8.2 unified Edit ModelPaid subscribers; Pro/Mega above $1M revenue $10 Basic / $30 Standard / $60 Pro / $120 Mega per month; 20% off with annual billing ($8/$24/$48/$96 per month equivalent). Billed in GPU hours, not per image.
Ideogram 4.0 Native 2K$0.03 Turbo / $0.06 Default / $0.10 QualityBest in class for layout-controlled typeYes — edit and reference toolsPaid API/subscription; weights are non-commercial API: $0.03 Turbo, $0.06 Default, $0.10 Quality per image (Ideogram 4.0). Character-reference generation on 3.0 raises this to $0.10/$0.15/$0.20. Subscriptions: Plus $20/mo ($15 annual) for 1,000 priority credits, Pro $60/mo ($42 annual) for 3,500, Team $20/user/mo. Weights are free to download.
Recraft V4.1 Resolution-independent for vector; ~2K raster$0.035 raster / $0.088 vector / $0.33 Pro vectorGood — strongest for design/typographic layoutsYes — raster $0.002–$0.25, vector ops $0.088Paid plans only; free-plan images owned by Recraft API units are prepaid and non-expiring at 1,000 units per $1. V4.1: $0.035 raster, $0.088 vector, $0.21 Pro, $0.33 Pro Vector, $0.035 Utility, $0.21 Utility Pro. V4 Styles: $0.035, $0.055 vector, $0.10 Pro, $0.12 Pro Vector. Raster edits $0.002–$0.25. Subscriptions: Free, Basic, Pro, Teams with top-up credits.
ByteDance Seedream 5.0 Pro ~2.7K long edge (4K 16:9 rolling out)$0.045 (≤2.36MP) / $0.09 (above)Strong — best-in-class Chinese, good LatinYes — refs +$0.003 each after the firstPermitted under BytePlus terms $0.045 per image at ≤2.36MP, $0.09 above that. First reference image free, each additional +$0.003. Seedream 5.0 Lite $0.035. Third-party routes: ~$0.035 via Vercel AI Gateway, from $0.045 on OpenRouter.
Alibaba Qwen-Image-3.0 Not published~¥0.18 (~$0.025) reported, unverifiedStrong — small-text legibility is a stated focusSupported; details not documented publiclyHosted API terms; no weights licence Reported at approximately ¥0.18 per image (~$0.025) for Standard on Alibaba Model Studio; no verified USD rate card for Pro. Treat pricing as unconfirmed — Alibaba published no English-language rate card alongside the GA release.
xAI Grok Imagine Image 2.0 2048x2048 (2K)$0.02 standard; $0.05–$0.07 quality tierAdequate — behind GPT Image 2 and IdeogramYes — image input at $0.002/imagePermitted under xAI API terms grok-imagine-image: $0.02 per output image at 1K or 2K, $0.002 per input image. grok-imagine-image-quality: $0.05 (1K) / $0.07 (2K) output, $0.01 input. Grok Imagine Video is $0.05/second.
Luma Uni-1.1 2048px (2K) — the only tier priced$0.0404 (Uni-1.1) / $0.10 (Max) at 2KUnremarkable — not a stated strengthYes — up to 8 reference imagesPermitted under Luma API terms At 2048px (2K) — Uni-1.1: $0.0404 text-to-image, $0.0434 edit or 1 reference, $0.0464 with 2 references, $0.0644 with 8. Uni-1.1 Max: $0.1000 / $0.1030 / $0.1060 / $0.1240. Luma notes prices are approximate and token-derived. Legacy Photon was $0.015/image and Photon Flash $0.002.
Stability AI Stable Image Ultra / SD 3.5 ~1MP native for SD 3.5; upscaling billed separately$0.08 Ultra / $0.065 SD3.5 Large / $0.03 CorePoor — weakest of the major familiesYes — 5 credits ($0.05) per editFree under $1M revenue; Enterprise licence above 1 credit = $0.01. Stable Image Ultra $0.08, SD 3.5 Large $0.065, SD 3.5 Large Turbo $0.04, SD 3.5 Medium $0.035, Stable Image Core $0.03. Editing 5 credits ($0.05), upscaling 2–60 credits. API membership $20/month for 6,000 credits (750 Ultra or 2,000 Core images); overage at the same rate. New accounts get 25 free credits.
Adobe Firefly (Image Model 4 / 5) Up to 4K in Firefly web (model-dependent)Credit-based; no public API rate cardAdequate — behind GPT Image 2 and IdeogramYes — generative fill/expand across Adobe appsPermitted; enterprise IP indemnification offered Free $0/mo (25 generative credits). Standard $9.99/mo (2,000 credits). Pro Plus $49.99/mo. Premium $199.99/mo. Standard features cost 1 credit, premium generation 10–20 credits. Enterprise API (Firefly Services) has no published rate card; community-reported figures are roughly $0.02–$0.10/image with an approximate $1,000/month starting commitment — treat as unverified.
Reve 2.1 Native 4K (16MP full render)No API rate; consumer plans $7.99–$19.99/moVery good — layout-first architectureYes — structured editable layout in web appPermitted on paid consumer plans Consumer plans: Free tier, $7.99/mo and $19.99/mo. API pricing is no longer applicable — the API sunset on 14 August 2026 and unused credits were refunded before 31 August 2026. Third-party resellers still listing Reve 2.1 quote roughly $0.20–$0.25/image.
Google Imagen 4 (retired) n/a — discontinuedn/a (was $0.02 Fast / ~$0.04–$0.06 Standard)n/a — was mid-packn/an/a No longer purchasable. Historic reference: Imagen 4 Fast was $0.02/image. Migration target gemini-3.1-flash-image is $0.045 (0.5K) to $0.151 (4K).

C Entries

OpenAI GPT Image 2

OpenAI's current flagship image model (snapshot gpt-image-2-2026-04-21), also shipped in ChatGPT as Images 2.0. It plans and self-checks before rendering, which is why it tops the Artificial Analysis text-to-image leaderboard and renders legible text including non-Latin scripts. Sizes are flexible up to a 3840px long edge (655,360–8,294,400 total pixels, both dimensions multiples of 16, max 3:1 aspect), with low/medium/high/auto quality tiers. Supports masked editing, though the mask is treated as prompt guidance rather than a hard geometric boundary.

Max resolution3840px long edge (up to ~8.3MP)
Price per image~$0.01 / $0.03 / $0.12 (low/med/high at 1024²)
Text renderingBest in class; multi-script
Editing / inpaintingYes — mask-based and instruction edits
Commercial licencePermitted; user owns outputs
API availabilityPublic REST API, batch 50% off

Watch out: Token metering makes per-image cost hard to predict and budget: cost scales with resolution and quality tier, and the same call can vary. High quality at 4K is an order of magnitude more expensive than a Nano Banana 2 frame for a quality gap most viewers will not see, so it is the wrong default for bulk or background imagery. Outputs above 2560x1440 are flagged experimental by OpenAI's own docs. Content policy is the strictest of the major providers — expect refusals on real public figures, brand marks and mild suggestiveness that competitors allow.

Metered on tokens, not flat per image: $5.00/M text input, $8.00/M image input, $30.00/M image output; batch API is 50% off. At 1024x1024 that works out to roughly $0.01 low, $0.03 medium, $0.12 high. gpt-image-1.5 ($32/M output) and gpt-image-1-mini ($8/M output) remain listed as cheaper siblings.

Google Nano Banana 2 (Gemini 3.1 Flash Image)

Google's GA workhorse image model, served through the Gemini API as gemini-3.1-flash-image. Priced by output resolution with a published rate card — $0.045 at 512px, $0.067 at 1024², $0.101 at 2K, $0.151 at 4K — and halved again in batch mode, which makes it the cheapest credible route to 4K output. Handles conversational multi-turn editing and video-to-image generation natively. A Lite variant (gemini-3.1-flash-lite-image) runs $0.0336 at 1K for latency-sensitive work.

Max resolution4096x4096 (4K)
Price per image$0.067 at 1K; $0.151 at 4K
Text renderingGood; weaker on dense small type
Editing / inpaintingYes — conversational multi-turn edits
Commercial licencePermitted; SynthID watermark applied
API availabilityGemini API + Vertex; batch 50% off

Watch out: Google deprecates aggressively — this is the model that replaced Imagen 4, which itself was killed on the Gemini API in August 2026 after roughly a year. Expect to re-pin model IDs annually. Text rendering is good but trails GPT Image 2 and Ideogram on dense small type. SynthID watermarking is applied to outputs and cannot be disabled. Content filters are configurable but the defaults refuse a lot of person-related generation, and the filter behaviour has changed between snapshots without notice.

$0.045 (0.5K) / $0.067 (1K) / $0.101 (2K) / $0.151 (4K) per image; input $0.50/M tokens; batch pricing is 50% off across all tiers. Lite: $0.0336 at 1K.

Google Nano Banana Pro (Gemini 3 Pro Image)

The Pro rung of Google's image ladder, served as gemini-3-pro-image. Flat $0.134 whether you render at 1K or 2K, and $0.24 at 4K, against $2.00/M input tokens. It buys better prompt adherence on long, compositional prompts and more reliable text than the Flash tier, and it is the model to use when a single hero asset justifies 2x the cost. Same editing and reference-image workflow as Nano Banana 2.

Max resolution4096x4096 (4K)
Price per image$0.134 at 1K/2K; $0.24 at 4K
Text renderingVery good; better than Flash tier
Editing / inpaintingYes — instruction edits, multi-reference
Commercial licencePermitted; SynthID watermark applied
API availabilityGemini API + Vertex; batch 50% off

Watch out: At $0.134 it is roughly 2x Nano Banana 2 and 13x Meta's Muse Image for a quality margin that is visible on hard prompts and invisible on easy ones — using it as a default is a straightforward waste of money. It is a Gemini 3 generation model while the Flash tier has already moved to 3.1, so the Pro rung is the older architecture and a successor is plausibly imminent. Same SynthID watermarking and content-filter caveats as the Flash tier. No self-hosting path at any price.

$0.134 per image at 1K or 2K; $0.24 at 4K; input $2.00/M tokens; batch 50% off.

Microsoft MAI-Image-2.6

Released 4 September 2026, MAI-Image-2.6 is Microsoft AI's own image generation and editing model — not an OpenAI relabel. It took #2 on Arena text-to-image out of 77 systems and #1 on Artificial Analysis image editing, which makes it the strongest option specifically for edit-heavy work. Accepts up to five reference images per request, supports web grounding for current visual context, and offers dynamic aspect ratios up to 1.5K resolution. A Flash variant generates ~2.8x faster than GPT-Image-2-Medium.

Max resolution~1.5K, dynamic aspect ratios
Price per image$38/M output tokens; no per-image rate card
Text renderingStrong; #1 on AA editing benchmark
Editing / inpaintingYes — up to 5 reference images
Commercial licencePermitted under Azure/Foundry terms
API availabilityMicrosoft Foundry (public preview)

Watch out: Capped at roughly 1.5K resolution, so it is the wrong choice if you need 4K deliverables — it loses to Nano Banana 2 and GPT Image 2 on that axis outright. Still in public preview on Microsoft Foundry as of early September 2026, meaning no stability guarantees and no SLA. Barely a day old at the time of writing, so third-party tooling, LoRA-equivalents and community prompt knowledge effectively do not exist yet. Per-token pricing is sourced from aggregators rather than a Microsoft rate card.

$5.00/M text input, $8.00/M image input, $38.00/M image output (as listed on aggregators; Microsoft's launch post cites price-per-Elo leadership but publishes no per-image rate card). Free to test in MAI Playground and on Arena. · beta

Meta Muse Image

Meta Superintelligence Labs' first image model (codenamed Mango), launched to consumers on 7 July 2026 and to developers on the Meta Model API on 25 August 2026 at $0.01 per image. It reasons about layout before drawing, exposes three primitives — generate, edit, compose — and supports anchored composition from reference images for character and style consistency across a series. Output tops out around 1600px. Also available through fal since 1 September 2026.

Max resolution~1600px
Price per image$0.01
Text renderingInconsistent — varies run to run
Editing / inpaintingYes — edit and multi-image compose
Commercial licencePermitted; terms not detailed in docs
API availabilityMeta Model API (dev.meta.ai) and fal

Watch out: 1600px ceiling rules it out for print or any 2K/4K deliverable; you are buying price, not resolution. Meta's own developer documentation concedes that exact text rendering varies run to run and advises keeping baked-in text short and re-running when labels are unclear — so it is a poor fit for typography-critical work. Multi-input composition can reflow the layout rather than reproduce inputs faithfully. Every call is non-deterministic. Meta drew immediate public pushback in July over training on users' photos, which is a reputational consideration for consumer-facing brands. Commercial terms and rate limits are not spelled out in the developer docs.

$0.01 per image at production volumes on Meta Model API. Free for consumers in the Meta AI app, WhatsApp DMs and Instagram Stories.

Black Forest Labs FLUX.2 Pro

The mainstream tier of BFL's FLUX.2 family and the reference point for European-hosted image generation. Billed strictly by megapixel — $0.03/MP for the first megapixel of text-to-image output, $0.015/MP thereafter, $0.045/MP for editing, and $0.015/MP for reference-image input — which makes cost fully predictable from output dimensions. Maximum output around 4MP with a ~47K context. Sits between Klein (cheap, open weights) and Max ($0.07/MP) in BFL's ladder, with Flex at $0.05/MP offering per-step control.

Max resolution~4MP (≈2048x2048)
Price per image$0.03 at 1MP; $0.045/MP for edits
Text renderingWeak — headlines only
Editing / inpaintingYes — FLUX Tools, Kontext-style edits
Commercial licencePermitted on hosted API
API availabilityBFL API, plus Replicate/fal/OpenRouter

Watch out: Text rendering is the family's persistent weakness — usable for a headline, unreliable for paragraphs, small labels or non-Latin scripts, and clearly behind GPT Image 2, Ideogram and Nano Banana. Megapixel billing means high-resolution output gets expensive fast and reference images are charged as input, so multi-reference edits cost noticeably more than the headline rate suggests. Routing through Replicate or fal roughly doubles the price. BFL's attention has visibly shifted to FLUX 3 video and robotics; FLUX 3 Image was promised for 'coming weeks' after the 23 July 2026 announcement and had not shipped as of early September.

$0.03/MP text-to-image (≈$0.03 for a 1024² frame), $0.045/MP editing, $0.015/MP reference-image input, $0.015/MP for each MP beyond the first. Flex $0.05/MP, Max $0.07/MP. 1 credit = $0.01. Roughly $0.055/image on Replicate and $0.073 on fal if you route through a reseller.

Black Forest Labs FLUX.2 Max

The highest-quality rung of FLUX.2, priced at $0.07/MP for both text-to-image and editing (unlike Pro, where edits carry a premium over generation). Worth the step up when Pro's prompt adherence breaks down on dense compositional prompts or when edit fidelity on complex scenes matters. Same API surface, same megapixel billing model, same ~4MP ceiling as the rest of the FLUX.2 family.

Max resolution~4MP (≈2048x2048)
Price per image$0.07 at 1MP (generation and editing)
Text renderingWeak — same family limitation
Editing / inpaintingYes — same price as generation
Commercial licencePermitted on hosted API
API availabilityBFL API, plus resellers

Watch out: At $0.07/MP it is more than 2x Pro and roughly the same as Nano Banana Pro at 1K, without matching Google's text rendering or 4K support — so it wins only if you are already committed to FLUX's aesthetic and tooling. Shares the family's text-rendering weakness, which is the single most common reason to leave FLUX, and paying 2x does not fix it. Hosted-only: there is no open-weight Max, so this tier gives you none of FLUX's self-hosting advantage. Not worth using for drafts or iteration.

$0.07/MP for text-to-image and $0.07/MP for editing (≈$0.07 for a 1024² frame). 1 credit = $0.01.

Black Forest Labs FLUX.2 Klein

The small, open-weight end of FLUX.2, shipped in 4B and 9B parameter variants. On BFL's hosted API it is $0.014/MP (4B) and $0.015/MP (9B) — under half the cost of Pro — and the same weights can be downloaded and run locally, which is the reason most teams choose FLUX over Google or OpenAI at all. Practical on a single consumer GPU at the 4B size. This is the model to reach for when data cannot leave your infrastructure or when per-image API economics stop working at your volume.

Max resolution~4MP (≈2048x2048)
Price per image$0.014–$0.015 at 1MP hosted; GPU cost if self-hosted
Text renderingPoor — worst of the FLUX tiers
Editing / inpaintingYes — same rate as generation
Commercial licencePaid BFL tier required for commercial self-hosting
API availabilityBFL API + downloadable weights

Watch out: 'Open weights' is not open source: commercial self-hosting needs a paid BFL licence with monthly image caps, so the free-download story does not survive contact with a revenue-generating product. Output quality at 4B is a real step down from Pro — noticeably weaker on hands, complex scenes and anything compositional — and it inherits the family's poor text rendering, only worse. Self-hosting means you own the inference stack, scaling, safety filtering and NSFW mitigation yourself, which is a meaningful engineering cost people routinely underestimate when comparing against a $0.03 API call.

Hosted: from $0.014/MP (Klein 4B) and $0.015/MP (Klein 9B), same rate for text-to-image and editing. Self-hosted: weights are free to download; you pay only GPU compute. Commercial self-hosting requires a BFL licence tier (Builder through Enterprise, with 10K–100K images/month allowances).

Midjourney V8.2

V8.0 launched as an alpha on 17 March 2026, V8.1 on 14 April, and V8.2 became the default on 24 July 2026. V8 delivers native 2K output and generates roughly 4–5x faster than V7. V8.2 focuses on aesthetics and personalisation and introduces a unified Edit Model that replaces the previous Omni Reference, Character Reference and Retexture tools. Sold purely as a subscription with GPU-hour allowances: Basic $10, Standard $30, Pro $60, Mega $120 per month, 20% off annually. Relax (unmetered slow) mode from Standard up; Stealth mode on Pro and Mega only.

Max resolutionNative 2K
Price per imageNo per-image rate — $10–$120/mo GPU-hour plans
Text renderingImproved in V8 but behind leaders
Editing / inpaintingYes — V8.2 unified Edit Model
Commercial licencePaid subscribers; Pro/Mega above $1M revenue
API availabilityNone official; ToS bans third-party wrappers

Watch out: There is still no official public API in September 2026, no REST endpoint, no SDK and no documented API key flow — and Midjourney's ToS explicitly prohibits automated access and reselling, which it enforces. Building a product on PiAPI or similar wrappers means building on a violation. GPU-hour billing makes per-image cost opaque compared to every metered competitor. On Basic and Standard plans your generations are public by default; privacy costs $60/month minimum. Text rendering, while improved in V8, is still behind GPT Image 2 and Ideogram. Capped at 2K, so no 4K deliverables.

$10 Basic / $30 Standard / $60 Pro / $120 Mega per month; 20% off with annual billing ($8/$24/$48/$96 per month equivalent). Billed in GPU hours, not per image.

Ideogram 4.0

Released 3 June 2026, Ideogram 4.0 is a 9.3B-parameter diffusion transformer trained from scratch — not a fine-tune — and the company's first open-weight release, with nf4 and fp8 quantisations on Hugging Face that fit a 24GB GPU. Its distinguishing feature is a structured JSON prompting interface with explicit bounding-box layout control and colour-palette specification, which makes it the only model here where you can place text and elements by coordinate rather than by describing them. Native 2K output, best-in-class multilingual text rendering. Topped the open-weight DesignArena leaderboard at launch and placed ninth overall in text-to-image.

Max resolutionNative 2K
Price per image$0.03 Turbo / $0.06 Default / $0.10 Quality
Text renderingBest in class for layout-controlled type
Editing / inpaintingYes — edit and reference tools
Commercial licencePaid API/subscription; weights are non-commercial
API availabilityPublic API; 10 in-flight request default limit

Watch out: The open-weight release is a trap if you skim the headline: the licence is non-commercial, so self-hosting for anything revenue-generating requires a separate paid licence from Ideogram, and the announcement's 'commercial license' framing conflicts with the model agreement's actual terms — verify before you build. Capped at native 2K, no 4K. Photorealism and general aesthetic quality trail Midjourney, FLUX and the frontier multimodal models — it is a design tool, not a photography tool. API pricing is not published on Ideogram's own developer docs (the pricing page renders prices client-side), so the figures above come from third-party aggregators. Default rate limit is only 10 in-flight requests.

API: $0.03 Turbo, $0.06 Default, $0.10 Quality per image (Ideogram 4.0). Character-reference generation on 3.0 raises this to $0.10/$0.15/$0.20. Subscriptions: Plus $20/mo ($15 annual) for 1,000 priority credits, Pro $60/mo ($42 annual) for 3,500, Team $20/user/mo. Weights are free to download.

Recraft V4.1

Recraft is the only model here that outputs real, editable SVG rather than a raster image traced afterwards, which is why it persists in brand and design workflows despite unremarkable photorealism. V4.1 shipped May 2026 alongside the V4 Styles family. Pricing runs $0.035 for V4.1 raster, $0.088 for V4.1 Vector, $0.21 for V4.1 Pro and $0.33 for V4.1 Pro Vector; the V4 Styles line is cheaper at $0.035–$0.12. Also supports custom style training, so you can lock a brand look and reproduce it across assets. Raster editing operations range $0.002–$0.25 per call.

Max resolutionResolution-independent for vector; ~2K raster
Price per image$0.035 raster / $0.088 vector / $0.33 Pro vector
Text renderingGood — strongest for design/typographic layouts
Editing / inpaintingYes — raster $0.002–$0.25, vector ops $0.088
Commercial licencePaid plans only; free-plan images owned by Recraft
API availabilityPublic API, prepaid non-expiring units

Watch out: Photorealism is well behind FLUX, Midjourney and the frontier models — do not use it for photographic imagery. Vector output at $0.088–$0.33 is 3–10x the cost of a raster frame from any competitor, and generated SVG path structures are frequently messy enough that a designer ends up cleaning them by hand, which undercuts the core value proposition. The free tier is a licensing hazard: Recraft owns those images outright, so anything a team member generates while evaluating on the free plan cannot be shipped. Prepaid unit model means you commit cash before you know your volume.

API units are prepaid and non-expiring at 1,000 units per $1. V4.1: $0.035 raster, $0.088 vector, $0.21 Pro, $0.33 Pro Vector, $0.035 Utility, $0.21 Utility Pro. V4 Styles: $0.035, $0.055 vector, $0.10 Pro, $0.12 Pro Vector. Raster edits $0.002–$0.25. Subscriptions: Free, Basic, Pro, Teams with top-up credits.

ByteDance Seedream 5.0 Pro

Released 12 August 2026 and served through ByteDance's ModelArk / BytePlus API, Seedream 5.0 Pro is a combined generation and editing model with two flat price tiers: $0.045 at or below 2.36 megapixels, $0.09 above. Reference images are nearly free — the first is included, each additional one is $0.003 — which makes multi-reference composition unusually cheap compared to megapixel-metered competitors. Maximum output at launch is roughly 2.7K on the long edge, with 4K at 16:9 arriving shortly after. A Lite variant runs $0.035.

Max resolution~2.7K long edge (4K 16:9 rolling out)
Price per image$0.045 (≤2.36MP) / $0.09 (above)
Text renderingStrong — best-in-class Chinese, good Latin
Editing / inpaintingYes — refs +$0.003 each after the first
Commercial licencePermitted under BytePlus terms
API availabilityModelArk/BytePlus; OpenRouter, Vercel, fal

Watch out: ByteDance ownership is a hard blocker for many Western enterprises on data-residency and geopolitical grounds — check with legal before you prototype, not after. Documentation and console are China-first; the BytePlus international surface lags the domestic Volcengine one on both features and docs quality. Resolution tops out around 2.7K long edge at launch, so it is not a 4K option today. The two-tier flat pricing means a 2.4MP image costs exactly double a 2.3MP one — a cliff worth designing around. Community tooling and prompt knowledge outside Chinese-language forums is thin.

$0.045 per image at ≤2.36MP, $0.09 above that. First reference image free, each additional +$0.003. Seedream 5.0 Lite $0.035. Third-party routes: ~$0.035 via Vercel AI Gateway, from $0.045 on OpenRouter.

Alibaba Qwen-Image-3.0

Announced 21 July 2026 and generally available from 5 August, Qwen-Image-3.0 ships in Pro and Standard editions through Alibaba's hosted API. Its two genuine strengths are handling very long prompts without dropping instructions and rendering small text legibly — the two failure modes most image models share. The significant story is licensing: Qwen-Image 1.0 shipped in August 2025 under Apache 2.0 with a same-day technical report, 2.0 followed with its own report, and 3.0 arrived with no weights, no licence, no technical report and no official benchmarks. Alibaba closed its flagship image line.

Max resolutionNot published
Price per image~¥0.18 (~$0.025) reported, unverified
Text renderingStrong — small-text legibility is a stated focus
Editing / inpaintingSupported; details not documented publicly
Commercial licenceHosted API terms; no weights licence
API availabilityAlibaba Model Studio hosted API only

Watch out: The open-weight regression is the headline problem: if you adopted Qwen-Image because it was Apache 2.0, 3.0 is not a continuation of that bargain and there is no self-hosting path. No technical report and no official benchmarks means quality claims are unverifiable — you are trusting vendor marketing. Pricing is not published in a form I could verify against an official English rate card, so budget with caution. Same China-vendor data-residency and procurement concerns as Seedream. Documentation is China-first and international access routing is inconsistent.

Reported at approximately ¥0.18 per image (~$0.025) for Standard on Alibaba Model Studio; no verified USD rate card for Pro. Treat pricing as unconfirmed — Alibaba published no English-language rate card alongside the GA release.

xAI Grok Imagine Image 2.0

xAI's image model, with Image 2.0 (grok-imagine-image-2.0) launched 7 August 2026 and API access the following day. The standard tier is a flat $0.02 per output image at either 1024² or 2048², with input images at $0.002 — no megapixel metering, no quality-tier pricing cliff, which makes cost modelling trivial. A grok-imagine-image-quality variant runs $0.05 at 1K and $0.07 at 2K. Its practical differentiator is a markedly more permissive content policy than OpenAI, Google or Adobe.

Max resolution2048x2048 (2K)
Price per image$0.02 standard; $0.05–$0.07 quality tier
Text renderingAdequate — behind GPT Image 2 and Ideogram
Editing / inpaintingYes — image input at $0.002/image
Commercial licencePermitted under xAI API terms
API availabilityPublic xAI API

Watch out: Output quality trails the leaderboard leaders — this is a value model, not a frontier one, and it shows on complex compositions and text. Capped at 2K, so no 4K path. The permissive content policy that makes it useful also makes it a brand-safety liability: it will generate likenesses and content that would get you refused elsewhere, and 'the API allowed it' is not a defence. xAI has a track record of shipping fast and revising policy publicly afterwards, so terms may shift under you. Documentation and SDK ecosystem are thinner than OpenAI's or Google's.

grok-imagine-image: $0.02 per output image at 1K or 2K, $0.002 per input image. grok-imagine-image-quality: $0.05 (1K) / $0.07 (2K) output, $0.01 input. Grok Imagine Video is $0.05/second.

Luma Uni-1.1

Luma has quietly replaced the Photon and Photon Flash line with Uni-1.1, a unified image generation and editing model that no longer appears alongside Photon on the current API pricing page. Priced at 2048px: $0.0404 text-to-image, $0.0434 for an edit or one reference image, scaling to $0.0644 with eight references. The Max variant runs $0.10 / $0.103 / $0.124 across the same operations. The reference-image pricing curve is explicit and shallow, which makes multi-reference workflows cheap to model. Sits alongside Ray3.2 for video in the same API.

Max resolution2048px (2K) — the only tier priced
Price per image$0.0404 (Uni-1.1) / $0.10 (Max) at 2K
Text renderingUnremarkable — not a stated strength
Editing / inpaintingYes — up to 8 reference images
Commercial licencePermitted under Luma API terms
API availabilityLuma API (also MCP server)

Watch out: Luma's centre of gravity is video, and image is the secondary product — the Photon-to-Uni transition happened with essentially no announcement, the changelog does not mention Uni-1.1 at all, and the docs still surface a two-year-old Photon post. That is a bad sign for a dependency. Pricing is published only at 2048px, so other resolutions require guesswork, and Luma itself calls the figures approximate. It does not appear near the top of any current text-to-image leaderboard. If you are not already buying Luma video, there is little reason to choose it over Nano Banana 2 at a comparable price with far better documentation.

At 2048px (2K) — Uni-1.1: $0.0404 text-to-image, $0.0434 edit or 1 reference, $0.0464 with 2 references, $0.0644 with 8. Uni-1.1 Max: $0.1000 / $0.1030 / $0.1060 / $0.1240. Luma notes prices are approximate and token-derived. Legacy Photon was $0.015/image and Photon Flash $0.002.

Stability AI Stable Image Ultra / SD 3.5

Stability's hosted API is straightforward flat-credit pricing at 1 credit = $0.01: Stable Image Ultra $0.08, SD 3.5 Large $0.065, SD 3.5 Large Turbo $0.04, SD 3.5 Medium $0.035, Stable Image Core $0.03, with no separate input/output charge and no resolution-based variation. Its enduring value is the SD 3.5 open-weight family (8B Large, 8B Turbo, 2.5B Medium) and the enormous ecosystem of LoRAs, ControlNets, ComfyUI workflows and fine-tunes built on Stable Diffusion — nothing else in this list has comparable community tooling depth. The company raised a Series B in August 2026 and is stable, but its 2026 model release was Stable Audio 3.0, not an image model.

Max resolution~1MP native for SD 3.5; upscaling billed separately
Price per image$0.08 Ultra / $0.065 SD3.5 Large / $0.03 Core
Text renderingPoor — weakest of the major families
Editing / inpaintingYes — 5 credits ($0.05) per edit
Commercial licenceFree under $1M revenue; Enterprise licence above
API availabilityPublic REST API + downloadable weights

Watch out: The flagship open image models are SD 3.5 from October 2024 — nearly two years stale, and it shows: prompt adherence, text rendering and anatomy are all clearly behind FLUX.2, let alone the frontier multimodal models. Stability's own news page lists no 2026 image release, and the widely circulated 'Stable Diffusion 4' claims are not corroborated by any Stability announcement, so do not plan around one. Stable Image Ultra at $0.08 is more expensive than Nano Banana 2 at 1K while producing weaker output — the hosted API is hard to justify on merit. Text rendering is poor. The $1M revenue threshold on the Community License is a cliff that catches growing companies by surprise.

1 credit = $0.01. Stable Image Ultra $0.08, SD 3.5 Large $0.065, SD 3.5 Large Turbo $0.04, SD 3.5 Medium $0.035, Stable Image Core $0.03. Editing 5 credits ($0.05), upscaling 2–60 credits. API membership $20/month for 6,000 credits (750 Ultra or 2,000 Core images); overage at the same rate. New accounts get 25 free credits.

Adobe Firefly (Image Model 4 / 5)

Firefly's differentiator is not quality — it is provenance. Adobe trains on Adobe Stock and licensed content and offers enterprise customers indemnification against IP claims arising from generated output, which no other vendor here matches contractually. As of August 2026 the app runs on Firefly Image Model 4 by default with Model 5 available as an option, and Image Model 3 was fully retired in August 2026. Firefly has also become an aggregator, exposing 30+ third-party image and video models inside one credit-based subscription, and added custom style models trained on your own work in March 2026.

Max resolutionUp to 4K in Firefly web (model-dependent)
Price per imageCredit-based; no public API rate card
Text renderingAdequate — behind GPT Image 2 and Ideogram
Editing / inpaintingYes — generative fill/expand across Adobe apps
Commercial licencePermitted; enterprise IP indemnification offered
API availabilityFirefly Services, enterprise plans only

Watch out: Output quality is consistently mid-pack — Firefly does not appear near the top of any current text-to-image leaderboard, and the licensed-data constraint is a real quality ceiling, not just a marketing position. The credit system is genuinely confusing: 25 free credits buys about six high-quality images, and the 10–20 credit cost of premium generation means headline credit counts overstate what you get. There is no public API rate card at all, and enterprise access reportedly requires a four-figure monthly commitment, which puts it out of reach for small teams and makes cost comparison impossible. Adobe retires models on its own schedule inside the apps (Image 3 was pulled from Photoshop's picker in April 2026 and killed in August), so reproducibility across time is not guaranteed.

Free $0/mo (25 generative credits). Standard $9.99/mo (2,000 credits). Pro Plus $49.99/mo. Premium $199.99/mo. Standard features cost 1 credit, premium generation 10–20 credits. Enterprise API (Firefly Services) has no published rate card; community-reported figures are roughly $0.02–$0.10/image with an approximate $1,000/month starting commitment — treat as unverified.

Reve 2.1

Reve 2.1 shipped 9 July 2026 and is genuinely excellent: it plans an image as a structured, editable layout before rendering, outputs native 4K at a full 16-megapixel render, and reached #2 on Arena's text-to-image board at 1306 Elo. On 27 July 2026 OpenAI announced an investment in Reve and members of the AI research team moved to OpenAI to lead multimodal research; Reve stated it would continue operating independently with its products intact. Two and a half weeks later, on 14 August 2026, the API was sunset and unused credits were refunded by 31 August. The consumer web product at reve.com remains live.

Max resolutionNative 4K (16MP full render)
Price per imageNo API rate; consumer plans $7.99–$19.99/mo
Text renderingVery good — layout-first architecture
Editing / inpaintingYes — structured editable layout in web app
Commercial licencePermitted on paid consumer plans
API availabilityDISCONTINUED — sunset 14 August 2026

Watch out: The API is gone. Do not start an integration — any third-party endpoint still advertising Reve 2.1 is reselling capacity that no longer has a first-party source, and pricing there is 5x what comparable models cost direct. The wider signal is worse than the single shutdown: an investment announcement promising continuity was followed by an API sunset with roughly three weeks' notice, so treat public reassurances from single-product image startups as non-binding. Even as a web tool, the departure of the research team to OpenAI makes future model releases uncertain, and there is no self-hosting fallback.

Consumer plans: Free tier, $7.99/mo and $19.99/mo. API pricing is no longer applicable — the API sunset on 14 August 2026 and unused credits were refunded before 31 August 2026. Third-party resellers still listing Reve 2.1 quote roughly $0.20–$0.25/image. · deprecated

Google Imagen 4 (retired)

Included here because Imagen was one of the most-integrated image APIs of 2025 and a large number of production systems still reference it. Google deprecated the Imagen 3 and 4 endpoints on Vertex AI on 24 March 2026 and shut down imagen-4.0-generate-001, imagen-4.0-ultra-generate-001 and imagen-4.0-fast-generate-001 on the Gemini API on 17 August 2026. The named replacement is gemini-3.1-flash-image (Nano Banana 2), and Google states that Imagen prompt structures transfer directly to the Nano Banana models, which use the same conventions and add editing on top. Imagen 4 Fast was $0.02; the closest successor tier is around $0.039–$0.067.

Max resolutionn/a — discontinued
Price per imagen/a (was $0.02 Fast / ~$0.04–$0.06 Standard)
Text renderingn/a — was mid-pack
Editing / inpaintingn/a
Commercial licencen/a
API availabilityDISCONTINUED — Gemini API 17 Aug 2026, Vertex 24 Mar 2026

Watch out: Completely unavailable — calls to the Imagen 4 model IDs fail. If you have Imagen endpoints in production they are already broken. The migration is not purely mechanical despite Google's claim that prompts transfer: Nano Banana applies SynthID watermarking, has different aspect-ratio and resolution handling, different content-filter behaviour, and costs roughly 2–3x Imagen 4 Fast per image, so budget and QA both need revisiting. The 17-month lifespan of a flagship Google image API is the single best argument in this chapter for putting an abstraction layer between your product and any image model.

No longer purchasable. Historic reference: Imagen 4 Fast was $0.02/image. Migration target gemini-3.1-flash-image is $0.045 (0.5K) to $0.151 (4K). · discontinued