FLUX Image API Cost Per Image: 2026 Provider Guide
Connor Blier
Founding GTM
FLUX Image API Cost Per Image: 2026 Provider Guide
FLUX image API cost per image runs from about $0.0005 for a distilled schnell/klein tier to $0.07 for a premium max tier at 1024x1024. Between them, dev models land near $0.009-$0.030 and pro near $0.015-$0.055, depending on whether the provider bills per image, per megapixel or per GPU-second.
The question sounds simple: what does one FLUX image cost through an API? The answer depends on which FLUX tier you call, which provider serves it, and how that provider meters the work. This guide normalizes everything to a single 1024x1024 image so you can compare like for like, then pairs each price with the latency you should expect before you commit a production workload.
This is a commercial, provider-selection guide for ML engineers choosing where to ship. It optimizes for two axes at once: cost per image and latency.
TL;DR
Cheapest tier (distilled schnell/klein): roughly $0.0005 (DeepInfra FLUX-1-schnell) to $0.003 (Replicate FLUX.1 [schnell]) per 1024x1024 image.
Dev tier (balanced quality): ~$0.009 (DeepInfra FLUX-1-dev) to $0.030 (Replicate FLUX.1 [dev]).
Pro tier: ~$0.015 (DeepInfra FLUX-2-pro) to $0.055 (Replicate FLUX.1 [pro]).
Premium/max tier: ~$0.07 (DeepInfra FLUX-2-max).
Latency: a dev-class 1024px image lands in ~3.3s on Modular Cloud and ~6.75s baseline with vanilla diffusers, optimizable to under 3s (Modal). FLUX.2 has been demonstrated in sub-second on Modular.
How FLUX API pricing works
There are three billing models you will meet, and they do not compare directly until you normalize them.
Flat per-image. You pay a fixed amount per generation regardless of resolution or step count. Replicate lists FLUX.1 [pro] at "$0.055 per image," FLUX.1 [dev] at "$0.030 per image," and FLUX.1 [schnell] at "0.003 per image." DeepInfra also uses flat rates for some tiers: FLUX-2-max at $0.07 and FLUX-2-pro at $0.015.
Metered per-megapixel / per-iteration. The price scales with the pixels you render and the diffusion steps you run. DeepInfra prices FLUX-2-dev as "$0.01 x (w / 1024) x (h / 1024) x (iters / 28)", FLUX-1-dev as "$0.009 x (w / 1024) x (h / 1024) x (iters / 25)", and FLUX-1-schnell as "$0.0005 x (w / 1024) x (h / 1024) x iters". The klein variants scale by resolution only: FLUX-2-klein-9b at "$0.015 x (w / 1024) x (h / 1024)" and FLUX-2-klein-4b at "$0.014 x (w / 1024) x (h / 1024)".
GPU-second (self-hosted / serverless). You rent the GPU and pay for the wall-clock time each image takes. Here your cost per image is a function of throughput, not a sticker price. We cover the general economics of this model in our LLM inference cost at scale guide and the 2026 GPU buyer's guide, and the underlying serverless pricing shapes apply directly to image generation.
Normalization convention used below: every metered price is evaluated at w=1024, h=1024 and the provider's default step count, so all figures are per single 1024x1024 image.
Cost per image by FLUX tier
| Tier | Model (provider) | Cost per 1024x1024 | Billing basis |
|---|---|---|---|
| Speed / cheap | FLUX-1-schnell (DeepInfra) | ~$0.0005 | Metered (res x iters) |
| Speed / cheap | FLUX.1 [schnell] (Replicate) | $0.003 | Flat per-image |
| Balanced / dev | FLUX-1-dev (DeepInfra) | ~$0.009 | Metered (res x iters/25) |
| Balanced / dev | FLUX-2-dev (DeepInfra) | ~$0.01 | Metered (res x iters/28) |
| Balanced / dev | FLUX.1 [dev] (Replicate) | $0.030 | Flat per-image |
| Compact | FLUX-2-klein-4b (DeepInfra) | ~$0.014 | Metered (res only) |
| Compact | FLUX-2-klein-9b (DeepInfra) | ~$0.015 | Metered (res only) |
| Pro | FLUX-2-pro (DeepInfra) | $0.015 | Flat per-image |
| Pro | FLUX.1 [pro] (Replicate) | $0.055 | Flat per-image |
| Premium / max | FLUX-2-max (DeepInfra) | $0.07 | Flat per-image |
The spread is nearly 140x from the cheapest distilled tier to the premium tier. The obvious takeaway: pick the lowest tier that clears your quality bar, because tier choice dominates every other cost lever.
Latency and throughput
Cost only tells half the story. An image that costs $0.0005 but returns in 10 seconds may be unusable in an interactive product, while a slightly pricier tier that returns in a couple of seconds pays for itself in conversion.
Modal established the interactive threshold plainly: "in order to be competitive with APIs serving FLUX.1-dev images, we needed to return results in under three seconds." Their baseline with the standard Hugging Face diffusers library was ~6.75 seconds for a 1024x1024 image, which they cut roughly 3x (1.5x from compiler/hardware optimization, 2x from First Block Caching) to reach that sub-3-second target.
On the production side, Modular Cloud's Flux2-dev endpoint generated a 1024px, 25-step illustration in ~3.3 seconds measured over 24 hours of real traffic, and Modular has demonstrated FLUX.2 producing "good quality images in sub-second" with configurable quality tolerance.
On the self-hosted path, your latency and cost per image both come down to platform roundtrip and throughput. In our testing of LLM serving on Cerebrium hardware we measured a lowest TTFT of 123ms (vLLM) and 460 tokens/sec throughput (SGLang, batch 64) on a single H100 as the total roundtrip to the Cerebrium platform, not a vendor spec sheet number. The same discipline of measuring end-to-end roundtrip rather than trusting published figures is what you should apply to any FLUX endpoint before you trust its cost-per-image math. See the full methodology in our vLLM/SGLang/TensorRT benchmark and our network latency in the inference pipeline guide, since network hops often dwarf model compute.
Total cost at scale
Here is what tier choice does to a monthly bill at 100,000 images/month, all at 1024x1024:
Distilled (schnell/klein): DeepInfra FLUX-1-schnell ~$50; Replicate FLUX.1 [schnell] ~$300.
Dev: DeepInfra FLUX-1-dev ~$900; DeepInfra FLUX-2-dev ~$1,000; Replicate FLUX.1 [dev] ~$3,000.
Pro: DeepInfra FLUX-2-pro ~$1,500; Replicate FLUX.1 [pro] ~$5,500.
Max: DeepInfra FLUX-2-max ~$7,000.
At 1,000,000 images/month, multiply by ten: the gap between a distilled tier and a premium tier is the difference between a $500 line item and a $70,000 one. At that volume the GPU-second/serverless path becomes worth modeling, because sustained utilization can undercut per-image APIs, an effect we quantify in our inference cost at scale analysis and the case study on DistilLabs' 50% lower inference costs, which sustained up to 150 requests per second per model at peak on autoscaling infrastructure.
Choosing a provider
Interactive consumer app (chat, editors): prioritize latency. A dev tier at ~3.3s (Modular) or an optimized sub-3s endpoint (Modal) beats a cheaper-but-slower option. Pair with a low-cold-start serverless platform.
High-volume batch (thumbnails, variations): prioritize cost. A distilled tier at ~$0.0005-$0.003 wins outright; latency is hidden behind a queue.
Quality-critical (marketing, hero assets): a pro or max tier justifies its price on output quality, not throughput.
Unpredictable/spiky traffic: a metered or serverless model that scales to zero avoids paying for idle GPUs; see serverless cost economics and, if you are leaving hyperscalers, our AWS alternatives for AI workloads.
Normalize every quote to per-1024x1024, measure real roundtrip latency, then let your use case pick the axis.
Frequently asked questions
- What is the cheapest FLUX image API per image?
- The cheapest published rate is DeepInfra's FLUX-1-schnell at roughly $0.0005 per 1024x1024 image (metered by resolution and iterations). Replicate's FLUX.1 [schnell] is a flat $0.003 per image.
- How much does FLUX pro cost per image?
- It varies by provider: DeepInfra prices FLUX-2-pro at a flat $0.015 per image, while Replicate lists FLUX.1 [pro] at $0.055 per image. The premium FLUX-2-max tier on DeepInfra is $0.07.
- How is FLUX API pricing calculated?
- Three ways: flat per-image (e.g. Replicate), metered per-megapixel and per-iteration normalized to 1024x1024 and default steps (e.g. DeepInfra's dev and schnell tiers), or GPU-second billing when you self-host on serverless infrastructure.
- How fast is a FLUX image generated?
- A 1024px, 25-step dev image was measured at ~3.3s on Modular Cloud in production. Vanilla diffusers baseline is ~6.75s (Modal), optimizable to under 3s. FLUX.2 has been shown generating good-quality images in sub-second on Modular.
- Should I use a per-image API or self-host FLUX?
- Per-image APIs are simplest and cheapest at low-to-moderate volume. At high sustained volume, GPU-second serverless can undercut per-image pricing because you pay for utilization rather than a fixed sticker price per generation.
Get started with Cerebrium
Deploy AI models on serverless GPUs in minutes, with no infrastructure to manage. Start for free and pay only for the compute you use.
Sources
- replicate.com
“FLUX.1 [pro] is $0.055 per image.”
Replicate flat per-image price for the pro tier.
- replicate.com
“FLUX.1 [dev] is $0.030 per image.”
Replicate flat per-image price for the dev tier.
- replicate.com
“FLUX.1 [schnell] is 0.003 per image.”
Replicate flat per-image price for the schnell tier.
- deepinfra.com
“FLUX-2-dev | $0.01 x (w / 1024) x (h / 1024) x (iters / 28)”
DeepInfra metered pricing formula for FLUX-2-dev.
- deepinfra.com
“FLUX-2-klein-9b | $0.015 x (w / 1024) x (h / 1024)”
DeepInfra metered pricing for the klein-9b variant.
- deepinfra.com
“FLUX-2-klein-4b | $0.014 x (w / 1024) x (h / 1024)”
DeepInfra metered pricing for the klein-4b variant.
- deepinfra.com
“FLUX-2-max | $0.07”
DeepInfra flat per-image price for the max tier.
- deepinfra.com
“FLUX-2-pro | $0.015”
DeepInfra flat per-image price for the pro tier.
- deepinfra.com
“FLUX-1-dev | $0.009 x (w / 1024) x (h / 1024) x (iters / 25)”
DeepInfra metered pricing formula for FLUX-1-dev.
- deepinfra.com
“FLUX-1-schnell | $0.0005 x (w / 1024) x (h / 1024) x iters”
DeepInfra metered pricing formula for the cheapest schnell tier.
- modular.com
“Illustration (1024 px, 25 steps) | Flux 2-dev on Modular Cloud | ~3.3 s”
Production-measured latency for a 1024px dev-tier image.
- modular.com
“The tolerance for image quality is configurable, and we are able to get good quality images in sub-second.”
FLUX.2 sub-second generation demonstration.
- modal.com
“we needed to return results in under three seconds.”
Modal's interactive latency threshold for competitive FLUX.1-dev serving.
- modal.com
“Averaging across a variety of inputs, we find that we can generate a 1024x1024 image in ~6.75 seconds.”
Baseline FLUX.1-dev latency with standard diffusers.
- cerebrium.ai
“the lowest TTFT of 123ms”
Cerebrium first-hand measured lowest TTFT.
- cerebrium.ai
“SGLang was a clear winner achieving a throughput of 460 tokens per second on a batch size of 64”
Cerebrium first-hand measured peak throughput.
- cerebrium.ai
“This was the total roundtrip time of a request made to the Cerebrium platform.”
Methodology: measurements are end-to-end platform roundtrip.
- cerebrium.ai
“Distil Labs runs up to 150 requests per second per model during high-traffic periods, with multiple models deployed concurrently.”
Production scale proof for the serverless/at-scale section.