FLUX Image API Cost Per Image: 2026 Provider Guide

Connor Blier
Founding GTM

A hexagonal frame like a coin or badge with a plus sign at its centre, symbolizing a per-image cost unit.

FLUX Image API Cost Per Image: 2026 Provider Guide

FLUX image API cost per image runs from about $0.0005 for a distilled schnell/klein tier to $0.07 for a premium max tier at 1024x1024. Between them, dev models land near $0.009-$0.030 and pro near $0.015-$0.055, depending on whether the provider bills per image, per megapixel or per GPU-second.

The question sounds simple: what does one FLUX image cost through an API? The answer depends on which FLUX tier you call, which provider serves it, and how that provider meters the work. This guide normalizes everything to a single 1024x1024 image so you can compare like for like, then pairs each price with the latency you should expect before you commit a production workload.

This is a commercial, provider-selection guide for ML engineers choosing where to ship. It optimizes for two axes at once: cost per image and latency.

TL;DR

  • Cheapest tier (distilled schnell/klein): roughly $0.0005 (DeepInfra FLUX-1-schnell) to $0.003 (Replicate FLUX.1 [schnell]) per 1024x1024 image.

  • Dev tier (balanced quality): ~$0.009 (DeepInfra FLUX-1-dev) to $0.030 (Replicate FLUX.1 [dev]).

  • Pro tier: ~$0.015 (DeepInfra FLUX-2-pro) to $0.055 (Replicate FLUX.1 [pro]).

  • Premium/max tier: ~$0.07 (DeepInfra FLUX-2-max).

  • Latency: a dev-class 1024px image lands in ~3.3s on Modular Cloud and ~6.75s baseline with vanilla diffusers, optimizable to under 3s (Modal). FLUX.2 has been demonstrated in sub-second on Modular.

How FLUX API pricing works

There are three billing models you will meet, and they do not compare directly until you normalize them.

  1. Flat per-image. You pay a fixed amount per generation regardless of resolution or step count. Replicate lists FLUX.1 [pro] at "$0.055 per image," FLUX.1 [dev] at "$0.030 per image," and FLUX.1 [schnell] at "0.003 per image." DeepInfra also uses flat rates for some tiers: FLUX-2-max at $0.07 and FLUX-2-pro at $0.015.

  2. Metered per-megapixel / per-iteration. The price scales with the pixels you render and the diffusion steps you run. DeepInfra prices FLUX-2-dev as "$0.01 x (w / 1024) x (h / 1024) x (iters / 28)", FLUX-1-dev as "$0.009 x (w / 1024) x (h / 1024) x (iters / 25)", and FLUX-1-schnell as "$0.0005 x (w / 1024) x (h / 1024) x iters". The klein variants scale by resolution only: FLUX-2-klein-9b at "$0.015 x (w / 1024) x (h / 1024)" and FLUX-2-klein-4b at "$0.014 x (w / 1024) x (h / 1024)".

  3. GPU-second (self-hosted / serverless). You rent the GPU and pay for the wall-clock time each image takes. Here your cost per image is a function of throughput, not a sticker price. We cover the general economics of this model in our LLM inference cost at scale guide and the 2026 GPU buyer's guide, and the underlying serverless pricing shapes apply directly to image generation.

Normalization convention used below: every metered price is evaluated at w=1024, h=1024 and the provider's default step count, so all figures are per single 1024x1024 image.

Cost per image by FLUX tier

TierModel (provider)Cost per 1024x1024Billing basis
Speed / cheapFLUX-1-schnell (DeepInfra)~$0.0005Metered (res x iters)
Speed / cheapFLUX.1 [schnell] (Replicate)$0.003Flat per-image
Balanced / devFLUX-1-dev (DeepInfra)~$0.009Metered (res x iters/25)
Balanced / devFLUX-2-dev (DeepInfra)~$0.01Metered (res x iters/28)
Balanced / devFLUX.1 [dev] (Replicate)$0.030Flat per-image
CompactFLUX-2-klein-4b (DeepInfra)~$0.014Metered (res only)
CompactFLUX-2-klein-9b (DeepInfra)~$0.015Metered (res only)
ProFLUX-2-pro (DeepInfra)$0.015Flat per-image
ProFLUX.1 [pro] (Replicate)$0.055Flat per-image
Premium / maxFLUX-2-max (DeepInfra)$0.07Flat per-image

The spread is nearly 140x from the cheapest distilled tier to the premium tier. The obvious takeaway: pick the lowest tier that clears your quality bar, because tier choice dominates every other cost lever.

Latency and throughput

Cost only tells half the story. An image that costs $0.0005 but returns in 10 seconds may be unusable in an interactive product, while a slightly pricier tier that returns in a couple of seconds pays for itself in conversion.

Modal established the interactive threshold plainly: "in order to be competitive with APIs serving FLUX.1-dev images, we needed to return results in under three seconds." Their baseline with the standard Hugging Face diffusers library was ~6.75 seconds for a 1024x1024 image, which they cut roughly 3x (1.5x from compiler/hardware optimization, 2x from First Block Caching) to reach that sub-3-second target.

On the production side, Modular Cloud's Flux2-dev endpoint generated a 1024px, 25-step illustration in ~3.3 seconds measured over 24 hours of real traffic, and Modular has demonstrated FLUX.2 producing "good quality images in sub-second" with configurable quality tolerance.

On the self-hosted path, your latency and cost per image both come down to platform roundtrip and throughput. In our testing of LLM serving on Cerebrium hardware we measured a lowest TTFT of 123ms (vLLM) and 460 tokens/sec throughput (SGLang, batch 64) on a single H100 as the total roundtrip to the Cerebrium platform, not a vendor spec sheet number. The same discipline of measuring end-to-end roundtrip rather than trusting published figures is what you should apply to any FLUX endpoint before you trust its cost-per-image math. See the full methodology in our vLLM/SGLang/TensorRT benchmark and our network latency in the inference pipeline guide, since network hops often dwarf model compute.

Total cost at scale

Here is what tier choice does to a monthly bill at 100,000 images/month, all at 1024x1024:

  • Distilled (schnell/klein): DeepInfra FLUX-1-schnell ~$50; Replicate FLUX.1 [schnell] ~$300.

  • Dev: DeepInfra FLUX-1-dev ~$900; DeepInfra FLUX-2-dev ~$1,000; Replicate FLUX.1 [dev] ~$3,000.

  • Pro: DeepInfra FLUX-2-pro ~$1,500; Replicate FLUX.1 [pro] ~$5,500.

  • Max: DeepInfra FLUX-2-max ~$7,000.

At 1,000,000 images/month, multiply by ten: the gap between a distilled tier and a premium tier is the difference between a $500 line item and a $70,000 one. At that volume the GPU-second/serverless path becomes worth modeling, because sustained utilization can undercut per-image APIs, an effect we quantify in our inference cost at scale analysis and the case study on DistilLabs' 50% lower inference costs, which sustained up to 150 requests per second per model at peak on autoscaling infrastructure.

Choosing a provider

  • Interactive consumer app (chat, editors): prioritize latency. A dev tier at ~3.3s (Modular) or an optimized sub-3s endpoint (Modal) beats a cheaper-but-slower option. Pair with a low-cold-start serverless platform.

  • High-volume batch (thumbnails, variations): prioritize cost. A distilled tier at ~$0.0005-$0.003 wins outright; latency is hidden behind a queue.

  • Quality-critical (marketing, hero assets): a pro or max tier justifies its price on output quality, not throughput.

  • Unpredictable/spiky traffic: a metered or serverless model that scales to zero avoids paying for idle GPUs; see serverless cost economics and, if you are leaving hyperscalers, our AWS alternatives for AI workloads.

Normalize every quote to per-1024x1024, measure real roundtrip latency, then let your use case pick the axis.

Frequently asked questions

What is the cheapest FLUX image API per image?
The cheapest published rate is DeepInfra's FLUX-1-schnell at roughly $0.0005 per 1024x1024 image (metered by resolution and iterations). Replicate's FLUX.1 [schnell] is a flat $0.003 per image.
How much does FLUX pro cost per image?
It varies by provider: DeepInfra prices FLUX-2-pro at a flat $0.015 per image, while Replicate lists FLUX.1 [pro] at $0.055 per image. The premium FLUX-2-max tier on DeepInfra is $0.07.
How is FLUX API pricing calculated?
Three ways: flat per-image (e.g. Replicate), metered per-megapixel and per-iteration normalized to 1024x1024 and default steps (e.g. DeepInfra's dev and schnell tiers), or GPU-second billing when you self-host on serverless infrastructure.
How fast is a FLUX image generated?
A 1024px, 25-step dev image was measured at ~3.3s on Modular Cloud in production. Vanilla diffusers baseline is ~6.75s (Modal), optimizable to under 3s. FLUX.2 has been shown generating good-quality images in sub-second on Modular.
Should I use a per-image API or self-host FLUX?
Per-image APIs are simplest and cheapest at low-to-moderate volume. At high sustained volume, GPU-second serverless can undercut per-image pricing because you pay for utilization rather than a fixed sticker price per generation.

Get started with Cerebrium

Deploy AI models on serverless GPUs in minutes, with no infrastructure to manage. Start for free and pay only for the compute you use.

Sign up free

Sources

  1. replicate.com
    “FLUX.1 [pro] is $0.055 per image.”

    Replicate flat per-image price for the pro tier.

  2. replicate.com
    “FLUX.1 [dev] is $0.030 per image.”

    Replicate flat per-image price for the dev tier.

  3. replicate.com
    “FLUX.1 [schnell] is 0.003 per image.”

    Replicate flat per-image price for the schnell tier.

  4. deepinfra.com
    “FLUX-2-dev | $0.01 x (w / 1024) x (h / 1024) x (iters / 28)”

    DeepInfra metered pricing formula for FLUX-2-dev.

  5. deepinfra.com
    “FLUX-2-klein-9b | $0.015 x (w / 1024) x (h / 1024)”

    DeepInfra metered pricing for the klein-9b variant.

  6. deepinfra.com
    “FLUX-2-klein-4b | $0.014 x (w / 1024) x (h / 1024)”

    DeepInfra metered pricing for the klein-4b variant.

  7. deepinfra.com
    “FLUX-2-max | $0.07”

    DeepInfra flat per-image price for the max tier.

  8. deepinfra.com
    “FLUX-2-pro | $0.015”

    DeepInfra flat per-image price for the pro tier.

  9. deepinfra.com
    “FLUX-1-dev | $0.009 x (w / 1024) x (h / 1024) x (iters / 25)”

    DeepInfra metered pricing formula for FLUX-1-dev.

  10. deepinfra.com
    “FLUX-1-schnell | $0.0005 x (w / 1024) x (h / 1024) x iters”

    DeepInfra metered pricing formula for the cheapest schnell tier.

  11. modular.com
    “Illustration (1024 px, 25 steps) | Flux 2-dev on Modular Cloud | ~3.3 s”

    Production-measured latency for a 1024px dev-tier image.

  12. modular.com
    “The tolerance for image quality is configurable, and we are able to get good quality images in sub-second.”

    FLUX.2 sub-second generation demonstration.

  13. modal.com
    “we needed to return results in under three seconds.”

    Modal's interactive latency threshold for competitive FLUX.1-dev serving.

  14. modal.com
    “Averaging across a variety of inputs, we find that we can generate a 1024x1024 image in ~6.75 seconds.”

    Baseline FLUX.1-dev latency with standard diffusers.

  15. cerebrium.ai
    “the lowest TTFT of 123ms”

    Cerebrium first-hand measured lowest TTFT.

  16. cerebrium.ai
    “SGLang was a clear winner achieving a throughput of 460 tokens per second on a batch size of 64”

    Cerebrium first-hand measured peak throughput.

  17. cerebrium.ai
    “This was the total roundtrip time of a request made to the Cerebrium platform.”

    Methodology: measurements are end-to-end platform roundtrip.

  18. cerebrium.ai
    “Distil Labs runs up to 150 requests per second per model during high-traffic periods, with multiple models deployed concurrently.”

    Production scale proof for the serverless/at-scale section.


Related resources

See all
A single-colour bar chart icon: five vertical bars of increasing height rising left to right along a baseline, with an upward diagonal arrow above the tallest bars pointing to the upper right, symbolizing rising inference throughput.
MI300X vs H200 LLM Inference Throughput
A single mechanical gear rendered as a settings-style cog with a ringed outer edge and a hollow circular center, symbolizing configurable infrastructure for running large language model inference.
Baseten Alternatives for Production LLM Inference
A simple bar chart of five vertical bars increasing in height from left to right, with the tallest bar near the right side slightly shorter than its neighbor, sitting on a horizontal baseline, all rendered in solid black.
H100 vs H200 LLM Inference Throughput: A Buyer's Guide