Skip to content

Pricing

View markdown

Every job is billed as a single product, chosen by the job type and a small number of config parameters (output resolution, duration, number of input images). Prompt, seed, steps and other settings do not change the price.

Applies to the job type inference.gemini-omni-flash.txt2vid.v1.

This job type uses a fixed upstream spec: the caller sends only a prompt and the upstream model decides the output resolution and length. It is priced as a single flat product per video.

Price
$0.50

The product field is inference-gemini-omni-flash.

Applies to the Grok Imagine job types:

  • inference.grok-imagine.txt2vid.v1
  • inference.grok-imagine.img2vid.v1

Priced per second of output video. Duration is billed in whole seconds (1 to 15).

Per second 5 seconds 10 seconds 15 seconds
$0.08 $0.40 $0.80 $1.20

The product field follows the pattern inference-grok-imagine-duration-N, where N is the billed duration in seconds (1 to 15). For example a 10-second video bills as inference-grok-imagine-duration-10.

Applies to the job type inference.happyhorse-1-1.img2vid.v1.

Priced per second of output video by resolution (720P or 1080P). duration runs from 1 to 15 seconds and is billed in whole seconds, rounded up.

Resolution Per second 5 seconds 10 seconds 15 seconds
720P $0.126 $0.63 $1.26 $1.89
1080P $0.224 $1.12 $2.24 $3.36

The product field follows the pattern inference-happyhorse-1-1-{720p,1080p}-duration-N, where N is the billed duration in seconds (1 to 15). For example a 5-second 1080P clip bills as inference-happyhorse-1-1-1080p-duration-5.

Applies to the Kling job types inference.kling.txt2vid.v1 and inference.kling.img2vid.v1.

Priced per video by the model version, mode (std or pro, where the model offers both) and duration ("5" or "10" seconds). Aspect ratio and prompt do not change the price, and image-to-video bills the same as text-to-video. model defaults to kling-v1 and mode to std when omitted.

Model Mode Resolution 5 seconds 10 seconds
kling-v1 Standard 720P $0.15 $0.30
kling-v1 Pro 1080P $0.55 $1.10
kling-v1-6 Standard 720P $0.25 $0.50
kling-v1-6 Pro 1080P $0.55 $1.10
kling-v2-master — 1080P $1.40 $2.80
kling-v2-1 Standard 720P $0.28 $0.56
kling-v2-1-master — 1080P $1.40 $2.80

kling-v2-master, kling-v2-1 and kling-v2-1-master bill the same regardless of mode.

The product field follows the pattern inference-kling-{v1,v1-6}-{std,pro}-duration-N, inference-kling-v2-1-standard-duration-N, or inference-kling-{v2-master,v2-1-master}-duration-N, where N is 5 or 10. For example a 10-second kling-v1-6 video in pro mode bills as inference-kling-v1-6-pro-duration-10.

Applies to the LTX 2 Pro job types:

  • inference.ltx-2-pro.txt2vid.v1
  • inference.ltx-2-pro.img2vid.v1

Priced per second of output video by output resolution tier. duration runs from 1 to 10 seconds (default 8) and is billed in whole seconds, rounded up. Output resolutions above roughly 11 MP bill at the 1440p tier and above roughly 6 MP at the 4k tier; everything else bills at 720p or 1080p by output size.

Resolution Per second 5 seconds 8 seconds 10 seconds
720p $0.12 $0.60 $0.96 $1.20
1080p $0.17 $0.85 $1.36 $1.70
1440p $0.25 $1.25 $2.00 $2.50
4k $0.32 $1.60 $2.56 $3.20

The product field follows the pattern inference-ltx-2-pro-{720p,1080p,1440p,4k}-duration-N, where N is the billed duration in seconds (1 to 10). For example an 8-second 1080p clip bills as inference-ltx-2-pro-1080p-duration-8.

Applies to the LTX 2.5 fast tier job types:

  • inference.ltx-2.5.fast.txt2vid.v0
  • inference.ltx-2.5.fast.img2vid.v0

Priced per clip by output duration. Fractional durations round up to the next whole second. Output resolution, aspect ratio and seed do not change the price, and image-to-video bills the same as text-to-video.

Duration Price
1 s $0.0152
2 s $0.0181
3 s $0.0231
4 s $0.0316
5 s $0.0346
6 s $0.0391
7 s $0.0445
8 s $0.0506
9 s $0.0581
10 s $0.0655

The product field follows the pattern inference-ltx-2-5-duration-N, where N is the billed duration in seconds (1 to 10). For example a 6-second clip bills as inference-ltx-2-5-duration-6.

Applies to the MiniMax H3 job types in five tiers:

  • Partner – inference.minimax.h3.txt2vid.v1, inference.minimax.h3.img2vid.v1 and inference.minimax.h3.ref2vid.v1 (the hosted MiniMax API)
  • Base – inference.minimax.h3.base.txt2vid.v1, inference.minimax.h3.base.img2vid.v1 and inference.minimax.h3.base.ref2vid.v1 (self-hosted, v0 equivalents also bill this way)
  • Fast – inference.minimax.h3.fast.txt2vid.v1, inference.minimax.h3.fast.img2vid.v1 and inference.minimax.h3.fast.ref2vid.v1 (self-hosted)
  • Turbo / Hyper – inference.minimax.h3.turbo.{txt2vid,img2vid,ref2vid}.v0 and inference.minimax.h3.hyper.{txt2vid,img2vid,ref2vid}.v0 (self-hosted)

All tiers bill per second of output video; output duration runs from 4 to 15 seconds (billed in whole seconds, a job with no duration set bills at the 6-second default). Only Partner takes a resolution of 768P (default) or 2K; Base, Fast, Turbo and Hyper use a fixed 768P-class output spec.

Resolution 4 s 5 s 6 s 7 s 8 s 9 s 10 s 11 s 12 s 13 s 14 s 15 s
768P $0.32 $0.40 $0.48 $0.56 $0.64 $0.72 $0.80 $0.88 $0.96 $1.04 $1.12 $1.20
2K $0.52 $0.65 $0.78 $0.91 $1.04 $1.17 $1.30 $1.43 $1.56 $1.69 $1.82 $1.95

Base is 768P only.

4 s 5 s 6 s 7 s 8 s 9 s 10 s 11 s 12 s 13 s 14 s 15 s
$0.32 $0.40 $0.48 $0.56 $0.64 $0.72 $0.80 $0.88 $0.96 $1.04 $1.12 $1.20

Fast is 768P only.

4 s 5 s 6 s 7 s 8 s 9 s 10 s 11 s 12 s 13 s 14 s 15 s
$0.32 $0.40 $0.48 $0.56 $0.64 $0.72 $0.80 $0.88 $0.96 $1.04 $1.12 $1.20

Both tiers share the same latency class and the same prices.

4 s 5 s 6 s 7 s 8 s 9 s 10 s 11 s 12 s 13 s 14 s 15 s
$0.0277 $0.0330 $0.0451 $0.0517 $0.0595 $0.0736 $0.0844 $0.1026 $0.1098 $0.1294 $0.1408 $0.1540

Reference videos on ref2vid jobs are billed at the same per-second rate as the output: the input video’s duration is added to the output duration before pricing (reference video is capped at 15 seconds, so billed durations run to 60 on Partner, Base and Fast, 30 on Turbo and Hyper). Billed durations above 15 seconds continue at the per-second rate, for example (Partner 768P and Base bill identically):

Billed duration Partner 768P / Base Partner 2K
20 s $1.60 $2.60
30 s $2.40 $3.90
45 s $3.60 $5.85
60 s $4.80 $7.80

Input images are counted across the reference images on ref2vid jobs and any first or last frame on img2vid jobs. Up to 5 input images are included in the per-second price. From the sixth image onwards a flat $0.04 surcharge per image is added to the whole video, regardless of duration or resolution:

Input images Surcharge per video
0 to 5 none
6 $0.04
7 $0.08
8 $0.12
9 $0.16

For example, a 6-second Partner ref2vid job with 7 reference images bills $0.48 + $0.08 = $0.56.

The product field is inference-minimax-h3[-2k][-M-in]-duration-N for Partner, where M is the number of input images (omitted when there are none) and N is the billed duration in seconds (4 to 60); for example a 6-second job with 7 reference images bills as inference-minimax-h3-7-in-duration-6. The self-hosted tiers carry their tier name in the code instead: inference-minimax-h3-{base,fast}[-M-in]-duration-N (4 to 60) and inference-minimax-h3-{turbo,hyper}[-M-in]-duration-N (4 to 30), for example inference-minimax-h3-fast-7-in-duration-6.

Applies to the P-Video job types:

  • inference.pruna.p-video.txt2vid.v1
  • inference.pruna.p-video.img2vid.v1
  • inference.pruna.p-video.aud2vid.v1

Priced per video by mode (draft: true for Draft, otherwise Standard), resolution and duration. Aspect ratio and prompt do not change the price. Text- and image-to-video take a duration of 5 or 10 seconds; audio-to-video follows the length of the input audio and bills at the 5, 10 or 15 second tier, rounded up.

Mode Resolution 5 seconds 10 seconds 15 seconds
Draft 720p $0.025 $0.050 $0.075
Draft 1080p $0.050 $0.100 $0.150
Standard 720p $0.100 $0.200 $0.300
Standard 1080p $0.200 $0.400 $0.600

The product field follows the pattern inference-pruna-p-video[-1080p][-draft]-duration-N, where N is 5, 10 or 15. 720p and Standard are the defaults and carry no suffix. For example a 10-second Draft video at 1080p bills as inference-pruna-p-video-1080p-draft-duration-10.

Applies to the job type inference.runway.gen4.vid2vid.v1.

Priced as a flat price per video-to-video generation. Input video length, aspect ratio, prompt and seed do not change the price.

Price
$0.75

The product field is inference-runway-gen4.

Applies to the Seedance 1.0 job types, in three tiers:

  • Lite – inference.seedance.lite.txt2vid.v1 and inference.seedance.lite.img2vid.v1
  • Pro – inference.seedance.pro.txt2vid.v1 and inference.seedance.pro.img2vid.v1
  • Pro Turbo – inference.seedance.proturbo.txt2vid.v1 and inference.seedance.proturbo.img2vid.v1

Priced per video by resolution and duration (5 or 10 seconds). Aspect ratio, prompt and seed do not change the price, and image-to-video bills the same as text-to-video.

Tier Resolution 5 seconds 10 seconds
Lite 480p $0.09 $0.18
Lite 720p $0.20 $0.40
Lite 1080p $0.44 $0.88
Pro 480p $0.12 $0.24
Pro 1080p $0.61 $1.22
Pro Turbo 480p $0.048 $0.096
Pro Turbo 1080p $0.244 $0.488

Pro and Pro Turbo do not offer 720p. When omitted, resolution defaults to 1080p and duration to 5 seconds.

The product field follows the pattern inference-seedance-{lite,pro,proturbo}-{480p,720p,1080p}-duration-{5,10}. For example a 10-second Pro Turbo video at 480p bills as inference-seedance-proturbo-480p-duration-10.

Applies to the Seedance 2.0 family of job types, in three tiers:

  • Seedance 2.0 – inference.seedance-2.txt2vid.v1 and inference.seedance-2.img2vid.v1
  • Seedance 2.0 Fast – inference.seedance-2.fast.txt2vid.v1 and inference.seedance-2.fast.img2vid.v1
  • Seedance 2.5 – inference.seedance-2-5.txt2vid.v1 and inference.seedance-2-5.img2vid.v1

Seedance 2.0 and 2.0 Fast are priced per second of output video by resolution. duration is 4 to 15 seconds (default 5); setting duration to -1 lets the model pick the length and the clip bills at its actual duration. Aspect ratio, prompt, seed and audio do not change the price, and image-to-video bills the same as text-to-video.

Tier Resolution Per second 5 seconds 10 seconds 15 seconds
2.0 480p $0.062 $0.31 $0.62 $0.93
2.0 720p $0.14 $0.70 $1.40 $2.10
2.0 1080p $0.31 $1.55 $3.10 $4.65
2.0 Fast 480p $0.05 $0.25 $0.50 $0.75
2.0 Fast 720p $0.07 $0.35 $0.70 $1.05

resolution defaults to 1080p on Seedance 2.0 and to 720p on Seedance 2.0 Fast. Seedance 2.0 Fast does not offer 1080p.

Seedance 2.5 uses a fixed upstream spec (prompt and optional input image only) and bills a flat price per video:

Tier Price per video
2.5 $1.05

Seedance 2.0 bills as inference-seedance-2-{480p,720p,1080p}-duration-N and Seedance 2.0 Fast as inference-seedance-2-fast-{480p,720p}-duration-N, where N is the billed duration in seconds. For example a 10-second Seedance 2.0 video at 1080p bills as inference-seedance-2-1080p-duration-10. Seedance 2.5 bills as inference-seedance-2-5.

Applies to the Sora 2 job types:

  • Sora 2 – inference.sora-2.txt2vid.v1 and inference.sora-2.img2vid.v1
  • Sora 2 Pro – inference.sora-2.pro.txt2vid.v1 and inference.sora-2.pro.img2vid.v1

Priced per second of output video. duration is 4, 8 or 12 seconds (default 4). Sora 2 renders at 720p; Sora 2 Pro takes a resolution of 720p (default) or 1080p. Aspect ratio, prompt and seed do not change the price, and image-to-video bills the same as text-to-video.

Model Resolution Per second 4 seconds 8 seconds 12 seconds
Sora 2 720p $0.10 $0.40 $0.80 $1.20
Sora 2 Pro 720p $0.30 $1.20 $2.40 $3.60
Sora 2 Pro 1080p $0.50 $2.00 $4.00 $6.00

The product field follows the pattern inference-sora-2-duration-N for Sora 2 and inference-sora-2-pro-{720p,1080p}-duration-N for Sora 2 Pro, where N is 4, 8 or 12. For example an 8-second Sora 2 Pro video at 1080p bills as inference-sora-2-pro-1080p-duration-8.

Applies to the job type inference.tenstorrent.wan2-2-lightning.img2vid.v1.

This model runs on Prodia’s own hardware rather than an upstream API, so it carries an internal price: a flat rate per video, independent of duration and settings.

Price
$0.084

The product field is inference-tenstorrent-wan2-2-lightning.

Applies to the Veo job types:

  • Veo 3 Fast – inference.veo.fast.txt2vid.v2 and inference.veo.fast.img2vid.v2 (and their v1 equivalents)
  • Veo 3 – inference.veo.txt2vid.v2 and inference.veo.img2vid.v2 (and their v1 equivalents)

Every Veo 3 video is 8 seconds long. Priced per video by tier and whether audio is generated (generate_audio: true). Output resolution (720p or 1080p), prompt and seed do not change the price, and image-to-video bills the same as text-to-video.

Model Length Without audio With audio
Veo 3 Fast 8 seconds $0.80 $1.20
Veo 3 8 seconds $1.60 $3.20

The product field is inference-veo or inference-veo-fast, with -audio appended when audio is generated. Jobs that set aspect_ratio explicitly carry it in the product as well, for example inference-veo-fast-16_9-audio.

Applies to the job types:

  • inference.veo.lite.txt2vid.v2
  • inference.veo.lite.img2vid.v2

Every Veo 3.1 Lite video is 8 seconds long. Priced per video by resolution (720p, the default, or 1080p) and whether audio is generated (generate_audio: true). Prompt, seed and image-to-video do not change the price.

Resolution Without audio With audio
720p $0.24 $0.40
1080p $0.40 $0.64

The product field follows the pattern inference-veo-lite-{720p,1080p}[-audio]-duration-8. For example an 8-second 1080p video with audio bills as inference-veo-lite-1080p-audio-duration-8.

Applies to the Vidu job types:

  • inference.vidu-q3-pro.txt2vid.v1
  • inference.vidu-q3-pro.img2vid.v1

Vidu Q3 Pro uses a fixed upstream spec (prompt and optional input image only) and bills a flat price per video.

Price
$0.50

The product field is inference-vidu-q3-pro.

Applies to the Wan 2.6 job types:

  • Video – inference.wan2-6.txt2vid.v1 and inference.wan2-6.img2vid.v1
  • Image – inference.wan2-6.txt2img.v1 and inference.wan2-6.img2img.v1

Video is priced per second of output by resolution. Text-to-video takes a size (width*height) and image-to-video takes a resolution (720P or 1080P, the default); sizes up to 1 MP bill at 720p and larger sizes bill at 1080p. duration runs from 2 to 15 seconds (default 5) and is billed in whole seconds, rounded up. Images are a flat price. Aspect ratio, prompt, seed and reference inputs do not change the price.

Resolution Per second 5 seconds 10 seconds 15 seconds
720P $0.084 $0.42 $0.84 $1.26
1080P $0.14 $0.70 $1.40 $2.10
Price per image
$0.028

Video bills as inference-wan2-6-{720p,1080p}-duration-N, where N is the billed duration in seconds (2 to 15). For example a 5-second 1080P clip bills as inference-wan2-6-1080p-duration-5. Images bill as inference-wan2-6-image.

Applies to the Wan 2.7 job types:

  • Video – inference.wan2-7.txt2vid.v1, inference.wan2-7.img2vid.v1 and inference.wan2-7.vid2vid.v1
  • Image – inference.wan2-7.txt2img.v1 and inference.wan2-7.img2img.v1

Video is priced per second of output by resolution (720P, the default, or 1080P). duration runs from 2 to 15 seconds (default 5) and is billed in whole seconds, rounded up. Video-to-video bills on the output duration. Images are a flat price. Aspect ratio, prompt, seed and reference inputs do not change the price.

Resolution Per second 5 seconds 10 seconds 15 seconds
720P $0.07 $0.35 $0.70 $1.05
1080P $0.105 $0.525 $1.05 $1.575
Output Example sizes Price
1K or 2K 1024×1024, 2048×2048 $0.021

Video bills as inference-wan2-7-{720p,1080p}-duration-N, where N is the billed duration in seconds (2 to 15). For example a 5-second 1080P clip bills as inference-wan2-7-1080p-duration-5. Images bill as inference-wan2-7-image.

Applies to the Wan 3.0 job types:

  • inference.wan3-0.txt2vid.v1
  • inference.wan3-0.img2vid.v1

Priced per second of output by resolution (480P, 720P, the default, or 1080P). duration runs up to 30 seconds (default 5) and is billed in whole seconds, rounded up. Aspect ratio, prompt, seed and audio do not change the price, and image-to-video bills the same as text-to-video.

Resolution Per second 5 seconds 10 seconds 15 seconds 30 seconds
480P $0.035 $0.175 $0.35 $0.525 $1.05
720P $0.07 $0.35 $0.70 $1.05 $2.10
1080P $0.14 $0.70 $1.40 $2.10 $4.20

The product field follows the pattern inference-wan3-0-{480p,720p,1080p}-duration-N, where N is the billed duration in seconds (1 to 30). For example a 10-second 1080P clip bills as inference-wan3-0-1080p-duration-10.

Applies to the FLUX.1 job types:

  • FLUX.1 [dev] – inference.flux.dev.txt2img.v2, inference.flux.dev.img2img.v2, inference.flux.dev.inpainting.v2 (and their v1 equivalents), plus inference.flux-fast.dev.txt2img.v1
  • FLUX1.1 [pro] – inference.flux.pro11.txt2img.v1
  • FLUX1.1 [pro] Ultra – inference.flux.pro11ultra.txt2img.v1 and inference.flux.pro11ultra.img2img.v1

FLUX.1 [dev] is priced per image by step count; the FLUX1.1 [pro] models are a flat price per image. Output resolution, prompt and seed do not change the price. Volume pricing applies from 1M generations.

Model Steps Price Volume discount
FLUX.1 [dev] Up to 28 $0.0200 25%
FLUX.1 [dev] 29 to 50 $0.0240 25%
FLUX1.1 [pro] — $0.0400 25%
FLUX1.1 [pro] Ultra — $0.0600 25%
Model Product
FLUX.1 [dev] inference-flux-dev-steps-28, inference-flux-dev-steps-50
FLUX1.1 [pro] inference-flux-pro11
FLUX1.1 [pro] Ultra inference-flux-pro11-ultra

Applies to the FLUX.1 Kontext job types:

  • FLUX.1 Kontext [dev] – inference.flux-fast.dev.kontext.img2img.v1
  • FLUX.1 Kontext [pro] – inference.flux-kontext.pro.txt2img.v2 and inference.flux-kontext.pro.img2img.v2 (and their v1 equivalents)
  • FLUX.1 Kontext [max] – inference.flux-kontext.max.txt2img.v2 and inference.flux-kontext.max.img2img.v2 (and their v1 equivalents)

Flat price per image. Output resolution, aspect ratio, prompt and seed do not change the price, and image-to-image bills the same as text-to-image. Volume pricing applies from 1M generations.

Model Example sizes Price Volume discount
FLUX.1 Kontext [dev] 1024×1024, 1568×672, 672×1568 $0.025 25%
FLUX.1 Kontext [pro] 1024×1024, 1568×672, 672×1568 $0.040 25%
FLUX.1 Kontext [max] 512×512, 512×768, 640×640 $0.080 25%
Model Product
FLUX.1 Kontext [dev] inference-flux-kontext-dev
FLUX.1 Kontext [pro] inference-flux-kontext-pro
FLUX.1 Kontext [max] inference-flux-kontext-max

Applies to the FLUX.2 job types:

  • FLUX.2 [dev] – inference.flux-2.dev.txt2img.v1 and inference.flux-2.dev.img2img.v1 (and their v0 equivalents)
  • FLUX.2 [pro] – inference.flux-2.pro.txt2img.v1 and inference.flux-2.pro.img2img.v1
  • FLUX.2 [flex] – inference.flux-2.flex.txt2img.v1 and inference.flux-2.flex.img2img.v1
  • FLUX.2 [max] – inference.flux-2.max.txt2img.v1 and inference.flux-2.max.img2img.v1

Priced per image by output resolution (in megapixels) and the number of input images (0 to 8). Steps, prompt and seed do not change the price. FLUX.2 [dev] is self-hosted; the [pro], [flex] and [max] tiers are proxied from the upstream API at pass-through rates.

Output resolution buckets (measured as width × height):

Bucket Pixel budget Example sizes
Up to 1 MP ≤ 1,048,576 1024×1024, 1536×683
Up to 2 MP ≤ 2,096,704 1448×1448, 2048×1024
Up to 3 MP ≤ 3,140,784 1772×1772, 2560×1229
Over 3 MP > 3,140,784 2048×2048, 3072×1536
Inputs Up to 1 MP Up to 2 MP Up to 3 MP Over 3 MP
Text-to-image (no inputs) $0.012 $0.024 $0.036 $0.048
1 input image $0.024 $0.036 $0.048 $0.06
2 input images $0.036 $0.048 $0.06 $0.072
3 input images $0.048 $0.06 $0.072 $0.084
4 input images $0.06 $0.072 $0.084 $0.096
5 input images $0.072 $0.084 $0.096 $0.108
6 input images $0.084 $0.096 $0.108 $0.12
7 input images $0.096 $0.108 $0.12 $0.132
8 input images $0.108 $0.12 $0.132 $0.144
Inputs Up to 1 MP Up to 2 MP Up to 3 MP Over 3 MP
Text-to-image (no inputs) $0.03 $0.045 $0.06 $0.075
1 input image $0.045 $0.06 $0.075 $0.09
2 input images $0.06 $0.075 $0.09 $0.105
3 input images $0.075 $0.09 $0.105 $0.12
4 input images $0.09 $0.105 $0.12 $0.135
5 input images $0.105 $0.12 $0.135 $0.15
6 input images $0.12 $0.135 $0.15 $0.165
7 input images $0.135 $0.15 $0.165 $0.18
8 input images $0.15 $0.165 $0.18 $0.195
Inputs Up to 1 MP Up to 2 MP Up to 3 MP Over 3 MP
Text-to-image (no inputs) $0.06 $0.12 $0.18 $0.24
1 input image $0.12 $0.18 $0.24 $0.30
2 input images $0.18 $0.24 $0.30 $0.36
3 input images $0.24 $0.30 $0.36 $0.42
4 input images $0.30 $0.36 $0.42 $0.48
5 input images $0.36 $0.42 $0.48 $0.54
6 input images $0.42 $0.48 $0.54 $0.60
7 input images $0.48 $0.54 $0.60 $0.66
8 input images $0.54 $0.60 $0.66 $0.72
Inputs Up to 1 MP Up to 2 MP Up to 3 MP Over 3 MP
Text-to-image (no inputs) $0.07 $0.10 $0.13 $0.16
1 input image $0.10 $0.13 $0.16 $0.19
2 input images $0.13 $0.16 $0.19 $0.22
3 input images $0.16 $0.19 $0.22 $0.25
4 input images $0.19 $0.22 $0.25 $0.28
5 input images $0.22 $0.25 $0.28 $0.31
6 input images $0.25 $0.28 $0.31 $0.34
7 input images $0.28 $0.31 $0.34 $0.37
8 input images $0.31 $0.34 $0.37 $0.40

The product field follows the pattern inference-flux-2-{dev,pro,flex,max}[-N-in]-M-mp, where N is the number of input images (omitted when there are none) and M is the 1 to 4 megapixel bucket. For example an image with 2 input images rendered at 1536×1536 (3 MP) bills as inference-flux-2-pro-2-in-3-mp.

Applies to the FLUX.2 Klein 4B distilled job types:

  • inference.flux-2.klein.4b.txt2img.v1 and inference.flux-2.klein.4b.img2img.v1
  • inference.flux-2.klein.txt2img.v1 and inference.flux-2.klein.img2img.v1 (the default aliases, currently routed to 4B distilled)

Priced per image by output resolution and the number of input images.

Inputs Up to 1 MP Up to 2 MP Up to 3 MP Over 3 MP
Text-to-image (no inputs) $0.0005 $0.0027 $0.0040 $0.0059
1 input image $0.0010 $0.0027 $0.0040 $0.0059
2 input images $0.0017 $0.0034 $0.0068 $0.0084
3 input images $0.0025 $0.0065 $0.0078 $0.0084
4 input images $0.0033 $0.0065 $0.0080 $0.0109
5 input images $0.0045 $0.0068 $0.0102 $0.0109
6 input images $0.0054 $0.0088 $0.0116 $0.0132
7 input images $0.0068 $0.0092 $0.0130 $0.0145
8 input images $0.0100 $0.0108 $0.0148 $0.0170

Applies to the Nano Banana and Gemini 3 image job types:

  • Nano Banana (Gemini 2.5 Flash) – inference.nano-banana.txt2img.v2, inference.nano-banana.img2img.v2 and inference.nano-banana.img2img.v1
  • Nano Banana Pro (Gemini 3 Pro) – inference.gemini-3-pro.txt2img.v1 and inference.gemini-3-pro.img2img.v1
  • Nano Banana 2 (Gemini 3.1 Flash) – inference.gemini-3-1-flash.txt2img.v1 and inference.gemini-3-1-flash.img2img.v1

Priced per image by the resolution setting (1K, 2K or 4K). Aspect ratio, prompt and input images do not change the price, and image-to-image bills the same as text-to-image.

Model Resolution Example sizes Price
Nano Banana (Gemini 2.5 Flash) 1K 1024×1024, 1568×672, 672×1568 $0.039
Nano Banana Pro (Gemini 3 Pro) 1K or 2K 1024×1024, 1568×672, 2560×1440 $0.150
Nano Banana Pro (Gemini 3 Pro) 4K 3840×2160 $0.300
Nano Banana 2 (Gemini 3.1 Flash) 1K 1024×1024 $0.080
Nano Banana 2 (Gemini 3.1 Flash) 2K 2048×2048 $0.120
Nano Banana 2 (Gemini 3.1 Flash) 4K 3840×2160 $0.160

Nano Banana (Gemini 2.5 Flash) only produces 1K output. resolution defaults to 1K when omitted.

Model Product
Nano Banana (Gemini 2.5 Flash) inference-nano-banana
Nano Banana Pro (Gemini 3 Pro) inference-gemini-3-pro, inference-gemini-3-pro-4k
Nano Banana 2 (Gemini 3.1 Flash) inference-gemini-3-1-flash, inference-gemini-3-1-flash-2k, inference-gemini-3-1-flash-4k

Applies to the job type inference.pruna.p-image-ideogram.txt2img.v1.

Priced per image by thinking level (high, the default, medium, low or very low) and output resolution. image_size is 1K (default) or 2K; setting explicit width/height instead bills by the requested pixel count (over 1 MP bills as 2K). Prompt and seed do not change the price.

Thinking 1K 2K
High (default) $0.015 $0.030
Medium $0.010 $0.020
Low $0.0075 $0.015
Very low $0.003 $0.006

The product field is inference-pruna-p-image-ideogram, with -medium, -low or -very-low for the lower thinking levels and -2k appended for 2K output. For example a 2K image at the low thinking level bills as inference-pruna-p-image-ideogram-low-2k.

Available in two tiers, chosen by job type:

  • Speed – inference.qwen.image-edit.plus.lightning.img2img.v2
  • Quality – inference.qwen.image-edit.plus.img2img.v1 and inference.qwen.image-edit.plus.lightning.img2img.v1

Priced per generation by output resolution and whether one or several input images are supplied. Steps, prompt and seed do not change the price.

Tier Inputs Up to 1 MP Up to 4 MP
Speed Single input image $0.0085 $0.0320
Speed Multiple input images $0.0170 $0.0650
Quality Single input image $0.0150 $0.0450
Quality Multiple input images $0.0300 $0.0900

Output resolution is measured in pixels (width × height, each 256 to 2048):

Bucket Pixel range Example sizes
Up to 1 MP ≤ 1,048,576 1024×1024, 1200×900, 1280×800
Up to 4 MP > 1,048,576 1600×1200, 1920×1080, 2048×2048

Input images are counted from the images array (1 to 3). Exactly one image bills as single input; two or three bill as multiple inputs.

The product field follows the pattern inference-qwen-image-edit-{speed,quality}-{small,large}-{single,multi}, where small is the up-to-1 MP bucket and large is the up-to-4 MP bucket. For example a Speed job with two input images rendered at 1920×1080 bills as inference-qwen-image-edit-speed-large-multi.

Applies to the Recraft V3 and Recraft V4 job types. Each is a flat price per generation; size, prompt and style do not change the price.

Model Output Job types Price
V3 Raster image inference.recraft.txt2img.v1, inference.recraft.img2img.v1 $0.04
V3 Vectors SVG inference.recraft.img2vec.v1 $0.08
V4 Raster image inference.recraft.v4.txt2img.v1 $0.04
V4 Pro Raster image inference.recraft.v4.pro.txt2img.v1 $0.25
V4 Vector SVG inference.recraft.v4.txt2vec.v1 $0.08
V4 Pro Vector SVG inference.recraft.v4.pro.txt2vec.v1 $0.30

V3 and V4 render up to about 1 MP (for example 1024×1024 or 2048×1024 for V3, 768×1536 or 1024×1024 for V4). V4 Pro renders up to about 4 MP (for example 2048×2048 or 1536×3072).

Model Product
V3 inference-recraft
V3 Vectors inference-recraft-img2vec
V4 inference-recraft-v4
V4 Pro inference-recraft-v4-pro
V4 Vector inference-recraft-v4-vector
V4 Pro Vector inference-recraft-v4-pro-vector

Applies to the SDXL job types inference.sdxl.txt2img.v1, inference.sdxl.img2img.v1 and inference.sdxl.inpainting.v1.

Priced per image by step count. Output resolution (512 to 1536 per side), prompt and seed do not change the price, and image-to-image and inpainting bill the same as text-to-image.

Steps Example sizes Price
Up to 25 1024×1024, 768×1024, 1024×768 $0.0020
26 to 50 1024×1024, 768×1024, 1024×768 $0.0025
51 and above 1024×1024, 768×1024, 1024×768 $0.0100

The product field is inference-sdxl-steps-25, inference-sdxl-steps-50 or inference-sdxl-steps-100 for the three step buckets.

Applies to the Seedream job types:

  • Seedream 4.0 – inference.seedream-4.txt2img.v1 and inference.seedream-4.img2img.v1
  • Seedream 4.5 – inference.seedream-4-5.txt2img.v1 and inference.seedream-4-5.img2img.v1
  • Seedream 5.0 Lite – inference.seedream-5-0.lite.txt2img.v1 and inference.seedream-5-0.lite.img2img.v1

Flat price per image at any supported resolution. Output size, prompt, seed and input images do not change the price, and image-to-image bills the same as text-to-image.

Model Example sizes Price
Seedream 4.0 1024×1024, 2048×2048, 4096×4096 $0.030
Seedream 4.5 1024×1024, 2048×2048, 4096×4096 $0.040
Seedream 5.0 Lite 1024×1024, 2048×2048 $0.035
Model Product
Seedream 4.0 inference-seedream-4
Seedream 4.5 inference-seedream-4-5
Seedream 5.0 Lite inference-seedream-5-0-lite

Applies to the job type inference.topaz.gigapixel-standard-2.upscale.v1.

Priced by the output resolution of the upscaled image. The caller may set width/height (up to 32,000 per side); when omitted, Topaz determines the output size. The price follows the actual output pixel count of the delivered image. Upstream this bills as 1 credit per 24 MP of output.

Output megapixels Price
Up to 24 MP $0.10
Up to 48 MP $0.20
Up to 72 MP $0.30
Up to 96 MP $0.40
Up to 128 MP $0.60
Up to 256 MP $1.00
Over 256 MP $1.60

The product field is inference-topaz-gigapixel-standard-2, with the suffix for the output megapixel tier appended: -48mp, -72mp, -96mp, -128mp, -256mp or -512mp. For example an upscaled image with 60 MP of output bills as inference-topaz-gigapixel-standard-2-72mp.

Applies to the job type inference.topaz.proteus.upscale.v1.

Priced per second of output video by the output resolution tier. The upscale scale is 2 or 4 (default 2), or an explicit width/height (up to 16,000 per side); the tier follows the actual output size of the delivered clip (up to ~1 MP bills as 720p, up to ~2.1 MP as 1080p, larger as 4k). Clip length is billed in whole seconds, rounded up, to 30 seconds.

Output resolution Per second 10 seconds 30 seconds
720p $0.01 $0.10 $0.30
1080p $0.02 $0.20 $0.60
4k $0.06 $0.60 $1.80

The product field follows the pattern inference-topaz-proteus-{720p,1080p,4k}-duration-N, where N is the billed clip length in seconds (1 to 30). For example a 10-second 1080p upscale bills as inference-topaz-proteus-1080p-duration-10.

Fixed-price utility jobs. Each bills a single flat product per call; input image size, prompt and other settings do not change the price. Volume pricing applies from 1M generations.

Utility Model Job type Price Volume discount
Background removal BiRefNet 2 inference.remove-background.v1 $0.0025 25%
NSFW image detection ViT inference.vit.img2label.v2 $0.0002 —
Upscale 2× R-ESRGAN inference.upscale.v1 with upscale: 2 $0.0010 25%
Upscale 4× R-ESRGAN inference.upscale.v1 with upscale: 4 $0.0020 25%
Upscale 8× R-ESRGAN inference.upscale.v1 with upscale: 8 $0.0030 25%
Upscale 2× HYPIR inference.hypir.upscale.v1 $0.0500 25%
Face restore GFPGAN inference.facerestore.v1 $0.0008 25%
Segmentation SAM 3 (Segment Anything 3) inference.sam3.segment.v1 $0.0050 25%
Utility Product
Background removal inference-background-removal
NSFW image detection inference-vit-nsfw
Upscale (R-ESRGAN) inference-upscale-2x, inference-upscale-4x, inference-upscale-8x
Upscale (HYPIR) inference-hypir-upscale-v1
Face restore inference-facerestore
Segmentation inference-sam3