Test fixed batching first when the workload is repeatable and offline processing windows are clear. Run the same representative image mix at progressively larger batch sizes, capturing peak memory, end-to-end p95, and the actual work required to process each batch. The useful comparison is not images per second in isolation; it is the improvement in GPU productivity after accounting for the full work associated with a larger batch. OpenRouter’s distinction between throughput and latency supports this separation: speed of completion and delay experienced by a request are different measures.
Use adaptive batching for bursty arrivals when the service has a defined latency target. Check whether the policy forms useful batches during bursts without leaving a long tail of undersized work during lighter traffic. Compare the measured outcome with the best compliant fixed configuration, and reject the policy if better utilization comes with a p95 breach, tighter memory pressure, or no measurable improvement in cost per million images.
For continuous batching, evaluate the complete request path rather than relying on model-level throughput. Measure queueing, execution, result delivery, memory, and end-to-end p95 latency under representative load, then apply the same passing-configuration rule used for fixed and adaptive batching.

Convert checkpoints into cost per million
Use a checkpoint worksheet to convert a measured configuration into a comparable unit cost. Record the batch size, images per million, measured elapsed time per batch, peak memory, p95 latency, quality result, GPU-hours per million, and assumed GPU-hour rate. Keep the assumptions separate from measurements: the elapsed time and latency come from the test, while the rate is an explicit cost assumption that should be replaced with the deployment’s actual rate.
| Checkpoint | Batch 8 | Batch 16 |
|---|---|---|
| Launches per million images | 125,000 | 62,500 |
| Measured seconds per batch | 0.40 | 0.56 |
| GPU-seconds per million | 50,000 | 35,000 |
| GPU-hours per million | 13.89 | 9.72 |
| Illustrative cost at $2 per GPU-hour | $27.78 | $19.44 |
This is an illustrative method example, not a reported benchmark. At batch 8, 125,000 launches multiplied by 0.40 seconds equals 50,000 GPU-seconds, or 13.89 GPU-hours after division by 3,600. At an explicitly assumed $2 per GPU-hour, that is $27.78 per million images. At batch 16, 62,500 launches multiplied by 0.56 seconds equals 35,000 GPU-seconds, or 9.72 GPU-hours, producing an illustrative $19.44 per million at the same assumed rate.
The checkpoint comparison shows a 30.0% reduction: the unrounded cost falls from $27.7778 to $19.4444 per million, and the difference divided by the batch-8 result is 0.3000. Accept that result only if both configurations pass the same quality and service-limit checks. The comparison is not valid if a faster checkpoint changes output quality, exceeds the p95-latency limit, or leaves insufficient peak-memory headroom.
Apply the worksheet stepwise as batch size increases. Compute cost per million from the measured batch duration, not from the batch size alone, and compare each result with the preceding configuration. Increase batch size stepwise and keep the largest configuration whose measured peak memory has headroom and whose p95 latency remains within the production limit; accept the change only if measured GPU work per batch grows slowly enough to offset the extra images.
Worked Example: Run the Numbers
This worked example is an illustration, not a benchmark: a GPU serving team runs a synthetic image-inference workload during 2026. The party is the deployment operator, and the comparison date basis is 2026. Each batch contains identical model, image format, and target sizes. The production limits are a p95 request latency of 200 milliseconds and peak GPU memory of 40 GB. Cost is represented in GPU-seconds per million images, avoiding an invented provider price while preserving the relevant economic comparison.
| Images per launch | Measured GPU-busy time | Images per second | GPU-seconds per million images | Peak memory | p95 latency |
|---|---|---|---|---|---|
| 1 | 0.120 s | 8.3333 | 120,000 | 9 GB | 55 ms |
| 8 | 0.550 s | 14.5455 | 68,750 | 18 GB | 95 ms |
| 16 | 0.850 s | 18.8235 | 53,125 | 28 GB | 135 ms |
| 32 | 1.450 s | 22.06897 | 45,312.5 | 45 GB | 220 ms |
The first calculation is images per second: images per launch divided by measured GPU-busy time. For example, the 16-image batch delivers 16 ÷ 0.850 = 18.8235 images per second. The second calculation converts that result into GPU-seconds per million images: 1,000,000 ÷ 18.8235, equivalently 1,000,000 × 0.850 ÷ 16, producing 53,125 GPU-seconds. Applying the same formula gives 120,000 at batch 1, 68,750 at batch 8, and 45,312.5 at batch 32.
The result illustrates why savings are workload-dependent. Moving from 8 to 16 images reduces GPU-seconds per million by 15,625, or 15,625 ÷ 68,750 = 22.727%, because batch time rises from 0.550 to 0.850 seconds while image count doubles. Moving from 16 to 32 still improves the measured efficiency, from 53,125 to 45,312.5 GPU-seconds per million, but the percentage improvement is smaller because batch time rises to 1.450 seconds.
Batch 16 is therefore the winner for this example: it uses 28 GB, leaving 12 GB of memory headroom, and its 135 ms p95 latency remains below the 200 ms production limit. Batch 32 is rejected because it reaches 45 GB and 220 ms. Against batch 16, a 32-image batch would need a measured batch time below 1.700 seconds to improve GPU-seconds per image, because 1.700 ÷ 32 equals 0.850 ÷ 16. It would still need to satisfy both the memory and latency limits. The operator should choose the largest measured configuration that passes those constraints and improves the measured efficiency result.

Know where the rule breaks
Batch size is not a free efficiency dial. Increasing it can improve GPU utilization, but the gain disappears when extra images add more work than the hardware can absorb efficiently. Treat the largest tested batch as a candidate only after checking both peak memory and end-to-end p95 latency; cost per million is the final decision metric, not a presumed percentage improvement.
| Edge case | When the rule breaks | When the largest passing batch still wins |
|---|---|---|
| VRAM saturation or out-of-memory | Peak allocation leaves no headroom, or one image changes the tensor shape and causes an allocation spike. | All tested inputs fit within a documented memory safety margin. |
| Latency-sensitive interactive requests | p95 request latency or queue time exceeds the service limit. | The batch increase remains inside the measured production SLA. |
| Variable image sizes or prompts | Padding makes a large batch perform substantial unnecessary compute. | Inputs are bucketed by shape or another stable workload property before batching. |
If measured peak GPU memory reaches the documented ceiling or an input can trigger an out-of-memory failure, then reject that batch size and return to the largest configuration with verified headroom. If a single image changes tensor shape, then test that input explicitly instead of extrapolating safety from smaller, similarly shaped inputs.
If p95 latency or queue time crosses the production limit, then reject the larger batch even when its throughput or cost per million looks better. If latency remains within the SLA under representative load, then keep the larger batch only if its measured GPU cost per million is lower.
If image sizes, prompt lengths, or aspect ratios vary widely, then bucket inputs by shape or workload class before measuring batch performance. If padding waste dominates a mixed batch, then compare like-with-like measurements and do not credit the larger batch with savings caused by an easier mix.
If a larger batch passes memory and latency checks but does not improve measured cost per million, then retain the smaller passing batch as the production default. If the larger batch improves cost per million while preserving the memory margin and SLA, then advance to the next tested size rather than assuming that the next increase will be equally beneficial.
Frequently Asked Questions
When should I test fixed batching instead of adaptive batching?
Test fixed batching first when the workload is repeatable and offline processing windows are clear.
Which metrics should I capture as I increase the batch size?
Capture peak memory, end-to-end p95 latency, and the actual work required to process each batch.
Why is images per second not enough for evaluating a larger batch?
The useful comparison is GPU productivity after accounting for the full work associated with a larger batch, not throughput in isolation.
What should I verify before using adaptive batching for bursty traffic?
Verify that the policy forms useful batches during bursts without leaving a long tail of undersized work during lighter traffic.
What should cause me to reject an adaptive batching policy?
Reject it if better utilization comes with a p95 breach, tighter memory pressure, or no measurable improvement in cost per million images.
What must be measured when evaluating continuous batching?
Measure queueing, execution, result delivery, memory, and end-to-end p95 latency under representative load rather than relying on model-level throughput.
Quick answers
| When should fixed batching be tested first? | Test fixed batching first when the workload is repeatable and offline processing windows are clear. |
| What should be measured while progressively increasing batch sizes? | Run the same representative image mix at progressively larger batch sizes, capturing peak memory, end-to-end p95, and the actual work required to process each batch. |
| What is the useful comparison when evaluating a larger batch? | The useful comparison is not images per second in isolation; it is the improvement in GPU productivity after accounting for the full work associated with a larger batch. |
| When should adaptive batching be used? | Use adaptive batching for bursty arrivals when the service has a defined latency target. |
| How should continuous batching be evaluated? | For continuous batching, evaluate the complete request path rather than relying on model-level throughput. |
Also worth reading: Mastering scalable growth strategies with generative AI: Mastering scalable growth strategies with · Transform your product images into professional lifestyle photos with AI: Transform your product images into · Image generation speed test 2026: Stable Diffusion XL 8 vs 32 cap, 32 cost wins: Image generation speed test 2026: