GPUDeck · Production goodput observatory

What an occupied GPU hour finishes

Hourly price alone cannot tell you what finishing the work costs. GPUDeck measures paid completions per occupied GPU-hour inside one exact workload cell, compares named deployments as ratios, and never aggregates unlike workloads into one number.

Latest published measurement: trailing 24 hours ending 24 Aug 2026 11:22 UTC. Page regenerated 2026-08-24 11:45 UTC; one operator; per-workload ratios only.

1 operator · 11 physical GPUs in the source cells · 11 workload-specific comparisons · no cross-workload aggregate

Latest evidence changes

Each row is one workload on one GPU pairing, compared with its own previous published observation — not with your last visit, and not with any other workload. Latest observation 2026-08-24. Of 11 comparisons, 11 moved and 0 appeared for the first time. There is deliberately no overall figure here: averaging these rows together is the error this record exists to refuse.

WorkloadComparisonNowPreviousChangeAttemptsCardsEvidence
dark_beast_krea2_fp8GeForce RTX 4080 vs GeForce RTX 50900.59x0.38x 2026-08-23+53.9%76 / 2942 / 6same operator · replication: none
krea2_turbo_fp8_scaledGeForce RTX 4080 vs GeForce RTX 50900.58x0.42x 2026-08-23+39.2%60 / 1922 / 6same operator · replication: none
krea2_identity_edit_sogni_v0_3_alphaGeForce RTX 4080 vs GeForce RTX 50900.51x0.37x 2026-08-23+38.3%18 / 342 / 5same operator · replication: none
krea2_identity_edit_v1_2GeForce RTX 4080 vs GeForce RTX 50900.56x0.42x 2026-08-23+33.2%41 / 1232 / 6same operator · replication: none
chroma1-hd_fp8_scaledGeForce RTX 4080 vs GeForce RTX 50900.41x0.53x 2026-08-23-21.6%63 / 702 / 4same operator · replication: none
z_image_bf16GeForce RTX 4080 vs GeForce RTX 50900.51x0.45x 2026-08-23+12.8%40 / 222 / 6same operator · replication: none
z_image_turbo_bf16GeForce RTX 4080 vs GeForce RTX 50900.28x0.31x 2026-08-23-9.6%32 / 572 / 4same operator · replication: none
dark_beast_z_image_turbo_v9_bf16GeForce RTX 4080 vs GeForce RTX 50900.30x0.33x 2026-08-23-6.9%95 / 2522 / 4same operator · replication: none
one_obsession_v22_fp16GeForce RTX 4080 vs GeForce RTX 50900.35x0.34x 2026-08-23+2.8%101 / 1242 / 4same operator · replication: none
dark_beast_krea2_identity_edit_v1_2GeForce RTX 4080 vs GeForce RTX 50900.44x0.43x 2026-08-23+2.5%144 / 582 / 6same operator · replication: none
chroma-v48-detail-svd_fp8GeForce RTX 4080 vs GeForce RTX 50900.48x0.48x 2026-08-23-1.5%40 / 912 / 4same operator · replication: none

A change here is a change in this fleet between two dated windows. It is not evidence that any GPU became faster: workload mix, customer cancellation behaviour and which physical cards were free all move these ratios, and none of them is controlled. Rows marked † pair a card running under virtualisation, so they carry a host difference as well as a hardware one.

Goodput

Paid completions per occupied GPU-hour, measured inside one workload at a time. Occupied time includes attempts that were cancelled before they finished, so this counts the resource a card actually consumes rather than only the work that survived.

Observed · single day2026-08-24 · 2,027 attempts

Replication. Every cell on this page comes from a single operator and a single day. 0 of the 11 rows below rest on one physical card on at least one side, marked ●. Across the whole table 11 distinct cards contributed. Attempt counts are not replication: a card that ran 600 attempts in a day is still one card, and confidence that treats those attempts as independent would be wrong by roughly the square root of that count.

The headline does not rest on a single card. Dropping every row that has only one physical card on a side leaves 11 rows and the same range, 1.70x to 3.60x. The thin rows sit inside the range, not at its edges, so they are not setting it.

WorkloadLarger-memory GPUSmaller-memory GPUAttemptsCards A/BSmaller-memory yieldGoodput ratio
dark_beast_krea2_fp8GeForce RTX 5090GeForce RTX 40803706/20.9081.70x
krea2_turbo_fp8_scaledGeForce RTX 5090GeForce RTX 40802526/20.9831.71x
krea2_identity_edit_v1_2GeForce RTX 5090GeForce RTX 40801646/20.9511.80x
krea2_identity_edit_sogni_v0_3_alphaGeForce RTX 5090GeForce RTX 4080525/20.7221.96x
z_image_bf16GeForce RTX 5090GeForce RTX 4080626/20.9501.96x
chroma-v48-detail-svd_fp8GeForce RTX 5090GeForce RTX 40801314/20.9252.10x
dark_beast_krea2_identity_edit_v1_2GeForce RTX 5090GeForce RTX 40802026/20.9242.27x
chroma1-hd_fp8_scaledGeForce RTX 5090GeForce RTX 40801334/20.8572.43x
one_obsession_v22_fp16GeForce RTX 5090GeForce RTX 40802254/20.9802.88x
dark_beast_z_image_turbo_v9_bf16GeForce RTX 5090GeForce RTX 40803474/20.9893.28x
z_image_turbo_bf16GeForce RTX 5090GeForce RTX 4080894/20.9693.60x

A deployment here is a GPU together with the host, virtualisation layer, software stack and dispatch it ran under — not the card in isolation. These ratios do not isolate a hardware effect, which is why the confounded rows are separated above rather than averaged in.

Same memory · different class2026-08-24 · 4,864 attempts

The pairings in this table

  • 32 GB — GeForce RTX 5090 against RTX PRO 4500 Blackwell, across 27 workloads
  • 16 GB † — GeForce RTX 4080 against GeForce RTX 5070 Ti, across 9 workloads, host-confounded

† one GPU in that pairing runs under a Windows virtualisation layer reserving roughly 2.4 GB, so those rows compare production deployments rather than silicon.

WorkloadGoodput ratio (A ÷ B)Attempts AAttempts BMemoryGPU AGPU B
32 GB — GeForce RTX 5090 (A) against RTX PRO 4500 Blackwell (B) · goodput A÷B, above 1.00 favours A · observed min–max 1.12x to 3.11x across 27 workloads · weakest support 8 attempts, worst A:B imbalance 5.4:1
one_obsession_v22_fp161.12x1246632 GBGeForce RTX 5090RTX PRO 4500 Blackwell
minimax-h3-fl2va-fp8_t2v1.28x331432 GBGeForce RTX 5090RTX PRO 4500 Blackwell
ltx25-22b-int8_i2v_distilled1.39x105432 GBGeForce RTX 5090RTX PRO 4500 Blackwell
ltx23-22b-10eros-v1.4-fp8mixed_i2v1.40x91732 GBGeForce RTX 5090RTX PRO 4500 Blackwell
dark_beast_krea2_fp81.50x29430732 GBGeForce RTX 5090RTX PRO 4500 Blackwell
minimax-h3-fl2va-fp8_i2v_turbo1.54x5912732 GBGeForce RTX 5090RTX PRO 4500 Blackwell
wan_v2.2-14b-fp8_animate-replace_lightx2v1.57x321432 GBGeForce RTX 5090RTX PRO 4500 Blackwell
qwen_image_edit_2511_fp8_lightning1.67x7011432 GBGeForce RTX 5090RTX PRO 4500 Blackwell
minimax-h3-fl2va-fp8_flf2v_turbo1.72x94132 GBGeForce RTX 5090RTX PRO 4500 Blackwell
krea2_identity_edit_v1_21.74x12329432 GBGeForce RTX 5090RTX PRO 4500 Blackwell
krea2_identity_edit_sogni_v0_3_alpha1.75x344732 GBGeForce RTX 5090RTX PRO 4500 Blackwell
dark_beast_krea2_identity_edit_v1_21.76x5812532 GBGeForce RTX 5090RTX PRO 4500 Blackwell
z_image_bf161.77x221132 GBGeForce RTX 5090RTX PRO 4500 Blackwell
minimax-h3-fl2va-fp8_i2v1.78x132832 GBGeForce RTX 5090RTX PRO 4500 Blackwell
chroma-v48-detail-svd_fp81.83x914932 GBGeForce RTX 5090RTX PRO 4500 Blackwell
ltx23-22b-fp8_i2v_distilled1.88x352432 GBGeForce RTX 5090RTX PRO 4500 Blackwell
ltx23-22b-fp8_t2v_distilled1.90x191932 GBGeForce RTX 5090RTX PRO 4500 Blackwell
dark_beast_z_image_turbo_v9_bf161.92x2527032 GBGeForce RTX 5090RTX PRO 4500 Blackwell
krea2_turbo_fp8_scaled1.94x19235932 GBGeForce RTX 5090RTX PRO 4500 Blackwell
minimax-h3-ref2va-fp8_r2v1.94x312932 GBGeForce RTX 5090RTX PRO 4500 Blackwell
wan_v2.2-14b-fp8_i2v_lightx2v2.05x393432 GBGeForce RTX 5090RTX PRO 4500 Blackwell
chroma1-hd_fp8_scaled2.15x702232 GBGeForce RTX 5090RTX PRO 4500 Blackwell
wan_v2.2-14b-fp8_t2v_lightx2v2.34x142632 GBGeForce RTX 5090RTX PRO 4500 Blackwell
minimax-h3-ref2va-fp8_r2v_turbo2.63x174532 GBGeForce RTX 5090RTX PRO 4500 Blackwell
qwen_image_edit_2511_fp82.72x261932 GBGeForce RTX 5090RTX PRO 4500 Blackwell
ltx25-22b-int8_t2v_distilled2.80x81232 GBGeForce RTX 5090RTX PRO 4500 Blackwell
minimax-h3-fl2va-fp8_t2v_turbo3.11x143832 GBGeForce RTX 5090RTX PRO 4500 Blackwell
16 GB † — GeForce RTX 4080 (A) against GeForce RTX 5070 Ti (B) · goodput A÷B, above 1.00 favours A · observed min–max 1.07x to 3.33x across 9 workloads · weakest support 11 attempts, worst A:B imbalance 12.0:1 · GPU and host changed together so the GPU effect cannot be isolated; descriptive only, excluded from the headline
one_obsession_v22_fp161.07x1017116 GBGeForce RTX 4080GeForce RTX 5070 Ti
chroma-v48-detail-svd_fp81.07x402316 GBGeForce RTX 4080GeForce RTX 5070 Ti
dark_beast_krea2_identity_edit_v1_21.08x1441216 GBGeForce RTX 4080GeForce RTX 5070 Ti
chroma1-hd_fp8_scaled1.11x632916 GBGeForce RTX 4080GeForce RTX 5070 Ti
dark_beast_krea2_fp81.23x763616 GBGeForce RTX 4080GeForce RTX 5070 Ti
dark_beast_z_image_turbo_v9_bf162.18x951116 GBGeForce RTX 4080GeForce RTX 5070 Ti
krea2_identity_edit_sogni_v0_3_alpha2.62x183216 GBGeForce RTX 4080GeForce RTX 5070 Ti
krea2_identity_edit_v1_22.78x4111316 GBGeForce RTX 4080GeForce RTX 5070 Ti
krea2_turbo_fp8_scaled3.33x6019616 GBGeForce RTX 4080GeForce RTX 5070 Ti

Hardware evidence here is 2 GPU-model pairings, not 36 comparisons. The 36 rows above are workload-level measurements nested inside those pairings. Excluding the pairing with a known host confound, equal-memory goodput ranges from 1.12x to 3.11x across 27 workloads on a single pairing — GeForce RTX 5090 against RTX PRO 4500 Blackwell.

That is one pairing profiled across many workloads, plus one configuration comparison that cannot isolate a GPU effect — not a replicated estimate of a general hardware effect. The remaining pairing is not demonstrated to be host-matched either; it is free of the one host confound identified here, which is a weaker and more accurate statement.

9 of these rows are not a card-isolated comparison. One GPU in that pairing runs under a Windows virtualisation layer that reserves roughly 2.4 GB and cannot be configured away, so those rows compare two production deployments and not two pieces of silicon. They are kept because the deployment is what a buyer actually rents, and marked because the distinction changes what they can be used to argue. The attempt counts in those rows are also strongly unequal, so the widest of them rests on far fewer observations on one side than the other — read it as descriptive, not as a stable magnitude.

The two GPUs in each row are listed in a fixed alphabetical catalogue order settled before any measurement was taken, and the ratio is always the first divided by the second. That order is arbitrary by design: it cannot be influenced by how the comparison turns out. It is deliberately not the larger figure over the smaller. That form cannot fall below 1.00, so it reports two equal-sized differences pointing in opposite directions as the same number, and the direction of the result cannot be recovered from the page. These 36 rows cover 2 distinct GPU-model pairings, not 36 independent ones — the same two GPU MODELS recur across many workloads, so the rows show how one comparison behaves across different work, not many separate comparisons. Each side pools every card of that model in the fleet, so a pairing is a model-to-model comparison and not two individual cards. Ratios are stated per workload and per memory size, and no row is combined with any other.

Goodput is measured on production hosts, and the host is inside the measurement. The class anchoring the widest ratios above runs under a virtualisation layer that reserves part of its memory and cannot be configured away, so treat those ratios as a result that includes a known hosting disadvantage rather than as a property of the hardware alone. The equal-memory table, where both classes sit on comparable hosts, is the cleaner contrast.

Stated per workload and never combined. A single figure across workloads would move with the mix of work each class happens to be sent, which is a fact about dispatch rather than about the hardware — the same substitution this record exists to refuse. Yield is recorded as production loss under the observed network regime, not as a property of any class.

Break-even price ratio

How much more per hour the larger-memory class can cost before the smaller one becomes the cheaper way to finish this workload. This is a ratio of two goodput figures measured on the same fleet in the same window — it contains no marketplace price, so nothing here depends on a listing being launchable, comparable, or representative. Bring your own quote and compare it against the threshold.

Derived from the goodput table above — no additional observations. This is the same set of measured pairs, reinterpreted as a price threshold rather than an output ratio, so it is a second reading of one body of evidence and not a second, independent result.

Measured · no external price11 workloads · 2,027 attempts across both GPUs
WorkloadLarger-memory GPUSmaller-memory GPUBreak-even price multiple
z_image_turbo_bf16GeForce RTX 5090 · 32 GBGeForce RTX 4080 · 16 GB3.60x
dark_beast_z_image_turbo_v9_bf16GeForce RTX 5090 · 32 GBGeForce RTX 4080 · 16 GB3.28x
one_obsession_v22_fp16GeForce RTX 5090 · 32 GBGeForce RTX 4080 · 16 GB2.88x
chroma1-hd_fp8_scaledGeForce RTX 5090 · 32 GBGeForce RTX 4080 · 16 GB2.43x
dark_beast_krea2_identity_edit_v1_2GeForce RTX 5090 · 32 GBGeForce RTX 4080 · 16 GB2.27x
chroma-v48-detail-svd_fp8GeForce RTX 5090 · 32 GBGeForce RTX 4080 · 16 GB2.10x
z_image_bf16GeForce RTX 5090 · 32 GBGeForce RTX 4080 · 16 GB1.96x
krea2_identity_edit_sogni_v0_3_alphaGeForce RTX 5090 · 32 GBGeForce RTX 4080 · 16 GB1.96x
krea2_identity_edit_v1_2GeForce RTX 5090 · 32 GBGeForce RTX 4080 · 16 GB1.80x
krea2_turbo_fp8_scaledGeForce RTX 5090 · 32 GBGeForce RTX 4080 · 16 GB1.71x
dark_beast_krea2_fp8GeForce RTX 5090 · 32 GBGeForce RTX 4080 · 16 GB1.70x

Read every row as a threshold, not a price. A multiple of 3.60x on a row says the larger-memory GPU is the cheaper way to finish that workload at any hourly rate below 3.60 times the smaller one; at exactly that multiple the two cost the same. The daggered rows are not usable this way, because part of their headroom is a hosting penalty specific to this fleet.

Read it as a threshold, not a price. A ratio of 4.00x says the larger class is the cheaper way to finish this workload at any hourly multiple below four times; above it, the smaller class wins. The comparison holds where both classes see the same share of billed time spent occupied — where utilisation differs, scale each side by its own occupied-to-billed fraction before comparing.

Pairs are fixed by largest memory, then lowest permanent class index, both chosen before any measurement was read. Selecting the fastest and slowest observed classes instead would pick each comparison on its own outcome. These 11 rows are drawn from 1 distinct GPU-model pairings, so they show how a small number of comparisons behave across many workloads rather than that many independent comparisons; the GPUs in each row are named so this is checkable.

Diagnostics — latency and completion yield, the two inputs goodput is built from

Completed-render latency ratios

A single fixed pair of GPUs — RTX 5090 at 32 GB against RTX 4080 at 16 GB — measured on the same completed workloads. Each row reports how much longer the RTX 4080 took on paid renders of that workload. Every row is this same one pairing, so the rows show how one comparison varies across workloads, not a set of independent comparisons.

Exploratory · single-day2026-08-23 · 6,434 completed renders
WorkloadCompleted rendersRTX 4080 took longer than RTX 5090 by
chroma1-hd_fp8_scaled851.60x
flux1-schnell-fp8192.07x
z_image_bf16702.30x
krea2_identity_edit_sogni_v0_3_alpha2272.44x
dark_beast_krea2_identity_edit_v1_25252.46x
krea2_identity_edit_v1_28322.47x
chroma-v48-detail-svd_fp82042.67x
krea2_turbo_fp8_scaled1,6812.88x
dark_beast_krea2_fp82,0273.01x
one_obsession_v22_fp162724.42x
z_image_turbo_bf16964.85x
dark_beast_z_image_turbo_v9_bf163966.20x

The observed latency gap runs from 1.60x to 6.20x. All workloads point in the same direction, but one day cannot establish that the multiplier itself varies by workload. That requires consecutive days and an explicit interaction test.

On the widest-gap workload, p90/p10 dispersion was 4.2x for the RTX 4080 and 9.4x for the RTX 5090. That comparison is descriptive, not a confidence interval. These are completed-render latency ratios, not throughput ratios; throughput would include all occupied time, including cancelled and failed attempts.

Completion yield

Of the attempts a class started, the fraction that finished and were paid. An attempt that is cancelled or fails still occupies the card, so yield is the difference between time bought and work delivered.

Observed · single closed day2026-08-23 · 8,839 attempts started
GPUMemoryAttempts startedCompleted and paidYield
GeForce RTX 5070 Ti16 GB6205400.871
GeForce RTX 308020 GB6155390.876
GeForce RTX 408016 GB1,7261,6050.930
GeForce RTX 509032 GB5,8785,6690.964
All GPUs8,8398,3530.945

Yield ran from 0.871 to 0.964 across GPUs on this day, and it does NOT order by memory capacity. The 16 GB RTX 4080 converts more of what it starts than the 20 GB RTX 3080. An earlier version of this table bucketed these figures into 16/20/32 GB classes and reported that yield ordered the same way latency does — but averaging the 4080 together with the 5070 Ti pulled the 16 GB bucket below the 20 GB one and manufactured that order. The bucketing produced the finding; the per-GPU data does not contain it.

What does survive is that the lowest yield belongs to the GPU running under the virtualisation layer, and the highest to the fastest card. Yield is still a real cost — an attempt that is cancelled or fails occupies the card and pays nothing — so a latency ratio alone understates the gap in delivered work. It is just not a property of how much memory a card has.

Read this as production loss recorded under the observed network regime, not as a property of any hardware class. An outcome here is produced by class, workload, customer behaviour, interface and elapsed exposure together — and a longer render simply offers a larger window for a cancellation to arrive in, which alone would produce part of this ordering. The census does not separate those causes.

Data, scope and citation

The download, method and boundary live together so the editorial argument cannot outrun the evidence behind it.

Scope. Every goodput cell comes from one operator’s production fleet. Hardware, workload and measurement window are named; operator identity, host identity and absolute throughput are not. Ratios are never aggregated across workloads.

Method. Goodput is paid completions per occupied GPU-hour inside one exact (GPU, workload) cell. Attempt counts are support, not replication; physical-card counts are published beside them. Latency and yield remain diagnostics because neither is goodput on its own.

Separate record. The longitudinal public-network census now lives at /census. It is context, not an input to the goodput ratios, and is never used to weight them.

Citation. Reuse is licensed under CC BY 4.0. The citation package records retrieval, attribution and what the series does not measure.