Hourly price alone cannot tell you what finishing the work costs. GPUDeck measures paid completions per occupied GPU-hour inside one exact workload cell, compares named deployments as ratios, and never aggregates unlike workloads into one number.
Latest published measurement: trailing 24 hours ending 24 Aug 2026 11:22 UTC. Page regenerated 2026-08-24 11:45 UTC; one operator; per-workload ratios only.
1 operator · 11 physical GPUs in the source cells · 11 workload-specific comparisons · no cross-workload aggregate
Each row is one workload on one GPU pairing, compared with its own previous published observation — not with your last visit, and not with any other workload. Latest observation 2026-08-24. Of 11 comparisons, 11 moved and 0 appeared for the first time. There is deliberately no overall figure here: averaging these rows together is the error this record exists to refuse.
| Workload | Comparison | Now | Previous | Change | Attempts | Cards | Evidence |
|---|---|---|---|---|---|---|---|
| dark_beast_krea2_fp8 | GeForce RTX 4080 vs GeForce RTX 5090 | 0.59x | 0.38x 2026-08-23 | +53.9% | 76 / 294 | 2 / 6 | same operator · replication: none |
| krea2_turbo_fp8_scaled | GeForce RTX 4080 vs GeForce RTX 5090 | 0.58x | 0.42x 2026-08-23 | +39.2% | 60 / 192 | 2 / 6 | same operator · replication: none |
| krea2_identity_edit_sogni_v0_3_alpha | GeForce RTX 4080 vs GeForce RTX 5090 | 0.51x | 0.37x 2026-08-23 | +38.3% | 18 / 34 | 2 / 5 | same operator · replication: none |
| krea2_identity_edit_v1_2 | GeForce RTX 4080 vs GeForce RTX 5090 | 0.56x | 0.42x 2026-08-23 | +33.2% | 41 / 123 | 2 / 6 | same operator · replication: none |
| chroma1-hd_fp8_scaled | GeForce RTX 4080 vs GeForce RTX 5090 | 0.41x | 0.53x 2026-08-23 | -21.6% | 63 / 70 | 2 / 4 | same operator · replication: none |
| z_image_bf16 | GeForce RTX 4080 vs GeForce RTX 5090 | 0.51x | 0.45x 2026-08-23 | +12.8% | 40 / 22 | 2 / 6 | same operator · replication: none |
| z_image_turbo_bf16 | GeForce RTX 4080 vs GeForce RTX 5090 | 0.28x | 0.31x 2026-08-23 | -9.6% | 32 / 57 | 2 / 4 | same operator · replication: none |
| dark_beast_z_image_turbo_v9_bf16 | GeForce RTX 4080 vs GeForce RTX 5090 | 0.30x | 0.33x 2026-08-23 | -6.9% | 95 / 252 | 2 / 4 | same operator · replication: none |
| one_obsession_v22_fp16 | GeForce RTX 4080 vs GeForce RTX 5090 | 0.35x | 0.34x 2026-08-23 | +2.8% | 101 / 124 | 2 / 4 | same operator · replication: none |
| dark_beast_krea2_identity_edit_v1_2 | GeForce RTX 4080 vs GeForce RTX 5090 | 0.44x | 0.43x 2026-08-23 | +2.5% | 144 / 58 | 2 / 6 | same operator · replication: none |
| chroma-v48-detail-svd_fp8 | GeForce RTX 4080 vs GeForce RTX 5090 | 0.48x | 0.48x 2026-08-23 | -1.5% | 40 / 91 | 2 / 4 | same operator · replication: none |
A change here is a change in this fleet between two dated windows. It is not evidence that any GPU became faster: workload mix, customer cancellation behaviour and which physical cards were free all move these ratios, and none of them is controlled. Rows marked † pair a card running under virtualisation, so they carry a host difference as well as a hardware one.
Paid completions per occupied GPU-hour, measured inside one workload at a time. Occupied time includes attempts that were cancelled before they finished, so this counts the resource a card actually consumes rather than only the work that survived.
Replication. Every cell on this page comes from a single operator and a single day. 0 of the 11 rows below rest on one physical card on at least one side, marked ●. Across the whole table 11 distinct cards contributed. Attempt counts are not replication: a card that ran 600 attempts in a day is still one card, and confidence that treats those attempts as independent would be wrong by roughly the square root of that count.
The headline does not rest on a single card. Dropping every row that has only one physical card on a side leaves 11 rows and the same range, 1.70x to 3.60x. The thin rows sit inside the range, not at its edges, so they are not setting it.
| Workload | Larger-memory GPU | Smaller-memory GPU | Attempts | Cards A/B | Smaller-memory yield | Goodput ratio |
|---|---|---|---|---|---|---|
| dark_beast_krea2_fp8 | GeForce RTX 5090 | GeForce RTX 4080 | 370 | 6/2 | 0.908 | 1.70x |
| krea2_turbo_fp8_scaled | GeForce RTX 5090 | GeForce RTX 4080 | 252 | 6/2 | 0.983 | 1.71x |
| krea2_identity_edit_v1_2 | GeForce RTX 5090 | GeForce RTX 4080 | 164 | 6/2 | 0.951 | 1.80x |
| krea2_identity_edit_sogni_v0_3_alpha | GeForce RTX 5090 | GeForce RTX 4080 | 52 | 5/2 | 0.722 | 1.96x |
| z_image_bf16 | GeForce RTX 5090 | GeForce RTX 4080 | 62 | 6/2 | 0.950 | 1.96x |
| chroma-v48-detail-svd_fp8 | GeForce RTX 5090 | GeForce RTX 4080 | 131 | 4/2 | 0.925 | 2.10x |
| dark_beast_krea2_identity_edit_v1_2 | GeForce RTX 5090 | GeForce RTX 4080 | 202 | 6/2 | 0.924 | 2.27x |
| chroma1-hd_fp8_scaled | GeForce RTX 5090 | GeForce RTX 4080 | 133 | 4/2 | 0.857 | 2.43x |
| one_obsession_v22_fp16 | GeForce RTX 5090 | GeForce RTX 4080 | 225 | 4/2 | 0.980 | 2.88x |
| dark_beast_z_image_turbo_v9_bf16 | GeForce RTX 5090 | GeForce RTX 4080 | 347 | 4/2 | 0.989 | 3.28x |
| z_image_turbo_bf16 | GeForce RTX 5090 | GeForce RTX 4080 | 89 | 4/2 | 0.969 | 3.60x |
A deployment here is a GPU together with the host, virtualisation layer, software stack and dispatch it ran under — not the card in isolation. These ratios do not isolate a hardware effect, which is why the confounded rows are separated above rather than averaged in.
The pairings in this table
† one GPU in that pairing runs under a Windows virtualisation layer reserving roughly 2.4 GB, so those rows compare production deployments rather than silicon.
| Workload | Goodput ratio (A ÷ B) | Attempts A | Attempts B | Memory | GPU A | GPU B |
|---|---|---|---|---|---|---|
| 32 GB — GeForce RTX 5090 (A) against RTX PRO 4500 Blackwell (B) · goodput A÷B, above 1.00 favours A · observed min–max 1.12x to 3.11x across 27 workloads · weakest support 8 attempts, worst A:B imbalance 5.4:1 | ||||||
| one_obsession_v22_fp16 | 1.12x | 124 | 66 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| minimax-h3-fl2va-fp8_t2v | 1.28x | 33 | 14 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| ltx25-22b-int8_i2v_distilled | 1.39x | 10 | 54 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| ltx23-22b-10eros-v1.4-fp8mixed_i2v | 1.40x | 9 | 17 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| dark_beast_krea2_fp8 | 1.50x | 294 | 307 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| minimax-h3-fl2va-fp8_i2v_turbo | 1.54x | 59 | 127 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| wan_v2.2-14b-fp8_animate-replace_lightx2v | 1.57x | 32 | 14 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| qwen_image_edit_2511_fp8_lightning | 1.67x | 70 | 114 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| minimax-h3-fl2va-fp8_flf2v_turbo | 1.72x | 9 | 41 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| krea2_identity_edit_v1_2 | 1.74x | 123 | 294 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| krea2_identity_edit_sogni_v0_3_alpha | 1.75x | 34 | 47 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| dark_beast_krea2_identity_edit_v1_2 | 1.76x | 58 | 125 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| z_image_bf16 | 1.77x | 22 | 11 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| minimax-h3-fl2va-fp8_i2v | 1.78x | 13 | 28 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| chroma-v48-detail-svd_fp8 | 1.83x | 91 | 49 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| ltx23-22b-fp8_i2v_distilled | 1.88x | 35 | 24 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| ltx23-22b-fp8_t2v_distilled | 1.90x | 19 | 19 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| dark_beast_z_image_turbo_v9_bf16 | 1.92x | 252 | 70 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| krea2_turbo_fp8_scaled | 1.94x | 192 | 359 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| minimax-h3-ref2va-fp8_r2v | 1.94x | 31 | 29 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| wan_v2.2-14b-fp8_i2v_lightx2v | 2.05x | 39 | 34 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| chroma1-hd_fp8_scaled | 2.15x | 70 | 22 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| wan_v2.2-14b-fp8_t2v_lightx2v | 2.34x | 14 | 26 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| minimax-h3-ref2va-fp8_r2v_turbo | 2.63x | 17 | 45 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| qwen_image_edit_2511_fp8 | 2.72x | 26 | 19 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| ltx25-22b-int8_t2v_distilled | 2.80x | 8 | 12 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| minimax-h3-fl2va-fp8_t2v_turbo | 3.11x | 14 | 38 | 32 GB | GeForce RTX 5090 | RTX PRO 4500 Blackwell |
| 16 GB † — GeForce RTX 4080 (A) against GeForce RTX 5070 Ti (B) · goodput A÷B, above 1.00 favours A · observed min–max 1.07x to 3.33x across 9 workloads · weakest support 11 attempts, worst A:B imbalance 12.0:1 · GPU and host changed together so the GPU effect cannot be isolated; descriptive only, excluded from the headline | ||||||
| one_obsession_v22_fp16 | 1.07x | 101 | 71 | 16 GB | GeForce RTX 4080 | GeForce RTX 5070 Ti |
| chroma-v48-detail-svd_fp8 | 1.07x | 40 | 23 | 16 GB | GeForce RTX 4080 | GeForce RTX 5070 Ti |
| dark_beast_krea2_identity_edit_v1_2 | 1.08x | 144 | 12 | 16 GB | GeForce RTX 4080 | GeForce RTX 5070 Ti |
| chroma1-hd_fp8_scaled | 1.11x | 63 | 29 | 16 GB | GeForce RTX 4080 | GeForce RTX 5070 Ti |
| dark_beast_krea2_fp8 | 1.23x | 76 | 36 | 16 GB | GeForce RTX 4080 | GeForce RTX 5070 Ti |
| dark_beast_z_image_turbo_v9_bf16 | 2.18x | 95 | 11 | 16 GB | GeForce RTX 4080 | GeForce RTX 5070 Ti |
| krea2_identity_edit_sogni_v0_3_alpha | 2.62x | 18 | 32 | 16 GB | GeForce RTX 4080 | GeForce RTX 5070 Ti |
| krea2_identity_edit_v1_2 | 2.78x | 41 | 113 | 16 GB | GeForce RTX 4080 | GeForce RTX 5070 Ti |
| krea2_turbo_fp8_scaled | 3.33x | 60 | 196 | 16 GB | GeForce RTX 4080 | GeForce RTX 5070 Ti |
Hardware evidence here is 2 GPU-model pairings, not 36 comparisons. The 36 rows above are workload-level measurements nested inside those pairings. Excluding the pairing with a known host confound, equal-memory goodput ranges from 1.12x to 3.11x across 27 workloads on a single pairing — GeForce RTX 5090 against RTX PRO 4500 Blackwell.
That is one pairing profiled across many workloads, plus one configuration comparison that cannot isolate a GPU effect — not a replicated estimate of a general hardware effect. The remaining pairing is not demonstrated to be host-matched either; it is free of the one host confound identified here, which is a weaker and more accurate statement.
† 9 of these rows are not a card-isolated comparison. One GPU in that pairing runs under a Windows virtualisation layer that reserves roughly 2.4 GB and cannot be configured away, so those rows compare two production deployments and not two pieces of silicon. They are kept because the deployment is what a buyer actually rents, and marked because the distinction changes what they can be used to argue. The attempt counts in those rows are also strongly unequal, so the widest of them rests on far fewer observations on one side than the other — read it as descriptive, not as a stable magnitude.
The two GPUs in each row are listed in a fixed alphabetical catalogue order settled before any measurement was taken, and the ratio is always the first divided by the second. That order is arbitrary by design: it cannot be influenced by how the comparison turns out. It is deliberately not the larger figure over the smaller. That form cannot fall below 1.00, so it reports two equal-sized differences pointing in opposite directions as the same number, and the direction of the result cannot be recovered from the page. These 36 rows cover 2 distinct GPU-model pairings, not 36 independent ones — the same two GPU MODELS recur across many workloads, so the rows show how one comparison behaves across different work, not many separate comparisons. Each side pools every card of that model in the fleet, so a pairing is a model-to-model comparison and not two individual cards. Ratios are stated per workload and per memory size, and no row is combined with any other.
Goodput is measured on production hosts, and the host is inside the measurement. The class anchoring the widest ratios above runs under a virtualisation layer that reserves part of its memory and cannot be configured away, so treat those ratios as a result that includes a known hosting disadvantage rather than as a property of the hardware alone. The equal-memory table, where both classes sit on comparable hosts, is the cleaner contrast.
Stated per workload and never combined. A single figure across workloads would move with the mix of work each class happens to be sent, which is a fact about dispatch rather than about the hardware — the same substitution this record exists to refuse. Yield is recorded as production loss under the observed network regime, not as a property of any class.
How much more per hour the larger-memory class can cost before the smaller one becomes the cheaper way to finish this workload. This is a ratio of two goodput figures measured on the same fleet in the same window — it contains no marketplace price, so nothing here depends on a listing being launchable, comparable, or representative. Bring your own quote and compare it against the threshold.
Derived from the goodput table above — no additional observations. This is the same set of measured pairs, reinterpreted as a price threshold rather than an output ratio, so it is a second reading of one body of evidence and not a second, independent result.
| Workload | Larger-memory GPU | Smaller-memory GPU | Break-even price multiple |
|---|---|---|---|
| z_image_turbo_bf16 | GeForce RTX 5090 · 32 GB | GeForce RTX 4080 · 16 GB | 3.60x |
| dark_beast_z_image_turbo_v9_bf16 | GeForce RTX 5090 · 32 GB | GeForce RTX 4080 · 16 GB | 3.28x |
| one_obsession_v22_fp16 | GeForce RTX 5090 · 32 GB | GeForce RTX 4080 · 16 GB | 2.88x |
| chroma1-hd_fp8_scaled | GeForce RTX 5090 · 32 GB | GeForce RTX 4080 · 16 GB | 2.43x |
| dark_beast_krea2_identity_edit_v1_2 | GeForce RTX 5090 · 32 GB | GeForce RTX 4080 · 16 GB | 2.27x |
| chroma-v48-detail-svd_fp8 | GeForce RTX 5090 · 32 GB | GeForce RTX 4080 · 16 GB | 2.10x |
| z_image_bf16 | GeForce RTX 5090 · 32 GB | GeForce RTX 4080 · 16 GB | 1.96x |
| krea2_identity_edit_sogni_v0_3_alpha | GeForce RTX 5090 · 32 GB | GeForce RTX 4080 · 16 GB | 1.96x |
| krea2_identity_edit_v1_2 | GeForce RTX 5090 · 32 GB | GeForce RTX 4080 · 16 GB | 1.80x |
| krea2_turbo_fp8_scaled | GeForce RTX 5090 · 32 GB | GeForce RTX 4080 · 16 GB | 1.71x |
| dark_beast_krea2_fp8 | GeForce RTX 5090 · 32 GB | GeForce RTX 4080 · 16 GB | 1.70x |
Read every row as a threshold, not a price. A multiple of 3.60x on a row says the larger-memory GPU is the cheaper way to finish that workload at any hourly rate below 3.60 times the smaller one; at exactly that multiple the two cost the same. The daggered rows are not usable this way, because part of their headroom is a hosting penalty specific to this fleet.
Read it as a threshold, not a price. A ratio of 4.00x says the larger class is the cheaper way to finish this workload at any hourly multiple below four times; above it, the smaller class wins. The comparison holds where both classes see the same share of billed time spent occupied — where utilisation differs, scale each side by its own occupied-to-billed fraction before comparing.
Pairs are fixed by largest memory, then lowest permanent class index, both chosen before any measurement was read. Selecting the fastest and slowest observed classes instead would pick each comparison on its own outcome. These 11 rows are drawn from 1 distinct GPU-model pairings, so they show how a small number of comparisons behave across many workloads rather than that many independent comparisons; the GPUs in each row are named so this is checkable.
A single fixed pair of GPUs — RTX 5090 at 32 GB against RTX 4080 at 16 GB — measured on the same completed workloads. Each row reports how much longer the RTX 4080 took on paid renders of that workload. Every row is this same one pairing, so the rows show how one comparison varies across workloads, not a set of independent comparisons.
| Workload | Completed renders | RTX 4080 took longer than RTX 5090 by |
|---|---|---|
| chroma1-hd_fp8_scaled | 85 | 1.60x |
| flux1-schnell-fp8 | 19 | 2.07x |
| z_image_bf16 | 70 | 2.30x |
| krea2_identity_edit_sogni_v0_3_alpha | 227 | 2.44x |
| dark_beast_krea2_identity_edit_v1_2 | 525 | 2.46x |
| krea2_identity_edit_v1_2 | 832 | 2.47x |
| chroma-v48-detail-svd_fp8 | 204 | 2.67x |
| krea2_turbo_fp8_scaled | 1,681 | 2.88x |
| dark_beast_krea2_fp8 | 2,027 | 3.01x |
| one_obsession_v22_fp16 | 272 | 4.42x |
| z_image_turbo_bf16 | 96 | 4.85x |
| dark_beast_z_image_turbo_v9_bf16 | 396 | 6.20x |
The observed latency gap runs from 1.60x to 6.20x. All workloads point in the same direction, but one day cannot establish that the multiplier itself varies by workload. That requires consecutive days and an explicit interaction test.
On the widest-gap workload, p90/p10 dispersion was 4.2x for the RTX 4080 and 9.4x for the RTX 5090. That comparison is descriptive, not a confidence interval. These are completed-render latency ratios, not throughput ratios; throughput would include all occupied time, including cancelled and failed attempts.
Of the attempts a class started, the fraction that finished and were paid. An attempt that is cancelled or fails still occupies the card, so yield is the difference between time bought and work delivered.
| GPU | Memory | Attempts started | Completed and paid | Yield |
|---|---|---|---|---|
| GeForce RTX 5070 Ti | 16 GB | 620 | 540 | 0.871 |
| GeForce RTX 3080 | 20 GB | 615 | 539 | 0.876 |
| GeForce RTX 4080 | 16 GB | 1,726 | 1,605 | 0.930 |
| GeForce RTX 5090 | 32 GB | 5,878 | 5,669 | 0.964 |
| All GPUs | — | 8,839 | 8,353 | 0.945 |
Yield ran from 0.871 to 0.964 across GPUs on this day, and it does NOT order by memory capacity. The 16 GB RTX 4080 converts more of what it starts than the 20 GB RTX 3080. An earlier version of this table bucketed these figures into 16/20/32 GB classes and reported that yield ordered the same way latency does — but averaging the 4080 together with the 5070 Ti pulled the 16 GB bucket below the 20 GB one and manufactured that order. The bucketing produced the finding; the per-GPU data does not contain it.
What does survive is that the lowest yield belongs to the GPU running under the virtualisation layer, and the highest to the fastest card. Yield is still a real cost — an attempt that is cancelled or fails occupies the card and pays nothing — so a latency ratio alone understates the gap in delivered work. It is just not a property of how much memory a card has.
Read this as production loss recorded under the observed network regime, not as a property of any hardware class. An outcome here is produced by class, workload, customer behaviour, interface and elapsed exposure together — and a longer render simply offers a larger window for a cancellation to arrive in, which alone would produce part of this ordering. The census does not separate those causes.
The download, method and boundary live together so the editorial argument cannot outrun the evidence behind it.
Scope. Every goodput cell comes from one operator’s production fleet. Hardware, workload and measurement window are named; operator identity, host identity and absolute throughput are not. Ratios are never aggregated across workloads.
Method. Goodput is paid completions per occupied GPU-hour inside one exact (GPU, workload) cell. Attempt counts are support, not replication; physical-card counts are published beside them. Latency and yield remain diagnostics because neither is goodput on its own.
Separate record. The longitudinal public-network census now lives at /census. It is context, not an input to the goodput ratios, and is never used to weight them.
Citation. Reuse is licensed under CC BY 4.0. The citation package records retrieval, attribution and what the series does not measure.