Hourly price alone cannot tell you what finishing the work costs. GPUDeck measures paid completions per occupied GPU-hour inside one exact workload cell, compares named deployments as ratios, and never aggregates unlike workloads into one number.
Latest published measurement: trailing 24 hours ending 4 Sep 2026 01:21 UTC. Page regenerated 2026-09-23 05:21 UTC; one operator; per-workload ratios only.
1 operator · 11 physical GPUs in the source cells · 12 workload-specific comparisons · no cross-workload aggregate
Each row is one workload on one GPU pairing, compared with its own previous published observation — not with your last visit, and not with any other workload. Latest observation 2026-09-04. Of 12 comparisons, 11 moved and 1 appeared for the first time. There is deliberately no overall figure here: averaging these rows together is the error this record exists to refuse.
| Workload | Comparison | Now | Previous | Change | Attempts | Cards | Evidence |
|---|---|---|---|---|---|---|---|
| one_obsession_v22_fp16 | GeForce RTX 4080 vs GeForce RTX 5090 | 0.84x | 0.44x 2026-08-31 | +90.6% | 180 / 57 | 2 / 4 | one operator · no independent replication |
| z_image_bf16 | GeForce RTX 4080 vs GeForce RTX 5090 | 0.67x | 0.37x 2026-08-29 | +79.2% | 12 / 18 | 2 / 4 | one operator · no independent replication |
| chroma-v48-detail-svd_fp8 | GeForce RTX 4080 vs GeForce RTX 5090 | 0.51x | 0.29x 2026-09-01 | +76.1% | 66 / 27 | 2 / 4 | one operator · no independent replication |
| krea2_identity_edit_sogni_v0_3_alpha | GeForce RTX 4080 vs GeForce RTX 5090 | 0.90x | 0.54x 2026-08-31 | +65.1% | 96 / 39 | 2 / 5 | one operator · no independent replication |
| krea2_turbo_fp8_scaled | GeForce RTX 4080 vs GeForce RTX 5090 | 0.55x | 1.32x 2026-09-03 | -58.3% | 33 / 43 | 2 / 5 | one operator · no independent replication |
| chroma1-hd_fp8_scaled | GeForce RTX 4080 vs GeForce RTX 5090 | 0.68x | 0.46x 2026-09-02 | +47.0% | 31 / 15 | 2 / 3 | one operator · no independent replication |
| dark_beast_krea2_fp8 | GeForce RTX 4080 vs GeForce RTX 5090 | 0.56x | 0.97x 2026-09-03 | -42.6% | 101 / 136 | 2 / 5 | one operator · no independent replication |
| dark_beast_krea2_identity_edit_v1_2 | GeForce RTX 4080 vs GeForce RTX 5090 | 0.51x | 0.76x 2026-09-03 | -32.1% | 88 / 38 | 2 / 5 | one operator · no independent replication |
| krea2_identity_edit_v1_2 | GeForce RTX 4080 vs GeForce RTX 5090 | 0.58x | 0.81x 2026-09-03 | -28.2% | 57 / 24 | 2 / 3 | one operator · no independent replication |
| z_image_turbo_bf16 | GeForce RTX 4080 vs GeForce RTX 5090 | 0.36x | 0.28x 2026-08-24 | +28.0% | 13 / 88 | 2 / 6 | one operator · no independent replication |
| dark_beast_z_image_turbo_v9_bf16 | GeForce RTX 4080 vs GeForce RTX 5090 | 0.27x | 0.29x 2026-08-27 | -6.5% | 100 / 50 | 2 / 4 | one operator · no independent replication |
| flux1-schnell-fp8 | GeForce RTX 4080 vs GeForce RTX 5090 | 0.25x | first observation | new | 18 / 58 | 2 / 5 | one operator · no independent replication |
A change here is a change in this fleet between two dated windows. It is not evidence that any GPU became faster: workload mix, customer cancellation behaviour and which physical cards were free all move these ratios, and none of them is controlled. Rows marked † pair a card running under virtualisation, so they carry a host difference as well as a hardware one.
Of the attempts a card started, the fraction that finished and were paid. An attempt that is cancelled or fails still occupies the card, so yield is the difference between time bought and work delivered.
Across 15 archived snapshots on this production fleet, cards finished 83.0–96.9% of the attempts they started. These are dated rolling windows, not closed calendar days: each row carries its own window_start_utc, window_end_utc and effective hours, some windows are shorter than others, and some overlap — where they overlap the same completions appear in two rows, so the rows are not independent observations. The gap between started and completed is observed non-completion. The record does not classify how much of it is a fleet fault GPUDeck could act on, so it is not a count of the failure GPUDeck attacks, and it is not a measurement of GPUDeck’s effect on anything.
| Day | RTX 5090 6 cards | RTX 4080 2 cards | RTX 3080 1 cards | RTX 5070 Ti 1 cards |
|---|---|---|---|---|
| 2026-08-20 | 94.7% 3,786 of 3,997 | 89.5% 1,904 of 2,127 | rate withheld 249 of 302 | rate withheld 533 of 718 |
| 2026-08-21 | 95.2% 4,776 of 5,017 | 85.0% 1,315 of 1,547 | rate withheld 425 of 544 | rate withheld 317 of 473 |
| 2026-08-22 | 95.0% 5,872 of 6,180 | 83.0% 1,251 of 1,507 | rate withheld 456 of 566 | rate withheld 378 of 557 |
| 2026-08-23 | 96.4% 5,669 of 5,878 | 93.0% 1,605 of 1,726 | rate withheld 539 of 615 | rate withheld 540 of 620 |
| 2026-08-24 | 93.4% 911 of 975 | 88.8% 214 of 241 | rate withheld 71 of 93 | rate withheld 55 of 62 |
| 2026-08-26 | 95.6% 7,648 of 8,000 | 92.2% 1,945 of 2,110 | rate withheld 709 of 836 | rate withheld 534 of 671 |
| 2026-08-27 | 95.2% 7,757 of 8,149 | 90.8% 1,924 of 2,118 | rate withheld 685 of 833 | rate withheld 538 of 704 |
| 2026-08-28 | 93.4% 423 of 453 | 94.7% 373 of 394 | rate withheld 129 of 149 | rate withheld 103 of 111 |
| 2026-08-29 | 96.9% 3,360 of 3,469 | 91.9% 1,581 of 1,721 | rate withheld 513 of 606 | rate withheld 713 of 791 |
| 2026-08-30 | 93.3% 336 of 360 | 86.6% 71 of 82 | rate withheld 9 of 14 | rate withheld 62 of 77 |
| 2026-08-31 | 95.1% 2,261 of 2,377 | 91.8% 1,171 of 1,275 | rate withheld 406 of 468 | rate withheld 544 of 606 |
| 2026-09-01 | 96.8% 842 of 870 | 90.3% 708 of 784 | rate withheld 195 of 233 | rate withheld 46 of 52 |
| 2026-09-02 | 89.5% 402 of 449 | 84.1% 122 of 145 | rate withheld 78 of 93 | rate withheld 163 of 180 |
| 2026-09-03 | 94.2% 2,306 of 2,447 | 95.1% 3,005 of 3,160 | rate withheld 991 of 1,087 | rate withheld 428 of 450 |
| 2026-09-04 | 95.7% 353 of 369 | 93.5% 260 of 278 | rate withheld 67 of 88 | rate withheld 23 of 23 |
No fleet-wide row appears in this table, and that is deliberate: yield differs by card, so a blended figure moves whenever the mix of work shifts between a faster and a slower card — which would read as reliability changing on a day when nothing about reliability changed. A fleet total is published separately below, as a portfolio outcome rather than a statement about any card.
Below the replication floor, and therefore unnumbered: RTX 3080 and RTX 5070 Ti. Each is backed by a single physical card, and a card that ran hundreds of attempts in a day is one observation of that card, not hundreds. The fleet does run them, and this record states no yield for them. Same two-card floor the goodput exports apply. The suppressed cells are listed in fleet-completion-yield.json.
2026-08-25 has no archived snapshot at all, and is shown as a gap rather than a zero — a date with no snapshot is not a claim that the fleet ran no work that day. The current UTC day is excluded throughout, because the archive file for it can still change.
GPUDeck was running on this fleet on every day in this table, so there is no untreated period to compare against and no claim of improvement is made or implied here. Nothing in this record says the software raised these numbers, held them anywhere, or prevented any particular loss; a change between two days is not evidence that anything caused it. What it establishes is narrower and checkable: unfinished paid work is a real, recurring, measurable cost on a production fleet, at a rate with a denominator rather than an anecdote. Whether the product reduces it would need a fleet running without it, and we do not have one.
Across 12 non-overlapping snapshots the fleet finished 92.5% of everything it started (50,746 of 54,855). That covers every card the fleet ran, including the ones whose per-card cells are suppressed above: the two-card floor governs what can be said about a card, not what the fleet did. This is a portfolio outcome, not a hardware one: it moves when work shifts between a higher- and a lower-yielding lane even though no card behaved differently, so it must not be read as any card’s reliability — the per-card cells above are the only hardware statement this record makes. 2026-08-21, 2026-08-27, 2026-08-29 is excluded because its measurement window overlaps a snapshot already counted, and the same completions must not be added twice. Computed over 2026-08-20, 2026-08-22, 2026-08-23, 2026-08-24, 2026-08-26, 2026-08-28, 2026-08-30, 2026-08-31, 2026-09-01, 2026-09-02, 2026-09-03, 2026-09-04.
Every published figure above is regenerable from fleet-completion-yield.csv, which carries the started and completed counts each percentage is computed from, plus the measurement window each row was observed over. Suppressed cells are listed with their started and completed counts and NO rate: the replication floor withholds the per-card claim, not the underlying evidence, so a reader can see exactly how much was set aside and recompute anything they like from it. What is withheld is the percentage, because one physical card cannot evidence a property of a card model. The lifecycle archive also records two occupancy counters that are deliberately not published: they cannot be reconciled against the started and completed counts in either direction, so any rate built on them would have a denominator we cannot define. The reason is recorded in the JSON beside the data.
Paid completions per occupied GPU-hour, measured inside one workload at a time. Occupied time includes attempts that were cancelled before they finished, so this counts the resource a card actually consumes rather than only the work that survived.
This table takes one day. Workload evidence carries a page per paid production workload, each with its own dated ratio cells over time.
The gap between started and completed is unclassified here. The incident record lists every detected operational incident, and the status of product-efficacy estimation for each incident class, with the reason no estimate exists.
Replication. Every cell on this page comes from a single operator and a single day. 11 of the 12 rows below rest on one physical card on at least one side, marked ●. Across the whole table 11 distinct cards contributed. Attempt counts are not replication: a card that ran 600 attempts in a day is still one card, and confidence that treats those attempts as independent would be wrong by roughly the square root of that count.
No range is claimed for this vintage. Only 1 of the 12 rows below is both free of the host confound and supported by at least two physical cards on each side. One row is not a range, so none is stated. The individual ratios stand as what they are — single cells, each reproducible from the dated open dataset — and the equal-memory table below is the cleaner contrast for this day.
| Workload | Larger-memory GPU | Smaller-memory GPU | Attempts | Cards A/B | Smaller-memory yield | Goodput ratio |
|---|---|---|---|---|---|---|
| wan_v2.2-14b-fp8_i2v_lightx2v | A100 80GB PCIe | GeForce RTX 5090 | 23 | 1/3 ● | 0.900 | 0.25x |
| minimax-h3-ref2va-fp8_r2v_turbo | A100 80GB PCIe | GeForce RTX 5090 | 63 | 1/4 ● | 1.000 | 0.43x |
| minimax-h3-fastvideo-int8_i2v_turbo | A100 80GB PCIe | GeForce RTX 5090 | 103 | 1/4 ● | 1.000 | 0.43x |
| ltx25-22b-int8_i2v_distilled | A100 80GB PCIe | GeForce RTX 5090 | 34 | 1/5 ● | 0.955 | 0.64x |
| minimax-h3-fl2va-fp8_i2v_turbo | A100 80GB PCIe | GeForce RTX 5090 | 22 | 1/4 ● | 1.000 | 0.64x |
| minimax-h3-fastvideo-int8_flf2v_turbo | A100 80GB PCIe | GeForce RTX 5090 | 25 | 1/4 ● | 1.000 | 0.74x |
| dark_beast_krea2_identity_edit_v1_2 | A100 80GB PCIe | GeForce RTX 4080 | 102 | 1/2 ● | 0.977 | 0.98x |
| minimax-h3-fastvideo-int8_t2v_turbo | A100 80GB PCIe | GeForce RTX 5090 | 19 | 1/2 ● | 1.000 | 1.02x |
| dark_beast_krea2_fp8 | A100 80GB PCIe | GeForce RTX 4080 | 259 | 1/2 ● | 0.950 | 1.03x |
| krea2_identity_edit_v1_2 | A100 80GB PCIe | GeForce RTX 4080 | 91 | 1/2 ● | 0.842 | 1.09x |
| chroma1-hd_fp8_scaled | A100 80GB PCIe | GeForce RTX 4080 | 19 | 1/2 ● | 0.909 | 1.50x |
| krea2_turbo_fp8_scaled | GeForce RTX 5090 | GeForce RTX 4080 | 76 | 5/2 | 0.909 | 1.82x |
A deployment here is a GPU together with the host, virtualisation layer, software stack and dispatch it ran under — not the card in isolation. These ratios do not isolate a hardware effect, which is why the confounded rows are separated above rather than averaged in.
Goodput is measured on production hosts, and the host is inside the measurement. The class anchoring the widest ratios above runs under a virtualisation layer that reserves part of its memory and cannot be configured away, so treat those ratios as a result that includes a known hosting disadvantage rather than as a property of the hardware alone. The equal-memory table, where both classes sit on comparable hosts, is the cleaner contrast.
Stated per workload and never combined. A single figure across workloads would move with the mix of work each class happens to be sent, which is a fact about dispatch rather than about the hardware — the same substitution this record exists to refuse. Yield is recorded as production loss under the observed network regime, not as a property of any class.
How much more per hour the larger-memory class can cost before the smaller one becomes the cheaper way to finish this workload. This is a ratio of two goodput figures measured on the same fleet in the same window — it contains no marketplace price, so nothing here depends on a listing being launchable, comparable, or representative. Bring your own quote and compare it against the threshold.
Derived from the goodput table above — no additional observations. This is the same set of measured pairs, reinterpreted as a price threshold rather than an output ratio, so it is a second reading of one body of evidence and not a second, independent result.
| Workload | Larger-memory GPU | Smaller-memory GPU | Break-even price multiple |
|---|---|---|---|
| krea2_turbo_fp8_scaled | GeForce RTX 5090 · 32 GB | GeForce RTX 4080 · 16 GB | 1.82x |
| chroma1-hd_fp8_scaled | A100 80GB PCIe · 80 GB | GeForce RTX 4080 · 16 GB | 1.50x |
| krea2_identity_edit_v1_2 | A100 80GB PCIe · 80 GB | GeForce RTX 4080 · 16 GB | 1.09x |
| dark_beast_krea2_fp8 | A100 80GB PCIe · 80 GB | GeForce RTX 4080 · 16 GB | 1.03x |
| minimax-h3-fastvideo-int8_t2v_turbo | A100 80GB PCIe · 80 GB | GeForce RTX 5090 · 32 GB | 1.02x |
| dark_beast_krea2_identity_edit_v1_2 | A100 80GB PCIe · 80 GB | GeForce RTX 4080 · 16 GB | 0.98x |
| minimax-h3-fastvideo-int8_flf2v_turbo | A100 80GB PCIe · 80 GB | GeForce RTX 5090 · 32 GB | 0.74x |
| minimax-h3-fl2va-fp8_i2v_turbo | A100 80GB PCIe · 80 GB | GeForce RTX 5090 · 32 GB | 0.64x |
| ltx25-22b-int8_i2v_distilled | A100 80GB PCIe · 80 GB | GeForce RTX 5090 · 32 GB | 0.64x |
| minimax-h3-fastvideo-int8_i2v_turbo | A100 80GB PCIe · 80 GB | GeForce RTX 5090 · 32 GB | 0.43x |
| minimax-h3-ref2va-fp8_r2v_turbo | A100 80GB PCIe · 80 GB | GeForce RTX 5090 · 32 GB | 0.43x |
| wan_v2.2-14b-fp8_i2v_lightx2v | A100 80GB PCIe · 80 GB | GeForce RTX 5090 · 32 GB | 0.25x |
Read every row as a threshold, not a price. A multiple of 1.82x on a row says the larger-memory GPU is the cheaper way to finish that workload at any hourly rate below 1.82 times the smaller one; at exactly that multiple the two cost the same. The daggered rows are not usable this way, because part of their headroom is a hosting penalty specific to this fleet.
Read it as a threshold, not a price. A ratio of 4.00x says the larger class is the cheaper way to finish this workload at any hourly multiple below four times; above it, the smaller class wins. The comparison holds where both classes see the same share of billed time spent occupied — where utilisation differs, scale each side by its own occupied-to-billed fraction before comparing.
Pairs are fixed by largest memory, then lowest permanent class index, both chosen before any measurement was read. Selecting the fastest and slowest observed classes instead would pick each comparison on its own outcome. These 12 rows are drawn from 3 distinct GPU-model pairings, so they show how a small number of comparisons behave across many workloads rather than that many independent comparisons; the GPUs in each row are named so this is checkable.
A single fixed pair of GPUs — RTX 5090 at 32 GB against RTX 4080 at 16 GB — measured on the same completed workloads. Each row reports how much longer the RTX 4080 took on paid renders of that workload. Every row is this same one pairing, so the rows show how one comparison varies across workloads, not a set of independent comparisons.
This table is a single day and measures latency. The same pairing measured on goodput, one workload at a time across every observation day, is at card comparisons — where the ratio changes with the workload, which is the finding, and there is no winner.
| Workload | Completed renders | RTX 4080 took longer than RTX 5090 by |
|---|---|---|
| dark_beast_krea2_identity_edit_v1_2 | 124 | 1.80x |
| dark_beast_krea2_fp8 | 229 | 2.04x |
| krea2_turbo_fp8_scaled | 73 | 2.22x |
| krea2_identity_edit_v1_2 | 72 | 2.52x |
The observed latency gap runs from 1.80x to 2.52x. All workloads point in the same direction, but one day cannot establish that the multiplier itself varies by workload. That requires consecutive days and an explicit interaction test.
On the widest-gap workload, p90/p10 dispersion was 2.8x for the RTX 4080 and 5.5x for the RTX 5090. That comparison is descriptive, not a confidence interval. These are completed-render latency ratios, not throughput ratios; throughput would include all occupied time, including cancelled and failed attempts.
The download, method and boundary live together so the editorial argument cannot outrun the evidence behind it.
Scope. Every goodput cell comes from one operator’s production fleet. Hardware, workload and measurement window are named; operator identity, host identity and absolute throughput are not. Ratios are never aggregated across workloads.
Method. Goodput is paid completions per occupied GPU-hour inside one exact (GPU, workload) cell. Attempt counts are support, not replication; physical-card counts are published beside them. Latency and yield remain diagnostics because neither is goodput on its own.
Separate record. The longitudinal public-network census now lives at /census. It is context, not an input to the goodput ratios, and is never used to weight them.
Citation. Reuse is licensed under CC BY 4.0. The citation package records retrieval, attribution and what the series does not measure.
REPLICATE THIS RECORD ON A SECOND FLEET
Publish daily workload, GPU, attempts, completions and occupied GPU-hours under your own fleet identifier. Fleets are never averaged; earnings and worker identity are not accepted. Protocol validation is the gate - submitting does not mean acceptance.
Publish daily workload, GPU, attempts, completions and occupied GPU-hours under your own fleet identifier. Fleets are never averaged; earnings and worker identity are not accepted. Submitting does not make data part of the record — protocol validation is the gate.