GPUDeck · Production goodput observatory

What an occupied GPU hour finishes

Hourly price alone cannot tell you what finishing the work costs. GPUDeck measures paid completions per occupied GPU-hour inside one exact workload cell, compares named deployments as ratios, and never aggregates unlike workloads into one number.

Latest published measurement: trailing 24 hours ending 4 Sep 2026 01:21 UTC. Page regenerated 2026-09-23 05:21 UTC; one operator; per-workload ratios only.

1 operator · 11 physical GPUs in the source cells · 12 workload-specific comparisons · no cross-workload aggregate

Latest evidence changes

Each row is one workload on one GPU pairing, compared with its own previous published observation — not with your last visit, and not with any other workload. Latest observation 2026-09-04. Of 12 comparisons, 11 moved and 1 appeared for the first time. There is deliberately no overall figure here: averaging these rows together is the error this record exists to refuse.

WorkloadComparisonNowPreviousChangeAttemptsCardsEvidence
one_obsession_v22_fp16GeForce RTX 4080 vs GeForce RTX 50900.84x0.44x 2026-08-31+90.6%180 / 572 / 4one operator · no independent replication
z_image_bf16GeForce RTX 4080 vs GeForce RTX 50900.67x0.37x 2026-08-29+79.2%12 / 182 / 4one operator · no independent replication
chroma-v48-detail-svd_fp8GeForce RTX 4080 vs GeForce RTX 50900.51x0.29x 2026-09-01+76.1%66 / 272 / 4one operator · no independent replication
krea2_identity_edit_sogni_v0_3_alphaGeForce RTX 4080 vs GeForce RTX 50900.90x0.54x 2026-08-31+65.1%96 / 392 / 5one operator · no independent replication
krea2_turbo_fp8_scaledGeForce RTX 4080 vs GeForce RTX 50900.55x1.32x 2026-09-03-58.3%33 / 432 / 5one operator · no independent replication
chroma1-hd_fp8_scaledGeForce RTX 4080 vs GeForce RTX 50900.68x0.46x 2026-09-02+47.0%31 / 152 / 3one operator · no independent replication
dark_beast_krea2_fp8GeForce RTX 4080 vs GeForce RTX 50900.56x0.97x 2026-09-03-42.6%101 / 1362 / 5one operator · no independent replication
dark_beast_krea2_identity_edit_v1_2GeForce RTX 4080 vs GeForce RTX 50900.51x0.76x 2026-09-03-32.1%88 / 382 / 5one operator · no independent replication
krea2_identity_edit_v1_2GeForce RTX 4080 vs GeForce RTX 50900.58x0.81x 2026-09-03-28.2%57 / 242 / 3one operator · no independent replication
z_image_turbo_bf16GeForce RTX 4080 vs GeForce RTX 50900.36x0.28x 2026-08-24+28.0%13 / 882 / 6one operator · no independent replication
dark_beast_z_image_turbo_v9_bf16GeForce RTX 4080 vs GeForce RTX 50900.27x0.29x 2026-08-27-6.5%100 / 502 / 4one operator · no independent replication
flux1-schnell-fp8GeForce RTX 4080 vs GeForce RTX 50900.25xfirst observationnew18 / 582 / 5one operator · no independent replication

A change here is a change in this fleet between two dated windows. It is not evidence that any GPU became faster: workload mix, customer cancellation behaviour and which physical cards were free all move these ratios, and none of them is controlled. Rows marked † pair a card running under virtualisation, so they carry a host difference as well as a hardware one.

Completion yield

Of the attempts a card started, the fraction that finished and were paid. An attempt that is cancelled or fails still occupies the card, so yield is the difference between time bought and work delivered.

Across 15 archived snapshots on this production fleet, cards finished 83.0–96.9% of the attempts they started. These are dated rolling windows, not closed calendar days: each row carries its own window_start_utc, window_end_utc and effective hours, some windows are shorter than others, and some overlap — where they overlap the same completions appear in two rows, so the rows are not independent observations. The gap between started and completed is observed non-completion. The record does not classify how much of it is a fleet fault GPUDeck could act on, so it is not a count of the failure GPUDeck attacks, and it is not a measurement of GPUDeck’s effect on anything.

Observed · 15 archived snapshots2026-08-20–2026-09-04 · one operator’s fleet
DayRTX 5090
6 cards
RTX 4080
2 cards
RTX 3080
1 cards
RTX 5070 Ti
1 cards
2026-08-2094.7%
3,786 of 3,997
89.5%
1,904 of 2,127
rate withheld
249 of 302
rate withheld
533 of 718
2026-08-2195.2%
4,776 of 5,017
85.0%
1,315 of 1,547
rate withheld
425 of 544
rate withheld
317 of 473
2026-08-2295.0%
5,872 of 6,180
83.0%
1,251 of 1,507
rate withheld
456 of 566
rate withheld
378 of 557
2026-08-2396.4%
5,669 of 5,878
93.0%
1,605 of 1,726
rate withheld
539 of 615
rate withheld
540 of 620
2026-08-2493.4%
911 of 975
88.8%
214 of 241
rate withheld
71 of 93
rate withheld
55 of 62
2026-08-2695.6%
7,648 of 8,000
92.2%
1,945 of 2,110
rate withheld
709 of 836
rate withheld
534 of 671
2026-08-2795.2%
7,757 of 8,149
90.8%
1,924 of 2,118
rate withheld
685 of 833
rate withheld
538 of 704
2026-08-2893.4%
423 of 453
94.7%
373 of 394
rate withheld
129 of 149
rate withheld
103 of 111
2026-08-2996.9%
3,360 of 3,469
91.9%
1,581 of 1,721
rate withheld
513 of 606
rate withheld
713 of 791
2026-08-3093.3%
336 of 360
86.6%
71 of 82
rate withheld
9 of 14
rate withheld
62 of 77
2026-08-3195.1%
2,261 of 2,377
91.8%
1,171 of 1,275
rate withheld
406 of 468
rate withheld
544 of 606
2026-09-0196.8%
842 of 870
90.3%
708 of 784
rate withheld
195 of 233
rate withheld
46 of 52
2026-09-0289.5%
402 of 449
84.1%
122 of 145
rate withheld
78 of 93
rate withheld
163 of 180
2026-09-0394.2%
2,306 of 2,447
95.1%
3,005 of 3,160
rate withheld
991 of 1,087
rate withheld
428 of 450
2026-09-0495.7%
353 of 369
93.5%
260 of 278
rate withheld
67 of 88
rate withheld
23 of 23

No fleet-wide row appears in this table, and that is deliberate: yield differs by card, so a blended figure moves whenever the mix of work shifts between a faster and a slower card — which would read as reliability changing on a day when nothing about reliability changed. A fleet total is published separately below, as a portfolio outcome rather than a statement about any card.

Below the replication floor, and therefore unnumbered: RTX 3080 and RTX 5070 Ti. Each is backed by a single physical card, and a card that ran hundreds of attempts in a day is one observation of that card, not hundreds. The fleet does run them, and this record states no yield for them. Same two-card floor the goodput exports apply. The suppressed cells are listed in fleet-completion-yield.json.

2026-08-25 has no archived snapshot at all, and is shown as a gap rather than a zero — a date with no snapshot is not a claim that the fleet ran no work that day. The current UTC day is excluded throughout, because the archive file for it can still change.

GPUDeck was running on this fleet on every day in this table, so there is no untreated period to compare against and no claim of improvement is made or implied here. Nothing in this record says the software raised these numbers, held them anywhere, or prevented any particular loss; a change between two days is not evidence that anything caused it. What it establishes is narrower and checkable: unfinished paid work is a real, recurring, measurable cost on a production fleet, at a rate with a denominator rather than an anecdote. Whether the product reduces it would need a fleet running without it, and we do not have one.

Realized network yield, whole fleet

Across 12 non-overlapping snapshots the fleet finished 92.5% of everything it started (50,746 of 54,855). That covers every card the fleet ran, including the ones whose per-card cells are suppressed above: the two-card floor governs what can be said about a card, not what the fleet did. This is a portfolio outcome, not a hardware one: it moves when work shifts between a higher- and a lower-yielding lane even though no card behaved differently, so it must not be read as any card’s reliability — the per-card cells above are the only hardware statement this record makes. 2026-08-21, 2026-08-27, 2026-08-29 is excluded because its measurement window overlaps a snapshot already counted, and the same completions must not be added twice. Computed over 2026-08-20, 2026-08-22, 2026-08-23, 2026-08-24, 2026-08-26, 2026-08-28, 2026-08-30, 2026-08-31, 2026-09-01, 2026-09-02, 2026-09-03, 2026-09-04.

Every published figure above is regenerable from fleet-completion-yield.csv, which carries the started and completed counts each percentage is computed from, plus the measurement window each row was observed over. Suppressed cells are listed with their started and completed counts and NO rate: the replication floor withholds the per-card claim, not the underlying evidence, so a reader can see exactly how much was set aside and recompute anything they like from it. What is withheld is the percentage, because one physical card cannot evidence a property of a card model. The lifecycle archive also records two occupancy counters that are deliberately not published: they cannot be reconciled against the started and completed counts in either direction, so any rate built on them would have a denominator we cannot define. The reason is recorded in the JSON beside the data.

Goodput

Paid completions per occupied GPU-hour, measured inside one workload at a time. Occupied time includes attempts that were cancelled before they finished, so this counts the resource a card actually consumes rather than only the work that survived.

This table takes one day. Workload evidence carries a page per paid production workload, each with its own dated ratio cells over time.

The gap between started and completed is unclassified here. The incident record lists every detected operational incident, and the status of product-efficacy estimation for each incident class, with the reason no estimate exists.

Observed · single day2026-09-04 · 836 attempts

Replication. Every cell on this page comes from a single operator and a single day. 11 of the 12 rows below rest on one physical card on at least one side, marked ●. Across the whole table 11 distinct cards contributed. Attempt counts are not replication: a card that ran 600 attempts in a day is still one card, and confidence that treats those attempts as independent would be wrong by roughly the square root of that count.

No range is claimed for this vintage. Only 1 of the 12 rows below is both free of the host confound and supported by at least two physical cards on each side. One row is not a range, so none is stated. The individual ratios stand as what they are — single cells, each reproducible from the dated open dataset — and the equal-memory table below is the cleaner contrast for this day.

WorkloadLarger-memory GPUSmaller-memory GPUAttemptsCards A/BSmaller-memory yieldGoodput ratio
wan_v2.2-14b-fp8_i2v_lightx2vA100 80GB PCIeGeForce RTX 5090231/3 ●0.9000.25x
minimax-h3-ref2va-fp8_r2v_turboA100 80GB PCIeGeForce RTX 5090631/4 ●1.0000.43x
minimax-h3-fastvideo-int8_i2v_turboA100 80GB PCIeGeForce RTX 50901031/4 ●1.0000.43x
ltx25-22b-int8_i2v_distilledA100 80GB PCIeGeForce RTX 5090341/5 ●0.9550.64x
minimax-h3-fl2va-fp8_i2v_turboA100 80GB PCIeGeForce RTX 5090221/4 ●1.0000.64x
minimax-h3-fastvideo-int8_flf2v_turboA100 80GB PCIeGeForce RTX 5090251/4 ●1.0000.74x
dark_beast_krea2_identity_edit_v1_2A100 80GB PCIeGeForce RTX 40801021/2 ●0.9770.98x
minimax-h3-fastvideo-int8_t2v_turboA100 80GB PCIeGeForce RTX 5090191/2 ●1.0001.02x
dark_beast_krea2_fp8A100 80GB PCIeGeForce RTX 40802591/2 ●0.9501.03x
krea2_identity_edit_v1_2A100 80GB PCIeGeForce RTX 4080911/2 ●0.8421.09x
chroma1-hd_fp8_scaledA100 80GB PCIeGeForce RTX 4080191/2 ●0.9091.50x
krea2_turbo_fp8_scaledGeForce RTX 5090GeForce RTX 4080765/20.9091.82x

A deployment here is a GPU together with the host, virtualisation layer, software stack and dispatch it ran under — not the card in isolation. These ratios do not isolate a hardware effect, which is why the confounded rows are separated above rather than averaged in.

Goodput is measured on production hosts, and the host is inside the measurement. The class anchoring the widest ratios above runs under a virtualisation layer that reserves part of its memory and cannot be configured away, so treat those ratios as a result that includes a known hosting disadvantage rather than as a property of the hardware alone. The equal-memory table, where both classes sit on comparable hosts, is the cleaner contrast.

Stated per workload and never combined. A single figure across workloads would move with the mix of work each class happens to be sent, which is a fact about dispatch rather than about the hardware — the same substitution this record exists to refuse. Yield is recorded as production loss under the observed network regime, not as a property of any class.

Break-even price ratio

How much more per hour the larger-memory class can cost before the smaller one becomes the cheaper way to finish this workload. This is a ratio of two goodput figures measured on the same fleet in the same window — it contains no marketplace price, so nothing here depends on a listing being launchable, comparable, or representative. Bring your own quote and compare it against the threshold.

Derived from the goodput table above — no additional observations. This is the same set of measured pairs, reinterpreted as a price threshold rather than an output ratio, so it is a second reading of one body of evidence and not a second, independent result.

Measured · no external price12 workloads · 836 attempts across both GPUs
WorkloadLarger-memory GPUSmaller-memory GPUBreak-even price multiple
krea2_turbo_fp8_scaledGeForce RTX 5090 · 32 GBGeForce RTX 4080 · 16 GB1.82x
chroma1-hd_fp8_scaledA100 80GB PCIe · 80 GBGeForce RTX 4080 · 16 GB1.50x
krea2_identity_edit_v1_2A100 80GB PCIe · 80 GBGeForce RTX 4080 · 16 GB1.09x
dark_beast_krea2_fp8A100 80GB PCIe · 80 GBGeForce RTX 4080 · 16 GB1.03x
minimax-h3-fastvideo-int8_t2v_turboA100 80GB PCIe · 80 GBGeForce RTX 5090 · 32 GB1.02x
dark_beast_krea2_identity_edit_v1_2A100 80GB PCIe · 80 GBGeForce RTX 4080 · 16 GB0.98x
minimax-h3-fastvideo-int8_flf2v_turboA100 80GB PCIe · 80 GBGeForce RTX 5090 · 32 GB0.74x
minimax-h3-fl2va-fp8_i2v_turboA100 80GB PCIe · 80 GBGeForce RTX 5090 · 32 GB0.64x
ltx25-22b-int8_i2v_distilledA100 80GB PCIe · 80 GBGeForce RTX 5090 · 32 GB0.64x
minimax-h3-fastvideo-int8_i2v_turboA100 80GB PCIe · 80 GBGeForce RTX 5090 · 32 GB0.43x
minimax-h3-ref2va-fp8_r2v_turboA100 80GB PCIe · 80 GBGeForce RTX 5090 · 32 GB0.43x
wan_v2.2-14b-fp8_i2v_lightx2vA100 80GB PCIe · 80 GBGeForce RTX 5090 · 32 GB0.25x

Read every row as a threshold, not a price. A multiple of 1.82x on a row says the larger-memory GPU is the cheaper way to finish that workload at any hourly rate below 1.82 times the smaller one; at exactly that multiple the two cost the same. The daggered rows are not usable this way, because part of their headroom is a hosting penalty specific to this fleet.

Read it as a threshold, not a price. A ratio of 4.00x says the larger class is the cheaper way to finish this workload at any hourly multiple below four times; above it, the smaller class wins. The comparison holds where both classes see the same share of billed time spent occupied — where utilisation differs, scale each side by its own occupied-to-billed fraction before comparing.

Pairs are fixed by largest memory, then lowest permanent class index, both chosen before any measurement was read. Selecting the fastest and slowest observed classes instead would pick each comparison on its own outcome. These 12 rows are drawn from 3 distinct GPU-model pairings, so they show how a small number of comparisons behave across many workloads rather than that many independent comparisons; the GPUs in each row are named so this is checkable.

Diagnostics — completed-render latency, the other input goodput is built from

Completed-render latency ratios

A single fixed pair of GPUs — RTX 5090 at 32 GB against RTX 4080 at 16 GB — measured on the same completed workloads. Each row reports how much longer the RTX 4080 took on paid renders of that workload. Every row is this same one pairing, so the rows show how one comparison varies across workloads, not a set of independent comparisons.

This table is a single day and measures latency. The same pairing measured on goodput, one workload at a time across every observation day, is at card comparisons — where the ratio changes with the workload, which is the finding, and there is no winner.

Exploratory · single-day2026-09-04 · 498 completed renders
WorkloadCompleted rendersRTX 4080 took longer than RTX 5090 by
dark_beast_krea2_identity_edit_v1_21241.80x
dark_beast_krea2_fp82292.04x
krea2_turbo_fp8_scaled732.22x
krea2_identity_edit_v1_2722.52x

The observed latency gap runs from 1.80x to 2.52x. All workloads point in the same direction, but one day cannot establish that the multiplier itself varies by workload. That requires consecutive days and an explicit interaction test.

On the widest-gap workload, p90/p10 dispersion was 2.8x for the RTX 4080 and 5.5x for the RTX 5090. That comparison is descriptive, not a confidence interval. These are completed-render latency ratios, not throughput ratios; throughput would include all occupied time, including cancelled and failed attempts.

Data, scope and citation

The download, method and boundary live together so the editorial argument cannot outrun the evidence behind it.

Scope. Every goodput cell comes from one operator’s production fleet. Hardware, workload and measurement window are named; operator identity, host identity and absolute throughput are not. Ratios are never aggregated across workloads.

Method. Goodput is paid completions per occupied GPU-hour inside one exact (GPU, workload) cell. Attempt counts are support, not replication; physical-card counts are published beside them. Latency and yield remain diagnostics because neither is goodput on its own.

Separate record. The longitudinal public-network census now lives at /census. It is context, not an input to the goodput ratios, and is never used to weight them.

Citation. Reuse is licensed under CC BY 4.0. The citation package records retrieval, attribution and what the series does not measure.

REPLICATE THIS RECORD ON A SECOND FLEET

Publish daily workload, GPU, attempts, completions and occupied GPU-hours under your own fleet identifier. Fleets are never averaged; earnings and worker identity are not accepted. Protocol validation is the gate - submitting does not mean acceptance.

Read CONTRIBUTING-DATA.md · Contact the record

Replicate this record on a second fleet.

Publish daily workload, GPU, attempts, completions and occupied GPU-hours under your own fleet identifier. Fleets are never averaged; earnings and worker identity are not accepted. Submitting does not make data part of the record — protocol validation is the gate.