GPUDECK · INCIDENT RECORD · ONE FLEET

What the incident record can and cannot say

Every operational incident detected on the reference fleet that cleared the admission rule, together with the status of product-efficacy estimation for each class — and why no estimate currently exists.

The current statement

This block is generated by gpudeck_incident_ledger.public_statement() and published verbatim. It is not a summary of the ledger written alongside it — it is the ledger’s own words, so this page cannot say something the method does not. A production check fetches this page and requires the served text to match that function’s output exactly.

129 operational incidents recorded. No GPUDeck remediation-effectiveness rate is currently published. CONFIG_DRIFT: not applicable - No automatic GPUDeck remediation in this class. LANE_DOWN: not yet estimated - 0 deck-attributed interventions since the definition froze at 2026-08-29T00:00:00Z; publication floor 10. OOM_RISK: not attributable to deck - Recovery performed by the container runtime, not GPUDeck, on 87 of 87 automatic row(s). Definition frozen 2026-08-29T00:00:00Z; analysis horizon 2026-10-10T00:00:00Z; definition hash a03c21b1ae0c34f3. One fleet - no claim of generality.

By incident class

ClassIncidentsProduct efficacy Median detectionMedian time to earning
CONFIG_DRIFT3not applicable
No automatic GPUDeck remediation in this class.
0s2821s
LANE_DOWN39not yet estimated
0 deck-attributed interventions since the definition froze at 2026-08-29T00:00:00Z; publication floor 10.
61s113s
OOM_RISK87not attributable to deck
Recovery performed by the container runtime, not GPUDeck, on 87 of 87 automatic row(s).
415s87s

Detection latency and time-to-earning are reported per class and never pooled: the classes do not measure the same thing with those fields. CONFIG_DRIFT’s zero detection is structural — the component that refuses a bad config is the same one that reports the refusal, so it is not a fast detection, it is not a detection. And its time-to-earning is not an outage duration: a held worker keeps earning on its previous version throughout, so that column measures how long a held version persisted and says nothing about lost earning. Read down a column at your peril; the classes only share it typographically.

Why no estimate exists

These are not three rates GPUDeck knows and declines to publish. CONFIG_DRIFT has no automatic remediation, so no such rate is applicable to it; OOM_RISK’s recoveries belong to the container runtime, so any rate over them would not be GPUDeck’s; and for LANE_DOWN the estimate does not exist yet. Calling them “withheld” would imply discretion — a number in hand, kept back — which is the impression the three-status vocabulary exists to remove. An estimate is stated only when four pre-registered conditions hold at once, each checked in code rather than promised in prose:

Deck-attributed The action must be one GPUDeck took. Of the first feed’s 107 automatic actions, 82 were the container runtime’s own restart policy — the deck detected nothing and did nothing. A row not reporting its source is excluded, never assumed to be ours.
Prospective The row must be observed after 2026-08-29T00:00:00Z, when the rule that scores it was frozen. The fleet operator reports that an earlier recovery definition scored 26 of 46 rows as unrecovered where the current one scores 0 on the same events — a figure GPUDeck cannot verify, because the feed carries only the current definition’s verdicts and no record of the earlier one. It is stated here as their account and not as a measurement. The principle does not depend on it: a rule chosen with the answers already in view measures the choice.
Falsifiable The denominator must be able to contain a failure. A ratio no possible input could have lowered reports the shape of the definition, not the state of the fleet.
Above the floor At least 10 qualifying observations, inside the accrual window closing 2026-10-10T00:00:00Z. At that instant, too few observations is published as a failure to estimate rather than becoming an indefinite wait.

Was every attempt counted?

Every guard above grades incidents that were recorded. None of them can see a remediation that never produced a record at all — and that gap was not hypothetical here. Before the fleet built an action journal, this ledger held 46 lane-down rows: 30 recovered, 16 false positives, and zero failures. A flawless record over a denominator no failure could enter, because the self-heal writes its log line after the restart call returns: an attempt that hung or was killed mid-flight left no line, and was classified “none recorded” rather than as a failed remediation. It did not score badly. It stopped being a candidate.

The fleet now records an immutable action id at the moment it decides to act, before the attempt is made, and reconciles that journal against what it publishes. An attempt with no confirmed recovery 2 hours after that decision is specified to become a real not recovered row in the denominator — pre-registered, hashed into the definition, and the first rule here that can make the published number worse. That path has not yet fired on a real event. The fleet has been healthy, so no attempt has expired; the rule is specified and reconciled, not demonstrated, and this page says so until one does.

NOT_YET_EXERCISED The action journal exists and is readable, and no remediation attempt has been recorded through it yet. It records from its own creation, so incidents published before it existed have no journalled attempt behind them - that is the gap, not a discrepancy.

A reconciliation that has audited nothing has established nothing, so an empty journal reports as not yet exercised rather than as proven; a journal the fleet cannot read reports as unproven rather than as zero missing attempts; and attempts falling outside the measurement window are counted separately rather than dropped, so a journal full of attempts this rate does not cover can never look like an empty one. No rate is published at all unless completeness reads PROVEN — a mismatched or unstated window, an unreadable journal, an unexercised one, terminal settings that differ from the pre-registered ones, or figures that do not add up each suppress the rate rather than merely annotating it.

Scope

One operator, one fleet. 129 admitted incidents; 14 candidate rows were rejected for carrying no affirmative evidence that work was available — a quiet period is not an incident. Demand evidence is that other lanes were completing paid work nearby, which establishes that work existed for this fleet and not that this lane would have been dispatched any. Nothing here says GPUDeck improved an outcome; no fleet ran without it, so there is no untreated comparison to make.

Definition hash a03c21b1ae0c34f3 · accrual window 42 days · regenerable from incident-record.json.