THE DECK · OPS LAYER FOR GPU OPERATORS

Your rigs earn while you sleep, or they don't.

If you supply GPUs to rental networks, you know the failure modes: a worker's memory climbs until the kernel kills it mid-render, destroying the paid job, and the container exits clean so nothing even tells you. A lane drops after an image update and earns nothing for a day. Dashboards say the process is healthy while the payout says otherwise. The deck is the operations layer for exactly this - built by an operator, for operators, running against a real production fleet every night.

What it does

Recycle before the memory cap

Workers restart in an idle window, before the memory cap - measured on anon-rss, not the cache-polluted number container stats show you. A recycle costs one idle restart; an OOM kill costs the restart plus the paid job it destroys. We do not claim a particular recycle prevented a particular kill — only that the restart happened on the measured condition, in an idle window.

Earning is the only truth

Checks verify at the payout layer, never the process layer. A green container that stopped earning is an incident; the deck treats it as one before you notice.

Incidents get acted on unattended

Lane down, model set wrong after a reboot, image update stalled - the recovery runs first and the report arrives after, with what failed and what fixed it.

Config changes on evidence

Memory limits, headroom, routing - rolled out as controlled comparisons with a pre-registered read-out, reverted automatically when the control wins.

OBSERVED lane-726 anonymous memory at 41% of its cgroup cap, 24 minutes into a paid shift — measured on anon-rss, not the cache-polluted number container stats report.

ACTION recycled at 01:23:57 inside a 51-second idle window between jobs — chosen, not lucky: the deck had been waiting for that window since 01:23:04, because restarting mid-render is the failure, not the fix.

OUTCOME the worker returned to an earning state and the job that render was carrying finished and paid.

NOT CLAIMED that this lane would otherwise have been killed. This case does not establish it: there is no growth slope here, no predicted exhaustion time, and no untreated control. GPUDeck refuses that inference on its own record and will not make an exception for its own product.

WHAT WOULD PROVE IT a shadow period where the same state is detected and not acted on, and the outcome recorded either way — then a rate, not an anecdote. That measurement is not done yet, and this page will say so until it is.

WHAT IS MEASURED how often that non-completion happens, which is a different question and the one we can answer. Completion yield is published per card per archived snapshot on the live fleet record, with the counts each figure is computed from. It establishes that unfinished paid work is real, recurring and material on a production fleet — a rate with a denominator rather than the anecdote above. It is not the measurement in the line before this one: the deck ran throughout, so there is no untreated period, and nothing there says the software moved those numbers.

EVIDENCE the run log, verbatim, lane ids relabelled:

01:23:04  memory: lane-11=58%  lane-726=41%  lane-9=24%  lane-725=22%
01:23:04  lane-726: 41% of cap (19.6 GiB anon, up 24 min) - waiting for an idle window
01:23:57  lane-726: RECYCLED before the cgroup could kill it (51s, between jobs)
01:26:05  nothing at threshold; no recycle needed

Real log lines from last night's production run, lane ids relabelled and otherwise unedited. The job that render was carrying finished and paid.

Read that third line critically — it says “before the cgroup could kill it”, and that is the deck asserting something it did not observe. The recycle happened; the kill did not, and no untreated run establishes that it would have. The log wording overstates and is being corrected at the source. It is reproduced here as it was written rather than quietly cleaned up, because the alternative is a page that claims a verbatim log and shows an edited one.

First cohort

The deck runs nightly on the fleet behind the GPUDeck record. Productizing it is next: self-serve, your machines, your keys, no agent phoning home. If that is your situation, say hello - cohort invitations go out in order.

Who this is for, precisely. Fleets that earn by completing work — job-settled inference and render networks, where a killed job is your lost revenue and "the worker is running" is not the same as "the worker is earning". Sogni is the network this was built on.

Who it is not for yet. If you rent whole machines or containers by the hour — Vast, Clore — the paid workload belongs to your renter, not to you, and the central move here (recycling a worker between jobs to save the job) is either meaningless or something a host should not do. Those platforms also already ship their own telemetry and rental alerts. We would rather say that than sell you a fit that is not there. If you run both kinds of supply, tell us — that is exactly the case we want to test.

No pricing is shown yet and nothing is charged. Joining does not qualify you: qualification is recorded by a human after a conversation, and is never inferred from fleet size or from this submission. There is no self-serve product to download until a pilot artifact actually exists.

Prefer mail? deck@gpudeck.com with subject “GPUDeck deck waitlist” works too.

Honesty, stated plainly: GPUDeck runs only on our reference fleet today. No outside operator has installed it yet — the founding cohort is for building and testing the first shadow-mode integrations on independent fleets. The tooling exists and runs nightly in production, on one fleet, entangled with that fleet's own scripts; what it is not yet is something a stranger can install alone. The first integration will be done with you, not handed to you — and what shadow mode would need from your fleet is written down already: what it reads, what it never does, what permissions it requires, what leaves your machine, and the questions we have not answered yet. No spam, no list resale — replies come from a person.

REPLICATE THIS RECORD ON A SECOND FLEET

Publish daily workload, GPU, attempts, completions and occupied GPU-hours under your own fleet identifier. Fleets are never averaged; earnings and worker identity are not accepted. Protocol validation is the gate - submitting does not mean acceptance.

Read CONTRIBUTING-DATA.md · Contact the record

Replicate this record on a second fleet.

Publish daily workload, GPU, attempts, completions and occupied GPU-hours under your own fleet identifier. Fleets are never averaged; earnings and worker identity are not accepted. Submitting does not make data part of the record — protocol validation is the gate.