Programmable RevenueBuilder Edition Subscribe
Issue 01 Reliability Field

Fail closed: why unknown never counts as healthy in fleet capacity accounting

Fleet capacity is a distributed-state problem: hundreds of domains whose status can be stale, missing, or wrong at read time. The design that keeps the number honest: explicit states, unknown never counts as clean, every rate ships its denominator, and sending halts when state cannot be read.

Thomas Cornelius August 22, 2026 9 min read
Technical schematic: four domain states, available, quarantined, blocked, unknown, with unknown highlighted and annotated holds the ramp.
The fourth state is the fix: unknown never counts as healthy.

How much can this fleet send today? For a system of hundreds of domains, that number is a distributed-state problem: every domain has a status, and at read time any given status can be stale, missing, or wrong. The design question is what the aggregate does when a status is unreadable.

The common default, inherited from application infrastructure, is to let missing data pass through as healthy. Our first fleet report had that default, and it taught us the discipline this article describes: blocked domains passed through as available capacity, and the total looked normal while it drifted from reality. Nothing here is specific to email; this is dashboard engineering for any fleet of anything.

Why silent drift is hard to catch

A report with that default cannot visibly fail. A domain with a missing status comes through as usable, so the total stays inside the range someone expects, and nobody goes looking. A report that fails silently does not look broken. It looks fine. That is the whole problem, and it is why the fix has to live in the data model rather than in vigilance.

A smoke detector with a dead battery does not show smoke. It shows nothing, which looks exactly like “no fire.” Any dashboard that treats “no data” and “all clear” as the same picture has the same flaw, and you will not find out which one you are looking at until it matters.

Engineers know this trade as fail-open versus fail-closed. Most application infrastructure fails open on purpose: when a cache or a feature flag service goes down, requests continue, because interrupting the product over a lookup is the worse outcome. That default is so common that it follows the data into places it does not belong. Reporting is one of those places. Sending is another.

The state machine that replaced the guess

A domain now has an explicit state: available, quarantined, blocked or unknown. Only available counts toward capacity.

available quarantined blocked unknown counts toward capacity held, under repair cannot deliver holds the ramp a failed lookup lands in unknown, so it REDUCES what we are willing to send instead of inflating it
Four explicit states. The bug was the missing fourth one.

Unknown is the load-bearing state: without it, anything the collector cannot read falls through to “usable.” With it, a failed lookup reduces the number we are willing to send rather than inflating it. The same rule reaches the reputation data. Every sending address is checked daily against nine blocklists, and a refused lookup is written down as unverified, never counted as clean.

What we changed in the report itself

Every rate now ships with the count and the denominator it was calculated from, because a percentage published without the counts behind it can hide this exact problem. A 95 percent health figure means something different over 40 domains than over 400, and the collapse from 400 to 40 is precisely what a bare percentage conceals.

The report also cannot outrun its own data. It reads a file written by the collection job rather than querying live when someone opens it. If collection fails, the report is visibly stale with a timestamp on it, instead of quietly rendering a clean page from nothing.

If you own a dashboard in any department, two questions catch most versions of this bug. What does this chart show when the data pipeline behind it is broken? And does every percentage on it say how many things it was counted from? “The same number as yesterday” is the wrong answer to the first question.

The same discipline, applied to sending

Daily sending limits are enforced by counters held in a fast in-memory datastore. In many systems, when that datastore becomes temporarily unavailable, requests are allowed to continue so an infrastructure problem does not interrupt the application. We use the opposite behaviour for sending. New sends halt until the state is readable again.

The two failure directions do not cost the same. Holding the batch drops a day’s volume, and the shortfall is made up over the following days. Carrying on sending can exceed a limit on a domain that took weeks to warm, and some of that reputation does not come back. The opt-out list follows the same rule: an opt-out that was recorded but is momentarily unreadable is treated as unread, never as absent. The counters and the opt-out list are both state, and neither has a safe default value when it cannot be read.

The gauntlet a batch runs before sending

The fail-closed rule is enforced as a chain of checks, and any check that cannot produce its evidence stops the batch at that point rather than passing it on.

evidence dry run approval recompute reserve send fleet state < 5 min old fingerprint of list, copy, opt-out file human, bound to that run, 72h expiry route rebuilt; fleet moved meanwhile all-or-nothing commit any check that cannot produce its evidence halts the batch here stops apply at six scopes and are enforced twice: before scheduling and again before submission
The gauntlet. No stage treats a missing answer as a passing one.

The system reads current fleet state and rejects anything older than five minutes. It fingerprints the exact list, the exact copy, and the exact opt-out file, and a person approves that exact run; the approval expires after 72 hours rather than staying valid indefinitely. The route is rebuilt before authorizing, because the fleet moves between approval and send. Then the whole run commits or none of it does, so a partially sent batch cannot happen. Stops apply at six scopes and are enforced twice, once before scheduling and again before submission. And because a provider accepting a message is not evidence that it arrived, acceptance and arrival are counted separately (the placement-test article is that distinction, measured).

There is no stage in that sequence where a missing answer is treated as a passing one. That sentence is the whole design.

Who this applies to

Anyone who owns a dashboard, in any department. The bug pattern needs no email and no fleet: a marketing attribution report, a sales pipeline count, and a finance rollup can all treat a broken pipeline as a healthy zero. The two questions in the callout above are the whole audit.

Still open

Fail-closed has a price we have not fully measured: the volume lost to held batches over a year, against what fail-open would have cost in burned reputation. We accept the trade on the asymmetry, lost volume returns in days while lost reputation takes weeks, rather than on a completed ledger of both sides.

The pattern, in code

Simplified to its logic, the rule that fixed our report fits in a case statement. The failure mode it removes is the empty string that means “clean” and the error that also looks like an empty string.

# A blocklist lookup has three outcomes, not two
answer=$(dig +short "$addr".zen.spamhaus.org) || answer="lookup-failed"

case "$answer" in
  127.*)           echo "listed"      ;;
  "")              echo "clean"       ;;
  *)               echo "unverified"  ;;  # holds the ramp, never counts as clean
esac

Whatever your stack, the test is the same: force the lookup to fail and watch what your dashboard reports. If the number goes up or stays flat, your report can lie to you the way ours did.

The principle

A missing answer is never a passing one. It is the cheapest reliability rule we know, it would have caught our own report on day one, and it closes this issue where it started: build so that the system tells you the truth, especially when part of it is down.

Build with graph8