What broke, and exactly why
A checklist of “trust us” claims is worth nothing. A track record of admitted, explained failures is worth something. Every entry here is real: found on this infrastructure, root-caused, and fixed — including the ones that were embarrassing to find. See also live status for current uptime.
The signed key directory silently served "do not trust" for days after the fix was deployed
high/.well-known/x402-receipt-keys — the governance-signed root that lets a third party verify our receipts without trusting our TLS — kept returning signature:null and an explicit "do not trust" note even after the signing secret was correctly configured on the production server and the service was restarted.
Impact
Anyone trying to independently verify a FractalAI x402 receipt during this window got a directory that explicitly told them not to trust it — the opposite of the feature’s purpose. No funds or signatures were affected; this was a trust-verification surface, not a payment path.
Root cause
The route had no `export const dynamic = "force-dynamic"`. Next.js statically pre-rendered it ONCE, at CI build time — on a machine that never had the governance seed (it only exists on the production server’s runtime environment, not the build environment). The unsigned response got baked into the build output and copied to production; restarting the server process afterward could never fix it, because the response was frozen at build time, not evaluated per request.
Fix
Added the missing dynamic export so the route re-evaluates on every request. Verified by reproducing the exact bug locally first (build without the seed, start with it — still unsigned, matching production), then confirming the fix (build manifest flips from static to dynamic; response signature verifies cryptographically with ml_dsa65.verify, correct key/signature byte lengths, and a tampered message is correctly rejected) before deploying.
What our monitoring missed
We had no test asserting this route is dynamic, and no alert on "governance directory response has been identical since before a config change." Both are now candidates for the monitor a validator business needs — tracked, not yet built.
The deploy pipeline was silently disabled for days — no deploy landed, nothing alerted
highThe CI/CD workflow that builds and ships the frontend and node binaries had its GitHub Actions monthly minutes budget exhausted, causing every run to fail instantly. Someone disabled the workflow to stop the noise. It was never re-enabled.
Impact
Every push to main during the outage window updated the repository but changed nothing in production — including, ironically, some of the fixes described elsewhere on this page. No commit was lost; the code was safe in git the whole time. The risk was purely "we believe we shipped a fix and we have not."
Root cause
GitHub confirmed the exact failure via its own API: "the job was not started because an Actions budget is preventing further use." That is a monthly quota, not a permanent block — but the workflow’s `state` had been flipped to `disabled_manually`, which does not reset with the billing cycle. Nothing in our monitoring watched the deploy workflow’s own enabled/disabled state or run history, only the production endpoints it deploys to — so a pipeline that stopped running looked identical, from the outside, to a pipeline that had nothing to ship.
Fix
Re-enabled the workflow, confirmed the monthly quota had rolled over, and manually triggered a real deploy to verify it — then verified the actual production effect of that deploy (see the receipt-key entry above) rather than trusting a green checkmark.
What our monitoring missed
We now know to check `gh run list` state periodically. The deeper fix — an alert when the deploy workflow itself goes quiet for longer than its own trigger frequency implies — is not built yet.
Our own uptime/discovery monitor went dark for days, and its own absence was invisible
mediumVISION, the automated system that watches node health, x402 endpoint health, and ecosystem discoverability every 30 minutes, stopped running entirely for roughly three days.
Impact
No functional impact on the live node or x402 endpoints, which kept running — VISION observes, it does not operate anything load-bearing. The impact was blindness: no fresh observations, no reasoning cycles, during the gap.
Root cause
The self-hosted runner that executes VISION’s scheduled jobs entered a crash-loop after its own diagnostic logs filled the disk (nothing rotates that directory by default), then stopped for good when the disk pressure cleared — because its systemd unit had no `Restart=` directive at all. A process that dies with no restart policy just stays dead until a human notices.
Fix
Restarted the runner, then added three independent layers so this specific failure mode cannot recur silently: `Restart=always` with burst limits on the unit itself; a separate watchdog that checks every 5 minutes and distinguishes "process is dead" (restart it) from "process is alive but stopped producing output" (alert instead of restarting — restarting a healthy process does not fix a broken workflow, it just hides the symptom); and pruning of the diagnostic-log directory that caused the original disk-full crash.
What our monitoring missed
Closed for this specific cause. The general pattern — "what watches the watcher" — remains true for anything we have not built a second layer for yet, and we are not claiming otherwise.
Two blocks the node mined itself tied on difficulty, and it stopped syncing to its own canonical head
criticalBlock production froze. Two locally-mined blocks at the same height ended up with equal cumulative difficulty, and after the tie, several code paths re-seeded the next mining attempt from the just-processed block’s own header instead of the node’s actual canonical head — so the miner kept building on a block that was not (or was no longer) the tip.
Impact
The chain stopped advancing until the fix was deployed. No funds were at risk (this is a liveness bug, not a signature or balance bug), but block production halting on a chain that exists to prove liveness and verifiability is a serious failure in its own right.
Root cause
A self-race: the node is both the block producer and its own consumer of "what is the current head." Three call sites in the mining/sync path read the just-mined or just-received block’s own header as the next parent, rather than re-querying the node’s actual canonical head after each import — fine until a tie made those two things briefly different.
Fix
Re-seeded the mining parent from the node’s real canonical head at every relevant call site, and changed the miner to never re-attempt on an unconfirmed parent (reset to none after submitting a block, and re-fetch fresh before the next attempt) while keeping the normal block-time pacing delay intact — an earlier version of this exact fix accidentally dropped that pacing delay, which we caught before it shipped broadly because it produced obviously-too-fast block times.
What our monitoring missed
We did not have an alert for "block height unchanged since the last check" until after this incident — that check is part of the eye’s regular cycle now.
No incidents are omitted because they were minor or because we found them ourselves. If you find one we have not listed, tell us.