Prasith Govin

← All notes

a check that couldn't run said red

2 min read

At 23:12 on a Saturday my memory health check reported RED. Zero of 26 probes had run. The safety monitor saw the RED and froze memory writes.

Memory was fine. Seven minutes earlier, a live recall test had passed.

I had just upgraded my agent runtime, OpenClaw, from 2026.9.2 to 2026.9.6, and then run the check from inside an agent's shell. The new version blocks the agent command from that context. Every probe failed to start, and the script counted each one as a failure.

That is the bug worth writing down. The check had two outcomes, pass and fail, and it needed three. "I couldn't run" and "the thing I measure is broken" produced the same RED, and the freeze logic downstream treated them the same way. A monitor with no way to report its own failure will eventually punish a healthy system for it.

The fixes went in the same night. The health script now checks where it is running before it starts, and exits with code 3 when the context is wrong, so a skipped run reads as skipped. This is the guard at the top of the script, trimmed of its override flag:

if (process.env.OPENCLAW_SHELL === "exec") {
  console.error("health.mjs: refusing to run from an agent shell");
  process.exit(3);
}

The nightly memory consolidation job had been running as an agent turn. It now runs as a plain command, so the runtime change can't reach it. The update runbook got a new line: after a runtime upgrade, run health checks from a normal shell.

The freeze clears after two clean runs in a row. I kicked the first one at 23:40 and expected the 05:20 scheduled cycle to be the second.

It wasn't clean. With the false alarm out of the way, the next run came back RED for a real reason. Seven of 12 fleet cases failed, all with the same error from the new runtime when a parent session spawned a child. Recall still passed 26 of 26. The upgrade had broken something, just not the thing the first RED claimed.

An agent worked the compatibility fix overnight. By midday Sunday the freeze had lifted through the normal two-pass rule. The daily health report now reads:

Freeze: inactive
Recall suite: 26/26 overall · 24/24 positives · 2/2 negative controls
Fleet matrix: 12/12

The health tooling lives in a private repo, so those two snippets, copied verbatim, are the receipts I can publish.

I took two things from it, and I trust the second one more.

First, a health check needs a separate answer for "I didn't measure." Mine is exit code 3. Which number you pick is arbitrary. Having one is the fix.

Second, a false alarm hides the next real one. If the check still couldn't say "couldn't run," every RED after the upgrade would have looked like the same shell problem, and I would have waved the real regression through. Fixing the instrument first is what made the second RED believable.

This is one incident on one machine, so I'm not calling it a law about monitoring. It is a rule in my runbook now.