<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Prasith Govin, notes</title>
  <subtitle>Field notes on running AI agents in production.</subtitle>
  <link href="https://prasithg.com/feed.xml" rel="self"/>
  <link href="https://prasithg.com/"/>
  <id>https://prasithg.com/</id>
  <updated>2026-09-27T12:00:00Z</updated>
  <author><name>Prasith Govin</name><uri>https://prasithg.com/</uri></author>
  <entry>
    <title>a check that couldn&#x27;t run said red</title>
    <link href="https://prasithg.com/notes/a-check-that-could-not-run-said-red.html"/>
    <id>https://prasithg.com/notes/a-check-that-could-not-run-said-red.html</id>
    <updated>2026-09-27T12:00:00Z</updated>
    <summary>My memory health check froze a healthy system after an upgrade, and fixing that false alarm is what made the next, real RED believable.</summary>
    <content type="html">&lt;p&gt;At 23:12 on a Saturday my memory health check reported RED. Zero of 26 probes had run. The safety monitor saw the RED and froze memory writes.&lt;/p&gt;&lt;p&gt;Memory was fine. Seven minutes earlier, a live recall test had passed.&lt;/p&gt;&lt;p&gt;I had just upgraded my agent runtime, OpenClaw, from 2026.9.2 to 2026.9.6, and then run the check from inside an agent&#x27;s shell. The new version blocks the agent command from that context. Every probe failed to start, and the script counted each one as a failure.&lt;/p&gt;&lt;p&gt;That is the bug worth writing down. The check had two outcomes, pass and fail, and it needed three. &quot;I couldn&#x27;t run&quot; and &quot;the thing I measure is broken&quot; produced the same RED, and the freeze logic downstream treated them the same way. A monitor with no way to report its own failure will eventually punish a healthy system for it.&lt;/p&gt;&lt;p&gt;The fixes went in the same night. The health script now checks where it is running before it starts, and exits with code 3 when the context is wrong, so a skipped run reads as skipped. This is the guard at the top of the script, trimmed of its override flag:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;if (process.env.OPENCLAW_SHELL === &quot;exec&quot;) {
  console.error(&quot;health.mjs: refusing to run from an agent shell&quot;);
  process.exit(3);
}&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt; The nightly memory consolidation job had been running as an agent turn. It now runs as a plain command, so the runtime change can&#x27;t reach it. The update runbook got a new line: after a runtime upgrade, run health checks from a normal shell.&lt;/p&gt;&lt;p&gt;The freeze clears after two clean runs in a row. I kicked the first one at 23:40 and expected the 05:20 scheduled cycle to be the second.&lt;/p&gt;&lt;p&gt;It wasn&#x27;t clean. With the false alarm out of the way, the next run came back RED for a real reason. Seven of 12 fleet cases failed, all with the same error from the new runtime when a parent session spawned a child. Recall still passed 26 of 26. The upgrade had broken something, just not the thing the first RED claimed.&lt;/p&gt;&lt;p&gt;An agent worked the compatibility fix overnight. By midday Sunday the freeze had lifted through the normal two-pass rule. The daily health report now reads:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Freeze: inactive
Recall suite: 26/26 overall · 24/24 positives · 2/2 negative controls
Fleet matrix: 12/12&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The health tooling lives in a private repo, so those two snippets, copied verbatim, are the receipts I can publish.&lt;/p&gt;&lt;p&gt;I took two things from it, and I trust the second one more.&lt;/p&gt;&lt;p&gt;First, a health check needs a separate answer for &quot;I didn&#x27;t measure.&quot; Mine is exit code 3. Which number you pick is arbitrary. Having one is the fix.&lt;/p&gt;&lt;p&gt;Second, a false alarm hides the next real one. If the check still couldn&#x27;t say &quot;couldn&#x27;t run,&quot; every RED after the upgrade would have looked like the same shell problem, and I would have waved the real regression through. Fixing the instrument first is what made the second RED believable.&lt;/p&gt;&lt;p&gt;This is one incident on one machine, so I&#x27;m not calling it a law about monitoring. It is a rule in my runbook now.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>picking the colors before the model does</title>
    <link href="https://prasithg.com/notes/picking-the-colors-before-the-model-does.html"/>
    <id>https://prasithg.com/notes/picking-the-colors-before-the-model-does.html</id>
    <updated>2026-09-27T12:00:00Z</updated>
    <summary>The Opus product videos I saw all looked alike, so for the Parker teaser I wrote the palette, the copy, and a banned list before the model wrote any code.</summary>
    <content type="html">&lt;p&gt;Nearly every Claude Opus 5.5 product video I saw in September had the same look. Purple glow on near-black. Tiny uppercase labels in the corners. &quot;01 / 03&quot; step counters, fake timestamps, scanlines, a particle globe, words flashing in one at a time.&lt;/p&gt;&lt;p&gt;That look is what you get when the model picks the colors and the words. One reply under a popular example asked the author to &lt;a href=&quot;https://x.com/ggsimm/status/2103495318117708218&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot;&gt;get rid of the &quot;small caps and 01/03 everywhere.&quot;&lt;/a&gt;&lt;/p&gt;&lt;p&gt;For the Parker teaser I wrote the brief before any code existed and split the job in two. I owned the words and the look. Opus owned the motion.&lt;/p&gt;&lt;p&gt;The brief was specific enough to check against the render:&lt;/p&gt;&lt;ul&gt;&lt;li&gt;Three colors: paper F6F1E8, ink 14161A, one amber accent E07A2F. One soft red, used once, for a strike-through.&lt;/li&gt;&lt;li&gt;Inter for everything. Headlines 120 to 180 pixels tall on a 1080p frame.&lt;/li&gt;&lt;li&gt;Every scene holds at least 2.4 seconds and changes on the music&#x27;s downbeat. Nothing changes faster than twice a second. One camera move per scene, at most.&lt;/li&gt;&lt;li&gt;A banned list: small-caps corner labels, step counters, fake coordinates and timestamps, scanlines, particles, neon, dark mode.&lt;/li&gt;&lt;li&gt;Every on-screen phrase written by me, and a last frame that has to say &quot;Research prototype. Synthetic demo. Not medical advice.&quot;&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;Opus built it in Remotion, rendered a 16:9 and a 9:16 cut at 60 frames per second, and checked its own contact sheet against the brief before handing it over.&lt;/p&gt;&lt;p&gt;The light theme started as a way to look different from the wave. It held up for a better reason. Parker is built for people with effortful speech, starting with my dad, and the viewer I care about can&#x27;t read tiny gray UI on black. Big dark type on warm paper is legible. The trend look would have been wrong for the audience before it was wrong for the brand.&lt;/p&gt;&lt;p&gt;What I&#x27;m keeping: when a model has a strong default, the brief has to name the default and ban it. A brief that says &quot;clean and modern&quot; leaves those choices to the model. A hex list, a timing rule, and a banned list got me something I could check line by line.&lt;/p&gt;&lt;p&gt;The teaser is 30 seconds and sits under Parker on the front page.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>the visual test passed and the page still looked wrong</title>
    <link href="https://prasithg.com/notes/the-visual-test-passed-and-the-page-still-looked-wrong.html"/>
    <id>https://prasithg.com/notes/the-visual-test-passed-and-the-page-still-looked-wrong.html</id>
    <updated>2026-09-06T12:00:00Z</updated>
    <summary>My site passed its first visual review because the checks measured fit, not the composition I had approved.</summary>
    <content type="html">&lt;p&gt;I told my agent it could manage my personal site. It changed the headline.&lt;/p&gt;&lt;p&gt;That was the wrong place to spend the autonomy. The site already had a headline and composition I liked. What had gone stale were the project statements, build log, and field notes. The agent reached for the most visible copy on the page because it had a new positioning sentence and permission to keep the site current. &lt;a href=&quot;https://github.com/prasithg/prasithg.github.io/pull/2&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot;&gt;The regression PR is public&lt;/a&gt;.&lt;/p&gt;&lt;p&gt;The replacement did not look catastrophically broken. Nothing overlapped or clipped, the mobile page had no horizontal overflow, and the links still worked. The first visual review passed.&lt;/p&gt;&lt;p&gt;The page still looked worse.&lt;/p&gt;&lt;p&gt;On desktop, the new headline was only ten characters longer, but the word lengths forced two extra lines inside a narrow column. The headline grew by 118.72 pixels. The portrait stayed the same height. The whole hero grew by the same 118.72 pixels, and the project section moved down with it. That left a dead column of space below the portrait and turned a balanced block into a long stack of type next to an image that had already ended.&lt;/p&gt;&lt;p&gt;The test was green because it measured the wrong property. It answered one question: does this fit? It never asked whether the page still looked like the site I approved.&lt;/p&gt;&lt;p&gt;I corrected the agent, restored the old hero exactly, and reran the comparison at 1440 by 1100 and 390 by 844. This time the check was differential. The restored screenshots matched the approved desktop and mobile baselines pixel for pixel. The element geometry matched too: headline height, portrait height, hero height, and the point where the project section began.&lt;/p&gt;&lt;p&gt;Then &lt;a href=&quot;https://github.com/prasithg/prasithg.github.io/pull/3&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot;&gt;the correction became a contract&lt;/a&gt;. Routine automation can update verified project copy, receipt-backed build-log rows, and evidence-backed notes. It cannot change the header, hero, portrait, navigation, CSS, or section order unless I explicitly ask for a redesign. CI now fails if the hero text, portrait, navigation, or CSS changes. Section order is still a review rule, not a check.&lt;/p&gt;&lt;p&gt;That boundary matters more than the wording mistake. Giving an agent authority over a product does not mean every surface should stay mutable. Some decisions are settled. The useful autonomy is inside the areas that benefit from continued judgment: facts that go stale, new work that deserves a receipt, and lessons that have earned an article.&lt;/p&gt;&lt;p&gt;A green visual test can be as misleading as green CI. It may prove only that nothing broke hard enough for the test to notice. If the product already has an approved composition, the next version should be judged against that composition, not against the bare minimum of rendering without overflow.&lt;/p&gt;&lt;p&gt;The human correction was the grader. The durable part was making sure the same correction would not be needed twice.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>my own repo disproved my readme</title>
    <link href="https://prasithg.com/notes/my-own-repo-disproved-my-readme.html"/>
    <id>https://prasithg.com/notes/my-own-repo-disproved-my-readme.html</id>
    <updated>2026-07-01T12:00:00Z</updated>
    <summary>I caught my own README overclaiming Parker, and the real eval number was better than the lie.</summary>
    <content type="html">&lt;p&gt;My README said Parker understood dysarthric speech &quot;at least 90% of the time.&quot; I wrote that line. I believed it when I wrote it. It was wrong.&lt;/p&gt;&lt;p&gt;Parker is a voice assistant I built for my dad, who has Parkinson&#x27;s. The whole point is that it hears him when other systems give up, so the comprehension number is the product. A round number like 90% should have made me suspicious of myself. Nobody measures a messy real-world thing and lands on a clean 90.&lt;/p&gt;&lt;p&gt;I hadn&#x27;t lied on purpose. I wrote the README early, before the corpus eval existed, picked a number that felt safe, and never went back to reconcile it against what the harness actually reported.&lt;/p&gt;&lt;p&gt;The last real-audio eval told a different story. On dysarthria corpora, baseline comprehension was 58%. With the repair loop and n-best rescoring, it came up to 82%. A 24-point recovery lift, measured on real audio and backed by a 656-test suite, not the flat 90 I&#x27;d rounded to somewhere in my head.&lt;/p&gt;&lt;p&gt;I found the gap by opening my own &lt;code&gt;benchmark/reports/&lt;/code&gt; the way a stranger would. Someone skeptical clicks into that folder, sees 82, scrolls up, sees me claiming 90 on the front page. Two minutes to catch me. The part that stings: Parker ships an overclaim-guard eval. I built a thing to flag inflated numbers and then inflated one on my own landing page.&lt;/p&gt;&lt;p&gt;The fix had nothing to do with the number. I deleted the claim the evidence didn&#x27;t support and wrote down the one it did.&lt;/p&gt;&lt;p&gt;What I keep noticing is that the true story sells harder than the round one. &quot;58 to 82 on real dysarthric audio&quot; says the repair loop does real work, that I measured it, that I&#x27;ll show you the tests. &quot;90%&quot; says trust me. A skeptic believes the first and waves off the second, and the skeptic is the reader I want.&lt;/p&gt;&lt;p&gt;So the rule now: if a headline number is round, I go find the eval that produced it before it ships. No eval, no number.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>a repo can commit daily and still rot</title>
    <link href="https://prasithg.com/notes/a-repo-can-commit-daily-and-still-rot.html"/>
    <id>https://prasithg.com/notes/a-repo-can-commit-daily-and-still-rot.html</id>
    <updated>2026-07-01T12:00:00Z</updated>
    <summary>My agent audited our open-source sync and found three months of drift behind green CI and a daily commit streak.</summary>
    <content type="html">&lt;p&gt;Green CI and a daily commit streak feel like health. They are activity. My agent audited our own open-source sync and found three months of quiet drift sitting behind both.&lt;/p&gt;&lt;p&gt;The repo was the public Clawrari sync. Every check passed. Commits landed most days. By the numbers everyone trusts, it looked alive. Then the audit opened the front door instead of the dashboard.&lt;/p&gt;&lt;p&gt;What it found: an orphaned sister repo nobody had touched, a flagship domain pointing at a parking page, and two different taglines describing the same project depending on where you read. None of that trips a test. CI doesn&#x27;t check whether your homepage loads. A commit streak doesn&#x27;t notice that the story in the README stopped matching the tree three months ago.&lt;/p&gt;&lt;p&gt;The useful part is who caught it. An agent ran the audit as a task, with no ego in the diff. I&#x27;d been the one committing daily, so I was the last person likely to notice the rot. I was reading the same signals that were hiding it. The agent had no attachment to the streak, so it read the domain, read the taglines, read the sibling repo, and reported the gap flat.&lt;/p&gt;&lt;p&gt;We shipped v0.5 the same afternoon. The parking page got a real landing. The orphan got folded in or archived. The taglines collapsed to one.&lt;/p&gt;&lt;p&gt;What got me is how good the false signal was. Commit cadence and green checks are exactly the metrics you point at when someone asks if a project is maintained. They&#x27;re also the metrics that stay green while the project quietly stops being true.&lt;/p&gt;&lt;p&gt;So I changed what &quot;healthy repo&quot; means to me. It&#x27;s no longer &quot;CI is green and I committed this week.&quot; Now it&#x27;s two questions: does the front door work, and does the story on the page still match what&#x27;s in the tree. An agent can check both on a schedule, without caring how the answer reflects on whoever wrote the code.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>guardrails, not prompts</title>
    <link href="https://prasithg.com/notes/guardrails-not-prompts.html"/>
    <id>https://prasithg.com/notes/guardrails-not-prompts.html</id>
    <updated>2026-07-01T12:00:00Z</updated>
    <summary>A prompt tweak fixes one bad output; a guardrail in the loop stops every one that fails the same way.</summary>
    <content type="html">&lt;p&gt;A prompt tweak fixes one bad output. It patches the instance in front of you and does nothing about the next one that fails the same way. A guardrail in the loop is a different tool: an executable rule that fires on every output of that shape, whether or not I&#x27;m watching.&lt;/p&gt;&lt;p&gt;The way this works for me is a tripwire register. Every recurring failure graduates from a note into a machine check. A note says &quot;remember to verify after you mutate something.&quot; A tripwire named &lt;code&gt;RT-VERIFY-AFTER-MUTATE&lt;/code&gt; refuses to let an agent claim &quot;done&quot; without a verification read first. The note decays the moment I stop reading it. The tripwire runs on its own.&lt;/p&gt;&lt;p&gt;There&#x27;s a small register of these now. &lt;code&gt;RT-MCPORTER-ARGS&lt;/code&gt; catches a specific call I kept getting wrong. &lt;code&gt;RT-EVAL-GATE&lt;/code&gt; blocks a skill change that has no linked eval. &lt;code&gt;RT-GENERATED-CODE&lt;/code&gt; gates code that came out of a generator. Each one started as the same mistake showing up twice, then a third time, until &quot;I&#x27;ll be careful next time&quot; stopped counting as a fix.&lt;/p&gt;&lt;p&gt;The step I used to skip is making the guardrail visible when it fires. A guardrail you hope is working is just a comment. The night builds bind theirs in code: a mutation cap, a stop if a change touches more than 5% of a corpus, a rollback path. When one trips, it writes down why it stopped. I read the stop reason in the morning instead of trusting that nothing went wrong.&lt;/p&gt;&lt;p&gt;I still use prompt tweaks. If a failure is genuinely a one-off, low blast radius, unlikely to come back, a better instruction is the right size of fix. The mistake is reaching for a prompt tweak when the failure is a class. You re-prompt the same shape of bug across weeks, apologizing to yourself each time, when one rule in the loop would have ended the series.&lt;/p&gt;&lt;p&gt;When an agent output is wrong, I treat it like a production bug: root-cause it, then leave a check behind so that exact failure can&#x27;t come back quietly. A prompt tweak would have fixed today&#x27;s output and left tomorrow&#x27;s alone.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>managing coding agents is managing juniors who forget everything overnight</title>
    <link href="https://prasithg.com/notes/managing-coding-agents-like-juniors.html"/>
    <id>https://prasithg.com/notes/managing-coding-agents-like-juniors.html</id>
    <updated>2026-07-01T12:00:00Z</updated>
    <summary>Coding agents are capable juniors with no overnight memory, so the work is in the brief, not the model.</summary>
    <content type="html">&lt;p&gt;A coding agent is a capable junior developer who forgot everything overnight. Sharp on the task in front of it, no memory of the last one. Once I started managing it that way, most of my agent problems turned into management problems I already knew how to handle.&lt;/p&gt;&lt;p&gt;With a human junior, I&#x27;d never hand over real work in a one-line Slack message. I&#x27;d write down what &quot;done&quot; looks like. With an agent I kept doing exactly that, then acting surprised when it wandered off. The agent was fine. My brief was the problem.&lt;/p&gt;&lt;p&gt;So every spawn now carries the same five sections. Objective: the one thing this run exists to produce. Output path: where the artifact goes, exactly. Write scope: what it may touch and nothing else. Verification: how it proves the work is real before claiming done. Context budget: what to read, what to skip, when to stop. That&#x27;s the whole contract. It fits on a screen. It exists because the agent starts every task with no memory of the last one, so anything I don&#x27;t write down, it doesn&#x27;t know.&lt;/p&gt;&lt;p&gt;Verification is the section I underweighted longest. When a junior says &quot;it&#x27;s done,&quot; that&#x27;s a claim you still verify, which is why you ask to see the diff. An agent saying &quot;done&quot; is the same claim with more confidence and less shame. So I don&#x27;t let &quot;done&quot; be a status the agent declares. It&#x27;s a read I require: show me the file exists, show me the test passed, show me the number.&lt;/p&gt;&lt;p&gt;The other half is matching the task to the harness. Some jobs shouldn&#x27;t go to a chat subagent at all. Heavy work like dist surgery goes to a supervised executor instead, because I watched two of those blow up in a single day when I sent them to the wrong kind of worker. You wouldn&#x27;t hand a delicate migration to the junior in their first week either, however sharp they are on small tasks.&lt;/p&gt;&lt;p&gt;The model keeps getting better. The brief is still where the work lives. A strong model with a vague brief gives you a confident junior building the wrong thing very fast. The test I use now: if I couldn&#x27;t hand this brief to a competent stranger and expect the right output, it isn&#x27;t ready to hand to an agent.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>a subagent dies quietly: the context-death contract</title>
    <link href="https://prasithg.com/notes/the-context-death-contract.html"/>
    <id>https://prasithg.com/notes/the-context-death-contract.html</id>
    <updated>2026-07-01T12:00:00Z</updated>
    <summary>Five subagents died in two weeks, all context blowouts, so the anti-death rules became a required contract on every spawn.</summary>
    <content type="html">&lt;p&gt;Five subagents died on me in two weeks. Three on one day, two a couple of weeks before. The cause was the same every time: they read too much.&lt;/p&gt;&lt;p&gt;A subagent death looks like this. It gets a task, opens a large file with a full read to &quot;understand the context,&quot; fills its window, keeps exploring past the point where it already had enough to act, then times out or runs out of room before it writes a single line to disk. Hours of a capable model, zero output. The plan was fine. The agent never executed it because it spent its whole budget looking.&lt;/p&gt;&lt;p&gt;Once I saw the same failure five times, it stopped being bad luck and became a class. So I wrote a contract every spawn has to follow. Five sections: objective, output path, write scope, verification, and the one that carries the weight, a context budget.&lt;/p&gt;&lt;p&gt;The context budget is a short list of rules that sound obvious and get ignored under pressure. Write the skeleton of your output file first, before reading anything, so partial progress survives. Read narrowly and on purpose, never a whole large file when a targeted slice answers the question. Route the genuinely heavy files to a supervised executor instead of a chat subagent, since that&#x27;s the tool built to chew through them. Stop exploring the moment you have enough to decide. Don&#x27;t re-read what you already read.&lt;/p&gt;&lt;p&gt;The one that saved the most agents is the first. Skeleton file first. An agent that writes its output structure before it starts reading will, at worst, leave a half-finished file I can inspect and salvage. An agent that reads first and writes last leaves nothing behind when it dies. I would rather have a rough draft on disk than a perfect plan inside a process that just ran out of memory.&lt;/p&gt;&lt;p&gt;There&#x27;s a lesson in here about my own instincts. I kept blaming the model for the deaths and shopping for a smarter one. A smarter model dies the same way if you point it at a huge file and tell it to understand everything before it acts. The problem was discipline, and discipline lives in the brief.&lt;/p&gt;&lt;p&gt;So the contract is the rule now. No spawn without those five sections, and the context budget is not the optional one, because the expensive failures on this stack have all been the same quiet death.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>metadata-only receipts: trace, then tripwire</title>
    <link href="https://prasithg.com/notes/metadata-only-receipts.html"/>
    <id>https://prasithg.com/notes/metadata-only-receipts.html</id>
    <updated>2026-07-01T12:00:00Z</updated>
    <summary>Every agent run leaves a metadata-only trace that a fail-closed tripwire gates on before the next run proceeds.</summary>
    <content type="html">&lt;p&gt;Every agent run in my stack leaves a trace. It doesn&#x27;t capture the content the agent read or wrote. It records metadata: the plan it ran, the context files it opened, the tool calls it made, the metrics, the exit status. A small JSON file per run, plus one append-only row in an index. That&#x27;s it.&lt;/p&gt;&lt;p&gt;It started as a way to stop scraping raw logs. When an overnight run did something surprising, I&#x27;d dig through pages of output trying to reconstruct what happened. The trace collapsed that into one file I read in ten seconds: here&#x27;s the plan, here&#x27;s what it touched, here&#x27;s how it exited. No archaeology.&lt;/p&gt;&lt;p&gt;Keeping it metadata-only is deliberate, and it buys two things at once. Because I don&#x27;t store what the agent read or generated, a trace can&#x27;t leak the contents of a private file or a customer record. And it still answers the questions that matter: did the run touch files it shouldn&#x27;t have, did it exit clean, did it stay inside its budget. None of those need the content.&lt;/p&gt;&lt;p&gt;Then the trace grew a second half. A tripwire reads the trace before the next run is allowed to proceed, and it fails closed. The pre-run gate is the load-bearing piece: &lt;code&gt;lifecycle_start&lt;/code&gt; runs first, and if the trace looks wrong, the run doesn&#x27;t start. Postflight can change how the next run proceeds too. The gate defaults to blocking, so a broken or missing trace stops the pipeline instead of waving it through.&lt;/p&gt;&lt;p&gt;I shipped v0.1 as a fail-closed CI gate. Zero dependencies, no content capture, wired into CI so it runs on every change and not only when I remember to look. That last part matters more than it sounds. A safety check I have to run by hand is a safety check I&#x27;ll skip on the busy day, which is exactly the day I need it running.&lt;/p&gt;&lt;p&gt;The loop I ended up with is short: trace every run, gate on the trace, block by default when the trace is bad. Three steps, no content capture, and it runs whether or not I&#x27;m paying attention. That last property is the whole reason it earns a place in the pipeline.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>the night build queue</title>
    <link href="https://prasithg.com/notes/the-night-build-queue.html"/>
    <id>https://prasithg.com/notes/the-night-build-queue.html</id>
    <updated>2026-06-01T12:00:00Z</updated>
    <summary>I set the plan at 10:30pm and agents build while I sleep; it only works if every run re-derives its numbers by morning.</summary>
    <content type="html">&lt;p&gt;I set the plan at 10:30 at night and go to sleep. The agents build until they&#x27;re done or blocked. My job starts in the morning, when I review what showed up.&lt;/p&gt;&lt;p&gt;On one long session, execution autonomy landed somewhere north of 95%. Plan set at 10:30pm, zero interventions overnight. I didn&#x27;t wake up to babysit a stuck run. That number sounds like the point, but it isn&#x27;t. Autonomy is easy to fake. You can hit 95% by lowering the bar for what counts as done.&lt;/p&gt;&lt;p&gt;What makes it real is whether I can trust the morning review. The build I hold up as the standard is one Claude Code finished at 01:12. By the time I looked, there was working code, a benchmark, three evidence files, and a script called &lt;code&gt;autoloop-smoke.sh&lt;/code&gt; that re-derives every number in the report when you run it. I didn&#x27;t have to trust the summary. I ran the script and watched the numbers come back the same.&lt;/p&gt;&lt;p&gt;That&#x27;s the line for me. A night build I can&#x27;t re-verify in the morning isn&#x27;t done, however clean the summary reads. If the only proof a thing worked is the agent telling me it worked, I&#x27;ve automated the work and kept none of the accountability. Re-derivation is how I keep the accountability without staying up for it.&lt;/p&gt;&lt;p&gt;So the queue has a shape. At night I set the plan, because taste doesn&#x27;t automate and I&#x27;m not going to pretend it does. The agents run. In the morning I read the evidence, rerun the smoke script, and either accept or reject. The accept and the reject are mine. Deciding what&#x27;s worth building is mine. The typing, more and more, is not.&lt;/p&gt;&lt;p&gt;The failure mode I watch for is the run that &quot;passed&quot; and can&#x27;t prove it. Green summary, no re-derivable numbers, nothing I can rerun. That build goes back in the queue, because a pass I can&#x27;t reproduce is a story, and I&#x27;ve been burned by my own stories before. What I want from every overnight run is that morning-me can re-derive the result without trusting night-me&#x27;s agent.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>two agents reviewing each other over a private channel</title>
    <link href="https://prasithg.com/notes/two-agents-reviewing-each-other.html"/>
    <id>https://prasithg.com/notes/two-agents-reviewing-each-other.html</id>
    <updated>2026-06-01T12:00:00Z</updated>
    <summary>Two agents on two machines reviewing each other produced our most novel work, and a six-hour deadlock that taught the most.</summary>
    <content type="html">&lt;p&gt;For a stretch this summer I ran two top-level agents on two different machines, reviewing each other&#x27;s work. Claw on one box, Hermes on another, talking async over a shared Dropbox folder and a Discord channel. Two peers on roughly equal footing, arguing, rather than one coordinator handing tasks down to helpers.&lt;/p&gt;&lt;p&gt;The output was the most novel thing either of them produced alone. One long session generated around eighteen design artifacts plus a working prototype: an agent-doorbell, a cron-health-sentinel with four passing unit tests, a handoff-lint. The interesting part is why two peers beat one agent with helpers. A solo agent context-switches across my whole queue. A peer agent can sit on one thread for eight hours and refuse to let go, which the solo agent structurally can&#x27;t.&lt;/p&gt;&lt;p&gt;Then the same setup deadlocked for six hours. One agent asked the other a blocking question and waited. The other never answered, because nothing told it the question was blocking. Both sat there, polite and stuck. We were building agent-coordination tooling at the time, and we&#x27;d just reproduced the exact failure that tooling exists to prevent.&lt;/p&gt;&lt;p&gt;I could have hidden that. It&#x27;s the best part of the story. The deadlock was the eval signal. It told us precisely which coordination primitive was missing, which is why the doorbell and the handoff-lint exist at all. The doorbell exists because of that morning: two agents ignored each other for six hours, and neither could ring the other to break it.&lt;/p&gt;&lt;p&gt;What I took from it is that peer review between agents is worth the coordination cost, as long as you treat the coordination failures as data. Two agents on two machines find bugs a single agent won&#x27;t, in the work and in the handoff between them. The handoff is where most of the surprises live.&lt;/p&gt;&lt;p&gt;The open question I&#x27;m still sitting with is how many peers is too many. Two found things one couldn&#x27;t. Two also spent six hours deadlocked on a question a human would have answered in ten seconds. I don&#x27;t know yet where the coordination cost overtakes the benefit, and I&#x27;d rather find that out on tooling than on anything that matters.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>cross-model review is the default</title>
    <link href="https://prasithg.com/notes/cross-model-review-is-the-default.html"/>
    <id>https://prasithg.com/notes/cross-model-review-is-the-default.html</id>
    <updated>2026-05-01T12:00:00Z</updated>
    <summary>Every change gets reviewed by a different model family before it ships: one writes, one reviews, single round.</summary>
    <content type="html">&lt;p&gt;Before code ships here, a different model family reviews it. Claude writes, Codex reads it and returns a verdict. One pass, and the reviewer has no authorship stake in what it&#x27;s grading. That asymmetry is the point.&lt;/p&gt;&lt;p&gt;I use different families on purpose. Two instances of the same model share the same blind spots. Ask a model to review its own work and it will confidently miss the same things it missed writing it, because the gap is in the training, not the effort. A model from another family has different priors, so it catches what the author&#x27;s priors hid.&lt;/p&gt;&lt;p&gt;One round only. I looked at multi-round setups where two models argue toward consensus, and on cost-adjusted quality, plain review beat debate cleanly. One writer, one reviewer, one round, a structured verdict. Extra rounds mostly added tokens. So we kept review and didn&#x27;t grow it into an argument that never settles.&lt;/p&gt;&lt;p&gt;It catches real problems well beyond style nits. Our task board has a whole &quot;In Review&quot; state that exists only because that&#x27;s where cross-model review and the accept-or-reject happen. One night it caught a genuine P1 on a Codex cross-review that would have shipped otherwise. That&#x27;s the bar I hold it to. If cross-review only ever flagged formatting, I&#x27;d drop it. It earns its slot by catching what the author was too close to see.&lt;/p&gt;&lt;p&gt;The cost is honest and small. One extra state on the board, one more model in the loop, a few minutes per change. For that I get a second set of priors on everything before it goes out, and a habit that treats &quot;I wrote it and it looks right to me&quot; as the least trustworthy signal in the building.&lt;/p&gt;&lt;p&gt;So the default here is blunt: nothing I care about ships on one model&#x27;s self-assessment. The author, human or agent, is the worst judge of whether the author got it right. A reviewer with different training and no stake in the diff beats one more pass of the writer checking their own work.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>memory is the product surface</title>
    <link href="https://prasithg.com/notes/memory-is-the-product-surface.html"/>
    <id>https://prasithg.com/notes/memory-is-the-product-surface.html</id>
    <updated>2026-05-01T12:00:00Z</updated>
    <summary>For an agent, memory isn&#x27;t a bolt-on feature; it gets a measured architecture with benchmarked recall and negative controls.</summary>
    <content type="html">&lt;p&gt;The memory plan I wrote opens with a flat claim: for this agent, memory is the product surface. I still think that&#x27;s the most important sentence in the doc, and everything else follows from taking it literally.&lt;/p&gt;&lt;p&gt;A smart model that forgets everything is expensive autocomplete. The intelligence you pay for resets every session unless something outside the context window remembers for it. So memory can&#x27;t be whatever the current plugin happens to do by default. It&#x27;s the substrate the whole system stands on, and it deserves the same rigor I&#x27;d give any load-bearing part.&lt;/p&gt;&lt;p&gt;Taking it seriously meant turning it into something measured. Local-first, so the source of truth is files I own and can read, not a service I hope is up. Benchmarked recall, so I know retrieval works instead of assuming it. Bounded latency, so remembering doesn&#x27;t cost more than it saves. Drift detection I can actually see. And explicit decision gates before anything gets swapped or disabled, so no backend change happens on a hunch.&lt;/p&gt;&lt;p&gt;The benchmark is where this got real. It has recall buckets, and it has negative controls: cases where the correct answer is to retrieve nothing. Ask what nineteen times twenty-three is, and the right move is to answer, not to search memory for it. Ask for a password, and it should refuse instead of inventing one. Ask whether we decided to move everything to the cloud, and the honest answer is no, with the reason why. A memory system that only passes direct keyword recall isn&#x27;t good enough, because the worst failures are confident retrieval of things that were never true.&lt;/p&gt;&lt;p&gt;The discipline underneath all of it is one habit: write everything down. Context is temporary and files persist, so the job every session is to move what mattered out of the window and onto disk before it evaporates. The architecture is the grown-up version of that habit.&lt;/p&gt;&lt;p&gt;One line from that plan still runs everything I build: treat memory like the product, keep measuring whether it recalls, and stay suspicious of any system that can&#x27;t correctly retrieve nothing but insists it can retrieve everything.&lt;/p&gt;</content>
  </entry>
</feed>
