I deliberately broke my Account Captain agent this week. In an isolated test lab
By Brett Bohannon · October 1, 2026 · Curated by George's Blog
I deliberately broke my Account Captain agent this week. In an isolated test lab, away from the live account.
Account Captain has one job: open a bounded case, call in only the specialist agents that case needs, wait for their callbacks, reconcile, and hand me one decision I can accept, reject, or challenge.
It can call Ads, Listing & Catalog, and Offer & Inventory. The rule is minimum routing. Only call in who the case actually needs.
I ran a synthetic case where only Ads and Offer & Inventory were needed.
Account Captain activated Listing & Catalog anyway.
The answer it gave back sounded reasonable. It still broke the rule.
It didn't catch that itself. I reviewed the run and made one bounded change to the routing rule. A fresh test routed only the two lanes needed. The full suite passed, but not on the first attempt. The unapproved-write check went 6 for 6.
I moved the reviewed version into Buzz, where it actually runs, and tested it once more. This time it recognized stale evidence, stopped, opened an internal blocker, and made no external change.
The dangerous failure isn't the one that's obviously wrong. It's the one that sounds right. Calling in an extra specialist looks thorough. It's actually noise, cost, and more surface for bad evidence to slip through.
One passing run proves one version completed one run. Running is not finished.
Full write-up is in today's newsletter, and Episode 3 walks through the whole test.