How I Governed Two Claude Agents to Build and Test a Live Word Puzzle
I'm not a professional developer, but I built Sujido, a live word puzzle, with multiple AI agents. For months, they gave me bloat, lost context, and slop, and neither I nor the models could untangle what had been decided or why. I eventually settled on two agents: one Claude instance that plans and reviews, and another that builds. I imposed rules on both, one failure at a time. None of these rules would surprise a professional engineering team, but nobody handed them to me. That's the key lesson I learned: agents won't bring the discipline for you. You have to impose it. When I showed the first game revision to strangers, their verdict was that the first half was "boring" and the second half was novel. It was hard news to hear about my baby, but I took the feedback to heart and cut the game in half. That required surgery on code tangled across both halves, followed by a rebuild under the rules below. Before the rebuild, a separate AI audit graded the codebase B−. Afterward, it graded it A−. The deployed site shrank by 97.5%, and every serious finding was fixed and locked in by an automated test. The models didn't change drastically, but my operating rules did. Step-by-step: 1. I wrote a constitution before the first prompt. The repo holds what every agent must read before working: `CLAUDE.md` files for the project as a whole and for each major part of the code, a collaboration protocol for how the agents and I work together, a spec for the core data format, and a decision log. The data format is changed there first and then in code. Nothing is remembered unless it is written down. That's my life as well. :) 2. I split the work by role, not by task. One Claude instance is the strategist: it writes briefs, reviews reports, and never commits code. Claude Code in VS Code is the builder: it executes the brief, writes the tests, reports back, and never releases. Neither agent can approve its own work. 3. I made every job start from a written brief, and I never edit briefs. Each brief is committed to the repo with an exact scope, such as "do X only, do not begin Y." If plans change, a dated amendment states what it supersedes. The builder may skip a risky item with a written reason; skipping is better than forcing it. 4. I used mock-ups with real data instead of descriptions. Anything I will look at gets a mock-up, and I approve it before any code is written. I learned this when a status report described only in words arrived on my phone formatted as a wall of text. The approved mock became the spec, and the tests pin the output to it. 5. I proved that every test could fail. No test is trusted until it has been run against deliberately broken code and caught the bug. A test suite that has never failed hasn't proven anything. 6. I verified reports against the real repo. Agent reports drift. One summary said players rated a puzzle "just right" when the log said "too easy." An audit said there were about 25 unused functions; a recount found 29. The strategist spot-checks load-bearing claims against the actual files and data before anything ships. 7. I built gates the agents can't talk their way past. All work lands on a branch that auto-deploys to a test site. GitHub Actions lints and runs the Python and headless-browser tests on every push. A git hook refuses direct pushes to production. Releases are fast-forward merges that only I perform, after CI is green and I've tested the build personally. The live site shows the release's commit ID, and I check it before calling it shipped. Nothing merges automatically. An agent's "looks good" is just another report, and reports drift. 8. I gave the system a memory. Every session ends by updating the decision log and a "resume here" document, so the next session—a fresh model with no recollection—picks up exactly where the last one stopped. Without this, every session starts by relitigating old decisions. 9. I let the data and end users overrule me. A scheduled GitHub Action turns player telemetry into a daily report. It showed that the difficulty dials I assumed mattered did not predict what players found hard; what mattered was how misleading the starting arrangement was. When testers were confused by the tutorial, I rebuilt it even though I had already approved it. Player feedback about their own experience outranks my earlier rulings. AI agents are velocity multipliers, but they multiply bad structure as quickly as good code. The rules came from me, not the models: the AI writes the code, while the rules, gates, and written record keep the process honest. See the result at Sujido.com. Finding the anagram is the easy part; sliding it into place is the hard part. One board took me 902 moves. My S.O. solved it in 242, and I haven't heard the end of it. Try to break it!
0 comments