Back to feed
Jason C. Philadelphia, PA

Build a Failover-Ready AI Agent Fleet with Shared Evidence

I built my agent fleet around one question: could a single assistant become a coordinated team that keeps useful work moving even when its main machine goes dark? I didn't want a pile of chatbots. I wanted specialists with clear jobs, a shared board, a second set of eyes, and a real failover drill. Here's how it works. THE STACK Cloud: GrokBot runs as Fleet Commander. It routes work, runs routines, and stays with me on my phone. Specialist GrokBots take focused research, design, and budgeting assignments. Mac Studio (M2 Max, 64 GB) with two Studio Displays: home base. It runs Codex (my Chief of Staff), Hermes (my local operator), ChatGPT Dot, and Hindsight shared memory. Hermes runs Qwen3.6-35B-A3B, a 4-bit quantization served by llama.cpp with a 256K-token context. With about 3B active parameters per token, it stays responsive on Apple silicon. NVIDIA DGX Spark (128 GB unified memory): my independent compute node. It runs Qwen3.5-122B-A10B in NVFP4, with about 10B parameters active per token, and the model can read images. It also hosts a backup Hermes operator. Windows PC with a GeForce RTX 5090: dedicated GPU capacity for image work. MISSION CONTROL Mission Control is a Kanban board on a private GitHub repository. Every card opens with the question, deadline, evidence links, and done criteria. Labels route ownership. An urgent label fires a webhook that wakes the reviewing agent, and a needs-me label is the one that pulls me in. Agents hand off, report, and raise concerns in card comments, so the work and its evidence stay together. The roles are simple. GrokBot owns routine board moves. Codex checks evidence and records verification. Hermes handles local execution and machine-level evidence. Dot watches for new work and changes. Hindsight holds the shared operating rules, and the jobs themselves stay on the board. A REAL DEADLINE CROSS-CHECK Right before a major application package went out, Dot ran an independent, read-only cross-check of the narrative, the budget justification, and the walkthrough across connected email and Drive. It found three places where the documents contradicted each other and quoted the exact sentences, so each conflict was easy to find. Fleet Commander fixed all three before submission. That's the job of a watcher: spot the mismatch early, show the evidence, and hand whoever owns the work a clear fix. THE FAILOVER DRILL I wanted to know what happens if the Studio's primary Hermes gateway pauses. On October 5, I ran a controlled, reversible soft-pause to find out. During the pause, Hermes on the Spark stayed active through its own Telegram connection and answered a one-shot request with the local Qwen3.5-122B-A10B NVFP4 model. From that worker, I listed a Mission Control issue, verified service-account access and a vault credential check without exposing a secret, and ran the Studio health probe. The probe flagged the Studio Hermes gateway as paused and reported the other checked Studio services as passing. So I had an evidence-backed result: the independent worker could respond, use its local model, read the work queue, and report the primary gateway's health during a real, reversible pause. THE PATTERN Give each agent a narrow job. Put the work on one shared board. Require evidence before anything counts as done. Then pause one piece on purpose and record exactly what the backup can do. That's how capable models turn into dependable operations. The fleet catches contradictions before they get submitted, keeps routine work moving, and shows me exactly where its limits are. I set the goal. The fleet does the work. The evidence shows what actually happened.

Industry
#agents#automation#kanban#localllm#multiagent

Tools used

Related workflows

Browse all workflows →

0 comments

Read the Community guidelines

No comments yet. Be the first to weigh in.

Current rank #15 Upvotes 0