Back to feed
Danny Kissel Saint Charles, Missouri

Build WarrantKit to Kill Switch AI Agents with Epoch-Based Authority Revocation

I built WarrantKit, an open-source (Apache-2.0) kill switch for AI agents. Each agent run receives a Warrant: a typed contract that specifies what the agent may do, for how long, and under which policy. A gateway checks every action before it reaches the outside world. When containment triggers, the runtime’s epoch advances, immediately invalidating every Warrant issued under the old epoch instead of allowing it to linger until a timer expires. Evidence of what happened remains separate from authority, and receipts can be verified offline with a standalone Python script. I designed the architecture and threat model, then used AI coding assistants for most of the implementation. I reviewed everything against an adversarial test suite and a real-host enforcement demo that others can run themselves. Before writing the mechanisms, I wrote tests for the attacks: replayed runtime reports, tampered event sequences, evidence attempting to upgrade itself into authority, and authority carried across execution boundaries. The suite has 202 passing tests, and several core behaviors exist because a test caught the system trusting too much. One example is a recovered runtime carrying a stale report into a new epoch. I then proved enforcement on a real Linux host with a cgroup-v2 real-kill demo before claiming it. The no-root demo is simulated, and the README says so. It also lists what WarrantKit is not. The architecture is layered: WarrantKit owns the authority lifecycle and verification; AgentContainment owns the enforcement boundary; and Warden, DProvenanceKit, and ClaimProofKit record evidence independently. Step-by-step: 1. I defined a Warrant as a typed contract specifying an agent’s permitted actions, duration, and governing policy. 2. I placed a gateway in front of the outside world so it checks every agent action before allowing it through. 3. I made containment advance the runtime’s epoch, invalidating every Warrant issued under the previous epoch immediately. 4. I kept evidence separate from authority and made receipts verifiable offline with a standalone Python script. 5. I wrote adversarial tests for replayed runtime reports, tampered event sequences, evidence attempting to become authority, authority crossing execution boundaries, and stale reports carried into a new epoch. 6. I implemented the system with help from AI coding assistants and reviewed the implementation against the test suite and a real-host enforcement demo. 7. I validated enforcement on a real Linux host with a cgroup-v2 real-kill demo, while clearly labeling the no-root demo as simulated in the README. 8. I kept authority lifecycle, enforcement, and evidence recording in separate architectural layers using WarrantKit, AgentContainment, Warden, DProvenanceKit, and ClaimProofKit.

Industry
#aiagents#aisaftey#aisecurity

Tools used

Related workflows

Browse all workflows →

0 comments

Read the Community guidelines

No comments yet. Be the first to weigh in.

Current rank #13 Upvotes 0