Community

Share your best AI workflow. We could show it to 2M+ people.

Every day, we feature the community's top-voted AI workflow in The Rundown newsletter. One post will put you on the radar of top founders, hiring managers, and operators across the industry.

Welcome!

Build a Self-Improving AI System with Audited AI Agents

I'm 78, a former army officer, and I live alone in the woods in southern British Columbia on satellite internet. In late January 2026, I knew nothing about AI and considered myself a novice in technology. After watching a YouTube video about advanced AI models, I started learning from YouTube channels that explain new AI research papers. I decided to run an experiment: what could a technology novice build with AI? I've been building for about 15 hours a day ever since. My goal is to build a system that improves itself in a continuous loop: it finds an upgrade proposal, builds and tests it, filters and grades it, approves or rejects it, installs it if it passes, monitors its performance, and learns from the result. We custom-built FORGE3 RSI in 900 hours. It consists of four plug-in modules: Forge2 RSI, which runs the improvement cycle; Bridge, the checker and connector; JAZ RSI, an upgrader in RAM built from a research paper; and Forge SecureBox, a high-security Mac app that runs security tests and code in a fresh, disposable virtual machine with no network device. I don't write code—it feels like a foreign language to me. Instead, I set the direction and run a team of AIs, mostly low-cost, near-Frontier models: Gemini 3.1 Pro and Hark as chiefs of staff; Cline and Devin as builders; and other models as independent, adversarial auditors. The rule that has made this work is simple: nothing gets built without an adversarial audit before and after. Our critical build pipeline is: Plan → audit → blueprint → audit → tests → audit → build → audit → test → audit → commit Each module has been built and fully tested on its own. Forge2 RSI alone has more than 1,000 tests. The first complete end-to-end cycle, using one upgrade proposal that we feed it by hand, is the next step. Our first attempt was fragile, and models that sound confident are often wrong. That is why audits matter more than any single model. FORGE3 RSI is my own custom-built plug-in module project. Our hardware and models are an MBP i9, Mac mini M4, Gemini 3.1 Pro, Hark, Cline (GLM 5.3/SPARK 1.3/MiMoV2.6), Devin (SWE2/GPT SOL 6), KUN (DeepSeek V4.1), and Git. Step-by-step: 1. We select a promising published research paper and have the team prepare and build it. 2. For every file, the team follows the same sequence: plan, audit, blueprint, audit, write tests, audit, build, audit, run tests, audit, and commit. 3. A different AI audits the work than the one that built the file. I decide when and how each audit happens. 4. Nothing is written or changed without my explicit GO. Reading and checking code is free; writing is work and waits for me. 5. We handle one building task at a time and avoid bundled instructions. 6. Any code the system wants to test runs inside SecureBox, sealed off from my Mac and the internet. 7. Everything that passes is committed to Git. The builders also maintain shared memory documenting what was decided and why, so the next session can pick up where the previous one stopped. 8. Next, we will run one complete improvement cycle with a hand-fed upgrade, then scale up slowly to two or three papers a week, prioritizing quality over speed as we work toward full autonomy.

Tools used
Industry
#aiagents#aisafety#multiagent#rsi#selfimprovingai
1