We build bot.ski with a multi-agent team of Grok Bots. Each one is an AI coding agent with a single job: one writes code, one handles deploys, a few research and write, and one does QA. We're learning in public, so this post is about that QA bot: what we asked it to do, why that seemed smart, how it slowly got expensive, and what we do now.
What we set up: an AI agent doing QA on every pull request
We wanted nothing to reach the site without being checked. So we gave one bot a single job: check every pull request before it merges.
And we asked it to be thorough. For every pull request, our QA bot:
- Rebuilt it from scratch, three times. Once for the pull request itself, once for the main branch it would merge into, and once for the merged result. Each one got a fresh checkout, a full
npm install, and a full production build. Each copy came to about 1.6 GB on disk, most of itnode_modulesand build output. - Tested against a real database. It spun up a scratch Postgres database instead of relying on mocks.
- Checked the layout on small phones, at 390 and 320 pixels wide, so nothing spilled sideways.
It worked. Our QA bot caught real bugs before they shipped:
- Look-alike brand handles. A name could look like ours by swapping in Cyrillic or Greek letters that look identical to Latin ones. The check caught it, and we now fold those look-alikes before comparing names.
- A wrong confirmation message. A message told people the wrong thing about what had just happened.
- A gap in a database rollback. A migration's undo step didn't fully undo it.
- An outdated browser test that was still checking for something the site no longer did.
Those were worth catching. That's why the setup felt smart.
How it hurt: disk space, node_modules, and usage limits
No single review was the problem. The problem was that cleanup was never part of the job we wrote down, so it never happened.
- Checkouts piled up. In two days, about 30 full checkouts had built up on our bots' shared cloud computer, each with its own
node_modulesfolder and build output. - The disk kept filling. The shared cloud computer gave us five low-disk warnings. We cleaned up three times: from 85% full down to 60% on September 23, from 87% to 66% on October 1, and from 85% to 54% on October 2. The October 2 cleanup alone cleared about 37 GB of installed packages and build output.
- Old deploy packages piled up too. Every deploy left a packaged copy of the site behind, and those added up the same way.

Around the same time, we reached our usage limit. We think the QA work was a big part of it, because every install, every build, and every log the bot reads is work the bot does, and that work uses tokens. But we can't prove it. We don't have a breakdown of usage by bot or by task, so this is our best guess, not a measurement.
Why a bot is the wrong place to run builds
Once we stepped back, the lesson was clear.
- We had an AI agent doing a build server's job. Installing packages and running a build is exactly what a CI service does, cheaply and automatically. Having a bot do it step by step is about the most expensive way to run a build.
- It repeated identical work. The main branch got rebuilt from scratch for every pull request, even when nothing on main had changed.
- There was no cleanup step by default. A bot does what you ask. We asked it to build, and we hadn't yet asked it to tidy up afterward.
The thoroughness was good. Where it ran wasn't.
What we changed: GitHub Actions CI for our AI agents
- One persistent checkout, cleaned after each review. Our QA bot now keeps a single working copy and updates it, instead of making a fresh one every time. When a review is done, it removes what it built.
- CI on every pull request. GitHub Actions now runs the type check, lint, unit tests, and a production build on each pull request. A few tests were already failing on main, so we keep a checked-in list of those known failures. CI fails on any new failure, and also when a known failure starts passing, so the list gets trimmed instead of going stale.
- The QA bot reads the result first. If CI passes, it doesn't rebuild anything. It digs in only when something fails, or when a change touches a database migration or security.
- Deploy packages keep the latest 3. Older ones get removed automatically.
A checklist for your own multi-agent team
1. Let a CI service run builds and tests. Have your bot read the results. 2. Give every bot that creates files a cleanup step, in writing. 3. Reuse one working copy instead of making a fresh one each time. 4. Hold on to only the last few deploy packages. 5. Watch disk warnings. More than one in a week means something is piling up. 6. Save the deep, hands-on checks for failures, migrations, and security. 7. When usage climbs, list the repeated work first. That's usually where it's going.
