Nine stages, two workers, one topic name

The factory that writes code for me is 2,274 lines of Python. Nine stages: draft a spec, plan it, implement it, prove it ran, review the diff, commit, push, open a pull request, then watch that pull request and act on whatever the human does to it. Nine stages, and the two files that hold all of them come to less than a small Rails controller’s worth of real logic each.

That number is the whole point of the experiment. Every one of those stages needs a durable queue behind it, because an agent turn can take twenty minutes and a laptop can close. It needs an HTTP surface, because I submit specs from a browser. It needs a scheduler, because the pull request gate is a poll. It needs traces, because when a turn goes quiet for ten minutes the first question is always “is it working or is it dead.” I wrote none of that. I wrote the loop and the turn, and I let iii be everything underneath.

I’d built this before. The earlier version is a hand-rolled Python service that owns its own queue and its own state handling. It works fine. It also meant every new stage cost me plumbing before it cost me thinking. So the question I actually wanted answered was narrow: can a dark factory be composed rather than built? One afternoon of mostly agent-driven work later, the answer is yes, with a couple of scars I’ll get to.

What the thing does

A spec goes in. A pull request comes out. Nothing merges itself.

The spec is a markdown file, written by hand or drafted by a read-only agent turn against the target repo, so the first version names real files instead of plausible ones. Submitting it creates a git worktree off the target repo’s default branch. A planning turn reads the code on a thinking model and writes an approach. An implementation turn carries that plan out on a cheaper model, which is where the tokens go. Then the factory tries to commit.

That commit is the interesting part. The target repo’s own pre-commit hook runs, and if it refuses, what it printed goes straight back to the agent as the brief for another turn. Verbatim, not summarized. The gate’s own words are the most useful thing anyone has said about the work, and paraphrasing them throws that away.

The loop is bounded at two revisions by default. It also exits early if the refusal comes back byte-for-byte identical, because a gate repeating itself has already proved its complaint isn’t about the diff.

After the commit lands: a prove turn runs the software against the spec’s acceptance criteria and pastes the commands it ran, a review turn reads the diff against the spec and stamps a verdict on the pull request, and then everything stops. A human merges, comments, or closes. A cron tick reads the pull request every minute and turns a merge into teardown, a comment into another turn on the same branch, a close into a close.

The factory never merges. In a dark factory a check reports; only the merge accepts.

Two workers and the seam between them

iii organizes everything into three primitives: a worker hosts work, a trigger causes it, a function does it. So the design question became “how few workers can this be,” and the answer was two.

Two topic names are the entire contract between the loop and the model; a reviewer’s comment re-enters at the same place a first draft does.

factory-worker owns the loop: the HTTP endpoints, the worktrees, git, the pull request, the job records. It never talks to a model. claude-harness owns the turn: one call into the Claude Agent SDK, pointed at the worktree. It holds no state and never touches git.

Every arrow between those two boxes is a durable queue message, and so is every stage inside the loop. A crash between two stages resumes instead of restarting, and each handler guards on the job’s own stage, so at-least-once delivery does the work once. I get that from declaring a subscriber trigger on a topic. The retries, the dead-letter queue, and the persistence to disk are the queue worker’s problem, and the queue worker is a line in a config file.

The seam between the two custom workers is a topic name. A job carries harness: "claude", and the factory publishes to claude.run. Nothing else in the factory knows what a model is. A second harness (one that owns its own turn loop instead of calling a stock SDK) is a worker that consumes .run and publishes factory.pr when it’s done. Everything a turn needs rides in the message, so a harness has no store to share and nothing to migrate.

That seam is the reason I built this as a proof of concept rather than the real thing. Phase two is a custom harness, and the point of phase one was to make phase two a topic name rather than a rewrite. That part is done and untested by use, which is an honest thing to say about a seam.

What I didn’t write

Registering a function and giving it a trigger is four lines. In iii’s own shape, minus my names:

worker.register_function("tick", handler)
worker.register_trigger(
{"type": "cron", "function_id": "tick",
"config": {"expression": "0 * * * * *"}}
)p

Swap cron for http and it’s an endpoint. Swap it for durable:subscriber and it’s a queue consumer. Three trigger types cover the entire factory: ten HTTP routes, eight durable topics, one cron. Same registration call every time, so adding a stage is writing the handler and naming the trigger.

The engine’s config file lists eighteen workers, and I use maybe six of them directly. The queue gives me durability. The HTTP worker serves the dashboard and the two JSON endpoints the browser talks to. Cron runs the gate poll. The console (iii’s own, on its own port) shows every function and trigger in the system, queue depth, the dead-letter topics, and a trace waterfall for a turn. I did not build an admin UI, and I have not once wanted one, because the console is the layer under my dashboard: mine shows a change moving, iii’s shows the machinery moving it.

The comparison that matters isn’t lines of code, it’s the shape of a new feature. Adding the review stage to the hand-rolled version would have meant a new state, a new persistence path, and a new way to run something in the background. On iii it was a handler, a topic name, and a verdict parser. Prove was the same shape. So was the improve lane. So was the plan/execute split, which was the largest behavioral change in the project and still didn’t touch the loop’s structure — plan is just another topic the harness consumes, on a different model.

The one stock piece I took back out

The stock state worker held my job records at first. Then it restarted mid-run and took a live job with it.

That’s not really a bug report. It’s a tradeoff I’d chosen without noticing: a store whose disappearance loses work is the wrong store for a factory that runs unattended. So job records became JSON files under state/, written by the factory, one per job. Boring, and cat state/jobs/.json turned out to be the cheapest observability in the project.

I want to be fair about what that episode shows, because it isn’t “the state worker is bad.” Composition made the wrong choice cheap to make and cheap to reverse. The whole change was a small commit in one worker. If the job store had been welded into the same process as the queue and the HTTP surface, the way it was in my hand-rolled version, that swap would have been its own project.

When the factory turned on itself

The stage I’m least able to be objective about is the improve lane. One endpoint hands a turn everything that went wrong recently and asks what would have prevented it. Every proposal names a lane: the target repo’s charter, the factory loop, or the harness.

Run against a job where the target repo’s hook refused every commit, it produced four proposals. Two of them indicted my own code by file and line.

factory — Stop revising when the gate repeats itself word for word. Revisions 1 and 2 got byte-identical refusals. It spent a second turn to learn what the first one already proved: the gate’s complaint is not a function of the diff.

Nothing is applied. Accepting a proposal writes a spec into specs/ and stops there. To become real it goes through the same pipeline as any other work, gated by the same pull request, because the improve lane may not edit the charter, the factory, or the harness on its own authority. That’s the one rule keeping it from being the single thing here that escapes the factory’s own gate.

So I accepted that proposal, sent the spec, and the factory opened a pull request against itself. I merged it. The early-exit on identical refusals that I described four sections ago is that pull request. It’s in the git history as a merge commit, and I did not write the code.

Feedback on iii, having actually shipped something on it

The good parts are the ones I’ve already described, so I’ll be specific about the friction instead. This is a 0.22 release. None of what follows surprised me for a pre-1.0 tool, and all of it cost me real minutes.

Configuration moves out from under you. My config.yaml still has a queue: block, and it is no longer read. The value now lives in a configuration worker, and the file has a comment saying so. That’s a reasonable design, and it means the file I edited and the file that’s in force are different objects. Expect to check iii config rather than trusting the file.

The console rewrites your config on boot. I run the console on 3123 rather than the stock 3113, because a second iii project already had 3113. The engine consumes that port as a startup seed and comments it out of config.yaml on every boot, so I have a nine-line script that puts it back before the engine starts, and a Makefile target that re-seeds before restarting the console. When the worker manager restarts the console on its own, it starts from the stripped config, comes back on 3113, and silently loses to whatever is already there. Silent, and the failure mode is a console that looks fine and isn’t yours.

Two projects on one machine means moving three ports. Factory, console, and worker manager all had to come off their stock numbers. Worth knowing on day one rather than day two.

Managed microVMs are the default, and my workers can’t use them. Both of my workers need git, gh, and the host’s Claude Code login, so they self-connect as local host processes. That’s supported and it works, but the tutorials aren’t written for it, and “which of these two hosting models am I in” was the single thing I was most confused about early.

None of that is a reason not to use it. I’d make the same call again, and the honest summary is that iii cost me a handful of pre-1.0 papercuts and saved me the entire distributed-systems half of the project.

If you’re trying this

Pick the seam before you pick the workers. The one design decision that paid for itself was deciding that a harness is a topic name. Everything else, down to which stock workers and how many stages, turned out to be cheap to change. That one wasn’t. It’s the only thing I’d have had to rewrite if I’d got it wrong.

Put everything a handler needs in the message. My harness holds no state at all, and that’s why a second harness needs no migration, no shared store, and no coordination with the first. Shared state between workers is the thing that turns composition back into integration.

Use the console before you build a dashboard. I built mine because I wanted a stage rail and a live log for a non-technical view of a job. Every debugging session I’ve actually had went through iii’s console instead.

Don’t assume a stock worker is the right worker. The state worker was the right shape and the wrong durability. Composition means swapping it costs a commit, so the mistake is cheap but only if you notice you made it, which took an unattended job dying to teach me.

Bound your loops in the spec, not in the prompt. Two revisions, then stop. Early exit on a repeated refusal. An agent that can’t satisfy a gate twice won’t satisfy it on the ninth try, and finding out costs a turn each time.

Phase two is the custom harness: my own turn loop, my own context handling, behind the same topic name. The factory shouldn’t notice. That’s the claim this proof of concept exists to test, and I’ll write up whether it survived contact.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论