Here's the Gist
This runs against most of the AI industry, so I expect some grief for it, but somebody should say it. In March I built Atelier, an assembly line of eight AI agents, each a digital coworker with a name, one job and a human sign-off before the next could start. By October I had deleted nearly all of it, and now I use what's built in, which comes down to one foreman session and four small helpers. Read on.
AI models in use: Opus and Sonnet 4.8 in March; Fable, Opus 5, then 5.5 in October.
Atelier
Atelier was my factory, and I built it like a human team because that felt right. Plenty of other people were doing the same, and whole companies had formed around giant AI workflows chasing the "dark factory," meaning software built with no humans in the loop.
There was even a character named "Robert" who wrote the product spec, because I had apparently decided the one thing this assembly line lacked was more of me. Since I also signed off on every phase by hand, I was technically middle management, and a single run took hours.
- Spec"Robert"
- OK?
- DesignSable
- OK?
- ArchitectureSarah
- OK?
- BuildColby
- OK?
- ReviewPoirot, blind
- OK?
- DocsAgatha
- OK?
- CommitEllis
20 Minutes To Meltdown
In one planning session I gave the agent a rule, and about 20 minutes later it proposed something that broke it, as cheerful as a golden retriever with a new idea. The AI hadn't made anything up and it hadn't wandered off topic. It had simply forgotten what I told it, the way you walk into the kitchen and forget why you came.
"Most 'the agent got confused' moments are not intelligence failures. They are working-memory failures."
What I wrote down at the time
That one sentence ended up costing me the team: on July 18 I uninstalled Atelier, and the memory loss was my reason.
What I took out
I took Atelier apart in six steps between July and early October, and the project's change history kept the receipts.
| When | What went | What it looked like |
|---|---|---|
| Jul 18 | RemovedAtelier itself: the manager program and its Eva character, the rulebook, every program that enforced it, the review scripts and 10 of the assembly-line agents. KeptSeven specialists that work on their own, with shared memory now required for every one. | 96 files changed, 15,875 lines deleted |
| Late Aug | RemovedRules that only existed on paper. The files said "enforced," and nothing enforced them. | 104 rules in 14 files became 10 expectations, each naming the program that actually enforces it |
| Sep 13 | RemovedTinker, my catalog of ready-made specialists: 19 packages holding 23 characters. KeptOne final check before work counts as done, four agents, and one short cheat sheet for each kind of technology. | 753 files changed, about 95,000 lines deleted |
| Sep 21 | RemovedA review "driver" (a program that runs reviews) that agent sessions built in one week. Nobody asked for it, and that includes me. KeptTwo blind reviews per change, until the new flow of Oct 3 (described below). Serious findings get fixed, and the rest get logged. | 29 files, about 5,100 lines, about 400 tests |
| Sep 25 | RemovedInstruction text a newer AI no longer needed. | The opening instructions every agent reads: 2,848 words down to 275. Reference docs: 23,233 words down to 6,366. |
| Early Oct | RemovedThe last of the catalog, which nobody on the live system had ever used. The to-do list moved from Linear to GitHub Issues. | 1,820 issues moved |
The test that ended the specialists
On Sep 13 I ran the test I should have run in March: twelve cases, each done with the specialist's instructions and again without them. In 11 of the 12 the specialist made no difference to the result, and on the one case I timed it only made everything slower.
The test wasn't perfect, and I wrote the flaws down at the time. It covered only four of the 19 packages, my grading had two bugs, and one case changed its answer when I changed the turn limit.
| With specialist | Without | |
|---|---|---|
| Back-and-forth turns | 7.2 | 3.6 |
| Cost | 20 cents | 12 cents |
| Time | 37 seconds | 24 seconds |
"I need an equivalent, I don't need to reintroduce specialized agents."
Something I said along the way
Capital letters don't enforce anything
I had named those 104 rules "Iron Laws," capital letters and all. The agent read "this is enforced" the way you read a wet paint sign on a dry bench. What really says no is a hook, a small program that runs by itself at a set moment, such as right before code is saved. If a rule is broken, it refuses to continue, and you can think of it as the bouncer at the door. I also wrote a test that fails whenever the instructions claim a block exists and nothing is wired up to do it.
The Foreman
This is the part I'm proudest of. I tell the foreman "here is Milestone 1," which is a batch of related issues (tasks on the to-do list). It reads the batch and starts anywhere from 5 to 15 separate Claude sessions in the cloud, one for each issue. Every session reads its issue and tells me its plan before it does anything, and only then does it get to work.
That's the right human in the middle and the human out of the way. Atelier made me a gate at every phase, but now I read the plans once and the sessions carry on to a pull request without waiting for my go-ahead. (A pull request is a finished piece of work submitted for review.)
Four helpers, each for a clean slate
"Asking you to do the pass yourself is like asking me to tell everyone how awesome I am."
Why the author shouldn't review its own work
| Helper | Why it's separate |
|---|---|
| Blind review | Sees the change, not the conversation that produced it, and takes one angle at a time |
| Security review | Starts clean, with a narrower job, and goes when the change calls for it |
| Bug detective | Starts clean, so it doesn't inherit my first guess about what went wrong |
| Reading | Reads something huge so the main session doesn't have to |
How a change gets reviewed now
Review is the part I put back. On Oct 3 I approved a new flow in which five blind reviewers each take one angle, and the security reviewer joins them when the change calls for it, which makes six. They are still the same four kinds of helper, with the reviewer simply sent out several times.
| Step | Who | What happens |
|---|---|---|
| 1. First review | Five blind reviewers, one for each angle: bugs, edge cases, contract (does it do what was promised), reuse and cleanup. The security reviewer joins them when needed. | Each one saves its findings as Important (could cause real harm) or Nit (minor). |
| 2. Verify | One more reviewer | Tries to prove each Important finding and drops any it can't. |
| 3. Fix | The session that wrote the code | Important findings get fixed, test first. Nits go in a log. |
| 4. Fix check | One reviewer for each angle that found something | Checks that the fix really fixed it. |
Before and after, counted from the project's history
These are file counts from the project's agent folder at three moments in time. The instructions shrank by almost 90%, the hook scripts fell by less than a fifth, and the tests that check the hooks grew from 1 to 23. That is exactly the point: the rules moved out of written instructions and into programs that refuse.
| Count | Atelier Jul 17 | Catalog peak Sep 12 | After the cuts Sep 29 | Change since Jul 17 |
|---|---|---|---|---|
| Agent files | 21 | 17 | 4 | down 81% |
| Skill files | 114 | 123 | 13 | down 89% |
| Reference docs | 36 | 32 | 13 | down 64% |
| Words in those docs | 52,008 | 48,803 | 6,569 | down 87% |
| Hook scripts (programs that refuse) | 42 | 50 | 34 | down 19% |
| Tests that check the hooks | 1 | 26 | 23 | from 1 to 23 |
Since Sep 30 the four agents, six skills and shared hooks live in my own user folder and work in every project, so no project keeps its own copy.
The numbers
Part of the drop in failed calls is simply fewer calls, since a deleted agent can't fail, but that is also the point: those calls weren't adding anything.
| Measure | Before | After | Source |
|---|---|---|---|
| Failed or refused agent calls (two weeks each side of Sep 13) | 130 of 576 23% | 41 of 542 8% | My records |
| Cost of handing off a job (950 jobs) | Inline in the main session: about 3x Copy of the session: about 2x | Fresh helper: 1x, the cheapest | My records |
| Agent teams, usage billed against one session | Anthropic's docs: about 7x when teammates are set to plan before acting | My four weeks: about the same per job as ordinary helpers | Anthropic docs, my records |
Use both agent-team numbers or neither, because the 7x on its own is a trap and so is mine.
Work completed per day
Each merged pull request counts as one piece of finished work. I split the timeline at Sep 20. The review driver came out on Sep 21, and on Sep 24 came the flow that starts one session per issue with a single command.
Aug 7 to Sep 19
Sep 21 to Oct 5
Is this just hindsight?
Partly, yes. Everything I kept, the required memory, the blind review and the programs that refuse, was born inside the big assembly line. I couldn't have subtracted my way here in March, because there was nothing yet to subtract from. Structure also goes stale, and how fast depends on how good the AI is that month.
If you're using a weaker AI, or one that runs on your own computer, some of what I threw out may still be right for you. The same goes for teams where the sign-off at each step is the actual product. The direction held, though: friends who copied the setup where every issue gets its own session are still running it, with results very similar to mine. Atelier, meanwhile, stopped growing the day I stopped feeding it.
What I'd tell you if you were starting today
Don't buy or build complex workflows. They'll cost you more and slow you down.
We've seen this before. The tools developers write code in kept getting fancier, and we spent years arguing over whose was best. Meanwhile a few developers just opened vi, a bare-bones text editor from the 1970s, and were coding faster than the rest of us before our splash screens finished loading.
Now we're doing it again with AI workflow products, lining one up against the next and comparing feature lists, and it's the same argument with a new splash screen.
A factory, whether the "dark" kind or any other, will not set you apart from anyone. What does is building something simple that improves how you build your product and the quality of what comes out.
What I'd do instead
- Use what Claude Code ships. Anthropic says that as of May 2026, more than 80% of the code merged into its own codebase was written by Claude. Boris Cherny at Anthropic posted in January that he had shipped 22 pull requests one day and 27 the next, each one written by Claude. I think the heavy workflow products sold on top of it do the same job with a lot more parts, and the parts are where the cost and the failures live.
- Start with one session, and add a helper only for work that needs a clean slate.
- Put your rules in programs that refuse, because written instructions are only a polite request.
- Run the with-and-without test on day one instead of waiting until day 180, as I did.
- Keep the human at the plan and out of every phase.
- Give every agent shared memory. When one gets confused, check what it forgot before you hire a whole new character to do the remembering.
The character named Robert didn't make the cut, which is the most honest performance review I've ever given myself.
Where the numbers come from. Counts, dates and line totals come from the git history of the project. Failed calls, the 12-case test, the 950 jobs and the agent-team comparison come from my own logs and tests. Pull requests per day count the commits on the main branch that carry a pull request number, starting from the first merged pull request on Aug 7.
Outside sources: When AI builds itself, Anthropic Institute and Fortune, January 29, 2026, reporting Boris Cherny's post.