Here's the Gist

This runs against most of the AI industry, so I expect some grief for it, but somebody should say it. In March I built Atelier, an assembly line of eight AI agents, each a digital coworker with a name, one job and a human sign-off before the next could start. By October I had deleted nearly all of it, and now I use what's built in, which comes down to one foreman session and four small helpers. Read on.

AI models in use: Opus and Sonnet 4.8 in March; Fable, Opus 5, then 5.5 in October.

Complex workflows cost moreEvery extra agent adds more back-and-forth, more waiting and a bigger bill.
They fail more oftenMore moving parts means more handoffs that fail or get refused.
They didn't give me better resultsI tested the specialists with and without them, and almost nothing changed.

Atelier

Atelier was my factory, and I built it like a human team because that felt right. Plenty of other people were doing the same, and whole companies had formed around giant AI workflows chasing the "dark factory," meaning software built with no humans in the loop.

There was even a character named "Robert" who wrote the product spec, because I had apparently decided the one thing this assembly line lacked was more of me. Since I also signed off on every phase by hand, I was technically middle management, and a single run took hours.

Eva ran the whole show
  1. Spec"Robert"
  2. OK?
  3. DesignSable
  4. OK?
  5. ArchitectureSarah
  6. OK?
  7. BuildColby
  8. OK?
  9. ReviewPoirot, blind
  10. OK?
  11. DocsAgatha
  12. OK?
  13. CommitEllis
A human sign-off, every phasePoirot saw the code, not the conversation behind it
Simplified. Eight characters, each with one job.

20 Minutes To Meltdown

In one planning session I gave the agent a rule, and about 20 minutes later it proposed something that broke it, as cheerful as a golden retriever with a new idea. The AI hadn't made anything up and it hadn't wandered off topic. It had simply forgotten what I told it, the way you walk into the kitchen and forget why you came.

"Most 'the agent got confused' moments are not intelligence failures. They are working-memory failures."

What I wrote down at the time

That one sentence ended up costing me the team: on July 18 I uninstalled Atelier, and the memory loss was my reason.

What I took out

I took Atelier apart in six steps between July and early October, and the project's change history kept the receipts.

WhenWhat wentWhat it looked like
Jul 18RemovedAtelier itself: the manager program and its Eva character, the rulebook, every program that enforced it, the review scripts and 10 of the assembly-line agents.
KeptSeven specialists that work on their own, with shared memory now required for every one.
96 files changed, 15,875 lines deleted
Late AugRemovedRules that only existed on paper. The files said "enforced," and nothing enforced them.104 rules in 14 files became 10 expectations, each naming the program that actually enforces it
Sep 13RemovedTinker, my catalog of ready-made specialists: 19 packages holding 23 characters.
KeptOne final check before work counts as done, four agents, and one short cheat sheet for each kind of technology.
753 files changed, about 95,000 lines deleted
Sep 21RemovedA review "driver" (a program that runs reviews) that agent sessions built in one week. Nobody asked for it, and that includes me.
KeptTwo blind reviews per change, until the new flow of Oct 3 (described below). Serious findings get fixed, and the rest get logged.
29 files, about 5,100 lines, about 400 tests
Sep 25RemovedInstruction text a newer AI no longer needed.The opening instructions every agent reads: 2,848 words down to 275. Reference docs: 23,233 words down to 6,366.
Early OctRemovedThe last of the catalog, which nobody on the live system had ever used. The to-do list moved from Linear to GitHub Issues.1,820 issues moved

The test that ended the specialists

On Sep 13 I ran the test I should have run in March: twelve cases, each done with the specialist's instructions and again without them. In 11 of the 12 the specialist made no difference to the result, and on the one case I timed it only made everything slower.

The test wasn't perfect, and I wrote the flaws down at the time. It covered only four of the 19 packages, my grading had two bugs, and one case changed its answer when I changed the turn limit.

With specialistWithout
Back-and-forth turns7.23.6
Cost20 cents12 cents
Time37 seconds24 seconds

"I need an equivalent, I don't need to reintroduce specialized agents."

Something I said along the way

Capital letters don't enforce anything

I had named those 104 rules "Iron Laws," capital letters and all. The agent read "this is enforced" the way you read a wet paint sign on a dry bench. What really says no is a hook, a small program that runs by itself at a set moment, such as right before code is saved. If a rule is broken, it refuses to continue, and you can think of it as the bouncer at the door. I also wrote a test that fails whenever the instructions claim a block exists and nothing is wired up to do it.

The Foreman

This is the part I'm proudest of. I tell the foreman "here is Milestone 1," which is a batch of related issues (tasks on the to-do list). It reads the batch and starts anywhere from 5 to 15 separate Claude sessions in the cloud, one for each issue. Every session reads its issue and tells me its plan before it does anything, and only then does it get to work.

That's the right human in the middle and the human out of the way. Atelier made me a gate at every phase, but now I read the plans once and the sessions carry on to a pull request without waiting for my go-ahead. (A pull request is a finished piece of work submitted for review.)

Me"Here is Milestone 1."
↓
ForemanReads the milestone and starts one cloud session per issue
issueissueissueissueissueissueissueissue
↓
Human in the middleEach session posts its planI read what every session is about to do
↓
Human out of the waySessions build to a pull requestTests, blind reviews by angle, then the pull request. No sign-off at each phase.
↓
Milestone doneReleased to the test site
One session works one issue. Every issue goes through a pull request.

Four helpers, each for a clean slate

"Asking you to do the pass yourself is like asking me to tell everyone how awesome I am."

Why the author shouldn't review its own work

HelperWhy it's separate
Blind reviewSees the change, not the conversation that produced it, and takes one angle at a time
Security reviewStarts clean, with a narrower job, and goes when the change calls for it
Bug detectiveStarts clean, so it doesn't inherit my first guess about what went wrong
ReadingReads something huge so the main session doesn't have to

How a change gets reviewed now

Review is the part I put back. On Oct 3 I approved a new flow in which five blind reviewers each take one angle, and the security reviewer joins them when the change calls for it, which makes six. They are still the same four kinds of helper, with the reviewer simply sent out several times.

StepWhoWhat happens
1. First reviewFive blind reviewers, one for each angle: bugs, edge cases, contract (does it do what was promised), reuse and cleanup. The security reviewer joins them when needed.Each one saves its findings as Important (could cause real harm) or Nit (minor).
2. VerifyOne more reviewerTries to prove each Important finding and drops any it can't.
3. FixThe session that wrote the codeImportant findings get fixed, test first. Nits go in a log.
4. Fix checkOne reviewer for each angle that found somethingChecks that the fix really fixed it.

Before and after, counted from the project's history

These are file counts from the project's agent folder at three moments in time. The instructions shrank by almost 90%, the hook scripts fell by less than a fifth, and the tests that check the hooks grew from 1 to 23. That is exactly the point: the rules moved out of written instructions and into programs that refuse.

CountAtelier
Jul 17
Catalog peak
Sep 12
After the cuts
Sep 29
Change since Jul 17
Agent files21174down 81%
Skill files11412313down 89%
Reference docs363213down 64%
Words in those docs52,00848,8036,569down 87%
Hook scripts (programs that refuse)425034down 19%
Tests that check the hooks12623from 1 to 23

Since Sep 30 the four agents, six skills and shared hooks live in my own user folder and work in every project, so no project keeps its own copy.

The numbers

Part of the drop in failed calls is simply fewer calls, since a deleted agent can't fail, but that is also the point: those calls weren't adding anything.

MeasureBeforeAfterSource
Failed or refused agent calls (two weeks each side of Sep 13)130 of 576
23%
41 of 542
8%
My records
Cost of handing off a job (950 jobs)Inline in the main session: about 3x
Copy of the session: about 2x
Fresh helper: 1x, the cheapestMy records
Agent teams, usage billed against one sessionAnthropic's docs: about 7x when teammates are set to plan before actingMy four weeks: about the same per job as ordinary helpersAnthropic docs, my records

Use both agent-team numbers or neither, because the 7x on its own is a trap and so is mine.

Work completed per day

Each merged pull request counts as one piece of finished work. I split the timeline at Sep 20. The review driver came out on Sep 21, and on Sep 24 came the flow that starts one session per issue with a single command.

Before Sep 20
Aug 7 to Sep 19
7.4325 pull requests in 44 days
After Sep 20
Sep 21 to Oct 5
20.9313 pull requests in 15 days
Average pull requests merged per day, counted from the project's history. The "after" average includes one 88-PR day on Sep 23, and without it the figure is 16.1 a day. If you compare only the 15 days just before (Sep 5 to 19), the "before" figure is 9.2 a day, so either way the pace is about double. The two windows covered different work, so read this as a direction rather than a multiplier.

Is this just hindsight?

Partly, yes. Everything I kept, the required memory, the blind review and the programs that refuse, was born inside the big assembly line. I couldn't have subtracted my way here in March, because there was nothing yet to subtract from. Structure also goes stale, and how fast depends on how good the AI is that month.

If you're using a weaker AI, or one that runs on your own computer, some of what I threw out may still be right for you. The same goes for teams where the sign-off at each step is the actual product. The direction held, though: friends who copied the setup where every issue gets its own session are still running it, with results very similar to mine. Atelier, meanwhile, stopped growing the day I stopped feeding it.

What I'd tell you if you were starting today

Don't buy or build complex workflows. They'll cost you more and slow you down.

We've seen this before. The tools developers write code in kept getting fancier, and we spent years arguing over whose was best. Meanwhile a few developers just opened vi, a bare-bones text editor from the 1970s, and were coding faster than the rest of us before our splash screens finished loading.

Now we're doing it again with AI workflow products, lining one up against the next and comparing feature lists, and it's the same argument with a new splash screen.

A factory, whether the "dark" kind or any other, will not set you apart from anyone. What does is building something simple that improves how you build your product and the quality of what comes out.

What I'd do instead

  1. Use what Claude Code ships. Anthropic says that as of May 2026, more than 80% of the code merged into its own codebase was written by Claude. Boris Cherny at Anthropic posted in January that he had shipped 22 pull requests one day and 27 the next, each one written by Claude. I think the heavy workflow products sold on top of it do the same job with a lot more parts, and the parts are where the cost and the failures live.
  2. Start with one session, and add a helper only for work that needs a clean slate.
  3. Put your rules in programs that refuse, because written instructions are only a polite request.
  4. Run the with-and-without test on day one instead of waiting until day 180, as I did.
  5. Keep the human at the plan and out of every phase.
  6. Give every agent shared memory. When one gets confused, check what it forgot before you hire a whole new character to do the remembering.

The character named Robert didn't make the cut, which is the most honest performance review I've ever given myself.


Where the numbers come from. Counts, dates and line totals come from the git history of the project. Failed calls, the 12-case test, the 950 jobs and the agent-team comparison come from my own logs and tests. Pull requests per day count the commits on the main branch that carry a pull request number, starting from the first merged pull request on Aug 7.

Outside sources: When AI builds itself, Anthropic Institute and Fortune, January 29, 2026, reporting Boris Cherny's post.