Moses had a bottleneck problem.
Exodus 18. The man is sitting from morning until evening judging every dispute in Israel by himself, and the line of people waiting on him never gets shorter. His father-in-law Jethro watches one full day of this and doesn't sugarcoat it.
"What you are doing is not good. You and the people with you will certainly wear yourselves out, for the thing is too heavy for you. You are not able to do it alone."
Catch what Jethro actually says there. It isn't just Moses who burns out. It's Moses AND everybody stuck in his line. The bottleneck breaks the man at the top and stalls everything downstream of him at the same time.
Then he hands Moses the fix. Appoint captains over thousands, over hundreds, over fifties, over tens. Let them handle the small cases. Only the hard ones come to you.
That's an org chart. And an org chart turns out to be the exact architecture that keeps AI agents from choking.
I've been running this pattern for months without teaching it start to finish. The Content Spy in Build Notes #017 was three separate workflows with a boss workflow sitting on top of them. The Dream 100 agent in #021 pulled the same trick. I showed you the machines and skipped the blueprint.
So here it is.
All of August, Build Notes is building ONE thing: a complete paid ads department that a single person can run. Every Thursday adds another piece.
This issue is the org chart that whole department sits on, so read it even if you never touch an ad account, because the pattern works on any job you're currently drowning in.
One Employee With Twenty Job Descriptions
Here's what almost everybody does when they start building agents.
They make one agent. Then they bolt another tool onto it. Then another. Then they paste in a fourth set of instructions, and a fifth, and pretty soon this poor thing has twenty tools, six jobs, and a set of directions longer than a mortgage.
Literally, right now there are people with Hermes and OpenClaw agents that are built like this.
And it gets DUMBER, not smarter.
Anthropic (Claude) shows us why...
Their engineering team found that an agent with too many tools, "often 10+," starts struggling to pick the right one for the job. Think about hiring one guy and handing him 10 job descriptions on his first day.
He's not going to be 10x more useful. He's going to have a breakdown, not if... WHEN!
The second failure is quieter, and it's the one that costs you money.
Every AI has a context window. That's the amount of stuff it can hold in its head at once, kind of like the desk space in front of a worker. Everything the agent reads, every tool result it gets back, every half-useful web page it opened by accident, all of it piles up on that desk.
And once the desk is buried, quality drops.
The agent starts forgetting what you asked it in the first place.
One agent doing six jobs drowns itself in context and gets stupider as the day goes on.
Clean Desks Beat More Brains
The fix is Jethro's fix.
You stop building one agent and start building a small department. One orchestrator (fancy word for the manager, the agent that takes the job, decides who does what, and merges the results) plus a handful of subagents (the interns, each one hired for exactly one task).
When we were hiring humans, it was expensive to add 10 more. With agents it's almost free to hire 10 more of them, and they perform better when each has a smaller job description.
Manager agent (orchestrator) takes the assignment and cuts it into pieces, one piece per intern (subagent). Then he waits. Then he takes what comes back and builds the final output.
Every subagent gets its own fresh context window. Its own desk. The intern goes off, reads forty pages of garbage, finds the three sentences that matter, and hands the manager a summary. The manager never sees the forty pages. His desk stays clear the whole time.
This is what makes pro agent builders, pro. It's the art of agentic architecture and it will be a multi-six figure role soon. THIS is the new role for hiring humans.
Humans build agents. Agents do the work.
Anthropic names the three benefits as context protection, parallelization, and specialization, which is a very polite way of saying: clean desks, they work at the same time, and each one only has to be good at one thing.
Parallelization is the part you FEEL.
Anthropic clocked the combination of running 3 to 5 subagents at once plus letting each of them call 3+ tools simultaneously, and it cut research time by up to 90% on complex queries.
90% that's not nuthin!
Five interns reading five sources at the same time beats one guy reading five sources one after the other.
(Hold onto that. It's about to matter a lot.)
The Number Everybody Quotes
Anthropic ran their multi-agent research system against a single agent on their own internal evaluation. Opus 4 as the manager with Sonnet 4 interns beat solo Opus 4 by 90.2%.
That's the stat going around in every AI newsletter this year.
I'm going to be straight with you about it: that's Anthropic testing Anthropic's product on Anthropic's own test. It's real, I believe it, but I still wouldn't build my business on one company's internal scorecard.
So let's go find a study that tested everybody's models instead of just its own.
The Study That Ruins The Party
MIT and Google published a paper called Towards a Science of Scaling Agent Systems. They ran 260 configurations across six benchmarks, five architectures, and three different model families (OpenAI, Google, and Anthropic all got run through the same grinder).
That's the kind of unglamorous work almost nobody does before writing a Twitter thread about agent swarms.
Four findings came out of it.
Two of them will save you real money.
One. On work that comes apart into independent pieces, multi-agent won big. Structured financial reasoning improved 80.8% when they routed it through a central coordinator.
Two. On work where step two needs step one's answer, multi-agent LOST. Every single multi-agent setup they tried came out worse than one agent working alone on sequential planning, somewhere between 39% and 70% worse, with the ugliest number belonging to agents running independently with nobody coordinating them.
Seventy percent worse. You think the guy selling you a 50-agent swarm ran that number?? Or that he'd tell you if he had???
That finding is the single most expensive thing in this article, because it's the exact opposite of what's being sold to you, and it's the reason I'm writing this instead of just showing you another cool build.
Three and four go together. Mistakes multiply when nobody's in charge: independent agents amplified errors 17.2x compared to a single agent, and routing those same agents through an orchestrator dropped it to 4.4x (the in-between setups they tested, decentralized and hybrid, landed at 7.8x and 5.1x, so it's a gradient, not a light switch).
And more agents stops helping fast.
They found the sweet spot at 3 to 4 agents, and once a single agent was already clearing 45% accuracy on a task, adding agents produced NEGATIVE returns.
Three to four. Not 10, 20, 50. Every "I built a swarm of 50 agents" post you scrolled past this year is a guy who never measured. He's just brute forcing and it will cost him.
Then There's The Bill
Agents burn tokens (tokens are how AI usage gets billed, roughly a chunk of a word, and every word in and out costs you). A single agent uses about 4x the tokens of a normal chat. A multi-agent system uses about 15x.
Anthropic's follow-up guidance says a multi-agent system typically runs 3x to 10x the tokens of a single agent doing the same job. Those two stats agree with each other, by the way, they're just measured against different baselines. One is versus chat, the other is versus one agent.
So before you hire a team of interns, do the math a business owner does. A $12 task becomes a $120 task. If the answer is worth $500, hire the team and smile. If you're generating social captions, one agent and a good prompt is the right call and the swarm is just an expensive hobby.
Know when you need multi-agent teams and when 1 single agent is fine.
The Guys Who Changed Their Minds
My favorite receipt in this whole article is a reversal.
In June 2025, Cognition (the team behind Devin) published a piece titled "Don't Build Multi-Agents." Blunt title, blunt argument.
Their example was perfect: the job was to build a Flappy Bird clone, one subagent misread its slice of the work and built a Super Mario style background, the other built the bird, and the two never compared notes. Mario's castle behind a flying bird.
Then in April 2026 they came back and published "Multi-Agents: What's Actually Working."
They refined it. And they landed on this:
Keep the writing single-threaded. Parallelize the reading.
Multiple agents reading, researching, checking, and analyzing at the same time? Great. Multiple agents WRITING to the same thing at the same time? That's how you get Mario's castle behind a flying bird.
LangChain (another fancy pants AI tech company) says the same thing in fewer words: read actions are inherently more parallelizable than write actions.
And OpenAI, in their own guide for building agents, tells you to maximize a single agent's capabilities FIRST before you split anything. The companies selling you the tools are more cautious about this than the influencers selling you the courses.
Funny how that works...
I trust a team that changed its mind with receipts a whole lot more than a guy who's never had to.
The Split Test (Not the A/B kind)
So here's the rule I use now, and I want you to steal it.
Before you hire a single intern, ask one question: does this job come apart?
If the pieces are independent (research five competitors, check four platforms, grade eight landing pages, summarize ten transcripts), it comes apart. Build a team. That's the +80.8% side.
If step two needs step one's answer (write the outline, then the draft, then the edit, then the format), it does NOT come apart. Keep one agent. That's the 39-to-70-percent-worse side, and splitting it makes your output worse AND costs you up to 10x more to get there.
I call it The Split Test, and it takes about four seconds to run. Four seconds has saved me from building at least three agent departments I had no business building.
Look at my Content Spy from #017 through that lens. One YouTube workflow watching channels, one X workflow watching profiles, both running at the same time, neither needing the other's answer, both reporting up to a boss workflow that wrote the final report.
Textbook "comes apart-ness". That's why it worked.
How To Actually Build One
Inside MindStudio use the Run Workflow block, which lets one workflow call another workflow like an employee calling in a specialist.
You pass information IN through Launch Variables (the fields you fill at the top of the workflow). The subagent does its job. Then it hands results back OUT through the Terminator Block, under Return Data, as JSON Output.
JSON is just a tidy way of labeling data so another machine can read it without guessing, like a form with named boxes instead of a paragraph. You don't need to know JSON, AI does, you just need to know that your AI likes it... ALOT!
In my builds the child workflow starts with a clean slate and only knows what I handed it through those Launch Variables, which is exactly the isolation you want. Set Execution Mode to Parallel and you get both tasks being done at the same time. That's your manager and your interns.
With Claude and Codex it's as simple as telling it you want to split the work between subagents, which models, which tools, etc.
With Hermes, I use their "profiles" feature which allows you to split one Hermes install into multiple specialist trained personas. Then you can have the main persona (in my case Jeffri) task all the profiles like they were the interns.
The Five Rules
- Give every intern an objective, an output format, tool guidance, and hard boundaries. A vague boundary is how you end up with two interns doing the same task and nobody doing the third. Anthropic hit exactly this: subagents misread the assignment and ran the identical searches as each other.
- Single writer. Many can read. ONE writes. No exceptions until you've been burned enough to know when to break it.
- Everything routes through the manager. Interns don't chat with each other unsupervised. 4.4x beats 17.2x.
- Cap the team at 3 or 4. If you think you need twelve, you've probably got a sequential job wearing a decomposable costume.
- Run The Split Test first. Every time.
Your Build Checklist
- Pick one job your current agent is doing
- Run The Split Test on it (does it come apart, or does step two need step one?)
- If it's sequential, stop here and go fix your prompt instead
- If it comes apart, list the independent pieces (aim for 3 to 4, not 12)
- Write one intern per piece: one objective, one output format, only the tools it needs
- Assign a cheap model to the grunt interns and your best model to the manager
- Decide which single agent is allowed to WRITE, and make sure it's only one
- Route every result back through the manager, never intern to intern
- Run it once and compare against your old single agent on the same job
- Check the token bill before you scale it up
Sources
- Anthropic, "How we built our multi-agent research system" (Jun 13, 2025): https://www.anthropic.com/engineering/multi-agent-research-system
- Anthropic, "Building multi-agent systems: when and how to use them" (Jan 23, 2026): https://claude.com/blog/building-multi-agent-systems-when-and-how-to-use-them
- Kim et al. (MIT + Google), "Towards a Science of Scaling Agent Systems," arXiv:2512.08296: https://arxiv.org/abs/2512.08296
- Google Research summary (Jan 28, 2026): https://research.google/blog/towards-a-science-of-scaling-agent-systems-when-and-why-agent-systems-work/
- Cognition, "Don't Build Multi-Agents" (Jun 12, 2025): https://cognition.com/blog/dont-build-multi-agents
- Cognition, "Multi-Agents: What's Actually Working" (Apr 22, 2026): https://cognition.com/blog/multi-agents-working
- LangChain, "How and when to build multi-agent systems" (Jun 16, 2025): https://www.langchain.com/blog/how-and-when-to-build-multi-agent-systems
- OpenAI, "A Practical Guide to Building Agents" (Apr 2025): https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf
- MindStudio Run Workflow block: https://university.mindstudio.ai/building-ai-agents/blocks-reference/run-workflow-block
- Exodus 18:17-18, 21-22 (ESV)
Continue with Build Notes+
This Monday deep dive is free in full. For complete Thursday builds and supporting files, join Build Notes+.
