Multi-Agent AI: Why the Creator Needs a Separate Critic
Brilliant Monday, mediocre Thursday — one agent holding writer and editor instincts at once. Split the roles, cap the loop, calibrate the critic.
A Bangkok content agency's AI wrote client newsletters that were brilliant on Monday and mediocre on Thursday — same prompt, same inputs, wildly different polish. The owner's fix was more instructions: 60 rules about tone, length, structure. Output got worse: formulaic, cautious, forgettable. The problem was never the rules. It was that one agent was being asked to be both the writer and its own editor — two jobs with opposite instincts, welded into a single prompt.
The fix is role separation: stop building one over-instructed generalist; build a small team of specialists. 
This is the design pattern that most reliably moves AI work from "sometimes great" to "dependably good" — and it costs less than people fear when managed right.
Reliable multi-agent AI comes from separating roles with conflicting incentives. A single agent told to "write it well and check it's good" faces a conflict of interest — completion instinct beats error-finding, so quality drifts. Instead, split the pipeline: a Creator optimized for tone and narrative, a separate Critic optimized for finding errors and checking constraints, plus upstream specialists (researcher, analyst) each with one objective. Constrain the loop (hard cap on Creator–Critic exchanges), define failure precisely so the Critic neither rubber-stamps nor rejects everything, and accept the modest extra token cost as the price of consistency. :::
The all-in-one trap
One agent, one long instruction list — the default architecture, and the source of three predictable failures:
| Failure mode | What you see | Why it happens |
|---|---|---|
| Prompt bloat | Rules ignored, priorities scrambled | 60 rules = no rules; attention fragments |
| Rigidity | Formulaic, soulless output | Over-instructed = checkbox writing |
| Quality whiplash | Gems Monday, duds Thursday | One agent juggling contexts it can't prioritize |
The cruel part: the instinctive fixes make it worse. Inconsistent? Add rules → bloat. Too rigid? Loosen rules → whiplash. You're tuning a single knob that controls two opposites. The escape is not better tuning — it's more knobs: different agents, different incentives.
Role separation: the team, not the generalist
Decompose by instinct, not just by step (this builds on the blobs-to-blocks pipeline rule — same decomposition, organized around conflicting goals):
RESEARCHER → find relevant data points. Optimized for:
recall, sourcing, no interpretation.
ANALYST → evaluate: which insights actually matter?
Optimized for judgment, ranking.
CREATOR → draft from the vetted insights.
Optimized for tone, narrative, brand voice.
CRITIC → attack the draft. Optimized for errors,
constraint violations, missing facts.
Explicitly NOT allowed to rewrite — only to flag.
Each role gets a short prompt (3–5 rules — the instruction-discipline limit) because each has exactly one objective. No juggling, no bloat.
Why Creator ≠ Critic matters most
"Write a draft and make sure it's good" is a conflict of interest in a single agent: the drive to complete the task outranks the drive to find its flaws (the same goal-momentum that leaks drafts past DO-NOT rules). Split them:
- Creator persona: "Your only job is a compelling draft in the client's voice, using ONLY the provided insights."
- Critic persona: "Your only job is to find problems. List every unsupported claim, constraint violation, and tone mismatch. Do not rewrite."
Creator wants to be read; Critic wants to be right. When the draft passes between them, each pass does its one job well — and the output that survives is dramatically more consistent than anything a single generalist produces.
{
"loop_contract": {
"max_exchanges": 2,
"critic_output": "issue list only, severity-tagged",
"creator_revision": "address issues; do not expand scope",
"escalation": "unresolved high-severity after 2 passes → human review"
}
}
That max_exchanges cap is non-negotiable. Without it, Creator and Critic can loop forever — each "fix" introducing a new flagged issue, tokens burning, nobody converging. Two passes, then escalate. Perfectionism is a failure mode, not a feature.
Managing the trade-offs honestly
Strictness calibration
A Critic too lenient rubber-stamps; too strict rejects good work and the Creator over-corrects into mush. The Goldilocks zone comes from defining failure precisely — "a missing price is a failure; a slightly long sentence is not" — and then iterating with real outputs until the rejection rate sits around a sane 10–20%. Anything higher means your definition of failure, not your drafts, is the problem.
Cost reality
Each extra agent and loop costs tokens — but the accounting surprises people:
- Specialized prompts are short; five small calls often cost near one bloated mega-prompt
- The loop cap bounds worst-case spend
- The real comparison is against the human re-work that inconsistency causes — the agency's Thursday duds took an hour of manual rewriting each
Maintenance
More roles = more prompts to maintain. Mitigate: shared context via the stateful table (each agent reads/writes rows, not chat history), versioned prompts, and an eval suite run on every change so you know which role drifted.
When NOT to multi-agent
Match architecture to stakes, as always:
- Solo quick drafts, internal notes → one agent, light prompt; role separation is ceremony
- Customer-facing output weekly volume → Creator + Critic is the highest-ROI pair; skip researcher/analyst unless data-heavy
- High-volume, multi-source work → full pipeline + evals + state table
Adding agents to a badly decomposed task makes it worse and more expensive — more steps only help when each step's logic is sound.
And when the reliable pipeline is feeding growth, check the demand side with the same rigor: a free geo-grid scan at https://gbppeak.com/free-maps shows where your Google Maps ranking stands across your service area — the map your content team is actually playing on.
Frequently Asked Questions
Why can't one agent check its own work?
Self-review creates a conflict of interest: the drive to complete the task outranks the drive to find flaws in it. Splitting Creator and Critic gives each a single, non-conflicting objective — and the critic persona instructed only to flag, never rewrite, stays aggressive instead of diplomatic.
Doesn't multiple agents cost a lot more?
Less than expected: specialized prompts are short, a hard loop cap bounds worst-case spend, and the comparison that matters is against human rework of inconsistent output. The expensive version is the unbounded feedback loop — cap exchanges at two and escalate.
How do I stop the Critic from rejecting everything?
Define failure precisely and narrowly — concrete, checkable conditions, not vibes — then calibrate against real drafts until rejections sit around 10–20%. A Critic rejecting half of everything is reporting a spec problem, not a quality problem.
How many roles should a small team start with?
Two: Creator and Critic. That pair delivers most of the reliability gain. Add a Researcher when inputs are scattered, and an Analyst when judgment over data matters — only when the simpler pipeline demonstrably fails.
Final note
The generalist agent was never going to be consistent — it has too many masters. Build the small team instead: one to make, one to check, a cap on the argument, and a table keeping the minutes. That's how "sometimes great" becomes the floor instead of the ceiling.