From Blobs to Blocks: Building AI Workflows That Run the Same Every Time
Same prompt, same data, three different results. Break the blob into single-purpose blocks — one verb per step, JSON contracts, checks per station.
A Bangkok fitness studio owner built a "magic" AI workflow: one prompt that read the week's member messages, spotted at-risk members, drafted win-back offers, and formatted the Monday report. In testing it was brilliant. In production it was chaos — Monday's run flagged three members, Tuesday's version invented a "bring a friend free" policy, and Thursday's output skipped the offers entirely. Same prompt, same data, three different results.
That's not bad luck; it's architecture. The workflow was a blob — one AI call juggling four mental jobs at once. The fix is decomposition: break the blob into blocks, single-purpose steps with checkable outputs. 
This guide shows the pattern, the trade-offs, and when you actually want an agent's improvisation instead.
Reliable AI workflows come from decomposing big prompts into single-purpose blocks. A blob asks one AI call to search, extract, decide, and write simultaneously — performance on each part degrades and results vary run to run. A block does one bounded act (extract quotes / classify each / summarize the labels) with a defined input and structured output, so each step is verifiable, re-runnable, and reusable. You trade some setup time and token cost for consistency — and for business automations, predictability beats improvisation every time. :::
Why blobs fail
"Read these messages, find the unhappy customers, figure out why, and write each one a win-back email." That's four reasoning modes in one pass: filtering, sentiment judgment, causal analysis, and persuasive writing. While composing the email, the model is still holding the filtering rules in attention — and dropping them. Multi-objective blobs produce the classic failure bouquet:
- Missed items (filtering degraded while writing)
- Invented policies (writing mode outran the facts — see our hallucination defenses guide)
- Run-to-run variance (each run balances the objectives differently)
| Blob (one mega-prompt) | Blocks (small pipeline) | |
|---|---|---|
| Consistency across runs | Low — varies every time | High — same steps, same order |
| Error rate | Compounds silently | Isolated per step |
| When it breaks | Rerun everything | Rerun one block |
| Verification | "Looks about right" | Automated checks per step |
| Reuse | None — welded to one task | Blocks reused across workflows |
| Cost | One call | More calls, smaller each |
The assembly-line principle
Think of each block as a station on an industrial line: the robot arm may perform complex movements, but it does exactly one job, every time, with a fixture that holds the part in place. Your workflow is the conveyor. The intelligence is in the segmentation.
A good block has:
- Bounded scope — one verb: extract, classify, summarize, compare
- A contract — defined input, structured output (JSON), so the next block knows exactly what it receives
- A check — validation the output must pass before moving on
- Re-run safety — failure means re-running this block, not the line
The pattern in practice
Case 1: The member-retention workflow (fixing our studio owner)
BLOB: "read messages, find unhappy members, draft offers" ❌
BLOCKS:
1. EXTRACT → every message quote expressing a sentiment
(output: list of verbatim quotes + member id)
2. CLASSIFY → each quote → risk: high|medium|none + reason tag
(output: JSON row per member)
3. DRAFT → for high-risk rows only, write a win-back message
using the ACTUAL quote and the real offer sheet
(output: draft per member, human approves)
4. REPORT → counts + trends table for the owner
Every Monday, identical output shape. Step 2 drifted? Re-run step 2 for the affected rows only. New offer policy? Update the sheet step 3 reads — nothing else changes.
Case 2: Research-then-write
The blob version ("search the web and write a sales email") hallucinates length. The block version separates concerns: search-and-filter block → insight-extraction block → writing block that only ever sees pre-vetted notes. The writer can't invent a fact it was never given, because it never sees the open web — only the extracted notes.
{
"block": "classify",
"input": "one verbatim quote",
"output_schema": {
"member_id": "string",
"risk": "high | medium | none",
"reason_tag": "price | results | schedule | trainer | other",
"evidence": "the quote itself"
},
"check": "reason_tag must exist in taxonomy; evidence must be verbatim"
}
That check line is what turns an AI call into infrastructure. Invalid output → rejected → retried or flagged. Nothing broken flows downstream.
Implementation principles
- One reasoning act per step — if your prompt contains "and then," it's probably two blocks
- Define contracts — structured outputs (JSON with a schema), not prose between steps
- Ensure re-run safety — persist each block's output so failures resume, not restart
- Build for composition — the classify block above works identically on reviews, tickets, and DMs
This is also the backbone that makes eval suites practical: you can only test a pipeline whose steps have checkable outputs.
The honest trade-offs
- Setup time — an afternoon instead of ten minutes; the blob was faster to write and slower to trust
- Token cost — more calls, but each is small; total cost is typically similar, since blobs re-read everything per attempt
- Reduced improvisation — a fixed path won't invent a clever move on weird inputs; it will flag them
That last point is a feature in business contexts. Most owners would happily trade a "creative" agent that works 60% of the time for a boring pipeline that works 99%.
When you DO want an agent
Autonomous, self-directing agents still earn their place:
- Human-in-the-loop sessions — a monitored conversation has a safety net; a scheduled workflow doesn't
- Open-ended exploration — brainstorming, research rabbit holes, no "right" output
- Prototyping — use an agent to discover the steps; then freeze the discovered path into blocks
The mature pattern: agents to explore, blocks to operate. Prototype with freedom, production with rails.
And if the workflow you're about to decompose feeds your growth decisions, start with clean input: a free geo-grid scan at https://gbppeak.com/free-maps gives you a structured, block-ready view of where your Google Maps ranking stands across your service area.
Frequently Asked Questions
Why can't a well-written prompt just handle everything?
Even perfect prompts produce run-to-run variance when the model juggles multiple objectives at once — attention is finite. Decomposition fixes variance structurally: fewer objectives per call means fewer failure modes, and each step's output becomes checkable instead of trust-me.
Isn't multiple AI calls more expensive than one?
Slightly more calls, but each is smaller and rarely needs retries; blobs often cost more in practice because failures force full re-runs. The real economics: a 99%-reliable pipeline replaces hours of manual checking every week.
How small should each block be?
One verb per block. If your instruction contains "and then also," split it. If you can't write a one-line automated check for the output, the block is still too big.
Should everything be blocks, then?
No — use agents for monitored conversations, open-ended exploration, and prototyping. The pattern that works: agents to discover the path, blocks to run it in production.
Final note
The question stopped being "how smart is the prompt" and became "how reliable is the system." Segment the work, contract the outputs, check each step — and your AI stops being a slot machine and starts being a machine.