Our fulfilment runs on an AI operations layer: agents draft, monitor and pattern-match at scale, and a human approves every send.
By Joel Wylie, Founder · Last updated 7 August 2026
Our agency's entire fulfilment operation runs on Claude Code. AI agents draft every prospect reply, QA every campaign before launch, sweep every client account nightly and prepare every weekly report, and a human approves every single action that touches a client or a prospect. We did not buy an AI SDR. We built an AI operations layer underneath a human operator, and that architectural difference is the whole reason it works.
Most "AI for agencies" content is written by people selling software. This is written by people running the thing. We operate cold email campaigns at up to 50,000 sends a month per client, and the honest version of our stack is not "AI does everything" or "AI is just a toy". It is a specific division of labour: AI does the reading, drafting, monitoring and pattern-matching at a scale one operator could never cover manually, and a human stays on every decision that leaves the building.
It means the operating manual of the business is written down as instructions an AI agent can execute, and the agent runs those instructions on a schedule or on demand. Our copywriting doctrine, launch checklist, reply-handling rules and diagnostic playbooks all exist as documents the system reads before it acts. When a prospect replies, when a campaign is ready to launch, when an account needs its nightly check, an agent picks up the relevant playbook and does the work.
The critical design choice is where the automation stops. The system produces drafts, findings and recommendations. It never produces an unsupervised send. Every reply to a prospect, every campaign launch, every report to a client arrives as a proposal with the work already done, and a human says yes or no. That single rule is what separates an operations layer from a liability.
The second design choice is that everything is written down. An instruction that lives in someone's head cannot be executed by an agent, audited after a mistake, or improved when it fails. Writing the operation down was painful and it is also why the business survived scaling: the doctrine is the product, and the AI is the engine that executes it consistently.
Every client has a knowledge base: their offer, their proof points, their case studies, their tone, the exact promises their campaigns made. When a prospect replies, an agent classifies the reply by intent, reads that client's knowledge base, and drafts a response in the client's voice in about 30 seconds. Objections get answered with real proof points. Interested replies get a clear next step. Wrong-person replies get routed, not pitched.
The draft then lands in front of a human for one-tap approval before anything sends. That review gate is not a temporary training wheel we plan to remove. High-stakes conversations always deserve a human eye, and the gate is also where the system learns: every edit a human makes to a draft is a signal about what the agent got wrong. We wrote up the full architecture in the AI reply agent playbook.
The result is that speed and quality stop trading off against each other. Reply handling used to be the fastest leak in outbound: a warm lead replies, the operator is busy, three days pass, the deal is gone. Now every reply across every client gets a considered, grounded draft within a minute of arriving, and the human's job shrinks to judgment.
Every launch, resume and lead upload passes through an automated pre-launch checklist that works as a hard gate. The agent verifies the list, checks that every personalisation variable resolves on every lead, scans all copy against a spam-word list, confirms sequence structure and send settings, and checks the sending infrastructure end to end. Anything that fails blocks the launch. No checklist, no launch, no exceptions.
We learned the value of the gate the hard way, twice. One campaign went out to a list where a small slice of recipients sat behind a strict corporate email security gateway; that slice was 7% of the list and produced 77% of all bounces. Another time a client asked us to swap one approved phrase back to the word "free", the edit went straight to the live campaign without a re-scan, and deliverability told on us within days. Both failure modes are now impossible, because the gate re-checks everything on every launch regardless of where the copy or the list came from.
That is the real argument for AI-run QA. A human running a 40-item checklist for the tenth time that week skips steps. An agent runs item one to item forty identically at midnight on a Friday, and it treats "the client wrote this bit" and "we already checked this list last month" as reasons to re-check, not reasons to skip.
Every night, an automated health check sweeps every client account: bounce rates and their multi-day trend, disconnected sender accounts, campaigns that have quietly stopped sending, replies that have not been handled, and approvals that were given but never executed. Problems surface as cards with a diagnosis and a proposed fix already attached. Healthy accounts surface as nothing at all, because a report that says "everything is fine" forty times is how real problems get missed.
On a weekly cadence, a per-client review cycle runs on schedule: it reads the week's numbers, compares variants, checks the account against our benchmarks, and drafts the client report and the next week's plan. A human reviews and approves before anything reaches the client. The judgment call stays human; the four hours of reading that used to precede the judgment call no longer exist.
This is the part of the system with no human-scale equivalent. One operator cannot read every reply thread, every bounce log and every campaign stat across a full client roster every single day. An agent can, and the practical effect is that problems get caught in hours instead of being discovered in the week's numbers.
Every human correction feeds back into the knowledge base. When an operator edits a drafted reply, the edit is captured and distilled into a voice rule for that client. When a campaign outcome validates or kills an offer, the learning is written into that client's file. When a mistake produces a new rule, the rule lands in the playbook the agents read, which means the mistake is structurally unrepeatable rather than just remembered.
This is the compounding loop most AI tooling misses. A generic AI tool is as good on day 200 as it was on day one. An operations layer with a feedback loop drafts better replies in month six than it did in month one, because it has absorbed hundreds of corrections, and every client inherits the cross-client lessons. The same pattern powers our sales engineer agent playbook, where founder-corrected answers feed the knowledge base until the agent handles most technical questions alone.
Because the AI SDR category gets the architecture backwards. An AI SDR tries to replace the salesperson: it owns the conversation, it owns the send, and the human is reduced to reading dashboards. That is exactly why prospects can smell AI SDR output a mile away, and why the category keeps disappointing buyers. The hard part of outbound was never typing the words. It is judgment: which offer, which list, when to push, when to walk away.
| AI SDR | AI operations layer (our model) | |
|---|---|---|
| What the AI owns | The conversation and the send | Drafting, monitoring, QA, pattern-matching |
| What the human owns | A dashboard | Every send, launch and client-facing decision |
| Grounding | Generic templates plus scraped personalisation | A per-client knowledge base that grows weekly |
| When it is wrong | The prospect finds out | The reviewer finds out, and the correction becomes a rule |
| How it improves | Vendor ships an update | Every human edit feeds the knowledge base |
| What it replaces | The salesperson | The reading, typing and checking around the salesperson |
The economics follow the architecture. Because the AI absorbs the hours and the human keeps the judgment, one operator supervises a workload that would otherwise need a team, without the quality collapse that full automation produces. We have broken down those numbers against the in-house alternative in our agency versus SDR cost comparison.
Our position is blunt: full automation of outbound is a quality disaster, and zero automation is a scale ceiling. The businesses that win with AI over the next few years will not be the ones that removed humans. They will be the ones that redesigned the operation so machines do everything below the judgment line and humans do everything above it, with a written doctrine connecting the two.
Yes. Our entire fulfilment operation runs on Claude Code: reply drafting, pre-launch QA, nightly account health checks, weekly review cycles and report drafting. The constraint is design discipline, not the model. Every action that touches a client or a prospect still passes through a human approval gate.
No. The AI drafts every prospect reply, but a human approves each one before it sends. The same gate applies to campaign launches, client reports and client messages. Automation covers the reading, drafting and monitoring; a person owns every decision that leaves the building.
An AI SDR replaces the salesperson and owns the send, which is why prospects smell it instantly. An AI operations layer sits under a human operator: it drafts, monitors and pattern-matches at machine scale, while the human keeps judgment, taste and every approve decision. Same technology, opposite architecture.
About 30 seconds from reply landing to draft ready for review, grounded in that client's knowledge base and written in their voice. A reply that used to take 30 minutes of context-gathering and writing now takes one human read and one tap to approve.
We run campaigns at up to 50,000 sends a month per client, with every account swept nightly for bounces, disconnected senders and stuck campaigns. One operator supervises the whole roster because the AI does the reading and the human only handles decisions.
Yes. Every edited draft, every approved reply and every campaign outcome feeds back into the client's knowledge base. The agent that drafts next week's replies has read everything that happened this week, so drafting quality compounds instead of resetting with each new conversation.
Want outbound like this run for you, end to end?
Book A Call
Three routes to booked B2B meetings, compared on cost structure, risk, and what happens when results are zero.

SaaS companies hold a structural advantage in cold email that most of them waste. Here is how to match the offer to your motion and stop leading with a demo ask.

Every agency inbox pitch sounds identical. The one that wins is the one that gives away the most real work for free. Here is the doctrine we run, on our clients and on ourselves.