GTM

How much of your outbound should AI actually write?

Written by
Pravin Kumar
Published on
Sep 24, 2026

How much of your outbound should AI actually write?

Let a model do the research and the first draft, and keep the claim, the ask, and the send decision human. In practice that means AI writes maybe half the words and none of the promises. The moment a model is deciding what you offer, you have automated the wrong half.

I get asked this by founders who are about to hire an agency or buy a sequencing tool, and the question underneath it is always the same. Can I send more without becoming the thing I delete from my own inbox?

The answer is yes, but not by generating more variants of the same weak email. Volume was never the constraint. Relevance was, and relevance is the part people try hardest to skip.

What does AI genuinely do well in outbound?

Three things. It reads faster than you, it drafts without ego, and it never gets bored on the two hundredth account. Research, summarisation, and first drafts are real wins. Deciding who to contact and what to promise them is not on that list.

The research win is the big one. Reading a prospect's pricing page, changelog, careers listings, and recent announcements to answer one question about whether they have the problem you solve is exactly the kind of tedious reading a model is good at. The output is a paragraph for your eyes, not an email.

The drafting win is smaller than people think, because a draft from thin research is fluent and empty. Fluency is not the scarce resource in outbound. A reason to write today is the scarce resource, and no amount of language quality manufactures one.

What happens to deliverability when you scale with AI?

Volume moves you into a stricter rulebook. Google's published sender requirements say that starting February 1, 2024, all senders to Gmail accounts must set up SPF or DKIM, keep valid forward and reverse DNS records, use a TLS connection, and keep spam rates reported in Postmaster Tools below 0.3 percent.

The page also sets a second tier. Senders of more than 5,000 messages per day to Gmail accounts have to meet additional requirements, including having the domain in the From header aligned with either the SPF domain or the DKIM domain to pass DMARC alignment, and marketing and subscribed messages have to support one click unsubscribe with a clearly visible unsubscribe link in the body.

Read those two paragraphs again with an AI tool in mind. Generation makes it trivially easy to cross a volume line that comes with obligations, and the spam rate threshold is the one that punishes bad targeting rather than bad writing. Recipients marking you as spam is the metric, and a model cannot talk its way out of it.

Check Google's current documentation before you build on any of this, since sender requirements are exactly the kind of thing that gets updated, and the same goes for other mailbox providers whose rules I am deliberately not summarising from memory here.

Which parts should stay human?

The list, the claim, and the send. Who is worth contacting, what you are asserting about their business, and whether this specific message should go out today. Those three decisions carry all the risk and almost none of the typing.

The list is first because everything downstream inherits it. A generated email to the wrong person is not a small waste, it is a spam complaint waiting to happen, and complaints cost you the channel rather than the deal. If you cannot describe why this account is on the list in one sentence, take it off.

The claim is second. Any sentence asserting what a prospect is currently doing, spending, or failing at needs to be true, and a model will happily generate a confident guess. I have seen drafts assert that a company was running a stack it had abandoned two years earlier, which is an instant credibility loss on a first touch.

The send decision is third and it is the cheapest safeguard. A person glancing at the final message before it goes catches the generated line that reads oddly, and the one sent to the wrong company after a bad merge. Machines draft, people release.

How do recipients tell a researched email from a generated one?

They look for something that could only have been written to them. A real detail, a real constraint, a real question. Generic personalisation tokens read as automation now, so the giveaway is not the writing quality, it is whether the message survives being sent to anyone else.

My own test before any outbound goes out is the swap test, the same one I use on positioning. Could this email be sent unchanged to the next company on the list? If yes, it is a broadcast wearing a first name, and it will be treated as one.

The second giveaway is asking for too much too early. A message that clearly took thirty seconds to make and asks for a thirty minute call has the ratio backwards. Making the ask proportional to the effort shown is basic manners and it works, which is part of why content that helps your champion sell internally often beats a meeting request as a first ask.

What does a sane AI assisted outbound motion look like?

Narrow list, deep research, human claim, drafted first pass, human send, small volumes, and a quality bar you can defend. Roughly twenty five well researched messages a week beat a thousand generated ones, and they leave your domain reputation intact for next quarter.

The stack matters less than the sequence. Enrichment and research can run in Clay, Apollo, or a scripted pipeline in Zapier or Make. Drafting can run through ChatGPT or Claude. Sending can go through a sequencer or plain Gmail. None of those choices fix a bad list, and all of them amplify one.

I would also keep outbound honest about where it fits. It is one channel among several, it suits a specific kind of offer, and it deserves the same patience and the same kill criteria as any other. Deciding how long to give a channel before killing it matters more than which tool you picked.

How do you measure whether this is working?

Measure replies that lead somewhere, not opens. Count positive replies, meetings held, and opportunities created per hundred messages, and watch your spam complaint rate as a hard constraint rather than a metric. Any volume increase that raises complaints is a loss dressed as growth.

I separate two numbers deliberately. Response quality tells you whether the list and the claim are right. Deliverability health tells you whether you are allowed to keep going. Teams that only watch the first number tend to discover the second one abruptly.

Attribute honestly too. Outbound often gets credit for deals that inbound warmed up first, and the reverse happens as well. If your reporting cannot survive that ambiguity, you will over invest in whichever channel logs the last touch.

Where is the ethical line?

Do not fake familiarity, do not fabricate details about a prospect, and do not pretend a machine drafted message came from a conversation that never happened. Everything else is a judgement call about relevance, and relevance is your responsibility rather than the model's.

The fabrication point is the one that ends relationships. Inventing a mutual connection, a shared event, or a problem you have not verified turns a cold email into a lie, and it is the kind of thing that gets screenshotted. A model will produce it on request without hesitation, which is exactly why the claim stays human.

My personal line is that I will send anything I would be comfortable having quoted back to me in a first call. If a sentence only works because the recipient will not check it, it does not go out.

Who should you actually be writing to?

The accounts that look like the ones you already serve well, described by the situation they are in rather than by their size or industry code. AI assistance makes a narrow list practical, because deep research on fifty accounts is now an afternoon instead of a week.

That is the real unlock and almost nobody uses it that way. Given a research assistant, most teams keep the same shallow list and multiply the sending. The better trade is to keep the volume and multiply the depth, and it takes a defined profile to know what to research for, which is why defining an ICP from a small customer base comes before any of this tooling.

If your profile is still a guess, do that work first. Outbound to a guessed audience is the most expensive way to discover your positioning is wrong.

What should you do next?

Take your next twenty five accounts, have a model research each one against a single yes or no question, write the claim yourself, let it draft the rest, and read every message before it sends. Then check your complaint rate and your reply quality before you increase volume.

If you want help deciding which parts of your go to market motion should be automated and which should stay in human hands, that is a conversation I have with founders most weeks. Tell me what you are sending today and I will tell you what I would automate first. Let's chat.

Get found, cited and the back office automated

Let's make your site the source AI engines quote and wire up the systems behind it.

Contact

Let's get your website found and cited by AI

Tell me what you're working on, whether AI search is skipping your product, your back office is buried in manual work, or you need a build that does both.

Got it, thanks. I read every message personally and reply within 1-2 business days.
Oops! Something went wrong while submitting the form.