Two subject lines, 150 sends each. Version A gets 6 replies, version B gets 3. It is tempting to crown A and move on.
Now picture one identical email going to both groups, at a true 3% reply rate. A gap of 3 or more replies still turns up about 39% of the time. The full arithmetic comes later.
How to A/B test cold emails, in one line: shuffle one list and split it in two, change one thing between the halves, write down the stop point before the first send, and score the result on replies, never opens.
The hard part is volume. Say 3 in every 100 prospects reply today and a new version would get 4.5. A standard test (95% confidence, 80% power) needs about 2,500 sends of each version before it reliably spots that gain.
On a list of a few hundred prospects, the only differences you can see are big ones: a different segment, a different offer, a different angle, a different sequence.
So on a small list, change the structure of the campaign. Word swaps can wait until the volume is there.
What should you test first in a cold email?
Sort every test idea into one of two kinds: structural changes (who gets the email, who sends it, what it offers) and wording changes (subject line, opening line, call to action).
In my Journal post on cutting cold email copy I described what we vary in live campaigns: "In a real campaign we are always running variants: different subject lines, different angles, different numbers of steps in the sequence".
The same post puts copy in its place: "Copy is worth maybe a few points of reply rate."
A few points is roughly the gap a list of a few hundred cannot see, as the tables further down show. So a small list should spend its tests on the structural kind.
Structural or wording: where a small list should spend its tests (guidance, based on the sample-size math below)
| What you change | Kind of change | When it is worth testing |
|---|---|---|
| Who receives it (titles, company size, country) | Structural | Before anything else |
| Who sends it (founder or rep) | Structural | From the first campaign |
| The offer (what the first call gives them) | Structural | From the first campaign |
| The angle (which problem you lead with) | Structural | From the first campaign |
| The proof (a real client result in the message, or none) | Structural | From the first campaign |
| The sequence (number of steps, email only or with LinkedIn) | Structural | Once a first version works |
| Opening line and call to action | Wording | A few thousand sends per version |
| Subject line wording, send time | Wording or timing | Several thousand sends per version |
Start with the segment, because it decides who is able to say yes at all. In my post on B2B lead lists I described companies that skip the basic list checks "and then spend three months A/B testing a subject line."
A segment test runs differently from the rest. The list is the thing you change, so keep the email, the sender and the sending days identical. A good first pair: companies that look like your past and current clients against the broader list you would otherwise send to.
A segment that replies more has a bonus: the next tests on it finish sooner. In the sample-size table below, a 50% lift needs 2,515 sends per version at a 3% baseline and 1,468 at 5%.
Proof is its own lever. My advice is to put past clients and their results inside the message, so one concrete client result against no proof at all is a fair structural test.
How many sends does each version need?
Four inputs set the number: your current reply rate, the size of lift worth catching, the confidence level (usually 95%) and the power, meaning how often the test should catch a real difference (usually 80%).
The standard formula for comparing two proportions is published with a free calculator by Select Statistical Consultants. With a as your current rate and b as your target: n = (1.96 + 0.84)² x (a(1 - a) + b(1 - b)) / (a - b)².
Their own worked example, a purchase rate moving from 15% to 25%, needs 248 people per group.
Their example is a supermarket promotion with rates in the teens. Here is the same formula run on cold email reply rates.
Sends needed per version, 95% confidence, 80% power (illustrative calculation)
| Current reply rate | Catch a 25% lift | Catch a 50% lift | Catch a doubling |
|---|---|---|---|
| 2% | 13,807 | 3,823 | 1,139 |
| 3% | 9,097 | 2,515 | 746 |
| 4% | 6,743 | 1,861 | 550 |
| 5% | 5,330 | 1,468 | 432 |
| 15% (for comparison) | 1,562 | 422 | 118 |
| 20% (for comparison) | 1,091 | 291 | 79 |
I ran the formula with unrounded z-values and rounded up; the same method gives Select's 248 for their example. Catching the same 50% lift takes more than 8 times as many sends at 3% as at 20%.
Low reply rates make small counts jumpy. A version sent to 150 people at a true 3% rate expects 4.5 replies, yet anything from 2 to 7 turns up about 86% of the time by chance alone (illustrative calculation).
Select also notes that this approximation stops being valid when a proportion is close to 0, so read every row from 2% to 5% as an order of magnitude.
If you want your own numbers, Evan Miller's free sample size calculator takes your baseline and the lift you care about.
What can a list of 500 prospects actually detect?
Turn the question around. How big does the difference have to be before your test can see it?
Reply rate the better version must reach to be caught 80% of the time (illustrative calculation, list split evenly)
| Prospects in the test | Baseline 2% | Baseline 3% | Baseline 5% |
|---|---|---|---|
| 300 | 9.4% | 11.2% | 14.5% |
| 500 | 7.2% | 8.9% | 11.9% |
| 1,000 | 5.3% | 6.8% | 9.6% |
| 2,000 | 4.2% | 5.5% | 8.1% |
| 4,000 | 3.4% | 4.7% | 7.1% |
With 500 prospects and a 3% baseline, the new version has to reach about 8.9% before the test will reliably show it. That is roughly three times the old reply rate.
A reworded subject line is unlikely to triple replies. A new offer or a new angle sometimes changes who answers at all. That is the whole case for testing structure first on a small list.
Why does the losing email sometimes win the rematch?
Send the same email twice. Two groups of 150, the same true reply rate of 3%, nothing different at all.
Worked out exactly with the binomial distribution, there is about a 39% chance that one group ends up at least 3 replies ahead of the other, and about a 23% chance of a lead of 4 or more. Identical emails. A founder reading that dashboard would still pick a winner.
Now the opening example: 6 replies against 3, from 150 sends each. That is exactly the 3-reply gap above, which identical emails produce about 39% of the time. A standard two-proportion z-test agrees: p = 0.31, far above the usual 0.05 bar (illustrative calculation).
A confidence interval says the same thing another way. The 95% range for the real difference runs from B being about 1.9 points better to A being about 5.9 points better. When that range includes zero, you have a tie, whatever color the dashboard shows.
This is why a "winner" from a small test so often loses the rematch. It never won in the first place.
The same effect works across tests. Pick the best of several small results and part of what you picked is luck, so the number tends to shrink when you send the version again.
A second trap has nothing to do with noise. Every test answers a question about one list, one sender and one month. Identical copy can land 7.5% replies at one company and only 2% at another, so a version that won on one list is only a well-informed guess on the next.
Why does checking results every morning break the test?
Your sending tool shows replies as they arrive, and it is tempting to call the test as soon as one version pulls ahead. Evan Miller showed what that does in a 2010 essay on stopping tests early.
His setup: a change that does nothing, a test stopped as soon as the result looks significant at the 5% level, checked after every observation, with a cap of 150 observations.
The real false positive rate comes out at 26.1%, more than five times the 5% the experimenter thinks he is running. In his words, "the more you peek, the more your significance levels will be off."
His essay points to two ways out.
Settle the sample size before the first send, and read the result once, when every prospect in both groups has finished the sequence.
Or use a rule that expects you to look. Miller's simple sequential test (2015) counts only successes, which in cold email means replies. He notes it works extremely well with low conversion rates, which is exactly where cold email sits:
- Before you start, pick N, the maximum total number of replies you will collect across both versions.
- Split prospects 50/50 at random and count replies for each version.
- If the new version's lead over the old one reaches 2 x √N replies, stop. The new version wins.
- If total replies reach N first, stop. There is no winner.
To catch a 50% lift at 5% significance and 80% power, his table gives N = 170 replies, with a winning lead of 26. The rule is one-sided: it tells you whether the new version beats the old one, and nothing else.
A 500-prospect list at a 3% reply rate produces about 15 replies in total. The rule is honest with you: that list cannot settle a 50% question.
With a small list, test the offer and leave the wording alone
What to test at each list size (guidance drawn from the tables above)
| Prospects per version | What to test | How to read it |
|---|---|---|
| Under about 500 | One structural change: segment, offer, angle, proof, sender or sequence shape | Replies, booked calls and what people actually say |
| About 500 to 2,500 | Structural changes, with the number of sends fixed in advance | Reply rate once, at the end, plus reply quality |
| Above about 2,500 | Wording: opening line, call to action, subject line | Reply rate with a fixed sample or the sequential rule |
The 2,500 line comes from the 3% row: it is roughly what a 50% lift needs. Smaller lifts need far more.
Back in 2023 I wrote that "increasing the volume is the last step, not the first!" If a campaign cannot win with 300 prospects, sending it to 30,000 will not make it win. Small numbers are where you find out whether the offer works at all.
A second rule of mine applies here: the more versions you try, the sooner a winner turns up. On a small list that means one structural test after another, each read properly. Ten tiny variants read on day three teach you very little.
If your company sells seven services, the first test is simpler than it looks. Pick one offer for one audience and test that against a second offer, rather than testing seven offers inside one email.
An offer test inside a live campaign
The case below comes from our own client work. Smirnov Consulting Group is a Prague-based B2B outbound lead generation agency that runs cold email and LinkedIn campaigns for founder-led B2B companies and books qualified sales calls.
One of the case studies on our site is about a US marketing agency. The campaign was LinkedIn outbound followed by warm emails (email sent after a LinkedIn first touch), with full profile optimization, lists from Sales Navigator and Apollo, and an A/B test of 2 offers.
Over 3 months the campaign sent 823 InMails and 2,099 emails and got 141 replies. Sales calls went from 17 in month 1 to 33 in month 2 and 49 in month 3, 99 in total.
The offer test ran alongside the profile optimization and the outreach itself, so read it as an example of an offer-level test inside a real campaign.
Notice the reply count: even all 141 replies together sit below the 170 that the sequential rule plans for when checking a 50% lift. At this scale, read an offer test on booked calls and on what people say in their replies, next to the raw counts.
What if one campaign is too small even for a structural test?
Two moves help. One adds campaigns together. The other checks the list itself, which takes a handful of lookups and no sends at all.
Run one test through several campaigns. Put the same two versions into every campaign you launch, split 50/50 at random inside each one, and keep one running tally per version.
Four campaigns of 500 prospects make a 2,000-prospect test. At a 3% baseline, the detectable-rate table above drops from about 8.9% to about 5.5%.
Keep the split even inside every campaign. If one version got more of an easier segment, the pooled total would reward the segment instead of the email. The sequential rule pools the same way: replies from every campaign count toward one N.
Check the list before you spend replies on it. When a real share of the list is wrong, a small sample finds it. In the list post I suggest pulling 20 random rows and looking each person up by hand.
Say one row in five has the wrong title or the wrong company size. The chance that 20 random rows show none of them is about 1% (illustrative calculation). Twenty lookups answer that question. At a 3% reply rate, 20 sends would not even produce one reply on average.
Addresses are cheaper still. Verification checks every address before the first send, so dead ones come out without costing a single email.
How long does a cold email A/B test take?
Three numbers set the calendar: the sends each version needs, the new prospects each version gets per day, and the length of your sequence.
An illustration: at a 3% baseline, a 50% lift needs 2,515 sends per version. At 60 new prospects per version per day, that is about 42 days of first sends. The test is not finished on day 42, because the prospects who entered last still have their follow-ups to come.
That tail matters more than people expect. In our campaigns, follow-up number three is typically the one that books the call, provided the gaps between messages are real (how I build the sequence).
Stop the test before the last group gets that third follow-up, and you cut off the replies that decide it.
So a fixed calendar rule like "two weeks" only works if two weeks happens to contain enough completed sequences. Count prospects who finished the sequence. Ignore the calendar.
A test plan for a list under 1,000 prospects
Steps you can run this week:
-
Write one sentence on what changes and why it could move replies a lot. If you cannot explain why the difference should be big, choose a bigger change.
-
Shuffle the list, then split it. Never send version A to one list source and version B to another. Do that and you are testing lists, which is fine only when the segment is the thing under test.
-
Change one thing and keep everything else identical. Same sending mailboxes, same days, same number of steps. Swap two things at once and you get a result you cannot carry into the next campaign.
-
Fix the stop point before the first send, and do not act on the dashboard before it. Either a fixed number of prospects per version who have finished the whole sequence, or N for the sequential rule.
Write down what counts as a reply at the same time. Looking is human. Stopping early is the mistake.
-
Read every reply. Who answered, what they asked, which version the booked calls came from: on a small list that is often the clearest signal you have.
-
Call ties ties. If neither version wins by your rule, keep the one that is easier to run or brought better conversations, then line up the next structural test.
-
Move down to wording tests only once each version gets thousands of sends. Before that, a subject line test mostly tells you about luck.
One more condition: testing only makes sense on a stable sending setup. In our own order of work, A/B tests and spintax come after the domains and mailboxes are warmed up and the volume has been raised.
FAQ
How many replies does a cold email A/B test need?
Far more than a small list produces. To check a 50% lift at 5% significance and 80% power, Evan Miller's sequential rule plans for up to 170 replies across both versions.
A list of 500 prospects at a 3% reply rate gives about 15. Below that range, any gap is a hint worth a proper test later.
My sending tool picked a winner after three days. Should I trust it?
Only once the test has reached the number of sends you fixed in advance and every prospect has finished the sequence.
A winner picked on early replies has the peeking problem described above, and the follow-ups that often decide the test have not gone out yet. Until then, treat it as an early read.
What counts as a reply in a cold email test?
A message a person actually wrote. Auto-replies and bounce notices stay out of the count, whichever version they land on.
Write that down together with the stop point, then mark which of the human replies were positive, because that split says more about the offer than the raw count does.
Should I test on open rate or reply rate?
Reply rate. An open is counted when a tracking image loads, while a reply is a person choosing to talk to you, and only replies turn into sales calls.
Opens need fewer sends because the rates are higher, but they measure the wrong thing. If volume allows, go one step further and count positive replies.
