What you need to know
- At a 2% reply rate, 150 emails per version can only reliably tell 2% apart from 9%.
- Test the offer before the wording. In thousands of public headline tests, the best beat the average by only about 1.18 times.
- A fair test changes one thing, splits by business and sends every version on the same days.
- Decide what counts as success before you send, and count meaningful replies, not opens.
- Most replies to cold email are machines. Read the words before you count them.
Most cold email A/B testing advice says 100 to 200 emails per version is enough. At a 2% reply rate, that many emails can only reliably tell 2% apart from 9%. Anything smaller than a four-fold difference disappears into luck.
We know this because we've just started testing our own outreach properly. Gibson Promotions has spent 20 years helping Sydney businesses find customers, and we've stopped trusting our own gut about what works. We're now sending 1,053 Sydney businesses one of three offers, split at random. The results come in after 29 October. This article is about how we set it up, what we believe, and what it has already taught us before a single result is in.
How many emails do you need for an A/B test?
More than most people think, and it depends on how often people reply. The lower your reply rate, the bigger the sample you need to see a real difference. Here is what a standard test can reliably detect (80% power, the usual 5% significance level).
| Emails per version | If 2% normally reply, it can spot | If 5% normally reply, it can spot |
|---|---|---|
| 150 | 2% vs 9% | 5% vs 14% |
| 500 | 2% vs 5.3% | 5% vs 9.6% |
| 1,000 | 2% vs 4.2% | 5% vs 8.1% |
So a test of 150 emails per version will only show a "winner" when one version is several times better than the other. Smaller wins are real, but a test that size can't see them. When it does crown one, it's often luck.
That luck is easy to underestimate. In a public archive of thousands of randomised headline tests, the best headline beat the average one by only about 1.18 times once you correct for chance. Most "winning" wording was noise that looked like a result.
To be fair to the usual advice: if you're testing open rates, where far more people take part, smaller samples see smaller differences. But opens aren't customers. For replies, meetings and quotes, you need hundreds per version, and you should only test changes big enough to matter.
What should you test first in cold email?
Test the offer before you test the words. Changing what you offer changes how people respond far more than rewording the same offer. That's why our test compares three different offers, not three subject lines:
- a letterbox drop for their business
- getting back the customers they've lost
- a free review of their own customer records: finding everyone they never got back to, cleaning up the data and sorting out who's worth contacting
We think most small businesses test the wrong thing. They split-test a subject line on 200 people, declare a winner, and never notice it was luck. A different offer is a decision worth testing. A different comma isn't.
“Most small-business A/B tests can't tell you anything because they're too small and they test the wrong thing. We test offers, not subject lines, and we decide what counts as success before we send a single email. That's how we're testing Reignite on 1,053 Sydney businesses right now.”
How do you make a marketing test fair?
Change one thing at a time, and let chance decide who gets what. Here is how we set ours up, and why:
- Split by business, not by email address. If a business has two branches with two inboxes, both get the same version. Otherwise one business sees both offers and the comparison is muddied.
- Send every version on the same days, alternating. If one version goes out on a Monday and the other on a Friday, you're testing the day, not the offer.
- Keep the sender the same where you can. A study of about 80,000 near-identical pitches found the sender alone changed how many people replied. In our test, two of the offers come from the same person, so comparing them changes only the offer. Our letterbox offer comes from a different sender, because that's how we'd really sell it. We say so plainly when we report it, and don't pretend it's a pure offer test.
- Write the plan down before you start: what you're testing, how many, and what counts as success. We froze ours before the first email went out.
Government teams work the same way. The NSW Government's Behavioural Insights Unit tests message wording with randomised trials, such as one at St Vincent's Hospital Sydney, so that the answer comes from the people, not the person who wrote the message.
How do you measure if outreach is working?
Decide what counts before you send, and count replies that mean something. We count the first positive reply within 15 business days: someone asking for a price, asking to talk, or booking a time. Two people check each one against that definition, so we can't quietly count the replies we like.
Clicks and opens are tempting because they're big numbers. But research on headlines found that attention tricks, like numbers and teasers, win clicks without telling you who'll buy. A click is curiosity. A "how much would that cost?" is interest.
What did we learn before the results came in?
The test is still running, but it has already taught us four things that have nothing to do with which offer wins:
- Most replies to cold email aren't people. About half of the replies on our first day were automatic acknowledgements: "we've received your email and will respond within 48 hours". If you count replies, you're counting machines. Read the words.
- Software can misread a "yes". Our email ends with "If it is not relevant, reply 'no thanks'". When someone replies and the reply quotes our email, a naive filter sees "no thanks" and marks an interested person as a no. We now read only the sender's own words.
- An automatic reply can quietly stop your follow-up. "We'll get back to you" isn't a person saying anything. If your system treats it as a reply, that business never gets your second email. We fixed ours so acknowledgements don't pause anything.
- "Random" isn't always random. When we audited public test datasets released by large companies, several weren't as evenly split as their documentation claimed. Check your own split before you trust the result.
“Before a single result came in, our own test showed us that about half of the replies to cold email were machines, and that software can file an interested person as a no. Any Australian business counting replies to judge its outreach should read the words first, then count.”
What we believe
Most small-business A/B tests are noise. Not because people are careless, but because the numbers are too small for the differences they're looking for.
Test the offer before the words. The offer decides whether anyone is interested at all.
Hold some people back. Contacting customers isn't always harmless. In one well-known randomised study, a proactive "better plan" offer raised customer churn from about 6% to 10%. When you re-contact your own customers, keep a small group you don't contact, so you can see what your outreach actually changed.
Publish what didn't work. We'll share what this test finds, including the offers that lose. A business that only ever reports wins isn't testing. It's advertising.
Where reasonable people disagree
"Testing is for big companies." If you send 50 emails a month, a formal test will take a long time. We'd still say: test only big changes, keep the plan written down, and pool results over a few months. You'll learn more than from guessing.
"Just trust your gut." Sometimes your gut is right. Our letterbox offer has done well for years without a test. But "it feels like it's working" and "it works" aren't the same thing, and a fair test is the only way to tell them apart.
If you'd like to see what's sitting in your own customer records, see how Reignite works.
Frequently asked questions
How many emails do you need for a cold email A/B test?
It depends on your reply rate. At a 2% reply rate, 150 emails per version can only reliably spot 2% against 9%. With 500 per version you can spot 2% against about 5%. Test only changes big enough to matter, and treat small "wins" as probably luck.
Should I A/B test subject lines or the offer?
Test the offer first. Changing what you offer usually moves response far more than rewording. In a large public archive of headline tests, the best wording beat the average by only about 1.18 times once chance was taken out. Subject-line tests on small lists mostly measure luck.
What should count as success in a cold email test?
Decide before you send, and count meaningful replies: someone asking a price, asking to talk, or booking a time. Opens and clicks are big numbers but weak signals. Have two people check each positive reply against the same definition.
What is Gibson Promotions testing with Reignite?
Gibson Promotions is running a randomised test across 1,053 Sydney businesses to see which offer gets more owners to put their hand up. That includes Reignite, its lead reactivation platform, and a free review of a business's own customer records. Results will be published after the test window closes.
Sources and evidence
- Matias et al. (2021), The Upworthy Research Archive, Scientific Data: 32,487 randomised headline tests. The 1.18 times figure is Gibson's re-analysis after correcting for chance.
- Gornall and Strebulaev (2025), Management Science: About 80,000 pitch emails with the sender randomised.
- Ascarza, Iyengar and Schleicher (2016), Journal of Marketing Research: A proactive plan offer raised churn from about 6% to 10%.
- NSW Behavioural Insights Unit (2015), St Vincent's Hospital Sydney: Randomised trial of reminder wording.

