Cold-email testing guide

A/B test cold email and find results you can trust.

When replies are rare, one early response can make a weak variant look decisive. Define the hypothesis, outcome, audience, and stopping rule before comparing sequence versions.

Reviewed September 8, 202611 minute readProduct behavior checked against current implementation
A useful cold-email test changes one meaningful variable, measures a preselected recipient outcome, and runs long enough to separate a plausible effect from ordinary variation.

Write the test down before it starts

State the hypothesis, primary metric, denominator, minimum worthwhile improvement, audience, and stopping rule in advance. This prevents the team from changing the question after seeing a convenient result.

DecisionExampleWhy it matters
HypothesisA more concrete problem statement will increase positive replies.Connects the copy change to a reason.
Primary outcomeInterested or meeting-booked replies.Avoids declaring victory on any reply, including objections.
DenominatorPositive replies ÷ delivered messages.Makes rates comparable and auditable.
Minimum useful effectThe smallest lift worth changing the default for.Separates practical value from a mathematical difference.
Stopping ruleRequired observations or a preplanned analysis date.Prevents stopping after the first favorable reply.

Change one material variable

Test an offer frame, opening, proof point, call to action, or follow-up path. Keep the audience definition, sender conditions, delivery window, and all unrelated copy as stable as practical. If several things change together, the result can select a complete version but cannot explain which change caused the difference.

Subject-line tests can put too much weight on opens. Open data is unreliable, and MailSequence does not currently list it as a campaign metric. Focus on reply quality and downstream conversations.

Use comparable groups

  • Draw both groups from the same audience definition and time period.
  • Balance important attributes such as role, company size, geography, and list source.
  • Avoid putting one variant on healthy inboxes and the other on newly connected senders.
  • Apply the same verification, suppression, daily caps, and reply-stopping rules.
  • Record operational incidents that could distort one side of the comparison.

Choose a sample size for the decision

For a reply outcome, sample size depends on the expected baseline rate, the smallest effect worth detecting, the chosen significance level, and the desired power. Rare outcomes and small expected improvements require more observations. A fixed internet rule such as “send 100 emails per variant” ignores those inputs.

Choose the significance threshold and analysis plan before the test. A small p-value is not the same as a valuable business result; a large sample can make a trivial difference look statistically significant, while a small sample can miss a useful effect.

If the test cannot collect enough observations, report it as directional evidence. Do not silently promote the leading variant to a proven winner.

Read the whole outcome

Compare delivered volume, replies, reply dispositions, bounces, complaints, stop requests, and operational health. More replies can still be a poor result when negative responses rise. Review the absolute effect and the confidence around the difference.

How to structure the comparison in MailSequence

MailSequence preserves published sequence versions and can assist with drafting alternative copy. Campaign reporting lets operators select a version and inspect sends, delivered messages, replies, steps, and reply dispositions. Label the denominator used in every analysis.

MailSequence does not currently list a random-assignment experiment engine as a shipped feature. For a controlled comparison, create stable versions and allocate comparable audience groups. Record the allocation and leave each version unchanged during the test.

Sources and product basis

Run tests that can change a real decision.

Define the outcome, publish stable versions, allocate comparable audiences, and interpret reply quality before choosing a default.