Direct answer: Small agencies should test one variable at a time, at the biggest lever first: the first line, then the ask, then the subject. Split each batch evenly, run until each side has meaningful volume, judge on replies rather than opens, and keep a log so wins accumulate instead of evaporating. You do not need enterprise volume to test. You need discipline about testing one thing and patience to let the numbers speak.
Key takeaways
- Test order matters: first line, ask, subject, length, in that order of leverage.
- One variable per test. Two changes means zero conclusions.
- Replies are the metric. Opens lie, and pixel tracking grows less reliable yearly.
- Small volume means slower verdicts, not no verdicts. Log everything.
What should a small agency test first?
The first line, because it decides the two second glance that everything else depends on. Test signal types against each other: role reference lines versus funding reference lines versus pain reference lines, same frame beneath. Next lever: the ask. Micro ask ("worth a look?") versus calendar ask ("15 minutes Thursday?"). Then subjects: plain ("sales roles") versus context ("your Manchester launch"). Length and send timing come last, since their effects are real but smaller. Skip cosmetics entirely: signature colors and greeting styles test noise.
How much volume does a verdict need?
Honest answer: more than one week's sends at small scale. A practical rule for agencies: run each side to at least 100 sends before leaning, 200 before declaring, and treat sub 50 send observations as anecdotes. Slow verdicts still compound: a small agency running one clean test a month has twelve accumulated wins a year later, which is how small senders end up out converting big lazy ones. Inside SDR GROW, the 16 touch flow reports replies per template variant, so the counting happens without spreadsheet archaeology, and winning lines roll into the drafts the lead engine and Industry Insight keep feeding.
What does the testing log contain?
Five columns: date range, variable tested, both versions verbatim, sends per side, replies per side. Plus one habit: the losing version gets archived, not deleted, because market shifts revive old losers and memory alone will re test them by accident. The log turns testing from vibes into an asset a new team member can read in ten minutes.
Checklist: clean test hygiene
- One variable changed, everything else frozen.
- Batches split randomly, same day, same list quality.
- Judged on replies, positive replies noted separately.
- Minimum volume reached before any verdict.
- Result logged with both versions in full.
Example
A three person agency tests role reference first lines against company news lines, 150 sends each over three weeks. Role lines pull nearly double the replies. Next month they test asks on the winning line: the named day calendar ask beats the open micro ask for their market. Two months, two clean wins, sequence reply rate up meaningfully, and both results written down. Their bigger rival, meanwhile, has "tested lots of stuff" and can name no numbers. Discipline is the whole edge.
Mistakes to avoid
- Testing subject and first line together, then crediting the wrong one.
- Calling winners after twenty sends because the early gap looked exciting.
- Testing during holiday weeks and shipping the polluted verdict.
- Winning a test and never deploying the winner. It happens more than anyone admits.
FAQ
Is open rate ever worth tracking??
As a deliverability smoke alarm, yes: a sudden drop flags spam placement. As a copy verdict, no.
Should each niche get its own tests??
If you serve genuinely different markets, yes. A line that wins with fintech founders can lose with construction directors.
How long before testing stops paying??
It slows after the big levers settle, then market drift restarts it. Yearly re validation of your champion lines is the maintenance dose.
Related reading
Ready to build predictable pipeline for your agency?
Book a Strategy Call →