Headlines experiment
Experiment: above-the-fold headline test
The headline test is the most over-run and least-honestly-analyzed indie SaaS experiment. Founders test headlines on 200 visitors and call it a result; the math says they need 2,000+ per variant for statistical reliability. The experiment below names what to test, when the test is honest, and the sample-size discipline most founders skip.
Min sample size: 1,000 visitors per variant for a 50%+ relative lift detection. 2,000+ for 25% lift. 5,000+ for 10% lift. Most indie SaaS sites cannot run honest headline tests below 1,000 visits per variant.
Duration: 7-21 days minimum, even at adequate traffic. Day-of-week effects matter; 7 days is the floor.
Verified · editorial policy
Hypothesis structure
Changing the headline from [CURRENT] to [VARIANT] will increase [PRIMARY METRIC] by at least [EXPECTED LIFT] because [SPECIFIC REASON tied to Wrong Person / Weak Offer / Weak Belief diagnosis].
If you cannot complete this template, you do not have an experiment — you have a guess.
Variant design
Single change: only the H1 sentence. Keep sub-hook, CTA, trust block, and all other page elements identical. Two-variable changes muddle attribution.
Primary metric
Click-through to checkout or signup. Click-rate is more sample-efficient than conversion-to-paid; conversion-to-paid requires 5-10x the sample size.
Secondary metrics (watch but do not decide on)
- Scroll-depth past the fold (does the new headline keep readers reading?).
- Bounce rate (does the new headline drive visitors away faster?).
- Time-on-page (proxy for engagement at the top).
Procedure
Step 1
Write the hypothesis explicitly
Using the template above. If you cannot complete it, you do not have an experiment — you have a guess.
Step 2
Set traffic split with an honest tool
PostHog, GrowthBook, or a server-side split. Cookie-only client-side splits leak across sessions and produce noisy data.
Step 3
Run for full week-cycles
Day-of-week effects are real. End the test at the same day-of-week boundary you started — never mid-week.
Step 4
Refuse to peek at results before the duration ends
Peeking shifts your mental model and biases the analysis. Pre-commit to the duration; check at the end.
Step 5
Run a chi-square or two-proportion test
Eyeballing 'this looks better' is confirmation bias in numbered form. Use a real statistical test or a calculator (evanmiller.org/ab-testing/sample-size.html).
Step 6
Decide and ship
If the variant wins reliably, ship the variant and start the next experiment. If the variant loses, keep the control. If the result is inconclusive (within noise), the experiment was insufficient — usually small sample size.
Self-deceptions to avoid
- Calling a 200-visitor test 'a result'. At that volume, the noise dominates the signal.
- Testing two changes at once (headline + sub-hook). Muddles attribution; cannot tell which moved the metric.
- Stopping the test early because the variant 'looks better'. Early stopping inflates winner detection by 30-50%.
- Treating peak-traffic days as representative. Weekday-weekend mix matters.
- Confirming a hypothesis-shaped expectation. The honest test is whether the variant DOES better, not whether it COULD be argued to.
What success looks like
A variant that wins by 25%+ on click-through with 1,000+ visitors per side, statistically significant at p<0.05. Replicable on a fresh cohort if you re-run.
Related benchmark
See the directional range for landing page conversion rate to calibrate the expected lift in your hypothesis.
Frequently asked
- What if I don't have 1,000 visitors per variant?
- Then you cannot run an honest A/B test on headlines. The right move at lower volume is qualitative testing: show both headlines to 10-20 people in your target audience, ask which they would click and why. Qualitative beats noisy quantitative at small scale.
- Can I test headlines on different traffic sources separately?
- Yes, and you should be aware sources behave differently. Twitter traffic, Google organic, and direct-link traffic respond differently to the same headline. Pool only if the sources behave similarly historically.
Other experiments
Test on a page that is already pointed in the right direction
A/B tests on a misaligned page produce two losing variants. The diagnostic labels the alignment problem first; the test optimizes within the right alignment.