Skip to content

Headlines experiment

Experiment: above-the-fold headline test

The headline test is the most over-run and least-honestly-analyzed indie SaaS experiment. Founders test headlines on 200 visitors and call it a result; the math says they need 2,000+ per variant for statistical reliability. The experiment below names what to test, when the test is honest, and the sample-size discipline most founders skip.

Min sample size: 1,000 visitors per variant for a 50%+ relative lift detection. 2,000+ for 25% lift. 5,000+ for 10% lift. Most indie SaaS sites cannot run honest headline tests below 1,000 visits per variant.

Duration: 7-21 days minimum, even at adequate traffic. Day-of-week effects matter; 7 days is the floor.

Verified · editorial policy

Hypothesis structure

Changing the headline from [CURRENT] to [VARIANT] will increase [PRIMARY METRIC] by at least [EXPECTED LIFT] because [SPECIFIC REASON tied to Wrong Person / Weak Offer / Weak Belief diagnosis].

If you cannot complete this template, you do not have an experiment — you have a guess.

Variant design

Single change: only the H1 sentence. Keep sub-hook, CTA, trust block, and all other page elements identical. Two-variable changes muddle attribution.

Primary metric

Click-through to checkout or signup. Click-rate is more sample-efficient than conversion-to-paid; conversion-to-paid requires 5-10x the sample size.

Secondary metrics (watch but do not decide on)

  • Scroll-depth past the fold (does the new headline keep readers reading?).
  • Bounce rate (does the new headline drive visitors away faster?).
  • Time-on-page (proxy for engagement at the top).

Procedure

  1. Step 1

    Write the hypothesis explicitly

    Using the template above. If you cannot complete it, you do not have an experiment — you have a guess.

  2. Step 2

    Set traffic split with an honest tool

    PostHog, GrowthBook, or a server-side split. Cookie-only client-side splits leak across sessions and produce noisy data.

  3. Step 3

    Run for full week-cycles

    Day-of-week effects are real. End the test at the same day-of-week boundary you started — never mid-week.

  4. Step 4

    Refuse to peek at results before the duration ends

    Peeking shifts your mental model and biases the analysis. Pre-commit to the duration; check at the end.

  5. Step 5

    Run a chi-square or two-proportion test

    Eyeballing 'this looks better' is confirmation bias in numbered form. Use a real statistical test or a calculator (evanmiller.org/ab-testing/sample-size.html).

  6. Step 6

    Decide and ship

    If the variant wins reliably, ship the variant and start the next experiment. If the variant loses, keep the control. If the result is inconclusive (within noise), the experiment was insufficient — usually small sample size.

Self-deceptions to avoid

  • Calling a 200-visitor test 'a result'. At that volume, the noise dominates the signal.
  • Testing two changes at once (headline + sub-hook). Muddles attribution; cannot tell which moved the metric.
  • Stopping the test early because the variant 'looks better'. Early stopping inflates winner detection by 30-50%.
  • Treating peak-traffic days as representative. Weekday-weekend mix matters.
  • Confirming a hypothesis-shaped expectation. The honest test is whether the variant DOES better, not whether it COULD be argued to.

What success looks like

A variant that wins by 25%+ on click-through with 1,000+ visitors per side, statistically significant at p<0.05. Replicable on a fresh cohort if you re-run.

Related benchmark

See the directional range for landing page conversion rate to calibrate the expected lift in your hypothesis.

Frequently asked

What if I don't have 1,000 visitors per variant?
Then you cannot run an honest A/B test on headlines. The right move at lower volume is qualitative testing: show both headlines to 10-20 people in your target audience, ask which they would click and why. Qualitative beats noisy quantitative at small scale.
Can I test headlines on different traffic sources separately?
Yes, and you should be aware sources behave differently. Twitter traffic, Google organic, and direct-link traffic respond differently to the same headline. Pool only if the sources behave similarly historically.

Test on a page that is already pointed in the right direction

A/B tests on a misaligned page produce two losing variants. The diagnostic labels the alignment problem first; the test optimizes within the right alignment.

🚀 Explore Our Network

Full disclosure: UnlockSaaS is one of ten small products built and run by one independent operator. These are the other nine.

60 days
To First Paying Customer
7 steps
Proven Playbook
100%
Money-Back Guarantee
$49
Founding Price /mo

You shipped. Nobody paid. The playbook breaks the pattern or the code refunds you automatically.

Get Free Diagnosis

Refund runs from your dashboard, not a support ticket — the server re-checks eligibility and issues it through Stripe automatically.