Marketing / Testing

A/B Test Significance Calculator (p-Value and Sample Size)

Enter the visitors and conversions for version A and version B to see whether the difference is statistically significant, with the p-value, the relative change, a confidence interval for the difference, and the visitors each group needs to detect the lift you care about.

A/B Test Significance Calculator (p-Value and Sample Size): A two-proportion z-test compares the rates. With 200 conversions from 5,000 visitors in A (4%) and 230 from 5,000 in B (4.6%), z = 1.48 and p = 0.139, so the 15% lift is not significant at 95%; the true difference could be anywhere from −0.20 to +1.40 percentage points. Detecting a 20% lift on a 4% rate needs 10,317 visitors per group at 80% power. Runs 100% locally in your browser with zero server file uploads.

Runs
In your browser
Cost
Free · no sign-up
Availability
Ready to use
A/B test significance calculatorLocal processing

Runs entirely in your browser

ResultNot significantp = 0.1392 at 95% confidence
Conversion rates4% → 4.6%relative change +15%
Difference, 95% interval-0.2% to 1.4%in percentage points
Visitors needed per group10,317to detect a 20% lift with 80% power

The test compares the two conversion rates with a two-proportion z-test: if the variant made no difference, a gap as large as this would appear with probability p. Below 0.05 is significant at 95% confidence. Decide the sample size before you start and do not stop the test as soon as it looks significant, which inflates false positives. The interval shows the plausible range of the true difference.

The formulas

z = (p_B − p_A) ÷ √(p̄(1 − p̄)(1/n_A + 1/n_B)) with the pooled rate p̄, and the interval uses the unpooled standard error. The example matches statsmodels' proportions_ztest and confint_proportions_2indep.

Surveys and polls

To size a survey for a margin of error rather than a test, use the sample size calculator.

How to use it

  1. Enter the visitors and conversions for A and B.
  2. Choose the confidence level.
  3. Read the result, and check the sample size needed for the lift you want to detect.

Privacy & limitations

Everything is calculated in your browser.

Related tools

Frequently asked questions

What does p = 0.139 mean?

If B were really no different from A, a gap at least this large would turn up about 14% of the time by chance, too often to rule chance out.

Can I stop the test as soon as it is significant?

No: checking repeatedly and stopping at the first significant result greatly raises the false-positive rate. Fix the sample size first, or use a method designed for continuous monitoring.

Why is my sample size so large?

Small lifts on low conversion rates are hard to tell from noise; halving the lift you want to detect roughly quadruples the visitors needed.

Free tool · runs in your browser · no account required