QR Code Generator
Generate QR codes for websites, Wi-Fi networks, text, email, phone numbers, SMS messages, contacts, events, and locations.
The inputted text is too long to include in a share link.
Compare two A/B test variants with conversion rates, lift, p-value, confidence, winner guidance, and distribution charts.
Statistical significance estimates whether the observed conversion-rate difference is larger than you would expect from random variation alone. It does not prove that a result will hold forever, and it does not measure whether the lift is valuable enough for the business.
Use this calculator after the test has collected enough clean, randomized traffic. Avoid repeatedly checking early results and stopping the test the moment the calculator shows a favorable number.
p = \frac{x}{n}
x is conversions and n is visitors for a variant.
\hat{p} = \frac{x_A + x_B}{n_A + n_B}
The pooled rate is used in the standard error for the null hypothesis.
z = \frac{p_B - p_A}{\sqrt{\hat{p}(1 - \hat{p})(\frac{1}{n_A} + \frac{1}{n_B})}}
Positive z-scores mean variant B converted better than variant A.
The calculator uses a two-sided two-proportion z-test for conversion counts. The p-value estimates how surprising the observed difference would be if both variants had the same true conversion rate. The confidence shown here is one minus the p-value, and the note names the variant with the higher observed conversion rate.
The result-panel range bars show an approximate 95% range around each observed conversion rate. The distribution chart uses each variant's standard error to show how much the estimates overlap. This method is appropriate for simple binary outcomes such as signup, purchase, lead, or click conversion.
This is not a substitute for revenue-per-user tests, sequential testing plans, or experiments with non-random traffic assignment.
The pooled proportion, standard error, z statistic, and two-sided p-value follow the National Institute of Standards and Technology's difference-of-proportions test. NIST's discussion of critical values and p-values provides interpretation context; this calculator does not correct for repeated peeking or multiple comparisons.
Define one primary outcome, the eligible population, assignment method, and stopping rule before looking at results. Changing the metric or audience after seeing the data increases the chance of finding a favorable story by accident. Record guardrail metrics such as errors, refunds, unsubscribes, or latency when a variant could improve conversion while harming the broader experience.
Statistical significance is not practical importance. A large sample can detect a tiny lift that is too small to cover engineering, media, or operational cost. Review the absolute conversion difference and confidence interval as well as the p-value. Decide the smallest effect worth acting on before the test begins.
Repeatedly checking and stopping when the p-value crosses a threshold changes the false-positive behavior of a fixed-horizon test. Let the planned sample and duration complete unless the design uses a valid sequential method. Include enough calendar time to cover normal weekday, weekend, promotion, and billing patterns.
The two groups must be comparable. Broken randomization, users appearing in both variants, bot traffic, missing events, and unequal exposure can create a precise answer to the wrong question. Check assignment counts, instrumentation, and baseline characteristics before interpreting the winner.
A conversion result assumes independent observations more strongly than many product datasets allow. Multiple sessions from one person, household clustering, marketplace interactions, or organization-level assignment can understate uncertainty. Analyze at the assignment unit or use a method that accounts for clustering.
Do not treat a non-significant result as proof that the variants are identical. The data may still allow effects that matter. Read the confidence interval and ask whether it rules out both meaningful benefit and harm. If not, the experiment may be inconclusive rather than negative.
Replicate important wins when novelty, seasonality, or implementation risk is high. A result from one audience and period may not generalize to another. Save the hypothesis, screenshots, sample rules, exclusions, raw counts, dates, and decision so later teams can distinguish evidence from a recycled headline.
Built and maintained by utilkit. Updated . Found an issue? Send corrections to contact@utilkit.com
Generate QR codes for websites, Wi-Fi networks, text, email, phone numbers, SMS messages, contacts, events, and locations.
Check SPF, DKIM, DMARC, and MX DNS records for a domain to identify email authentication, alignment, and deliverability issues.
Build and review robots.txt rules for documented AI training crawlers, AI search crawlers, and user-triggered retrieval clients.