How the math works
Every number on this site comes from a formula you can check. This page explains each one and what it assumes.
Paywall A/B tests
Every paywall view ends in one outcome: no purchase, or a purchase of one of the plans. For each variant we model those outcomes as a multinomial with unknown probabilities, and put a Dirichlet prior on them. The prior has weight 1 on "no purchase" and 1/K on each of the K plans. That works out to a flat prior on overall conversion, worth one imaginary view, so it stops mattering once a test has a few hundred views.
Revenue per view is the sum over plans of the probability of buying that plan times the plan's value after the store fee. Because plan values are known, the uncertainty in revenue per view comes entirely from the plan probabilities. We draw 20,000 samples from each variant's posterior with a fixed random seed, so the same inputs always produce the same answer.
From those draws we report:
- Chance it earns most: the share of draws in which the variant has the highest revenue per view.
- Range: the middle 95% of draws (a 95% credible interval).
- Revenue vs control: the relative difference to the first variant, as a median and a 95% range.
- Expected cost of being wrong: the average revenue per view you give up by shipping a variant, counting zero in the draws where it is the best.
We check the single-plan case against the exact closed-form probability for two Beta distributions in our test suite.
When we say "Ship it"
We recommend a variant when it has at least a 95% chance of earning the most and its expected cost of being wrong is under 1% of average revenue per view. If every variant's expected cost is under that 1% line, we say there's no difference worth chasing. Before that, we say keep running. Tests with fewer than 200 views or 10 purchases in a variant are always "too early".
Frequentist checks
For people who need a p-value, we also show a two-sided pooled z-test for conversion and a Welch z-test for revenue per view. Each view's revenue is either zero or a plan value, so its variance follows from the plan counts.
Sample ratio mismatch
We compare the users in each group with the counts your planned split predicts, using a chi-square goodness-of-fit test with one fewer degree of freedom than there are groups. We flag a mismatch when p is below 0.001. That threshold is the common industry convention: a real mismatch nearly always comes from a bug, and a strict cutoff avoids false alarms when you check every day.
Sample size and duration
For conversion, views per variant are
n = (zα/2 √(2 p̄ (1 − p̄)) + zβ √(p1(1 − p1) + p2(1 − p2)))² / (p2 − p1)²
where p1 is the current conversion, p2 is p1 times one plus the change you want to detect, and p̄ is their average. A 5% baseline and a 20% change at 95% confidence and 80% power needs 8,158 views per variant.
For revenue per view we use n = 2 (zα/2 + zβ)² σ² / δ², with σ² the variance of one view's revenue under your current plan mix and δ the change in revenue per view. Revenue varies much more between people than a yes or no purchase does, so revenue tests need more traffic.
Duration is total views needed divided by daily paywall views times the share of traffic in the test, rounded up to whole days.
Lifetime value and payback
For each plan we assume payments at the start of every billing period from day 0 until the end of the time frame. A share of first payments is refunded, and refunded subscribers never renew. Each renewal happens with the same probability. Lifetime value per subscriber is then
LTV = price × (1 − store fee) × (1 − refund rate) × (1 − rN) / (1 − r)
where r is the renewal rate and N the number of billing periods inside the time frame. Per install, we multiply by the share of installs that subscribe. Payback is the first payment day on which cumulative revenue per install covers the cost per install.
What these calculators can't tell you
- They assume each view is independent. If the same person sees the paywall many times, count people, not impressions.
- They don't model free trials converting later. Wait until most trials have ended, or test on trial starts and track trial-to-paid separately.
- A constant renewal rate is a simplification. Real retention usually drops fastest after the first renewal.
- Novelty effects, seasonality and marketing pushes during a test can all move results. Run tests over full weeks.