How the math works

Every number on this site comes from a formula you can check. This page explains each one and what it assumes.

Paywall A/B tests

Every paywall view ends in one outcome: no purchase, or a purchase of one of the plans. For each variant we model those outcomes as a multinomial with unknown probabilities, and put a Dirichlet prior on them. The prior has weight 1 on "no purchase" and 1/K on each of the K plans. That works out to a flat prior on overall conversion, worth one imaginary view, so it stops mattering once a test has a few hundred views.

Revenue per view is the sum over plans of the probability of buying that plan times the plan's value after the store fee. Because plan values are known, the uncertainty in revenue per view comes entirely from the plan probabilities. We draw 20,000 samples from each variant's posterior with a fixed random seed, so the same inputs always produce the same answer.

From those draws we report:

We check the single-plan case against the exact closed-form probability for two Beta distributions in our test suite.

When we say "Ship it"

We recommend a variant when it has at least a 95% chance of earning the most and its expected cost of being wrong is under 1% of average revenue per view. If every variant's expected cost is under that 1% line, we say there's no difference worth chasing. Before that, we say keep running. Tests with fewer than 200 views or 10 purchases in a variant are always "too early".

Frequentist checks

For people who need a p-value, we also show a two-sided pooled z-test for conversion and a Welch z-test for revenue per view. Each view's revenue is either zero or a plan value, so its variance follows from the plan counts.

Sample ratio mismatch

We compare the users in each group with the counts your planned split predicts, using a chi-square goodness-of-fit test with one fewer degree of freedom than there are groups. We flag a mismatch when p is below 0.001. That threshold is the common industry convention: a real mismatch nearly always comes from a bug, and a strict cutoff avoids false alarms when you check every day.

Sample size and duration

For conversion, views per variant are

n = (zα/2 √(2 p̄ (1 − p̄)) + zβ √(p1(1 − p1) + p2(1 − p2)))² / (p2 − p1)²

where p1 is the current conversion, p2 is p1 times one plus the change you want to detect, and p̄ is their average. A 5% baseline and a 20% change at 95% confidence and 80% power needs 8,158 views per variant.

For revenue per view we use n = 2 (zα/2 + zβ)² σ² / δ², with σ² the variance of one view's revenue under your current plan mix and δ the change in revenue per view. Revenue varies much more between people than a yes or no purchase does, so revenue tests need more traffic.

Duration is total views needed divided by daily paywall views times the share of traffic in the test, rounded up to whole days.

Lifetime value and payback

For each plan we assume payments at the start of every billing period from day 0 until the end of the time frame. A share of first payments is refunded, and refunded subscribers never renew. Each renewal happens with the same probability. Lifetime value per subscriber is then

LTV = price × (1 − store fee) × (1 − refund rate) × (1 − rN) / (1 − r)

where r is the renewal rate and N the number of billing periods inside the time frame. Per install, we multiply by the share of installs that subscribe. Payback is the first payment day on which cumulative revenue per install covers the cost per install.

What these calculators can't tell you