Statistics & methodology
How MaxLift computes results — and how to audit it yourself.
MaxLift is deliberately transparent about its statistics. Every method traces to published literature, and you can inspect the exact SQL we run against your warehouse.
The readout
For each experiment we compute per-variant aggregates in your warehouse (users, conversions, sums), then: a two-proportion z-test for binomial metrics (Welch t for continuous), a relative-lift confidence interval via the delta method, and a sample-ratio-mismatch (SRM) chi-square check weighted by your design split.
Sequential testing (always-valid p-values)
Because MaxLift recomputes results continuously, a fixed-horizon p-value alone would inflate false positives under repeated looks. Every 'winner' verdict must also clear an always-valid p-value from a mixture sequential probability ratio test (mSPRT) — the same family GrowthBook and Optimizely use — so you can watch a running experiment without peeking penalties. (Johari, Koomen, Pekelis, Walsh — KDD 2017.)
Bayesian readout
Alongside the frequentist test, every variant carries a Bayesian beta-binomial readout: probability it beats control, expected loss if you ship it, and a 95% credible interval on the conversion rate — the questions a growth team actually asks, in plain probability.
CUPED variance reduction
When you map a pre-experiment covariate (e.g. a unit's prior 30-day spend), MaxLift applies CUPED to cut variance and reach significance faster — often halving the required sample at a covariate correlation of 0.7. The covariate is summed inside your warehouse as sufficient statistics; no row-level data ever leaves it. (Deng, Xu, Kohavi, Walker — WSDM 2013.)
Sample-ratio mismatch
If observed traffic split deviates from expected at p<0.001, results are quarantined — a bucketing or logging bug, not a real effect. This is one of the most common ways teams fool themselves; MaxLift blocks it rather than showing a corrupt verdict. (Fabijan et al., KDD 2019.)
The horizon rule
In fixed-horizon mode, no decision statistic is shown until the planned sample size or end date. Peeking at p-values and stopping on a 'win' produces false positives ~20%+ of the time; the Launch Gate computes the horizon up front so you don't.
Roadmap
- •Holdouts / mutual exclusion (namespaces)
- •Multi-metric false-discovery-rate correction
Audit it yourself
Open any experiment's results and click 'View SQL' to see the exact query. Run it in your warehouse and you get the same numbers — no second source of truth.