Unbalanced A/B Tests: The Smaller Group Decides Everything
95/5 holdouts, 90/10 canary tests and 5% lift studies are legitimate designs. Here is how to size them, check them and analyse them without inventing precision you do not have.
Same data, opposite conclusions
A team runs an unbalanced A/B test: 10,000 users see a new checkout, and a 5% holdout of 500 users keeps the old one. The variant converts at 5%, the holdout at 4%. Someone points out that the groups are “not comparable” because one is 20 times bigger, so they scale the holdout up to 10,000 users at the same 4% rate and run the significance test. Result: p = 0.0006. A clear winner.
Run the same test on the real counts, 500 against 10,000, and you get p = 0.31. Not significant. Nothing in the user behaviour changed between the two analyses; only the arithmetic did. One version ships a feature on evidence that does not exist.
This article covers how unequal traffic splits actually work: why teams use them, what they cost in statistical power, how small the small group can safely be, how to check that the split held, and how to analyse the result honestly. The core message fits in one sentence. In an unbalanced test, the smaller group sets the limit on what you can learn.
Key takeaways
- 1Unequal splits are legitimate. Use them to limit risk (a 5% variant) or cost (a 5% holdout).
- 2They cost power: a 95/5 split needs about 5.3× the total traffic of a 50/50 split.
- 3The small group needs at least about half of what each group would need in a 50/50 test, however large the other group is.
- 4Beyond roughly 4:1, extra traffic in the large group adds almost nothing.
- 5Check sample ratio mismatch against the planned ratio (e.g. 95/5), not 50/50.
- 6Never scale the small group up before testing. It creates false confidence and false winners.
- 7“Not significant” from an unbalanced test often means “couldn't tell”. Report the MDE.
Why teams run 95/5 and 90/10 splits
A 50/50 split gives the most statistical power for a given amount of traffic. Teams still choose unequal splits, and for two good reasons: the change might hurt users, or withholding it costs money. Both are legitimate. Neither changes the statistics.
Small variant, e.g. 95/5
Limit risk
Canary rollouts of risky changes: a new ranking model, new pricing, a major redesign, a payment provider migration. If the change is broken, only a few users notice. Kohavi, Henne and Sommerfield recommend this as treatment ramp-up: for example, start at 0.1%, then step through 0.5%, 2.5% and 10% before reaching 50%, checking for egregious problems at each step.
Small holdout, e.g. 95/5
Limit cost
Ad lift studies and feature holdouts, where every held-out user is a user who does not get the ad or the feature. Google's user-based Conversion Lift lets advertisers set the holdback anywhere from 1% to 50%, and its documentation notes that larger holdouts increase the sample but also have a higher opportunity cost.
Infrastructure teams have a related habit worth copying. Google and Waze recommend comparing a canary not against the whole production fleet but against a fresh baseline deployment that matches the canary in deployment time, number of instances and type of traffic. Their reason is operational (cache warm-up and load balancing differ between old and new instances), but the lesson carries over: the cleanest comparison is between groups that were created the same way, at the same time.
Geo experiments are a different design. When whole regions are chosen deliberately as test and control markets, users are not randomised individually, and the rules below do not apply directly. This article is about user-level randomisation with a fixed, unequal ratio.
The cost of an unequal split: statistical power
An unequal split needs more total traffic than a 50/50 split to detect the same effect. The reason is intuitive once you see where uncertainty comes from. An A/B test estimates the difference between two groups, and that difference inherits the noise of both. The large group's rate is measured precisely; the small group's rate is not. The comparison can never be more precise than its noisiest part, and in an unbalanced test the small group contributes most of the noise.
The penalty grows slowly at first and then very quickly. An 80/20 split costs 60% more traffic; a 95/5 split costs more than five times as much. Kohavi, Henne and Sommerfield give the same rule of thumb and note that a 99/1 split must run about 25 times longer than a 50/50 one.
Traffic needed for the same power
| Split | Small share | Total traffic vs 50/50 |
|---|---|---|
| 50/50 | 50% | 1× |
| 80/20 | 20% | 1.6× |
| 90/10 | 10% | 2.8× |
| 95/5 | 5% | 5.3× |
| 99/1 | 1% | 25.3× |
Same significance level, same power, same effect. Assumes similar per-user variance in both groups.
For the curious: the penalty formula
If w is the small group's share of traffic, the total sample size relative to a 50/50 design is:
At w = 0.5 the denominator is 1, so the penalty is 1. At w = 0.05 it is 1 / (4 × 0.05 × 0.95) = 1 / 0.19 = 5.26.
Where it comes from: the variance of the difference is proportional to 1/nsmall + 1/nlarge. With N total users that equals 1/(wN) + 1/((1 − w)N) = 1 / (N × w × (1 − w)). Setting this equal to the 50/50 value, 4/N50/50, gives N = N50/50 / (4w(1 − w)).
How much traffic does an unequal split cost?
Drag the slider to change the small group's share. The curve shows total traffic needed relative to a 50/50 test with the same power.
A 95/5 split needs 5.26× the total traffic of a 50/50 test to reach the same power. If a 50/50 test would take 2 weeks, this one takes about 10.5 weeks.
Clinical trialists have weighed this trade-off for decades. Torgerson and Campbell, writing in the BMJ, argue that a 2:1 allocation can save substantial money when one treatment is expensive, with only a modest loss of power. The formula agrees: 2:1 needs 12.5% more participants than 1:1. The lesson for growth teams is the same. A mild imbalance is cheap; an extreme one is not.
The half rule: how small can the small group be?
The small group needs at least about half the per-group sample size of a 50/50 test, however large the other group is. If a balanced test needs 40,000 users per group, your holdout needs at least 20,000, even if the treated group has ten million. Martin Bland makes the same point in his teaching notes on clinical trials: half the original group size is the lowest possible limit if you want to preserve power.
This is the most useful rule in this article because it reverses how most teams plan. They pick a percentage first (“let's hold out 5%”) and hope the numbers work. The half rule says to start from the small group's absolute size and let the percentage follow.
The rest of this section uses one piece of shorthand: the ratio k, which is simply how many times bigger the large group is than the small one. Divide the large share by the small share: an 80/20 split has k = 4 (80 ÷ 20, written 4:1), a 90/10 split has k = 9, and a 95/5 split has k = 19. A 50/50 split has k = 1.
For the curious: where the half rule comes from
Let n be the per-group size of the 50/50 design. Precision depends on the variance of the difference, which is proportional to:
To match the 50/50 design, this must be no bigger than the balanced value:
As nlarge grows towards infinity, 1/nlarge shrinks to zero and the condition becomes 1/nsmall ≤ 2/n, which means nsmall ≥ n/2. No amount of traffic in the large group can push the small group below that floor.
For a finite ratio k = nlarge / nsmall (19 for a 95/5 split), solving the same equation gives the standard unequal-allocation formula found in clinical-trial sample size texts:
Diminishing returns from the large group
| Ratio k (split) | Small group | Large group |
|---|---|---|
| 1 (50/50) | 100% of n | 1.0 n |
| 4 (80/20) | 62.5% of n | 2.5 n |
| 9 (90/10) | 55.6% of n | 5.0 n |
| 19 (95/5) | 52.6% of n | 10.0 n |
| ∞ | 50% of n | ∞ |
n = per-group sample size of the equivalent 50/50 test.
Small group size needed as the large group grows
Shown as a percentage of the per-group sample size n from a 50/50 plan. The ratio k is large group ÷ small group: 4:1 is an 80/20 split, 9:1 is 90/10, 19:1 is 95/5. The small group never drops below 50% of n.
Read the table from top to bottom. Going from 1:1 (50/50) to 4:1 (80/20) shrinks the small group by more than a third. Going from 4:1 to 19:1 (95/5) quadruples the large group and saves only another 10 percentage points in the small group. Beyond roughly 4:1, extra traffic in the large group buys almost nothing. Epidemiologists reached the same conclusion long ago, which is why case-control studies rarely use more than four or five controls per case.
Be honest about the assumption
The half rule assumes both groups have similar per-user variance. For conversion rates and small effects that is close enough: at a 4% baseline with a +10% lift, the exact unequal-allocation calculation, which uses each group's own variance, lands within about 2% of the simple rule, depending on which group is small. For averages with very different spreads, such as revenue per user after a pricing change, use each group's own variance, the way Welch's T-test does. In that situation the most efficient split gives more users to the noisier group, in proportion to its standard deviation.
How to plan an unbalanced test, step by step
Here is a worked example you can reproduce with our sample size calculator. A subscription product wants to test a new Pro-upgrade page. Pricing changes are risky, so only 5% of users will see it. The baseline upgrade rate is 4%, and the smallest lift worth acting on is +10% relative (4.0% to 4.4%). Significance is 95% two-sided and power is 80%.
- 1
Find the 50/50 sample size
A balanced test needs 39,476 users per group, 78,952 in total.
- 2
Apply the 95/5 penalty
78,952 × 5.26 ≈ 415,500 users in total.
- 3
Split it
About 20,800 users in the small group (5%) and about 394,800 in the large group (95%).
- 4
Sanity-check with the half rule
20,800 is above half of 39,476 (19,738). Consistent: at 19:1 the small group needs 52.6% of n.
The number that matters is the 20,800. If the product gets 20,000 new users a day, the test runs for about three weeks, not the four days a 50/50 version would take. That is the real price of protecting 95% of users from a risky change, and it is better to know it before launch than to discover it in an inconclusive readout. Try your own numbers below, or set the Traffic Split in the sample size calculator.
Interactive planner
Size an unequal split test
At 20,000 users a day, the 50/50 version finishes in about 4 days; this split needs about 21 days. Round up to whole weeks to cover day-of-week effects.
Assumes a two-sided test at 95% confidence and 80% power, using the same formula as our sample size calculator, then the unequal-allocation adjustment nsmall = n × (1 + 1/k) / 2.
Running it: check that the split held
Sample ratio mismatch (SRM) is a significant difference between the share of users you configured for each group and the share you actually observe. Fabijan and colleagues at Microsoft found it in about 6% of experiments, and in most cases it invalidates the result, because whatever removed users from one group (a redirect, a crash, a bot filter) usually removed a non-random slice of them.
SRM applies to unequal splits exactly as it does to 50/50 ones, with one change: test against the planned ratio. If you configured 95/5, the expected counts are 95% and 5% of the total. Lukas Vermeer's SRM checker FAQ states it plainly: the expected split does not need to be equal.
Why it matters more with a small group
Suppose 200,000 users are assigned at 95/5, and a tracking bug in the variant silently drops 300 of them. That is only 0.15% of all traffic, invisible on any dashboard. But it is 3% of the small group, and the planned-ratio test catches it (p = 0.003). The same 300 users lost from the large group would be a 0.16% dent that barely moves anything (p = 0.88). The small group is where a leak does the most damage, and also where the SRM check is most sensitive.
Why practitioners learn to ignore SRM
Many tools assume 50/50 and test every split against it. A deliberate 95/5 split then fails the check every time, so teams learn to click past the warning, and the one time it signals a real bug, nobody looks. Our SRM checker and significance calculator both have a Traffic Split setting, so they test against the split you actually configured. Try the scenario below.
Interactive check
Sample ratio mismatch against the planned split
Observed small-group share: 4.86% (planned 5%). Chi-square goodness-of-fit test with 1 degree of freedom; flagged when p < 0.01, the same threshold as our SRM checker.
Lift studies: assigned users, not exposed users
In an ad lift study the 95/5 ratio applies to users assigned to each group, not to users who actually saw an ad. Many users in the test group never get an impression, and the ones who do are a non-random subset chosen by the auction. Checking SRM, or computing lift, on “exposed test users vs all control users” breaks the randomisation.
Platforms handle this by comparing like with like. Meta's lift API randomises accounts into test and control percentages that sum to 100. The ghost-ads method described by Johnson, Lewis and Nubbemeyer logs the control users who would have seen the ad, so exposed users can be compared with their true control counterparts.
Analysing it: never rescale the small group
Comparing rates across groups of different sizes is perfectly fine; a conversion rate is already normalised. Inventing sample size is not. When you scale a 500-user holdout to “10,000 users at 4%”, the rate stays the same but the standard error shrinks as if you had observed 20 times more people. The test then reports a confidence it has not earned.
The demo below runs both versions through the same two-proportion Z-test our significance calculator uses. Switch between the two cases and watch the intervals.
Same rates, different sample sizes: real vs scaled control
Variant: 10,000 users at 5%. Control: 500 real users, or the same rate “scaled” to 10,000. Two-sided Z-test at 95% confidence.
| Correct (control = 500) | Scaled (control “= 10,000”) | |
|---|---|---|
| Result | Not significant | Significant |
| p-value | 0.3145 | 0.0006 |
| Z-score | 1.01 | 3.41 |
| 95% CI of the difference | -0.8 to +2.8 pts | +0.4 to +1.6 pts |
Control conversion rate, 95% CI
Difference (variant − control), 95% CI
In the 2% vs 5% case scaling is “only” overconfident: both versions are significant, but the scaled interval (+2.5 to +3.5 points) is less than half as wide as the honest one (+1.7 to +4.3). In the 4% vs 5% case it is worse. The honest difference interval runs from about −0.8 to +2.8 points and includes zero. The scaled one runs from +0.4 to +1.6 and does not. Scaling turned “we can't tell” into a false winner.
The fix is simple: enter the real counts. The Z-test and T-test handle unequal groups natively, because their standard error already contains the 1/n1 + 1/n2 term. There is nothing to correct.
The legitimate kind of scaling
There is one place where scaling belongs: projecting business impact to 100% of users, after the test, using the interval from the real counts. Meta's lift reporting, for instance, returns control conversions both raw and scaled to the test group's size, so incremental conversions can be read on a common footing. The uncertainty still has to come from the raw data.
In the 4% vs 5% example, rolling out to one million users projects 10,000 extra conversions (1 point × 1,000,000). Carrying the honest interval through gives a range from about 7,700 fewer to 27,700 more conversions. That is what the test actually tells you. If you want to translate it into money, our ROI calculator keeps the range attached.
After the test: what could you have detected?
A “not significant” result from an unbalanced test is often not evidence of no effect. It is evidence that the small group was too small to see one. The way to tell the difference is to report the minimum detectable effect (MDE) the test actually had.
For an unequal split, first convert the two group sizes into an equivalent number of users per group: the balanced test that would have the same precision.
For the curious: equivalent users per group
This is the harmonic mean of the two group sizes. For 500 and 10,000 users: 2 / (1/500 + 1/10,000) = 2 / 0.0021 ≈ 952. The 10,500 users in this test carry the same information as a 952 vs 952 balanced test.
The usual quick MDE formula for a conversion rate p, at 95% confidence and 80% power, is:
At a 2% baseline that gives 2.80 × √(2 × 0.02 × 0.98 / 952) ≈ 1.8 points. Treat it as a floor, not an answer. It assumes both groups keep the baseline's variance, which stops being true when the effect is large relative to the baseline. Solving the power equation for the actual two-proportion Z-test (pooled standard error under “no effect”, each group's own variance under the effect) gives +2.4 points, about +118% relative. We checked both by simulating 400,000 tests: +2.4 points is detected 80% of the time, +1.8 points only 58% of the time.
For the small effects most tests are planned around, the two agree closely: in the 95/5 Pro-upgrade example above (+10% relative) the quick formula gives +0.39 points and the exact calculation +0.40. Our MDE calculator uses the exact calculation for conversion rates.
At a 2% baseline, this test could only reliably detect a change to about 4.4%. Go back to the opening example: at a 4% baseline the same 500 vs 10,000 design has an MDE of about +3.0 points (+76% relative). The observed +1 point was never going to be detectable, so p = 0.31 says nothing about whether the new checkout works. If the MDE is bigger than any realistic effect, “not significant” means the test couldn't tell.
The reverse also holds. When an underpowered unbalanced test does come out significant, the estimate is likely to be inflated; see the winner's curse for why, and practical significance for how to judge whether an effect is worth acting on at all.
Interactive check
What could this test have detected?
Solved for the two-proportion Z-test our significance calculator runs, at 95% confidence and 80% power. The quick formula (z0.975 + z0.80) × √(2p(1 − p) / neq) gives +1.80 pts; it assumes both groups keep the baseline's variance, so it is optimistic when the effect is large relative to the baseline.
Checklist for unbalanced A/B tests
Before, during and after
- 1.Decide the smallest effect that matters before the test starts.
- 2.Size the small group first. It must hold at least about half the 50/50 per-group number.
- 3.Avoid going beyond about 4:1 unless risk or cost genuinely requires it.
- 4.Check SRM against the planned ratio, on assigned users.
- 5.Analyse the real counts. Never rescale a group.
- 6.Report the minimum detectable effect alongside any “not significant” result.
- 7.Scale up only to project business impact, after testing, with the uncertainty carried through.
The small group decides what you can learn
Unequal splits are a sensible answer to real constraints. A 5% canary protects users from a broken release; a 5% holdout keeps an ad budget working. What they cannot do is make a small group behave like a big one. Plan around the small group's absolute size, protect it with an SRM check against the planned ratio, and analyse it with its real counts. Do that, and a 95/5 test is as trustworthy as any 50/50 test. It just takes longer.
If you are designing the test from scratch, our A/B test design template has a place to record the split, the MDE and the SRM plan before launch.
References
- Kohavi, R., Henne, R. M., & Sommerfield, D. (2007). Practical Guide to Controlled Experiments on the Web: Listen to Your Customers not to the HiPPO. KDD 2007. Source of the 1/(4p(1 − p)) running-time rule and the treatment ramp-up example. DOI
- Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press. Chapter 21: Sample Ratio Mismatch and Other Trust-Related Guardrail Metrics. DOI
- Fabijan, A., Gupchup, J., Gupta, S., Omhover, J., Qin, W., Vermeer, L., & Dmitriev, P. (2019). Diagnosing Sample Ratio Mismatch in Online Controlled Experiments: A Taxonomy and Rules of Thumb for Practitioners. KDD 2019. DOI
- Bland, M. Allocation ratios in clinical trials. University of York, MSc teaching notes. Derivation of unequal allocation and the half-group-size limit. University of York
- Torgerson, D. J., & Campbell, M. K. (2000). Use of unequal randomisation to aid the economic efficiency of clinical trials. BMJ, 321(7263), 759. PMC
- Chow, S.-C., Shao, J., Wang, H., & Lokhnygina, Y. (2017). Sample Size Calculations in Clinical Research (3rd ed.). Chapman & Hall/CRC. Two-sample formulas with an allocation ratio.
- Johnson, G. A., Lewis, R. A., & Nubbemeyer, E. I. (2017). Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness. Journal of Marketing Research, 54(6). DOI
- Google Ads Help. Set up Conversion Lift based on users (holdback from 1% to 50%). Google
- Meta for Developers. Conversion Lift Measurement, Marketing API (test and control percentages, raw and scaled control results). Meta
- Davidovich, E., & Chamley, T. (2019). Canary analysis: lessons learned and best practices from Google and Waze. Google Cloud Blog. Google Cloud
- Vermeer, L. Sample Ratio Mismatch (SRM) checker: frequently asked questions. lukasvermeer.nl
Found this useful?
Share it with the team that owns your holdouts and canary rollouts.
Related Resources
Plan experiments with proper power analysis.
Analyze your experiment results.
Determine your minimum detectable effect.
Why significant A/B test wins overestimate impact.
Learn when a result actually matters.
Avoid costly errors with charts and checklist.
Frequently Asked Questions
Size the small group first
Get the 50/50 sample size, halve it for your holdout floor, and analyse the real counts when the data comes in.