← L3vlup Labs
Free · no sign-up

A/B Test Designer

The analytical round in product, APM, BizOps and data interviews nearly always lands here. The failure mode is rarely arithmetic: most candidates can compute a p-value and cannot say what would change their mind. Size a test, watch the sample explode as the effect you are chasing shrinks, then read a result and say honestly what it entitles you to claim.

Includes the question everyone gets asked — a 2% lift at p = 0.06, what do you do.

Sample needed per arm
39,473
78,946 total · 7 days at 12,000/day
Detecting
+10%
4.0% → 4.40%
Absolute effect
0.40pp
what you are chasing
Power
80%
α 5%

Design the test

7 days, which is 1 full week. Run whole weeks — weekday and weekend users are different populations.

Why the sample explodes

Sample size scales with the inverse square of the effect. Halve the lift you want to detect and you need roughly four times the users. This is the single most useful thing to know in the room, because it turns “let’s just test it” into a costed decision.

DetectPer armDaysvs current
+2%950,88415924.1×
+3%424,6177110.8×
+5%154,302263.9×
+7.5%69,377121.8×
+10%39,47371.0×
+15%17,94030.5×
+20%10,31420.3×
+30%4,78010.1×

Read a result

Control
Variant
Relative lift+10.0%
Control rate4.00%
Variant rate4.40%
Absolute lift+0.40pp
p-value0.0585
95% interval-0.01 to 0.81pp
Achieved power47%
Not significant, at p = 0.059. This is the case interviewers love, and the wrong answer is “run it longer until it crosses” — that is peeking, and it inflates your false-positive rate badly. The right answer names the decision: was the test powered for an effect you would act on, and is the interval narrow enough to rule out a lift worth having?

Four ways a test lies to you

Peeking

Checking daily and stopping when it crosses 5% turns a 5% false-positive rate into something closer to 30%. Fix the sample size before you start, or use a sequential test designed for it.

Multiple comparisons

Twenty metrics at 5% means one lights up by chance. Name your primary metric before the test runs; everything else is a guardrail or a hypothesis for next time.

Novelty and primacy

Users react to change itself. Early lifts fade; early drops recover. A week is usually the minimum, and a returning-user cut tells you which one you are seeing.

Sample ratio mismatch

If the split is meant to be 50/50 and it is 51/49 at scale, something is broken in the assignment and the result cannot be trusted. Check it first, every time.

Share your result 🧪

𝕏in💬🤖

Instagram and TikTok have no desktop share link, so copy the caption and paste it into the app.