An experiment readout that turns two conversion counts into a relative lift, its confidence interval, a significance verdict and a sample-size progress line — with every degenerate input handled as its own branch.
Preview in your theme
Loading preview…
This component is part of Pro
The install command, prompt and concept map are included with a Pro license — one purchase unlocks every Pro component, forever.
ready · the primary metric is up 6.7% but its 95% interval still crosses zero, while a guardrail moved 29% the wrong way and its interval does not — the inconclusive win and the conclusive regression sit in the same panel
Checkout redesign
Web · 12–26 March 2026 · 50/50 split
Inconclusive so far1 guardrail breached95% confidence
Planned sample48,282 of 60,000 exposures · 80%
About 11.7K more exposures are needed before the planned 30,000 per arm is reached.
Primary metric
Checkout completionSessions that reached the order confirmation page
Not significant
Control
4.98%
1,204 of 24,180 converted
Variant B
5.31%
1,281 of 24,102 converted
+6.7%relative to Control
95% CI -1.2% to +14.7%
+0.34 pp absolute · p = 0.095
no change
The interval still crosses zero (95% CI -1.2% to +14.7%): the true effect could be anywhere inside it, in either direction. Not significant is not the same as no effect — about 11.7K more exposures are planned before this is due to be read.
Guardrails
Refund rateRefund requested within 14 days, per exposed session
Breached
Control
1.25%
302 of 24,180 converted
Variant B
1.61%
389 of 24,102 converted
+29.2%relative to Control
95% CI +12.3% to +46.2%
+0.37 pp absolute · p < 0.001
no change
The whole interval (95% CI +12.3% to +46.2%) stays on the worse side of zero, so a change this size is unlikely to be noise.
Support contact rateSessions that opened a support conversation
Holding
Control
6.24%
1,510 of 24,180 converted
Variant B
6.18%
1,489 of 24,102 converted
-1.1%relative to Control
95% CI -8.0% to +5.8%
-0.07 pp absolute · p = 0.761
no change
The interval still crosses zero (95% CI -8.0% to +5.8%): the true effect could be anywhere inside it, in either direction. Not significant is not the same as no effect — about 11.7K more exposures are planned before this is due to be read.
Payment failuresSessions with at least one declined payment
Holding
Control
0.71%
171 of 24,180 converted
Variant B
0.65%
156 of 24,102 converted
-8.5%relative to Control
95% CI -29.2% to +12.2%
-0.06 pp absolute · p = 0.422
no change
The interval still crosses zero (95% CI -29.2% to +12.2%): the true effect could be anywhere inside it, in either direction. Not significant is not the same as no effect — about 11.7K more exposures are planned before this is due to be read.
An interval that includes zero also includes every effect it fails to exclude: “not significant” means undecided, not proven equal.
ready · every degenerate input in one panel — 0/0 exposures, zero conversions in both arms, a 0% baseline, byte-identical arms, 0% vs 100% separation, a 2-event sample and impossible counts. No NaN, no Infinity, no −100% invented from a missing denominator; with no planned sample on file the progress bar is absent rather than fabricated
Arithmetic edges
Each row is an input that turns the textbook formula into a division by zero, a negative variance or a confident lie
Primary metric up95% confidence
0 exposures in Control and 0 in Variant. No planned sample size is on file, so this panel cannot say whether the test is due to be read.
Primary metrics
Just started0 / 0 in both arms — no rate exists yet, and 0/0 is not 0%
No exposures
Control
—
0 of 0 converted
Variant
—
0 of 0 converted
—nothing to estimate yet
Nobody has been exposed in Control or Variant yet, so there is no rate to compare.
No conversions yetNobody converted anywhere — the standard error is exactly 0, so no interval is drawn
No interval
Control
0.00%
0 of 8,400 converted
Variant
0.00%
0 of 8,390 converted
0.00 ppabsolute — Control never converted, so there is no baseline to divide by
Both arms sit at exactly the same boundary rate, so there is no spread to build an interval from — nothing has been measured yet.
Feature nobody used before0% baseline — a relative lift would divide by zero, so the absolute change is reported instead
Significant lift
Control
0.00%
0 of 5,000 converted
Variant
0.94%
47 of 5,010 converted
+0.94 ppabsolute — Control never converted, so there is no baseline to divide by
95% CI +0.67 pp to +1.21 pp
p < 0.001 · under 5 events in a cell — the approximation is rough
no change
The whole interval (95% CI +0.67 pp to +1.21 pp) stays on the better side of zero, so a change this size is unlikely to be noise.
Identical armsByte-identical arms — a real 0.0% with a real interval around it, not an error
Not significant
Control
5.00%
600 of 12,000 converted
Variant
5.00%
600 of 12,000 converted
0.0%relative to Control
95% CI -11.0% to +11.0%
0.00 pp absolute · p = 1.000
no change
The interval still crosses zero (95% CI -11.0% to +11.0%), and no planned sample size is on file, so nothing says when this run becomes conclusive. Not significant is not the same as no effect.
Complete separation0% vs 100% — a zero-width interval would otherwise be reported as a significant 100 pp gap
No interval
Control
0.00%
0 of 400 converted
Variant
100.00%
400 of 400 converted
+100.00 ppabsolute — Control never converted, so there is no baseline to divide by
One arm is at 0% and the other at 100%. A normal approximation has no spread to work with here, so treat the gap as unverified until the counts leave the boundary.
Tiny sampleA tripling with only 2 and 6 events behind it — huge point estimate, useless interval
Not significant
Control
3.33%
2 of 60 converted
Variant
10.34%
6 of 58 converted
+210.3%relative to Control
95% CI -61.4% to +482.1%
+7.01 pp absolute · p = 0.129 · under 5 events in a cell — the approximation is rough
no change
The interval still crosses zero (95% CI -61.4% to +482.1%), and no planned sample size is on file, so nothing says when this run becomes conclusive. Not significant is not the same as no effect.
Impossible countsMore conversions than exposures — rejected by the contract, and guarded again at render time
Invalid counts
Control
—
120 of 100 converted
Variant
—
40 of 100 converted
—nothing to estimate yet
These counts cannot be right — a cell is negative or reports more conversions than exposures, so nothing is computed from them.
An interval that includes zero also includes every effect it fails to exclude: “not significant” means undecided, not proven equal.
ready · confidenceLevel={0.9} narrows every interval and relabels the pill — one number drives the CI text, the bars and the verdicts together. 12 metrics in means 12 metrics out (no cap, no fixed-height clipping), and both arms are already past the planned sample
Bulk readout
12 metrics, deterministic counts derived from the row index
Primary metric down90% confidence
Planned sample12,040 of 10,000 exposures · 120%
Both arms passed the planned 5,000 exposures, so the result is due to be read.
Primary metrics
Tracked metric 01
Significant drop
Baseline
4.00%
240 of 6,000 converted
Candidate
3.39%
205 of 6,040 converted
-15.1%relative to Baseline
90% CI -29.3% to -1.0%
-0.61 pp absolute · p = 0.078
no change
The whole interval (90% CI -29.3% to -1.0%) stays on the worse side of zero, so a change this size is unlikely to be noise.
Tracked metric 02
Not significant
Baseline
5.01%
313 of 6,250 converted
Candidate
5.50%
346 of 6,290 converted
+9.8%relative to Baseline
90% CI -3.2% to +22.9%
+0.49 pp absolute · p = 0.216
no change
The interval still crosses zero (90% CI -3.2% to +22.9%), and the planned sample is already in: any real effect is smaller than this test can resolve. That is still not the same as no effect.
Tracked metric 03
Not significant
Baseline
6.00%
390 of 6,500 converted
Candidate
6.50%
425 of 6,540 converted
+8.3%relative to Baseline
90% CI -3.3% to +19.9%
+0.50 pp absolute · p = 0.240
no change
The interval still crosses zero (90% CI -3.3% to +19.9%), and the planned sample is already in: any real effect is smaller than this test can resolve. That is still not the same as no effect.
Tracked metric 05
Not significant
Baseline
8.00%
560 of 7,000 converted
Candidate
8.49%
598 of 7,040 converted
+6.2%relative to Baseline
90% CI -3.4% to +15.7%
+0.49 pp absolute · p = 0.287
no change
The interval still crosses zero (90% CI -3.4% to +15.7%), and the planned sample is already in: any real effect is smaller than this test can resolve. That is still not the same as no effect.
Tracked metric 06
Not significant
Baseline
4.00%
290 of 7,250 converted
Candidate
4.50%
328 of 7,290 converted
+12.5%relative to Baseline
90% CI -1.3% to +26.2%
+0.50 pp absolute · p = 0.136
no change
The interval still crosses zero (90% CI -1.3% to +26.2%), and the planned sample is already in: any real effect is smaller than this test can resolve. That is still not the same as no effect.
Tracked metric 07
Significant drop
Baseline
5.00%
375 of 7,500 converted
Candidate
4.40%
332 of 7,540 converted
-11.9%relative to Baseline
90% CI -23.3% to -0.6%
-0.60 pp absolute · p = 0.084
no change
The whole interval (90% CI -23.3% to -0.6%) stays on the worse side of zero, so a change this size is unlikely to be noise.
Tracked metric 09
Not significant
Baseline
7.00%
560 of 8,000 converted
Candidate
7.50%
603 of 8,040 converted
+7.1%relative to Baseline
90% CI -2.5% to +16.8%
+0.50 pp absolute · p = 0.222
no change
The interval still crosses zero (90% CI -2.5% to +16.8%), and the planned sample is already in: any real effect is smaller than this test can resolve. That is still not the same as no effect.
Tracked metric 10
Not significant
Baseline
8.00%
660 of 8,250 converted
Candidate
7.39%
613 of 8,290 converted
-7.6%relative to Baseline
90% CI -16.1% to +1.0%
-0.61 pp absolute · p = 0.144
no change
The interval still crosses zero (90% CI -16.1% to +1.0%), and the planned sample is already in: any real effect is smaller than this test can resolve. That is still not the same as no effect.
Tracked metric 11
Not significant
Baseline
4.00%
340 of 8,500 converted
Candidate
4.50%
384 of 8,540 converted
+12.4%relative to Baseline
90% CI -0.3% to +25.1%
+0.50 pp absolute · p = 0.108
no change
The interval still crosses zero (90% CI -0.3% to +25.1%), and the planned sample is already in: any real effect is smaller than this test can resolve. That is still not the same as no effect.
Guardrails
Tracked metric 04
Holding
Baseline
7.01%
473 of 6,750 converted
Candidate
6.41%
435 of 6,790 converted
-8.6%relative to Baseline
90% CI -18.7% to +1.5%
-0.60 pp absolute · p = 0.162
no change
The interval still crosses zero (90% CI -18.7% to +1.5%), and the planned sample is already in: any real effect is smaller than this test can resolve. That is still not the same as no effect.
Tracked metric 08
Holding
Baseline
6.00%
465 of 7,750 converted
Candidate
6.50%
506 of 7,790 converted
+8.3%relative to Baseline
90% CI -2.4% to +18.9%
+0.50 pp absolute · p = 0.202
no change
The interval still crosses zero (90% CI -2.4% to +18.9%), and the planned sample is already in: any real effect is smaller than this test can resolve. That is still not the same as no effect.
Tracked metric 12
Holding
Baseline
5.01%
438 of 8,750 converted
Candidate
5.49%
483 of 8,790 converted
+9.8%relative to Baseline
90% CI -1.3% to +20.8%
+0.49 pp absolute · p = 0.146
no change
The interval still crosses zero (90% CI -1.3% to +20.8%), and the planned sample is already in: any real effect is smaller than this test can resolve. That is still not the same as no effect.
An interval that includes zero also includes every effect it fails to exclude: “not significant” means undecided, not proven equal.
loading · the skeleton keeps the real anatomy (header, sample bar, metric cards) so nothing jumps when the counts land
Loading experiment results
empty · the experiment exists but no exposures have been logged yet
No experiment results yet
Once the first exposures are logged, this panel reports the lift, its interval and how much of the planned sample is in.
error · onRetry is wired to real local state — requested 0 times. No partial numbers are shown, because a half-loaded experiment reads as a real result
Couldn't load this experiment
The results service didn't respond. No partial numbers are shown, because a half-loaded experiment reads as a real result.