Push the gate
Run A stays fixed at counts 1 / 2 / 17 (mean 0.900), the trusted baseline. Drag run B's buckets, its 20 items redistribute across wrong, partial and correct, and all three checks recompute live.
Check 1 · CI threshold gate (bar = 0.70)
Check 3 · drift (PSI against run A)
This panel resamples run B's current 20 items 2,500 times per update to build a live 95% bootstrap interval, so its numbers land close to, but not always bit-identical to, the companion script's fixed-seed n_boot=10,000 run (which reads [0.550, 0.875] for the 3/5/12 example). PSI floors every bucket proportion at 1e-6 so an empty bucket never sends ln(pA/pB) to negative infinity. Drift reads stable below 0.10, moderate from 0.10 to 0.25, and significant above 0.25; the gate blocks whenever the CI check fails or drift is significant.