Drag to feel the tradeoff: lower confidence calls tests sooner and is wrong more often. 95% is the convention for a reason.
Method: pooled two-proportion z-test (two-sided) for the verdict gate, Wilson score intervals for the ranges, Bayesian probability-to-beat with flat Beta(1,1) priors, and the classic two-proportion sample-size formula for the finish line. All computed in code; verified against scipy and statsmodels. Nothing you send is stored.