Field notes from real experiment reviews and postmortems — the ambiguous rollout calls, the metric traps, and the reasoning that actually resolved them.
A redesigned placement's own click-through rate dropped, but every platform-level guardrail went up. Ship it or roll it back? A look back at how that call actually got made.
When nobody commits to a single goal metric before launch, every hypothesis becomes another roll of the dice. Why that habit is mathematically identical to p-hacking, and how to close it off.
Replacing a 'see more' page with an inline carousel cut a navigation step — and made browsing worse. Why fewer clicks isn't automatically better UX, and why click and like aren't the same signal.
Planning docs, Slack threads, and results all live in different tools with no canonical write-up. An idea for an agent that drafts the one-pager the day an experiment starts, and closes it out when it ends.
A standard significance test and a non-inferiority test both get called 'the guardrail test,' but they compound differently as guardrails multiply — one needs Bonferroni, the other needs a power correction.
A side-project decision rule for rollout calls, using posterior expected value instead of a p-value — and the four separate ideas (decision rule, dev cost, risk attitude, two metrics) I'd left tangled together in it.
Premium. Dropping non-purchasers and clipping the top of a revenue distribution feel like the same cleanup step. They're not — and doing them in the wrong order quietly damages the exact high-value tail the metric is supposed to capture.
Premium. The posterior-positive rollout rule maximizes expected return by construction. Working out why the standard significance test and the non-inferiority test both fall short of it, and by how much.
Premium. Choosing a simple 'probability of winning' launch rule over frequentist testing, why that simplicity actually raises the realized success rate, and a surprising correction once a second metric joins the decision.
Premium. Checking whether a Bayesian rollout rule holds up across a whole experimentation program — the statistics engine and multiple-hypothesis testing fit cleanly, but sequential testing turns out to be a much harder translation.
The Bayesian literature confirms a simple posterior-positive launch rule but has almost nothing to say about goal and guardrail metrics. Working through why that split exists, and what happens once there's more than one guardrail.