Retrospective

Field notes from real experiment reviews and postmortems — the ambiguous rollout calls, the metric traps, and the reasoning that actually resolved them.

Worse on the Surface, Better for the Platform: A Rollout Call Worth Revisiting

A redesigned placement's own click-through rate dropped, but every platform-level guardrail went up. Ship it or roll it back? A look back at how that call actually got made.

RetrospectiveGuardrail MetricsProduct Analytics
Read article →

Hunting for a Goal Metric After the Fact Is Just Multiple Testing in Disguise

When nobody commits to a single goal metric before launch, every hypothesis becomes another roll of the dice. Why that habit is mathematically identical to p-hacking, and how to close it off.

RetrospectiveMultiple TestingGoal Metrics
Read article →

Not Every CTR Drop Means "Exploring Elsewhere" — A Case Where It Didn't

Replacing a 'see more' page with an inline carousel cut a navigation step — and made browsing worse. Why fewer clicks isn't automatically better UX, and why click and like aren't the same signal.

RetrospectiveProduct AnalyticsUX Measurement
Read article →

Every Experiment Doc Lives in Three Different Tools. What If an Agent Just Wrote the One-Pager?

Planning docs, Slack threads, and results all live in different tools with no canonical write-up. An idea for an agent that drafts the one-pager the day an experiment starts, and closes it out when it ends.

RetrospectiveExperiment OpsTooling
Read article →

Two Kinds of Guardrail Tests, Two Different Corrections (and Where I Mixed Them Up)

A standard significance test and a non-inferiority test both get called 'the guardrail test,' but they compound differently as guardrails multiply — one needs Bonferroni, the other needs a power correction.

RetrospectiveGuardrail MetricsStatistics
Read article →

Posterior Expected Value as a Rollout Rule: Untangling What I Was Actually Building

A side-project decision rule for rollout calls, using posterior expected value instead of a p-value — and the four separate ideas (decision rule, dev cost, risk attitude, two metrics) I'd left tangled together in it.

RetrospectiveBayesian StatisticsDecision Rules
Read article →

Winsorization and "Ignore Zero" Are Not the Same Correction

Premium. Dropping non-purchasers and clipping the top of a revenue distribution feel like the same cleanup step. They're not — and doing them in the wrong order quietly damages the exact high-value tail the metric is supposed to capture.

RetrospectiveWinsorizationMetric Design
Read article →

Maximizing Expected Return Isn't the Same Policy as Proving "Not Bad"

Premium. The posterior-positive rollout rule maximizes expected return by construction. Working out why the standard significance test and the non-inferiority test both fall short of it, and by how much.

RetrospectiveBayesian StatisticsDecision Rules
Read article →

Chance to Win ≥ 50% Maximizes Expected Value — Until You Add a Second Metric

Premium. Choosing a simple 'probability of winning' launch rule over frequentist testing, why that simplicity actually raises the realized success rate, and a surprising correction once a second metric joins the decision.

RetrospectiveBayesian StatisticsDecision Rules
Read article →

Going Bayesian for A/B Testing: Two Easy Fits and One That Isn't

Premium. Checking whether a Bayesian rollout rule holds up across a whole experimentation program — the statistics engine and multiple-hypothesis testing fit cleanly, but sequential testing turns out to be a much harder translation.

RetrospectiveBayesian StatisticsSequential Testing
Read article →

Bayesian A/B Testing Doesn't Have Goal and Guardrail Metrics — Should It?

The Bayesian literature confirms a simple posterior-positive launch rule but has almost nothing to say about goal and guardrail metrics. Working through why that split exists, and what happens once there's more than one guardrail.

RetrospectiveBayesian StatisticsGuardrail Metrics
Read article →