Maximizing Expected Return Isn't the Same Policy as Proving "Not Bad"

I keep coming back to the same question about the posterior-expected-value rollout rule from a couple of weeks ago: if the goal is genuinely to maximize expected return across every experiment you run, how does that rule actually compare to the two testing regimes everyone already uses? I can state the mechanism now. I still haven't found the clean, one-line way to say it out loud.

Three Policies, Side by Side

If the objective is purely to maximize expected return, the best policy is: for every experiment, ship it whenever its posterior expected return is positive...

💎

Premium membership required

Upgrade to premium to access the full article.