I keep coming back to the same question about the posterior-expected-value rollout rule from a couple of weeks ago: if the goal is genuinely to maximize expected return across every experiment you run, how does that rule actually compare to the two testing regimes everyone already uses? I can state the mechanism now. I still haven't found the clean, one-line way to say it out loud.
Three Policies, Side by Side
If the objective is purely to maximize expected return, the best policy is: for every experiment, ship it whenever its posterior expected return is positive...