| Abstract: |
A/B tests are standard in firm decision making. In the standard pipeline,
experimental data is converted to a deployment decision by applying a t-test
of the difference in means (the lift) and deploying the treatment if lift is
positive and statistically significant. This common workflow answers the wrong
question. We argue that firms need a decision rule for economic payoffs in the
future deployment environment, not a test of equality in the experimental
sample. We develop an ambiguity-averse decision framework in which each arm is
evaluated by its ambiguity-penalized value over distributions close to the
experimental outcome distribution. The resulting rule has a simple closed form
thanks to the Donsker-Varadhan representation and it requires only the outcome
data from a standard A/B test plus one interpretable parameter governing trust
in the experiment. Our rule is thus no more difficult to implement than a
t-test. A mean-variance approximation shows how the rule penalizes
variability, while a connection to utility maximization shows it to be a
certainty equivalent. We are able to perform a real-world evaluation of our
proposed rule in the context of digital marketing using an archive of 552
advertising experiments from an anonymous US-based online platform. The
proposed rule substantially reduces regret relative to conventional hypothesis
testing. The results show that economically conservative, distribution-aware
deployment rules can outperform statistical-significance rules in digital
experimentation. |