Decision Making
-

Research Card: Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection, and Learning
We discuss a new method for off-policy evaluation, selection, and learning in interactive systems. Title: Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection, and Learning Short Title: A New Principled Method for Improved Evaluation, Selection, and Learning in Interactive Systems Authors: Otmane Sakhi (Criteo AI Lab, France), Imad Aouali (CREST, ENSAE; Criteo AI Lab, France), Pierre…
