An Expert Community Where Every Voice Matters


Join Us

Research Card: Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection, and Learning


We discuss a new method for off-policy evaluation, selection, and learning in interactive systems.


Why did we work on this topic (the problem we want to solve)?

In decision-making under uncertainty, offline contextual bandit is a framework that help in making decisions based on historical data without needing real-time feedback. This is particularly relevant for Criteo, where past interactions shape recommendations and advertising strategies. The main challenge is to evaluate and find new, improved policies only using logged data collected under the current policy (the current recommender system).
Traditional methods focus on point estimates of policy performance, which are single values predicting how well a policy might perform. This may be insufficient in critical applications where errors are costly. Thus, there is a pressing need for pessimistic off-policy evaluation (OPE) methods, which are techniques used to evaluate the performance of a policy’s worst-case performance. This enables us to:

  • Safely select and deploy new policies that are guaranteed to meet some performance thresholds.
  • Mitigate risks associated with policy changes, especially in high-stakes environments, where decisions have significant consequences and errors can be costly.

What did we find? What did we achieve?

  • Our method provides the most precise and reliable estimates compared to existing techniques, enabling better prediction of new policy performance before deployment. By consistently reducing uncertainty and variability, it leads to more trustworthy evaluations and more informed, confident decisions, which is crucial for effectively refining policies and minimizing risks when implementing new strategies.
  • Using advanced mathematical frameworks, we’ve established solid theoretical backing for our method.
  • Our extensive experiments show that our method outperforms existing ones across a variety of tasks and situations, confirming its flexibility and practical utility in enhancing decision-making systems in real-world applications.

How did we proceed?

To enhance the reliability of off-policy evaluation, selection, and learning, we undertook the following steps:

  1. Unified Theoretical Framework: We analyzed a broad class of OPE methods within the same unified framework.
  2. Designing Logarithmic Smoothing: Aiming for the estimator with the best theoretical properties, we introduced a new estimator called Logarithmic Smoothing.
  3. Policy Evaluation, Selection and Learning Strategies: We used the new estimator to develop improved strategies for off-policy evaluation (OPE), selection (OPS) and learning (OPL).
  4. Experimental Validation: We conducted extensive experiments on OPE, OPS, and OPL tasks to validate the practical effectiveness and versatility of Logarithmic Smoothing.

In simple terms, we carried out a series of detailed tests to check if our theoretical ideas were correct. We looked at our new methods and compared them to existing ways of evaluating performance, known as empirical evaluation bounds. By using different sets of data and various situations, we showed that our new approach was better than the older methods. This testing is very important because it helps prove that our proposed methods are trustworthy and can be relied upon in real-world applications.

What is the originality here?

  • Novel Estimator Design.
  • Unified Theoretical Framework.
  • Precise and Rigorous Analysis.
  • Practical Impact on Critical Applications.

Check all our research cards 👇
https://medium.com/criteo-engineering/research-cards/home