An Expert Community Where Every Voice Matters


Join Us

Highlights from NeurIPS 2022


Criteo was pleased to send seven researchers to NeurIPS 2022 in the beautiful and entertaining city of New Orleans. It was the first NeurIPS in person in 3 years and was now running in hybrid mode. There were some modifications to the program, which was now a single track with keynotes, but all accepted papers were presented by posters only.

This year seven papers were accepted with Criteo co-authors:

  1. Alexandre Ramé (Sorbonne Université), Matthieu Kirchmeyer (Sorbonne Université & Criteo AI Lab), Thibaud Rahier (Criteo AI Lab), Alain Rakotomamonjy (Université de Rouen LITIS & Criteo AI Lab), Patrick Gallinari (Sorbonne Université & Criteo AI Lab), Matthieu Cord (Sorbonne Université & Valeo.ai) “Diverse weight averaging for out-of-distribution generalization” Diverse Weight Averaging for Out-of-Distribution Generalization
  2. V. Cabannes, F. Bach, V. Perchet, A. Rudi “Active Labeling: Streaming Stochastic Gradients” Active Labeling: Streaming Stochastic Gradients
  3. N. Kotelevskii (Skoltech), M. Vono (Criteo AI Lab), E. Moulines (Polytechnique), A. Durmus (ENS Paris Saclay), “FedPop: A Bayesian Approach for Personalised Federated Learning”, link: FedPop: A Bayesian Approach for Personalised Federated Learning
  4. Angeliki Giannou, Kyriakos Lotidis, Panagiotis Mertikopoulos, and Emmanouil Vasileios Vlatakis-Gkaragkounis, “On the convergence of policy gradient methods to Nash equilibria in general stochastic games”.
  5. Yu-Guan Hsieh, Kimon Antonakopoulos, Volkan Cevher, and Panagiotis Mertikopoulos, “No-regret learning in games with noisy feedback: Faster rates and adaptivity via learning rate separation”. No-Regret Learning in Games with Noisy Feedback: Faster Rates and…
  6. T.Moreau, M. Massias, A. Gramfort, et al. (with A. Rakotomamonjy), “Benchopt: Reproducible, efficient and collaborative optimization benchmarks”, Benchopt: Reproducible, efficient and collaborative optimization benchmarks
  7. NeurIPS 2022 — Systems Datasets and Benchmarks Track: Florent Bonnet (Extrality and Sorbonne University), Jocelyn Ahmed Mazari (Extrality), Paola Cinnella (Sorbonne University), Patrick Gallinari (Sorbonne Université & Criteo AI Lab), AirfoilRANS: High Fidelity Computational Fluid Dynamics Dataset for Approximating Reynolds-Averaged-Navier–Stokes Solutions https://openreview.net/pdf?id=Zp8YmiQ_bDC

We also had one paper in the AI4Science workshop: Continuous PDE Dynamics Forecasting with Implicit Neural… with Yuan Yin, Matthieu Kirchmeyer, Jean-Yves Franceschi, Alain Rakotomamonjy, and Patrick Gallinari.

David Rohde was also pleased to be one of the co-organizers of The third I can’t believe it’s not better Workshop 2022, which focuses on learning via falsificationism.

Finally, we would like to highlight the work of Giulia Romano from Politecnico di Milano also presented a very interesting paper, A Unifying Framework for Online Optimization with Long-Term Constraints, which has lots of applications for optimizing display advertising. Criteo is very pleased to welcome Giulia for an internship during her Ph.D., which we have very high expectations for.

In what follows, we present some papers we liked and that are relevant to some of Criteo’s research and R&D projects.

Papers on PAC-Bayesian Learning Theory

PAC-Bayesian learning is a non-Bayesian theory of generalization relevant to bidding and recommendation problems at Criteo.

NeurIPS 2022 saw some exciting developments in the theory as a way to explain the generalization ability of deep neural networks. This allows practitioners to have certifications on the performance of their favorite model before deployment. A line of work focused on providing tight bounds (Biggs et al. 2022; Get et al. 2022) for particular use cases, and another interesting development was extending the application of PAC-Bayesian learning to the online setting, which can be of great interest to all practitioners (Haddouche and Guedj 2022).

Papers on Federated Learning and Incentiveness

Federated Learning (FL) is a framework for learning while keeping data on each client and helps in coping with privacy issues. However, one crucial question about FL, especially when clients have disparities in the volume of data they share, is incentiveness. Why would one client share their data, and what would they gain in doing so? Several papers have tried to address this question and have looked at the importance of some samples through the notion of Shapley values. (Stoch et. al. 2022; Chau et al 2022)

Other works try to model incentiveness in the learning process directly and propose a theoretical framework or a loss function for incentivizing clients to share data (Karimireddy et al., 2022; Cho et al., 2022).

Papers on Deep Learning and Generative Models for Text Generation

At Criteo, text generation models are considered for extracting and generating information about products or for generating audio ads. However, those models tend to hallucinate, i.e., produce texts that are incorrect or contain incorrect facts. Controlling the quality of those texts is of primary importance in products. (Lee et. al. 2022; Su et. al. 2022). Steering the generation towards a specific domain or style is crucial for audio ads. These works highlight models capable of controlling text generation under given constraints (Li et al., 2022; Qin et al., 2022)

Controlling the length of the output is also critical for us since our displayed text has to fit in given banner dimensions or (if converted to speech) in a given time window. Although there are already strategies that guide the decoder or truncate the generation, this paper tackles the problem from the character level, which is helpful if, for instance, you want to display your text in a small banner, or a LED display (Liu et al. 2022)

Lastly, this paper proposes a new approach to generating text out of the left-to-right dominant strategy, using a different and parallelizable way of generating text. A direct application of this strategy is to enrich product descriptions in a catalog, which are often too succinct and robotic to be read by an end user (Lu et al., 2022).

Out-of-Distribution Generalization

The lack of robustness to out-of-distribution (OOD) data is a major challenge for safely deploying deep learning models in the real world. At NeurIPS, there have been a variety of papers proposed to improve the robustness of OOD data, as well as a Workshop on Distribution Shifts. One promising approach leverages ensembling to combine several diverse models for better OOD generalization. This was the idea behind our paper Diverse Weight Averaging for Out-of-Distribution Generalization, where we were able to achieve the state-of-the-art on the challenging and reference DomainBed benchmark. Our paper is based on an approximation to ensembling based on weight averaging. There have been a variety of papers at NeurIPS on the topics of weight averaging and weight averaging for OOD generalization (Ilharco et al. 2022; Arpit et al. 2022).

Model Compression

Criteo is a company that productionizes large deep learning models as such recent advances that show how to make models more compact are of great interest. Savarese et al. (2022) was one paper that learned the quantization representation during the training process.

Spatiotemporal Forecasting

We were also happy to present our paper (Yin et al. 2022) at the AI4Science workshop. We proposed a new generative method based on Implicit Neural Representations for spatiotemporal forecasting applied here to physical phenomena modeling. Dynamics forecasting is a key problem behind modeling temporal data such as those underlying video or time series.

Matthieu presenting his poster
New Orleans also offered an interesting city away from the conference.