Infrastructure
-

How We Cut HAProxy Fleet GCP Cost in Half by Moving from N2D to C4D
When you run a large production footprint in Google Cloud, changing a VM family is never just a hardware refresh. In our case, HAProxy sits on a critical path of the platform, serving as part of the traffic layer that hundreds of downstream systems quietly depend on every day. That means even a seemingly straightforward…
-

Insights from Kernel Recipes 2025
Supporting tech communities should be a cornerstone of the software industry. The digital products we use daily have been created by software engineers who are very passionate about what they build, but especially about how they build it. The Linux kernel is, without a doubt, one of those essential pieces of software that enable the…
-

Highlights from SRECon 2023
SRECon EMEA Africa 2023 took place in Dublin from October 10th to October 12th. It attracted over 650 participants from all over the world. The conference was the perfect moment for keynote speakers, technical sessions, workshops, and networking opportunities for developers, Site Reliability Engineers, and other professionals interested in the world of Site Reliability Engineering.…
Arthur Calas & 3 others -

Monitoring microservices — Central Monitoring: A tool for a global view of things
A bit of history Some years ago, Criteo switched from monolithic applications to microservices. With this new architecture comes challenges like monitoring hundreds of applications, all interacting with each other. At Criteo, there are several ways to introduce innovation. One of them is the yearly Hackathon! At the 2020 event, the fantastic Firewatch team aimed to…
-

Highlights from SRE Con 2022
We were excited to participate in the sold-out SREcon EMEA 2022 in Amsterdam with our team! We enjoyed meeting & conversing with many interesting people in the SRE community from all over the world. We were happy to have one of our own Bo Hou talk about “A better way to manage Command Line Tools”…
Daphnée Bestel & 4 others -

Scheduling Data Pipelines at Criteo — Part 3
The Proven Model in Production Building a successful Platform is a quest of the good abstraction level. If you’ve missed it, check out the previous articles in this series: Scheduling Data Pipelines at Criteo — Part 2 This week we deep dive into the key ideas leveraged by BigDataFlow medium.com Scheduling Data Pipelines at Criteo — Part 1 Introducing…
-

Highlights of KubeCon + CloudNativeCon Europe 2022
An article by Daphnée Bestel, Harold Dost, Khalil Kooli, Thomas Langé & Thibaut Sarion The global cloud-native developer population has grown by 1 million in the last 12 months, according to the 2021 State of Cloud-Native Development Report developed for CNCF by SlashData. The estimated number of cloud-native developers is about 7.1 million worldwide. That represents…
-

5 Do’s and Don’ts to restart a Hadoop cluster with no downtime
Our Hadoop cluster plays a pivotal role in our business operations. Operating a network of 3,000 advertisers, 5,000 publishers and 1.5 billion active shoppers simply wouldn’t be possible without a storage and compute environment that scales. Since beginning our Hadoop journey in 2012 we’ve continued to build on our big data operations, with a strong…
-

Serving growing user needs with automated tooling
Running a Big Data platform can offer a great deal of value to users, but only if that value is simple and easy to access. At Criteo, we’ve found that making our massive data sets accessible to our growing and geographically-distributed sales user base has come with certain challenges; not least the ability to scale…
