Sre

Importance of Graceful Shutdown in Kubernetes
Have you ever deployed a new version of your app in Kubernetes and noticed errors briefly spiking during rollout? Many teams do not even realize this is happening, especially if they are not closely monitoring their error rates during deployments. There is a common misconception in the Kubernetes world that bothers me. The official Kubernetes…

Enhancing Service Reliability with Feature-Based SLOs: A Comprehensive Approach
In recent years, our focus has been on defining Service Level Objectives (SLOs) for HTTP API endpoints. While this approach has proven beneficial for development teams, it has also had less impact on external customers and product teams. Users typically expect SLOs to be associated with features rather than individual endpoints or services. Communicating the…

Monitoring microservices — Central Monitoring: A tool for a global view of things
A bit of history Some years ago, Criteo switched from monolithic applications to microservices. With this new architecture comes challenges like monitoring hundreds of applications, all interacting with each other. At Criteo, there are several ways to introduce innovation. One of them is the yearly Hackathon! At the 2020 event, the fantastic Firewatch team aimed to…

Scheduling Data Pipelines at Criteo — Part 3
The Proven Model in Production Building a successful Platform is a quest of the good abstraction level. If you’ve missed it, check out the previous articles in this series: Scheduling Data Pipelines at Criteo — Part 2 This week we deep dive into the key ideas leveraged by BigDataFlow medium.com Scheduling Data Pipelines at Criteo — Part 1 Introducing…




