An Expert Community Where Every Voice Matters


Join Us

Highlights from SRECon 2023


SRECon EMEA Africa 2023 took place in Dublin from October 10th to October 12th. It attracted over 650 participants from all over the world.

The conference was the perfect moment for keynote speakers, technical sessions, workshops, and networking opportunities for developers, Site Reliability Engineers, and other professionals interested in the world of Site Reliability Engineering.

As a sponsor of SRECon EMEA Africa 2023 , CRITEO’s SRE team attended many high-quality and highly technical talks. We enjoyed all our conversations with the people who joined us at our booth.

During the talks, the attendees had the opportunity to learn about the latest developments in site reliability, systems engineering, and teams working with complex distributed systems at scale.

They also met many experts in this industry, such as representatives from leading technology companies, startups, and open-source communities, including people from Meta Ai Research Infra, Google/GCP SRE team, Slack, Grafana Labs or Datadog.

Criteos’ Takeaways

Jérémy’s insights

It was my first SRECon, and I have not been disappointed at all! I have been able to discover new tools, how other companies practice reliability, and learn from others’ experiences.

The presentations were diverse, from incident management to new technologies. With two tracks for three days, you could find many topics close to your interests.

The benefits of participating in such a conference are multiple: it helps you to stay up to date with the current state of the community practices, it allows you to discover technologies and services you may not have heard of, it provides interesting experience feedback and of course, you get an opportunity to exchange with people from different companies.

It is hard to select one talk out of all of the ones I listened to as I took pieces of information out of many of them, but one that was quite interesting is how a platform like Wikipedia went from a single datacenter to a multi-datacenter architecture with a seamless failover procedure From Exceptional Maintenance to Automated Routine Operation: A Story of the Datacenter Switchover for Wikipedia.

Ilyas’ takeaways

Attending SRECon 2023 was an inspiring experience for me. The organization provided the ideal backdrop for advancing our knowledge and connections in Reliability Engineering and it was also the first time for me to represent Criteo in a conference.

The event took place at The Convention Centre Dublin, situated on the banks of the River Liffey. The choice of Dublin as the location added a unique international flavour to the conference. The venue’s spacious and well-designed conference rooms provided the perfect setting for the technical discussions that unfolded over the event.

The organization of SRECon 2023 in Dublin was nothing short of outstanding. Attendees were welcomed with Irish hospitality, and the registration process was seamless. The conference Slack space played a pivotal role in connecting attendees, making it easy to ask questions to speakers during the sessions. The conference program was clear, even if we had to switch rooms during sessions to attend all the talks of interest.

Networking opportunities were abundant, with a social event that fostered engaging discussions and knowledge sharing. The presence of leading sponsors added value, offering innovative product demonstrations and interactive sessions that complemented the conference content.

The content was a true highlight, featuring an array of international experts who shared their insights into the evolving landscape of reliability engineering. I felt both inspired and equipped with new knowledge on tools used by other companies, understanding how they manage global infrastructure and their best practices on observability, reliability, and chaos engineering. I was particularly interested in the talks where folks presented how they implement SRE practices in particular cases, their outage management, and how they deal with SLA/SLO.

It’s not easy to choose one particular talk as many were interesting and diverse, but, being part of the Stream Processing team, I would partially keep, the presentation delved into the unanticipated challenges faced when a lone Kafka broker, not operating at full health, triggered cascading disruptions on the whole Kafka infrastructure and disproportionate issues on many producers, even minutes of full-service outage. The presentation highlighted many scenarios that caused processing outage and deeply dived into the configuration that created the cascading issues. The speaker shared a comprehensive set of modifications made to enhance Kafka’s resilience, ensuring it remained aptly configured to cater to business demands.

In summary, SRECon 2023 was a memorable experience. The venue’s stunning location, flawless organization, and rich content provided the ideal atmosphere for advancing our knowledge and connections in SRE. Looking forward to the next editions, it could be a great opportunity to present one of Criteo’s technologies in the next one.

Baptiste’s notes

I personally spent most of my time at the Criteo booth, presenting and evangelizing Criteo’s infrastructure. This gave me a better perspective of the exhibition area compared to my colleagues.

My first observation was the attendees’ level of seriousness: during the scheduled talks, the booth space was nearly empty! Outside of the scheduled talks, attendees navigated the exhibition area as expected, exploring demos and collecting swag. The individuals manning the booth were exceptionally welcoming and friendly, showcasing their technology brilliantly. Kudos to SquaredUp, our neighbors at the Criteo booth, for their impressive setup featuring a complete supply chain of custom Lego characters integrated with their observability solution.

However, conferences offer more than just booth swag, and I had the opportunity to attend some highly engaging talks that delved into subjects that Criteo encounters daily.

At Criteo, we configure and provision our fleet of over 40,000 bare-metal servers using Chef. As a DevLead of Criteo’s Infrastructure Systems & Services team, responsible for providing Chef internally, on day one, I was eager to attend the “Scaling Chef Emotionally” presentation by Brett Pemberton, Slack. In this era dominated by cloud and containers, Chef receives less attention. Most users primarily interact with Cloud providers and Kubernetes, making it intriguing to witness a public discussion about Chef. Furthermore, given Slack’s renowned status as a Tech company with a substantial infrastructure, the challenges they encounter couldn’t be anything less than fascinating!
I left with a strong sense of satisfaction. Not only did we make similar decisions when it came to testing and scaling our Chef infrastructure, but we also encountered identical challenges related to scaling and version updates. The presentation was truly engaging, but the icing on the cake was the response to the following question: “Have you considered moving entirely away from Chef in favour of a containerized system?”
Just like Criteo, Slack is using containerized systems, but it’s worth noting that these systems still require provisioning and operation on a properly configured infrastructure, hence the need for Chef — or equivalent.

On day 2, I attended two presentations on incident response and management, The World Blew Up but We’re All Okay: How We Managed a Massive-scale Incident at Datadog and When One Line Took Thousands of Websites Offline. These presentations were well-structured and dynamic, providing a clear understanding of the impact and triggering memories of past incidents. Once again, this allowed me to compare Criteo’s SRE practices with those of other tech industry players. It also served as a reminder to stay humble in the face of scale. Despite having built highly resilient and automated systems, the entire house of cards can collapse with just “a line of code” or an automatic update feature not disabled.

On the third day, I concluded the conference by attending How to Make Your Automation a Better Team Player, by Laura Nolan from Stanza at the closing plenary session, which, I believe, encapsulated the key takeaways I brought home with me. As SREs, observability, and automation are key aspects. However, to derive the maximum benefit or simply to avoid unnecessary headaches, one should:

  • avoid automation surprises — such as unattended upgrades that impacted Datadog.
  • clearly display intended actions — to clearly identify the impacted resources & services.

Enriched and inspired by all the experiences shared by presenters and visitors at the Criteo booth, I write down these notes on SRECON 2023. I look forward to the next opportunity to return to SRECon, and I am convinced that certain Criteo anecdotes could captivate many visitors!


In conclusion, SRECon 2023 proved to be an exhilarating and enlightening experience for our team. The captivating presentations, engaging discussions, and invaluable networking opportunities reaffirmed our commitment to keep working on improving our SRE practices.

Engineering

We are creators! From designing ground-breaking products to finding unique ways to solve technical challenges at an…

bit.ly