Author: Criteo Tech

A summary of Scala Days 2019
The Scala language is used extensively at Criteo and the Recommendation codebase makes no exception to this rule. Our Spark jobs, which compute recommendations from catalogues of 6 billion products and logs of user actions coming in at 1 million entries per second, are written in Scala. We also use it for a number of…

Data Science: real jobs beyond the buzz
It has been now almost 4 years since I joined Criteo and I still keep a vivid memory of my first day on the job. Hired to lead a team of 4 Data Scientists and being “given access to one of the world’s biggest datasets, and the computing power to exploit it” (as per the…

Hackathon at scale: 3 perspectives
On March 2019, more than 600 Criteo employees joined our worldwide hackathon for 48 hours. From Palo Alto to Paris and Singapore, our… 649 participants, 96 projects, 3 perspectives. Hackathon at scale: 3 perspectives. On March 2019, more than 600 Criteo employees joined our worldwide hackathon for 48 hours. From Palo Alto to Paris and…

5 Do’s and Don’ts to restart a Hadoop cluster with no downtime
Our Hadoop cluster plays a pivotal role in our business operations. Operating a network of 3,000 advertisers, 5,000 publishers and 1.5 billion active shoppers simply wouldn’t be possible without a storage and compute environment that scales. Since beginning our Hadoop journey in 2012 we’ve continued to build on our big data operations, with a strong…

My year(s) at Criteo as a visiting professor
After working on theoretical questions in high-dimensional statistics as an academic at UC, Berkeley, I felt I wanted to see what statistics and machine learning were at industrial scale. This is why I joined Criteo in September 2017 to contribute to large-scale machine learning in industry. Friends in France told me that Criteo had positions…

Serving growing user needs with automated tooling
Running a Big Data platform can offer a great deal of value to users, but only if that value is simple and easy to access. At Criteo, we’ve found that making our massive data sets accessible to our growing and geographically-distributed sales user base has come with certain challenges; not least the ability to scale…

Career tracks and leveling in Criteo R&D
One of the questions we hear a lot from engineers in our recruiting process is what kind of career path exists for technical specialists within Criteo. Their concern stems from the fact that they have already reached the top of the technical path in their current company and their only option for progressing is to…

Monitor Finalizers, contention and threads in your application
This post of the series details more complicated CLR events related to finalizers and threading. Part 1: Replace .NET performance counters by CLR event tracing. Part 2: Grab ETW Session, Providers and Events. Introduction In the previous post, you saw how the TraceEvent nuget helps you deciphering simple ETW events such as the one emitted when…

The need to scale bigger
At Criteo, hypergrowth is something we have gotten used to. Our business has grown from a single data center in Paris back in 2005, to a globally distributed, multi data center infrastructure, with over 35 000 servers and what is one of the largest Hadoop clusters in Europe. This blogpost is the story of how…

How to beat !dumpheap -stat?… with ClrMD
When you are dealing with large memory dumps, figuring out what instances of which types (i.e. the list of types sorted by size of their instances with their count) are stored in memory takes time. Sos in WinDBG provides the dumpheap -stat command that can take minutes. For example, on a 16.7 GB production dump,…










