An Expert Community Where Every Voice Matters


Join Us

ICCV 2023 recap


Why ICCV?

The Conference

The International Conference on Computer Vision (ICCV) is one of the most renowned conferences in the field of computer vision. This conference serves as a hub for researchers, academics, and professionals for the exchange of cutting-edge research and state-of-the-art ideas in the field of computer vision. This year, over 7000 researchers were registered. A total of 56 workshops and 10 tutorials were organized, and over 2000 papers were accepted at the conference. As the field continues to evolve rapidly, ICCV remains a major and important event for staying at the forefront of the latest research and development in artificial intelligence applied to computer vision.

Criteo @ ICCV

Historically, Criteo’s research was focused on traditional ML and recommendation conferences such as ICML, AISTATS, NeurIPS or RecSys. Since the creation of the Criteo AI Lab (CAIL) in 2018, our research has now diversified and the number of papers we publish tackling deep learning has grown dramatically. This scientific production has positioned Criteo for the development of deep learning and generative AI for visual data well before their boom from 2022.

Now that these techniques have become more mature, we felt that this was the right time to apply our knowledge to more specific tasks for innovative visual ads and technologies, such as 3D asset generation, virtual try-on, image processing, object recognition, 3D vision, augmented reality, and others. This is why ICCV 2023, held in Paris, was a special opportunity for our researchers to explore the latest state-of-the art research in these fields, grasp new trends and techniques that will soon be used in production in the industry, and promote the exchange of ideas that drive forward the field in general and our own research in particular.

General Takeaways from the Conference

You will find below the conference’s highlights of our attendees, Karim Kassab, Simon Lepage, Antoine Schnepf (CAIL PhD students) and Jean-Yves Franceschi (CAIL researcher) 👇

Photo by Dim Hou on Unsplash

Antoine’s notes

ICCV 2023 marked my very first experience attending a conference. It was an exceptional week during which I discovered a new facet of the research world. I had the opportunity to witness how researchers present and communicate their work and, perhaps most excitingly, put faces to names I had only encountered in research papers.

The first two days were dedicated to workshops. Typically, researchers present their work at workshops to obtain initial feedback on their research ideas without going into extensive and comprehensive comparisons with similar research. They might also present a mature work at the workshop, before publishing it at a different conference, creating more opportunities for discussions and recognition within the community.

I found this part of the conference truly enjoyable. Workshops are structured around casual poster sessions, offering an opportunity to engage with researchers presenting their work using technical posters as their sole support. It’s a much quicker way to grasp the core concepts behind a paper and understand how it works compared to reading it.

Each day, there were around a dozen different workshops, and I was particularly captivated by the one on 3D computer vision, where groundbreaking generative techniques for producing 3D content were showcased. During these workshops, renowned keynote speakers in their respective fields provided insights about the discipline, current trends, recent major developments, limitations of the best methods, and future research directions for the community. It’s a unique juncture where technical details and a broader vision of the field converge, shedding light on the current state of a specific research area. I was especially impressed by Ben Poole’s presentation. He is a Google researcher who developed DreamFusion, a revolutionary technique for generating 3D assets from textual descriptions. Understanding the origin of such groundbreaking ideas and gaining insights into the “how” behind their development was invaluable for a young researcher like me.

During the following three days, the main part of the conference is held, and the accepted papers are presented. It alternates between poster sessions, where you can meet and talk with the authors, short presentations, where the most groundbreaking or exciting research advancements are presented by their authors, and keynote talks where more fundamental and general aspects of the field are presented.

My personal favourite moment of the conference was one of the very last talks by Google DeepMind where more general aspects of AI were discussed and put into perspective. Notably, advances in protein folding for new medications, faster matrix multiplication for quicker computation, and reinforcement learning for self-learning agents left a lasting impression. These breakthroughs have the potential for significant positive impacts on our lives in the future.

Karim’s takeaways

Attending ICCV for the first time as a PhD student was an eye-opening experience. As someone currently immersed in the field, the conference offered me a rare opportunity to witness the magnitude of international expertise in computer vision, as well as the latest state-of-the-art research. The days were filled with rich information presented in the form of keynote lectures, paper presentations, workshops, and poster sessions. Renowned researchers such as Pushmeet Kohli from Google DeepMind and Dorsa Sadigh from Stanford University shared their insights, providing me with a broader perspective on the challenges and opportunities within the field. Additionally, the exchange of ideas during poster sessions was particularly enlightening. This not only allowed me to discuss with authors of my favorite papers, but also enabled me to discover novel ideas and a fresh perspective on some of the many sub-domains in computer vision.

On a more personal level, what stood out the most for me from this experience is the impact the conference had on my own research. As I prepare for a submission to another top-tier computer vision conference, ICCV broadened my understanding of the field and provided a fresh lens through which to view my own work. The exposure to high-quality research has undoubtedly raised the bar for my own contributions.

Simon’s highlights

ICCV was also my first conference, and it was nothing short of impressive. While understanding the sheer magnitude of the Computer Vision field is one thing, witnessing the tangible representation of this popularity in real life, with countless researchers and papers, is an entirely different experience.

Throughout the week, there was an extensive array of presentations, including both oral sessions and poster exhibitions, creating an uninterrupted flow of ideas and concepts. Exploring areas beyond my research focus proved to be incredibly interesting, as I gleaned inspiration from diverse sources. The conference expanded my exposure to these ideas in a way that surpasses the time-consuming process of reading papers. Engaging directly with authors allowed me to pose questions about specific aspects of their work, gaining valuable intuitions and putting faces to names I had previously encountered. Additionally, observing how researchers presented their work to a large audience added another layer to my experience.

In essence, I was deeply impressed by the vast array of research subjects, at times remarkably specific to practical applications, and the sheer abundance of innovative ideas presented to tackle them.

Jean-Yves’ points

While used to ML conferences, ICCV was my first CV conference. I was impressed to witness how the CV community is organized, with so many papers tackling so many specific vision tasks. While some of them are standard even in the ML community (e.g., image generation, object recognition and segmentation), other are more niche and domain-specific (e.g., removing weather effects on images or extracting hand positions for autonomous driving). This high diversity helps find out what techniques work best in a majority of applied domains close to production and not just on standard public benchmarks.

Three clear trends emerged: the training and fine-tuning of large-scale image generation models for various downstream tasks, automatic 3D scene modeling, and multimodality (models considering multiple image and non-image, such as textual, data modalities). All these trends are especially relevant to our applications at Criteo, as much as it is to other e-commerce / retail media actors like Amazon and Alibaba who were showcased at the workshop 3D Vision and Modeling Challenges in eCommerce.

Overall, ICCV in Paris was an exciting and well-organized event with a high-quality program. The presentations, papers and trends we observed will help us shape our future research and its applications. The only criticism I would have: extremely crowded poster sessions!

Our Scientific Highlights

Karim

While navigating the conference, many papers caught my attention. Below is a brief list of papers that I found most intriguing, as they are closely related to my work on 3D representations.

In addition, some workshops were more finely specialized in the domain of 3D representations:

Simon

My primary focus was on papers and presentations centered on large vision models applied to extensive sets of noisy or unlabelled data. Interestingly, the challenges addressed in these works closely resonate with those we grapple with in our work with Criteo’s data.

  • Towards Universal Image Embeddings: A Large-Scale Dataset and Challenge for Generic Image Representations. Large-scale embedding applications frequently deal with highly heterogeneous data. However, existing research on instance-level recognition tasks predominantly addresses small, homogeneous domains, neglecting the challenges tied to establishing a singular universal embedding. This paper introduces a novel benchmark tailored to address these issues.
  • Noise-aware Learning from Web-crawled Image-Text Data for Image Captioning. Given a very large amount of noisy image-text pairs, what is the best approach to create a captioning model ? Training on the entire dataset yields subpar results, and indiscriminately removing noisy pairs risks losing valuable information. This paper introduces a novel method that is sensitive to noise levels, enabling the utilization of the entire dataset while gaining control over captioning model quality.
  • Referring Image Segmentation Using Text Supervision. Some tasks, such as Referring Image Segmentation (RIS) or Detection, typically require extensive and costly pixel-wise annotations. This paper introduces a method to address RIS using solely image-text pairs.
  • Read-only Prompt Optimization for Vision-Language Few-shot Learning. Efficient prompt-tuning methods have been suggested for fine-tuning pretrained models on small datasets. Nevertheless, the introduction of new tokens in the model can alter its behavior by influencing attention patterns. This paper introduces read-only prompts, demonstrating improved generalization results in data-deficient settings.

Some workshops were specifically dedicated to these themes as well.

Antoine

I primarily focused on papers about 3D computer vision, and more specifically those at the intersection of generative models and neural radiances fields. I listed below some of the work that interested me the most.

  • LERF: Language Embedded Radiance Field. This paper simultaneously trains NeRF with new volumetric CLIP embeddings, making it possible to know which objects are in the NeRF.
  • Dreambooth3D: Subject Driven Text-to-3d Generation. Given a few images of a subject (person, animal or any object), Dreambooth 3d allows you to generate 3d assets of this subject in varied environments and scenarios. To do so, you should simply provide the model with a textual description of the desired scene. Simply amazing.
  • GenNVS: Generative Novel View Synthesis with 3D-Aware Diffusion Models. A new diffusion-based model that produces Neural Radiance Fields from a single image on the complex CO3D dataset. It is the first time a model can produce high quality NeRF from a singe a few on such a complex real-world dataset.
  • HoloDiffusion and HoloFusion. Two papers from the same authors (only the latter was presented at ICCV) that propose a new approach for training a 3d generative diffusion model only from image data.
  • DiT: Scalable Diffusion Models with Transformers. This paper introduces a new type of diffusion models, based on the transformer architecture, which show competitive performance compared to the standard and more mature U-Net based diffusion models.

Jean-Yves

There were many interesting papers and presentations at the conference. I am highlighting some of them below, but it was hard selecting them!

  • 3D-aware Image Generation using 2D Diffusion Models. Given an image generative model like Stable Diffusion, how could we generate different views of the same object? This is an active area of research that is particularly relevant in our use cases. The key idea is to predict the depth map for each generated images, which gives a strong conditioning on the image that would be obtained from a close camera angle. Missing details are then simply filled with the image generation model. Simple but effective!
  • Efficient Video Prediction via Sparsely Conditioned Flow Matching. Video generation might be the new generative AI revolution, but it comes at even greater computational costs than for images. This paper greatly reduces the costs of such models by finding a clever way to condition the generation of the future of a video on the previous frames in a diffusion model.
  • Erasing Concepts from Diffusion Models and Ablating Concepts in Text-to-Image Diffusion Models. Image generation models like Stable Diffusion are now performant enough for real-world applications, but their prodification remains a challenge because such models are hard to control. In particular, they might generate copyrighted images or unwanted concepts (nudity, violence, fake news, etc.). These two papers propose methods to prevent such content to be produced by these models. While they both require fine tuning, this is a nascent area of research that we should follow.
  • Workshop on Uncertainty Quantification for Computer Vision. This workshop particularly stood out with interesting talks on a topic that has been often been overlooked by the community. Yet, uncertainty quantification is an essential tool to improve the reliability of AI models.

Outside these papers, I was struck by the exponentially growing popularity of mixture of experts, enabling the training of even more performant models while keeping reasonable computational costs.


Engineering

We are creators! From designing ground-breaking products to finding unique ways to solve technical challenges at an…

bit.ly