LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Schuhmann, Christoph; Vencu, Richard; Beaumont, Romain; Kaczmarczyk, Robert; Mullis, C. T.; Aarush, Katta,; Coombes, Theo; Jitsev, Jenia; Komatsuzaki, Aran

doi:10.48550/arxiv.2111.02114

Cited by 91 publications

(136 citation statements)

References 3 publications

Supporting

Mentioning

135

Contrasting

Order By: Relevance

“…We pre-train the model for 20 epochs using a batch size of (Changpinyo et al, 2021), SBU captions (Ordonez et al, 2011)). We also experimented with an additional web dataset, LAION (Schuhmann et al, 2021), which contains 115M images with more noisy texts 1 . More details about the datasets can be found in the appendix.…”

Section: Pre-training Detailsmentioning

confidence: 99%

BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Li¹,

Li²,

Xiong³

et al. 2022

Preprint

139

View full text Add to dashboard Cite

Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based tasks or generation-based tasks. Furthermore, performance improvement has been largely achieved by scaling up the dataset with noisy image-text pairs collected from the web, which is a suboptimal source of supervision. In this paper, we propose BLIP, a new VLP framework which transfers flexibly to both vision-language understanding and generation tasks. BLIP effectively utilizes the noisy web data by bootstrapping the captions, where a captioner generates synthetic captions and a filter removes the noisy ones. We achieve state-of-the-art results on a wide range of vision-language tasks, such as image-text retrieval (+2.7% in average recall@1), image captioning (+2.8% in CIDEr), and VQA (+1.6% in VQA score). BLIP also demonstrates strong generalization ability when directly transferred to videolanguage tasks in a zero-shot manner. Code, models, and datasets are released.

show abstract

Section: Pre-training Detailsmentioning

confidence: 99%

BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Li¹,

Li²,

Xiong³

et al. 2022

Preprint

139

View full text Add to dashboard Cite

show abstract

“…We observe that the model performs best with a cosine similarity threshold of 0.3 as we achieve a 0.84, 0.47, 0.80, 0.58 for accuracy, precision, recall, and f1 score, respectively. Indeed, the 0.3 threshold is also used by previous work by Schuhmann et al [55] that inspected CLIP's cosine similarities between text and images and determined that 0.3 is a suitable threshold.…”

Section: Identifying Antisemitic and Islamophobic Imagesmentioning

confidence: 99%

Understanding and Detecting Hateful Content using Contrastive Learning

González-Pizarro¹,

Zannettou²

2022

Preprint

View full text Add to dashboard Cite

The spread of hate speech and hateful imagery on the Web is a significant problem that needs to be mitigated to improve our Web experience. This work contributes to research efforts to detect and understand hateful content on the Web by undertaking a multimodal analysis of Antisemitism and Islamophobia on 4chan's /pol/ using OpenAI's CLIP. This large pre-trained model uses the Contrastive Learning paradigm. We devise a methodology to identify a set of Antisemitic and Islamophobic hateful textual phrases using Google's Perspective API and manual annotations. Then, we use OpenAI's CLIP to identify images that are highly similar to our Antisemitic/Islamophobic textual phrases. By running our methodology on a dataset that includes 66M posts and 5.8M images shared on 4chan's /pol/ for 18 months, we detect 573,513 posts containing 92K Antisemitic/Islamophobic images and 246K posts that include 420 hateful phrases. Among other things, we find that we can use OpenAI's CLIP model to detect hateful content with an accuracy score of 0.84 (F1 score = 0.58). Also, we find that Antisemitic/Islamophobic imagery is shared in 2x more posts on 4chan's /pol/ compared to Antisemitic/Islamophobic textual phrases, highlighting the need to design more tools for detecting hateful imagery. Finally, we make publicly available a dataset of 420 Antisemitic/Islamophobic phrases and 92K images that can assist researchers in further understanding Antisemitism/Islamophobia and developing more accurate hate speech detection models.

show abstract

“…Besides CC12M, there also exist some other large- scale image-text datasets, such as WIT [41], WenLan [19], LAION-400M [39], and the datasets used in CLIP [38] and ALIGN [20]. More detailed discussions on them are provided in Appendix.…”

Section: Pre-training Datasetmentioning

confidence: 99%

“…• LAION-400M [39] has 400 million image-text pairs, and is recently released to public. Instead of applying human designed heuristics in data cleaning, this dataset relies on the CLIP [38] model to filter image-text pairs, where the cosine similarity scores between image and text embeddings are calculated and filtered by 0.3.…”

Section: B Comparison Of Image-text Datasetsmentioning

confidence: 99%

See 1 more Smart Citation

Scaling Up Vision-Language Pre-training for Image Captioning

Hu¹,

Gan²,

Wang³

et al. 2021

Preprint

View full text Add to dashboard Cite

In recent years, we have witnessed significant performance boost in the image captioning task based on visionlanguage pre-training (VLP). Scale is believed to be an important factor for this advance. However, most existing work only focuses on pre-training transformers with moderate sizes (e.g., 12 or 24 layers) on roughly 4 million images. In this paper, we present LEMON , a LargE-scale iMage captiONer, and provide the first empirical study on the scaling behavior of VLP for image captioning. We use the state-of-the-art VinVL model as our reference model, which consists of an image feature extractor and a transformer model, and scale the transformer both up and down, with model sizes ranging from 13 to 675 million parameters. In terms of data, we conduct experiments with up to 200 million image-text pairs which are automatically collected from web based on the alt attribute of the image (dubbed as ALT200M). Extensive analysis helps to characterize the performance trend as the model size and the pre-training data size increase. We also compare different training recipes, especially for training on large-scale noisy data. As a result, LEMON achieves new state of the arts on several major image captioning benchmarks, including COCO Caption, nocaps, and Conceptual Captions. We also show LEMON can generate captions with long-tail visual concepts when used in a zero-shot manner.

show abstract

LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Cited by 91 publications

References 3 publications

BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Understanding and Detecting Hateful Content using Contrastive Learning

Scaling Up Vision-Language Pre-training for Image Captioning

Contact Info

Product

Resources

About