🤖 LAION releases open dataset of 80 million videos

The non-profit group LAION has released the video dataset LAION-BVD: 80 million videos totaling 10 million hours were collected from 1.3 billion CommonCrawl links, cut into 55 million clips with synthetic descriptions, plus 300 million frames and audio of 1.7 million and 10 million tracks. According to arXiv:2608.24845, ViCLIP and CLAP on BVD outperform models on InternVid by up to +2.1%.

🌍 The largest video datasets for multimodal pre-training are concentrated at proprietary companies. BVD for the first time opened 10 million hours of video to the scientific community in three modalities — video-text, audio-text, image-text.

👤 You can fine-tune your video or audio model now: URL versions of all sub-samples are freely available on Hugging Face. The license is for research, commercial use is prohibited, and the full corpus with media files is provided upon application.

Source 1: https://arxiv.org/abs/2608.24845

Source 2: https://projects.laion.ai/bvd/index.html