Laion Announces Big Video Dataset
Laion, the non‑profit AI research organization, has announced the release of its BVD (Bilingual Visual Dataset), a large‑scale collection of images paired with captions in multiple languages. The dataset, which contains over 400 million image‑caption pairs, spans 20 languages and is designed to support multilingual vision‑language models. Laion has made the data publicly available under an open‑source license, encouraging researchers and developers to use it for training and benchmarking new AI systems.
The launch follows a growing demand for diverse, multilingual training data that can improve the performance of models on non‑English tasks. By providing a resource that covers a broad linguistic range, Laion aims to reduce language bias in vision‑language research and enable more inclusive AI applications. The release has already attracted attention on Hacker News, where the discussion thread has garnered 67 points and 18 comments, reflecting the community’s interest in the dataset’s potential to advance multilingual AI research.