WanJuan 1.0 Dataset
Shusheng·Wanjuan 1.0 is the first open-source version of the Shusheng·Wanjuan multimodal corpus, which includes three parts: text datasets, image-text datasets, and video datasets, with a total data volume exceeding 2TB. Based on the corpus constructed by the Large Model Data Alliance, the Shanghai AI Laboratory has performed fine-grained cleaning, deduplication, and value alignment on some of the data, resulting in Shusheng·Wanjuan 1.0, which features four main characteristics: diverse integration, meticulous processing, value alignment, and ease of use and efficiency.