[NeurIPS 2024] ๐ธ GlotCC Dataset and Pipline
-
Updated
Apr 6, 2025 - Jupyter Notebook
[NeurIPS 2024] ๐ธ GlotCC Dataset and Pipline
MMT: A Multilingual and Multi-Topic Indian Social Media Dataset (C3NLP / EACL 2023). Samples for multilingual topic modeling, language identification, code-mixing and Hinglish.
Multilingual dataset for principal parts detection in inflectional morphology (CoNLL 2025)
118 public A1-B2 graded stories in German, Spanish, French, Italian, Korean and Russian with aligned English translations, vocabulary and questions.
Multilingual emotional speech datasets for TTS training
The first open-source ๐บ๐๐น๐๐ถ๐น๐ถ๐ป๐ด๐๐ฎ๐น (5 languages) corpus for low-resource NLP, boldly bridging three distinct language branches. Built by a ๐ป๐ฎ๐๐ถ๐๐ฒ ๐ฆ๐๐น๐ต๐ฒ๐๐ถ Linguistics undergrad at ๐๐๐ฟ๐๐ธ ๐ฆ๐๐ฎ๐๐ฒ ๐จ๐ป๐ถ๐๐ฒ๐ฟ๐๐ถ๐๐, Russia. Targeting a 10K+ sentence dataset for MT/ASR training to computationally revitalize Sylheti.
Multilingual dataset of world cities with English and Arabic names, population, and country info. Provided in JSON, CSV, SQL, Excel formats. This will provide enriched information of countries, states and their capitals translate these in Arabic and show population of the city
Parallel Literary Corpora: Fiction and Poetry Translations
To associate your repository with the multilingual-dataset topic, visit your repo's landing page and select "manage topics."