Skip to content
#

multilingual-dataset

Here are 8 public repositories matching this topic...

The first open-source ๐—บ๐˜‚๐—น๐˜๐—ถ๐—น๐—ถ๐—ป๐—ด๐˜‚๐—ฎ๐—น (5 languages) corpus for low-resource NLP, boldly bridging three distinct language branches. Built by a ๐—ป๐—ฎ๐˜๐—ถ๐˜ƒ๐—ฒ ๐—ฆ๐˜†๐—น๐—ต๐—ฒ๐˜๐—ถ Linguistics undergrad at ๐—ž๐˜‚๐—ฟ๐˜€๐—ธ ๐—ฆ๐˜๐—ฎ๐˜๐—ฒ ๐—จ๐—ป๐—ถ๐˜ƒ๐—ฒ๐—ฟ๐˜€๐—ถ๐˜๐˜†, Russia. Targeting a 10K+ sentence dataset for MT/ASR training to computationally revitalize Sylheti.

  • Updated Sep 12, 2026
  • HTML

Multilingual dataset of world cities with English and Arabic names, population, and country info. Provided in JSON, CSV, SQL, Excel formats. This will provide enriched information of countries, states and their capitals translate these in Arabic and show population of the city

  • Updated May 23, 2025
  • Python

Add this topic to your repo

To associate your repository with the multilingual-dataset topic, visit your repo's landing page and select "manage topics."

Learn more