Bangla Wikipedia dataset

Aisha Khatun
A subset of the Bangla version of the Wikipedia text. To create the Wikipedia dataset, we collected the Bangla wiki-dump of 10th June, 2019. The files are then merged and each article is selected as a sample text. All HTML tags were removed and the title of the page was stripped from the beginning of the text. This dataset contains 70377 samples with a total number of words being 18229481. The entire dataset has 1289249...
This data repository is not currently reporting usage information. For information on how your repository can submit usage information, please see our documentation.