Natural Language Processing
chiphuyen/lazynlp
Library to scrape and clean web pages to create massive datasets.
๐ค Language Models๐ Text Mining
Pythonno licenseupdated Nov 11, 2020
List of textual data sources to be used for text mining in R
$ git clone https://github.com/EmilHvitfeldt/R-text-data.gitLibrary to scrape and clean web pages to create massive datasets.
Fast, Consistent Tokenization of Natural Language Text
Search, download, and process public domain texts from Project Gutenberg
Text preprocessing, representation and visualization from zero to hero.
Python package for Korean natural language processing.
Starter code to solve real world text data problems. Includes: Gensim Word2Vec, phrase embeddings, Text Classification with Logistic Regression, word count with pyspark, simple text preprocessing, pre-trained embeddings and more.
Data from GitHub ยท snapshot Sep 24, 2026