Natural Language Processing
ropensci/gutenbergr
Search, download, and process public domain texts from Project Gutenberg
๐ Text Mining
Rno licenseupdated Aug 8, 2026
Text mining using tidy tools :sparkles::page_facing_up::sparkles:
$ git clone https://github.com/juliasilge/tidytext.gitSearch, download, and process public domain texts from Project Gutenberg
R package to Embed All the Things! using StarSpace
extract text from any document. no muss. no fuss.
Library to scrape and clean web pages to create massive datasets.
Starter code to solve real world text data problems. Includes: Gensim Word2Vec, phrase embeddings, Text Classification with Logistic Regression, word count with pyspark, simple text preprocessing, pre-trained embeddings and more.
A collection of notebooks for Natural Language Processing from NLP Town
Data from GitHub ยท snapshot Sep 24, 2026