Natural Language Processing
ropensci/tokenizers
Fast, Consistent Tokenization of Natural Language Text
โ๏ธ Tokenization & Preprocessing๐ Text Mining
ROtherupdated Mar 27, 2024
Search, download, and process public domain texts from Project Gutenberg
$ git clone https://github.com/ropensci/gutenbergr.gitFast, Consistent Tokenization of Natural Language Text
R package for Tokenization, Parts of Speech Tagging, Lemmatization and Dependency Parsing Based on the UDPipe Natural Language Processing Toolkit
R package to Embed All the Things! using StarSpace
Library to scrape and clean web pages to create massive datasets.
Text mining using tidy tools :sparkles::page_facing_up::sparkles:
Starter code to solve real world text data problems. Includes: Gensim Word2Vec, phrase embeddings, Text Classification with Logistic Regression, word count with pyspark, simple text preprocessing, pre-trained embeddings and more.
Data from GitHub ยท snapshot Sep 24, 2026