Tokenization & Preprocessing
๐ #2,313 overall#84 of 148 in Tokenization & Preprocessing
bnosac /udpipe
R package for Tokenization, Parts of Speech Tagging, Lemmatization and Dependency Parsing Based on the UDPipe Natural Language Processing Toolkit
$ git clone https://github.com/bnosac/udpipe.gitStars
224
224
Forks
35
35
Language
C++
License
MPL-2.0
Created
Aug 25, 2017
9.1 years old
Last push
Mar 2, 2026
More in Tokenization & Preprocessing
Tokenization & Preprocessing
mit-ccc/TweebankNLP
[LREC 2022] An off-the-shelf pre-trained Tweet NLP Toolkit (NER, tokenization, lemmatization, POS tagging, dependency parsing) + Tweebank-NER dataset
๐ท๏ธ Named Entity Recognition๐ง Natural Language Processing
PythonApache-2.0updated Jan 24, 2024
Tokenization & Preprocessing
ropensci/tokenizers
Fast, Consistent Tokenization of Natural Language Text
๐ง Natural Language Processing๐ Text Mining
ROtherupdated Mar 27, 2024
Tokenization & Preprocessing
taishi-i/nagisa๐ฅ active
A Japanese tokenizer based on recurrent neural networks
๐ง Natural Language Processing
PythonMITupdated Sep 22, 2026
Tokenization & Preprocessing
janlukasschroeder/nlp-cheat-sheet-python
NLP Cheat Sheet, Python, spacy, LexNPL, NLTK, tokenization, stemming, sentence detection, named entity recognition
๐ท๏ธ Named Entity Recognition๐ง Natural Language Processing
Jupyter Notebookno licenseupdated Feb 11, 2023
Tokenization & Preprocessing
jfilter/clean-text
๐งน Python package for text cleaning
๐ง Natural Language Processing
PythonOtherupdated May 15, 2026
Data from GitHub ยท snapshot Sep 24, 2026