πŸ† #2,469 overall#96 of 148 in Tokenization & Preprocessing

ropensci /tokenizers

Fast, Consistent Tokenization of Natural Language Text

$ git clone https://github.com/ropensci/tokenizers.git
GitHub social preview for ropensci/tokenizers
Stars
188
188
Forks
23
23
Language
R
License
Other
Created
Mar 25, 2016
10.5 years old
Last push
Mar 27, 2024

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

bnosac/udpipe

R package for Tokenization, Parts of Speech Tagging, Lemmatization and Dependency Parsing Based on the UDPipe Natural Language Processing Toolkit

🧠 Natural Language ProcessingπŸ’Ž Text Mining
C++MPL-2.0updated Mar 2, 2026
GitHub β†—β˜… 224β‘‚ 35
Tokenization & Preprocessing

ropensci/hunspell

High-Performance Stemmer, Tokenizer, and Spell Checker for R

C++Otherupdated Feb 12, 2026
GitHub β†—β˜… 116β‘‚ 47
Tokenization & Preprocessing

jbesomi/texthero

Text preprocessing, representation and visualization from zero to hero.

🧬 Embeddings🧠 Natural Language ProcessingπŸ’Ž Text Mining
PythonMITupdated Aug 29, 2023
GitHub β†—β˜… 2.9Kβ‘‚ 236
Tokenization & Preprocessing

marcelroed/gigatoken

Language model tokenization at GB/s

🧠 Natural Language Processing
RustMITupdated Sep 2, 2026
GitHub β†—β˜… 4.1Kβ‘‚ 220
Tokenization & Preprocessing

natasha/natasha

Solves basic Russian NLP tasks, API for lower level Natasha projects

🏷️ Named Entity Recognition🧠 Natural Language Processing
PythonMITupdated Apr 13, 2026
GitHub β†—β˜… 1.4Kβ‘‚ 121
Tokenization & Preprocessing

lovit/soynlp

ν•œκ΅­μ–΄ μžμ—°μ–΄μ²˜λ¦¬λ₯Ό μœ„ν•œ 파이썬 λΌμ΄λΈŒλŸ¬λ¦¬μž…λ‹ˆλ‹€. 단어 μΆ”μΆœ/ ν† ν¬λ‚˜μ΄μ € / ν’ˆμ‚¬νŒλ³„/ μ „μ²˜λ¦¬μ˜ κΈ°λŠ₯을 μ œκ³΅ν•©λ‹ˆλ‹€.

🧠 Natural Language Processing
PythonOtherupdated Mar 10, 2026
GitHub β†—β˜… 993β‘‚ 183

Data from GitHub Β· snapshot Sep 24, 2026