๐Ÿ† #2,332 overall#85 of 148 in Tokenization & Preprocessing๐Ÿ”ฅ active this week

adbar /simplemma

Simple multilingual lemmatizer for Python, especially useful for speed and efficiency

$ git clone https://github.com/adbar/simplemma.git
GitHub social preview for adbar/simplemma
Stars
219
219
Forks
17
17
Language
Python
License
MIT
Created
Jan 18, 2021
5.7 years old
Last push
Sep 21, 2026
๐Ÿ”ฅ this week

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

daac-tools/vibrato๐Ÿ”ฅ active

๐ŸŽค vibrato: Viterbi-based accelerated tokenizer

๐Ÿง  Natural Language Processing
RustApache-2.0updated Sep 19, 2026
GitHub โ†—โ˜… 423โ‘‚ 26
Tokenization & Preprocessing

daac-tools/vaporetto

๐Ÿ›ฅ Vaporetto: Very accelerated pointwise prediction based tokenizer

๐Ÿง  Natural Language Processing
RustApache-2.0updated Jul 20, 2026
GitHub โ†—โ˜… 299โ‘‚ 12
Tokenization & Preprocessing

adbar/trafilatura๐Ÿ”ฅ active

Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML

๐Ÿง  Natural Language Processing๐Ÿ’Ž Text Mining
PythonApache-2.0updated Sep 21, 2026
GitHub โ†—โ˜… 6.9Kโ‘‚ 436
Tokenization & Preprocessing

marcelroed/gigatoken

Language model tokenization at GB/s

๐Ÿง  Natural Language Processing
RustMITupdated Sep 2, 2026
GitHub โ†—โ˜… 4.1Kโ‘‚ 220
Tokenization & Preprocessing

roshan-research/hazm

Persian NLP Toolkit

๐Ÿง  Natural Language Processing
PythonMITupdated Apr 1, 2026
GitHub โ†—โ˜… 1.4Kโ‘‚ 206
Tokenization & Preprocessing

cbaziotis/ekphrasis

Ekphrasis is a text processing tool, geared towards text from social networks, such as Twitter or Facebook. Ekphrasis performs tokenization, word normalization, word segmentation (for splitting hashtags) and spell correction, using word statistics from 2 big corpora (english Wikipedia, twitter - 330mil english tweets).

๐Ÿง  Natural Language Processing
PythonMITupdated Jun 2, 2025
GitHub โ†—โ˜… 673โ‘‚ 92

Data from GitHub ยท snapshot Sep 24, 2026