๐Ÿ† #2,066 overall#69 of 148 in Tokenization & Preprocessing

natasha /razdel

Rule-based token, sentence segmentation for Russian language

$ git clone https://github.com/natasha/razdel.git
GitHub social preview for natasha/razdel
Stars
288
288
Forks
34
34
Language
Python
License
MIT
Created
Nov 10, 2018
7.9 years old
Last push
Apr 13, 2026

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

natasha/natasha

Solves basic Russian NLP tasks, API for lower level Natasha projects

๐Ÿท๏ธ Named Entity Recognition๐Ÿง  Natural Language Processing
PythonMITupdated Apr 13, 2026
GitHub โ†—โ˜… 1.4Kโ‘‚ 121
Tokenization & Preprocessing

marcelroed/gigatoken

Language model tokenization at GB/s

๐Ÿง  Natural Language Processing
RustMITupdated Sep 2, 2026
GitHub โ†—โ˜… 4.1Kโ‘‚ 220
Tokenization & Preprocessing

jfilter/clean-text

๐Ÿงน Python package for text cleaning

๐Ÿง  Natural Language Processing
PythonOtherupdated May 15, 2026
GitHub โ†—โ˜… 1Kโ‘‚ 82
Tokenization & Preprocessing

cbaziotis/ekphrasis

Ekphrasis is a text processing tool, geared towards text from social networks, such as Twitter or Facebook. Ekphrasis performs tokenization, word normalization, word segmentation (for splitting hashtags) and spell correction, using word statistics from 2 big corpora (english Wikipedia, twitter - 330mil english tweets).

๐Ÿง  Natural Language Processing
PythonMITupdated Jun 2, 2025
GitHub โ†—โ˜… 673โ‘‚ 92
Tokenization & Preprocessing

yooper/php-text-analysis

PHP Text Analysis is a library for performing Information Retrieval (IR) and Natural Language Processing (NLP) tasks using the PHP language

๐Ÿง  Natural Language Processing๐Ÿ’Ž Text Mining
PHPMITupdated Dec 28, 2024
GitHub โ†—โ˜… 534โ‘‚ 91
Tokenization & Preprocessing

NLPOptimize/flash-tokenizer

EFFICIENT AND OPTIMIZED TOKENIZER ENGINE FOR LLM INFERENCE SERVING

๐Ÿง  Natural Language Processing
C++no licenseupdated Feb 2, 2026
GitHub โ†—โ˜… 458โ‘‚ 11

Data from GitHub ยท snapshot Sep 24, 2026