๐Ÿ† #2,594 overall#104 of 148 in Tokenization & Preprocessing

lyeoni /prenlp

Preprocessing Library for Natural Language Processing

$ git clone https://github.com/lyeoni/prenlp.git
GitHub social preview for lyeoni/prenlp
Stars
164
164
Forks
12
12
Language
Python
License
Apache-2.0
Created
Nov 12, 2019
6.9 years old
Last push
Dec 6, 2022

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

jfilter/clean-text

๐Ÿงน Python package for text cleaning

๐Ÿง  Natural Language Processing
PythonOtherupdated May 15, 2026
GitHub โ†—โ˜… 1Kโ‘‚ 82
Tokenization & Preprocessing

jbesomi/texthero

Text preprocessing, representation and visualization from zero to hero.

๐Ÿงฌ Embeddings๐Ÿง  Natural Language Processing๐Ÿ’Ž Text Mining
PythonMITupdated Aug 29, 2023
GitHub โ†—โ˜… 2.9Kโ‘‚ 236
Tokenization & Preprocessing

roshan-research/hazm

Persian NLP Toolkit

๐Ÿง  Natural Language Processing
PythonMITupdated Apr 1, 2026
GitHub โ†—โ˜… 1.4Kโ‘‚ 206
Tokenization & Preprocessing

explosion/spacy-streamlit

๐Ÿ‘‘ spaCy building blocks and visualizers for Streamlit apps

๐Ÿท๏ธ Named Entity Recognition๐Ÿ—‚๏ธ Text Classification๐Ÿง  Natural Language Processing
PythonMITupdated Jul 29, 2024
GitHub โ†—โ˜… 863โ‘‚ 113
Tokenization & Preprocessing

cbaziotis/ekphrasis

Ekphrasis is a text processing tool, geared towards text from social networks, such as Twitter or Facebook. Ekphrasis performs tokenization, word normalization, word segmentation (for splitting hashtags) and spell correction, using word statistics from 2 big corpora (english Wikipedia, twitter - 330mil english tweets).

๐Ÿง  Natural Language Processing
PythonMITupdated Jun 2, 2025
GitHub โ†—โ˜… 673โ‘‚ 92
Tokenization & Preprocessing

open-korean-text/open-korean-text

Open Korean Text Processor - An Open-source Korean Text Processor

๐Ÿง  Natural Language Processing
ScalaApache-2.0updated Mar 12, 2024
GitHub โ†—โ˜… 669โ‘‚ 95

Data from GitHub ยท snapshot Sep 24, 2026