πŸ† #1,355 overall#31 of 148 in Tokenization & Preprocessing

cbaziotis /ekphrasis

Ekphrasis is a text processing tool, geared towards text from social networks, such as Twitter or Facebook. Ekphrasis performs tokenization, word normalization, word segmentation (for splitting hashtags) and spell correction, using word statistics from 2 big corpora (english Wikipedia, twitter - 330mil english tweets).

$ git clone https://github.com/cbaziotis/ekphrasis.git
GitHub social preview for cbaziotis/ekphrasis
Stars
673
673
Forks
92
92
Language
Python
License
MIT
Created
Feb 7, 2017
9.6 years old
Last push
Jun 2, 2025

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

marcelroed/gigatoken

Language model tokenization at GB/s

🧠 Natural Language Processing
RustMITupdated Sep 2, 2026
GitHub β†—β˜… 4.1Kβ‘‚ 220
Tokenization & Preprocessing

daac-tools/vibratoπŸ”₯ active

🎀 vibrato: Viterbi-based accelerated tokenizer

🧠 Natural Language Processing
RustApache-2.0updated Sep 19, 2026
GitHub β†—β˜… 423β‘‚ 26
Tokenization & Preprocessing

daac-tools/vaporetto

πŸ›₯ Vaporetto: Very accelerated pointwise prediction based tokenizer

🧠 Natural Language Processing
RustApache-2.0updated Jul 20, 2026
GitHub β†—β˜… 299β‘‚ 12
Tokenization & Preprocessing

clipperhouse/uax29

A tokenizer based on Unicode text segmentation (UAX #29), for Go. Split graphemes, words, sentences.

🧠 Natural Language Processing
GoMITupdated Feb 16, 2026
GitHub β†—β˜… 123β‘‚ 7
Tokenization & Preprocessing

natasha/natasha

Solves basic Russian NLP tasks, API for lower level Natasha projects

🏷️ Named Entity Recognition🧠 Natural Language Processing
PythonMITupdated Apr 13, 2026
GitHub β†—β˜… 1.4Kβ‘‚ 121
Tokenization & Preprocessing

lovit/soynlp

ν•œκ΅­μ–΄ μžμ—°μ–΄μ²˜λ¦¬λ₯Ό μœ„ν•œ 파이썬 λΌμ΄λΈŒλŸ¬λ¦¬μž…λ‹ˆλ‹€. 단어 μΆ”μΆœ/ ν† ν¬λ‚˜μ΄μ € / ν’ˆμ‚¬νŒλ³„/ μ „μ²˜λ¦¬μ˜ κΈ°λŠ₯을 μ œκ³΅ν•©λ‹ˆλ‹€.

🧠 Natural Language Processing
PythonOtherupdated Mar 10, 2026
GitHub β†—β˜… 993β‘‚ 183

Data from GitHub Β· snapshot Sep 24, 2026