Tokenization & Preprocessing
marcelroed/gigatoken
Language model tokenization at GB/s
๐ง Natural Language Processing
RustMITupdated Sep 2, 2026
๐ Fast token estimation with 95%+ average accuracy in a 2kB bundle
$ git clone https://github.com/johannschopplich/tokenx.gitLanguage model tokenization at GB/s
Ekphrasis is a text processing tool, geared towards text from social networks, such as Twitter or Facebook. Ekphrasis performs tokenization, word normalization, word segmentation (for splitting hashtags) and spell correction, using word statistics from 2 big corpora (english Wikipedia, twitter - 330mil english tweets).
Ungreedy subword tokenizer and vocabulary trainer for Python, Go, C++ & Javascript
๐ค vibrato: Viterbi-based accelerated tokenizer
Fast and customizable text tokenization library with BPE and SentencePiece support
Data from GitHub ยท snapshot Sep 24, 2026