Tokenization & Preprocessing
marcelroed/gigatoken
Language model tokenization at GB/s
๐ง Natural Language Processing
RustMITupdated Sep 2, 2026
A tokenizer based on Unicode text segmentation (UAX #29), for Go. Split graphemes, words, sentences.
$ git clone https://github.com/clipperhouse/uax29.gitLanguage model tokenization at GB/s
Ekphrasis is a text processing tool, geared towards text from social networks, such as Twitter or Facebook. Ekphrasis performs tokenization, word normalization, word segmentation (for splitting hashtags) and spell correction, using word statistics from 2 big corpora (english Wikipedia, twitter - 330mil english tweets).
๐ค vibrato: Viterbi-based accelerated tokenizer
Fast and customizable text tokenization library with BPE and SentencePiece support
๐ฅ Vaporetto: Very accelerated pointwise prediction based tokenizer
Tokenizers and lemmatizers for Go
Data from GitHub ยท snapshot Sep 24, 2026