marcelroed/gigatoken
Language model tokenization at GB/s
Ekphrasis is a text processing tool, geared towards text from social networks, such as Twitter or Facebook. Ekphrasis performs tokenization, word normalization, word segmentation (for splitting hashtags) and spell correction, using word statistics from 2 big corpora (english Wikipedia, twitter - 330mil english tweets).
$ git clone https://github.com/cbaziotis/ekphrasis.gitLanguage model tokenization at GB/s
π€ vibrato: Viterbi-based accelerated tokenizer
π₯ Vaporetto: Very accelerated pointwise prediction based tokenizer
A tokenizer based on Unicode text segmentation (UAX #29), for Go. Split graphemes, words, sentences.
Solves basic Russian NLP tasks, API for lower level Natasha projects
νκ΅μ΄ μμ°μ΄μ²λ¦¬λ₯Ό μν νμ΄μ¬ λΌμ΄λΈλ¬λ¦¬μ λλ€. λ¨μ΄ μΆμΆ/ ν ν¬λμ΄μ / νμ¬νλ³/ μ μ²λ¦¬μ κΈ°λ₯μ μ 곡ν©λλ€.
Data from GitHub Β· snapshot Sep 24, 2026