๐Ÿ† #2,510 overall#98 of 148 in Tokenization & Preprocessing

johannschopplich /tokenx

๐Ÿ“ Fast token estimation with 95%+ average accuracy in a 2kB bundle

$ git clone https://github.com/johannschopplich/tokenx.git
GitHub social preview for johannschopplich/tokenx
Stars
178
178
Forks
7
7
Language
TypeScript
License
MIT
Created
Nov 27, 2023
2.8 years old
Last push
Sep 3, 2026

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

marcelroed/gigatoken

Language model tokenization at GB/s

๐Ÿง  Natural Language Processing
RustMITupdated Sep 2, 2026
GitHub โ†—โ˜… 4.1Kโ‘‚ 220
Tokenization & Preprocessing

dqbd/tiktokenizer

Online playground for OpenAPI tokenizers

TypeScriptMITupdated Apr 24, 2025
GitHub โ†—โ˜… 1.7Kโ‘‚ 187
Tokenization & Preprocessing

cbaziotis/ekphrasis

Ekphrasis is a text processing tool, geared towards text from social networks, such as Twitter or Facebook. Ekphrasis performs tokenization, word normalization, word segmentation (for splitting hashtags) and spell correction, using word statistics from 2 big corpora (english Wikipedia, twitter - 330mil english tweets).

๐Ÿง  Natural Language Processing
PythonMITupdated Jun 2, 2025
GitHub โ†—โ˜… 673โ‘‚ 92
Tokenization & Preprocessing

alasdairforsythe/tokenmonster

Ungreedy subword tokenizer and vocabulary trainer for Python, Go, C++ & Javascript

GoMITupdated Jul 10, 2026
GitHub โ†—โ˜… 627โ‘‚ 23
Tokenization & Preprocessing

daac-tools/vibrato๐Ÿ”ฅ active

๐ŸŽค vibrato: Viterbi-based accelerated tokenizer

๐Ÿง  Natural Language Processing
RustApache-2.0updated Sep 19, 2026
GitHub โ†—โ˜… 423โ‘‚ 26
Tokenization & Preprocessing

OpenNMT/Tokenizer

Fast and customizable text tokenization library with BPE and SentencePiece support

๐ŸŒ Machine Translation๐Ÿง  Natural Language Processing
C++MITupdated Jan 10, 2026
GitHub โ†—โ˜… 340โ‘‚ 84

Data from GitHub ยท snapshot Sep 24, 2026