๐Ÿ† #1,170 overall#25 of 148 in Tokenization & Preprocessing

niieani /gpt-tokenizer

The fastest JavaScript BPE Tokenizer Encoder Decoder for OpenAI's GPT models (gpt-5, gpt-o*, gpt-4o, etc.). Port of OpenAI's tiktoken with additional features.

$ git clone https://github.com/niieani/gpt-tokenizer.git
GitHub social preview for niieani/gpt-tokenizer
Stars
849
849
Forks
59
59
Language
TypeScript
License
MIT
Created
Mar 22, 2023
3.5 years old
Last push
Aug 16, 2026

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

dmitry-brazhenko/SharpToken

SharpToken is a C# library for tokenizing natural language text. It's based on the tiktoken Python library and designed to be fast and accurate.

C#MITupdated Mar 25, 2026
GitHub โ†—โ˜… 262โ‘‚ 19
Tokenization & Preprocessing

explosion/spaCy

๐Ÿ’ซ Industrial-strength Natural Language Processing (NLP) in Python

๐Ÿท๏ธ Named Entity Recognition๐Ÿ—‚๏ธ Text Classification๐Ÿง  Natural Language Processing
PythonMITupdated Aug 24, 2026
GitHub โ†—โ˜… 33.9Kโ‘‚ 4.7K
Tokenization & Preprocessing

jbesomi/texthero

Text preprocessing, representation and visualization from zero to hero.

๐Ÿงฌ Embeddings๐Ÿง  Natural Language Processing๐Ÿ’Ž Text Mining
PythonMITupdated Aug 29, 2023
GitHub โ†—โ˜… 2.9Kโ‘‚ 236
Tokenization & Preprocessing

explosion/spacy-streamlit

๐Ÿ‘‘ spaCy building blocks and visualizers for Streamlit apps

๐Ÿท๏ธ Named Entity Recognition๐Ÿ—‚๏ธ Text Classification๐Ÿง  Natural Language Processing
PythonMITupdated Jul 29, 2024
GitHub โ†—โ˜… 863โ‘‚ 113
Tokenization & Preprocessing

nlp-uoregon/trankit

Trankit is a Light-Weight Transformer-based Python Toolkit for Multilingual Natural Language Processing

๐Ÿค– Language Models๐Ÿง  Natural Language Processing
PythonApache-2.0updated Jul 22, 2025
GitHub โ†—โ˜… 799โ‘‚ 106
Tokenization & Preprocessing

zurawiki/tiktoken-rs

Ready-made tokenizer library for working with GPT and tiktoken

RustMITupdated Jul 1, 2026
GitHub โ†—โ˜… 411โ‘‚ 75

Data from GitHub ยท snapshot Sep 24, 2026