๐Ÿ† #2,671 overall#113 of 148 in Tokenization & Preprocessing

THUDM /icetk

A unified tokenization tool for Images, Chinese and English.

$ git clone https://github.com/THUDM/icetk.git
GitHub social preview for THUDM/icetk
Stars
152
152
Forks
16
16
Language
Python
License
None
Created
Dec 22, 2021
4.8 years old
Last push
Mar 23, 2023

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

ImadSaddik/Train_Your_Language_Model_Course

Train a language model to chat like you using your personal conversations from WhatsApp, Telegram, Signal, or other platforms.

Jupyter Notebookno licenseupdated Sep 26, 2025
GitHub โ†—โ˜… 287โ‘‚ 157
Tokenization & Preprocessing

toon-format/toon

๐ŸŽ’ Token-Oriented Object Notation (TOON) โ€“ compact, human-readable serialization of JSON data for LLM prompts. TypeScript SDK, CLI, benchmarks.

TypeScriptMITupdated Sep 3, 2026
GitHub โ†—โ˜… 25.4Kโ‘‚ 1.1K
Tokenization & Preprocessing

marcelroed/gigatoken

Language model tokenization at GB/s

๐Ÿง  Natural Language Processing
RustMITupdated Sep 2, 2026
GitHub โ†—โ˜… 4.1Kโ‘‚ 220
Tokenization & Preprocessing

AgentOps-AI/tokencost

Easy token price estimates for 400+ LLMs. TokenOps.

๐Ÿค– Language Models
PythonMITupdated Sep 5, 2025
GitHub โ†—โ˜… 2Kโ‘‚ 106

Data from GitHub ยท snapshot Sep 24, 2026