๐Ÿ† #2,068 overall#70 of 148 in Tokenization & Preprocessing

ImadSaddik /Train_Your_Language_Model_Course

Train a language model to chat like you using your personal conversations from WhatsApp, Telegram, Signal, or other platforms.

$ git clone https://github.com/ImadSaddik/Train_Your_Language_Model_Course.git
GitHub social preview for ImadSaddik/Train_Your_Language_Model_Course
Stars
287
287
Forks
157
157
Language
Jupyter Notebook
License
None
Created
Feb 12, 2025
1.6 years old
Last push
Sep 26, 2025

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

toon-format/toon

๐ŸŽ’ Token-Oriented Object Notation (TOON) โ€“ compact, human-readable serialization of JSON data for LLM prompts. TypeScript SDK, CLI, benchmarks.

TypeScriptMITupdated Sep 3, 2026
GitHub โ†—โ˜… 25.4Kโ‘‚ 1.1K
Tokenization & Preprocessing

marcelroed/gigatoken

Language model tokenization at GB/s

๐Ÿง  Natural Language Processing
RustMITupdated Sep 2, 2026
GitHub โ†—โ˜… 4.1Kโ‘‚ 220
Tokenization & Preprocessing

AgentOps-AI/tokencost

Easy token price estimates for 400+ LLMs. TokenOps.

๐Ÿค– Language Models
PythonMITupdated Sep 5, 2025
GitHub โ†—โ˜… 2Kโ‘‚ 106
Tokenization & Preprocessing

ScrapeGraphAI/toonify

Toonify: Compact data format reducing LLM token usage by 30-60%

Pythonno licenseupdated Feb 6, 2026
GitHub โ†—โ˜… 362โ‘‚ 29
Tokenization & Preprocessing

OpenNMT/Tokenizer

Fast and customizable text tokenization library with BPE and SentencePiece support

๐ŸŒ Machine Translation๐Ÿง  Natural Language Processing
C++MITupdated Jan 10, 2026
GitHub โ†—โ˜… 340โ‘‚ 84
Tokenization & Preprocessing

alpkeskin/gotoon

Token-Oriented Object Notation for Go โ€“ JSON for LLMs at half the token cost

GoMITupdated Nov 24, 2025
GitHub โ†—โ˜… 241โ‘‚ 14

Data from GitHub ยท snapshot Sep 24, 2026