๐Ÿ† #1,848 overall#59 of 148 in Tokenization & Preprocessing

ScrapeGraphAI /toonify

Toonify: Compact data format reducing LLM token usage by 30-60%

$ git clone https://github.com/ScrapeGraphAI/toonify.git
GitHub social preview for ScrapeGraphAI/toonify
Stars
362
362
Forks
29
29
Language
Python
License
None
Created
Nov 11, 2025
0.9 years old
Last push
Feb 6, 2026

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

toon-format/toon

๐ŸŽ’ Token-Oriented Object Notation (TOON) โ€“ compact, human-readable serialization of JSON data for LLM prompts. TypeScript SDK, CLI, benchmarks.

TypeScriptMITupdated Sep 3, 2026
GitHub โ†—โ˜… 25.4Kโ‘‚ 1.1K
Tokenization & Preprocessing

marcelroed/gigatoken

Language model tokenization at GB/s

๐Ÿง  Natural Language Processing
RustMITupdated Sep 2, 2026
GitHub โ†—โ˜… 4.1Kโ‘‚ 220
Tokenization & Preprocessing

AgentOps-AI/tokencost

Easy token price estimates for 400+ LLMs. TokenOps.

๐Ÿค– Language Models
PythonMITupdated Sep 5, 2025
GitHub โ†—โ˜… 2Kโ‘‚ 106
Tokenization & Preprocessing

ImadSaddik/Train_Your_Language_Model_Course

Train a language model to chat like you using your personal conversations from WhatsApp, Telegram, Signal, or other platforms.

Jupyter Notebookno licenseupdated Sep 26, 2025
GitHub โ†—โ˜… 287โ‘‚ 157
Tokenization & Preprocessing

alpkeskin/gotoon

Token-Oriented Object Notation for Go โ€“ JSON for LLMs at half the token cost

GoMITupdated Nov 24, 2025
GitHub โ†—โ˜… 241โ‘‚ 14
Tokenization & Preprocessing

ash-01xor/bpe.c

Simple Byte pair Encoding mechanism used for tokenization process . written purely in C

CMITupdated Nov 11, 2024
GitHub โ†—โ˜… 155โ‘‚ 4

Data from GitHub ยท snapshot Sep 24, 2026