๐Ÿ† #2,948 overall#133 of 148 in Tokenization & Preprocessing

frothywater /kanade-tokenizer

Kanade is a single-layer disentangled speech tokenizer that extracts compact tokens suitable for both generative and discriminative modeling.

$ git clone https://github.com/frothywater/kanade-tokenizer.git
GitHub social preview for frothywater/kanade-tokenizer
Stars
116
116
Forks
16
16
Language
Python
License
None
Created
Oct 30, 2025
0.9 years old
Last push
Jul 18, 2026

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

guillaume-be/rust-tokenizers

Rust-tokenizer offers high-performance tokenizers for modern language models, including WordPiece, Byte-Pair Encoding (BPE) and Unigram (SentencePiece) models

RustApache-2.0updated Jan 22, 2026
GitHub โ†—โ˜… 344โ‘‚ 34
Tokenization & Preprocessing

sugarme/tokenizer

NLP tokenizers written in Go language

๐Ÿง  Natural Language Processing
GoApache-2.0updated May 26, 2026
GitHub โ†—โ˜… 335โ‘‚ 69
Tokenization & Preprocessing

amazon-far/deltatok

[CVPR 2026 Highlight] A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens

PythonApache-2.0updated Jul 17, 2026
GitHub โ†—โ˜… 261โ‘‚ 7
Tokenization & Preprocessing

OpenMOSS/MOSS-Audio-Tokenizer

A 1.6B causal Transformer audio tokenizer with streaming, variable bitrates, and semantic alignment across speech, sound, and music

PythonApache-2.0updated Jun 16, 2026
GitHub โ†—โ˜… 258โ‘‚ 19
Tokenization & Preprocessing

rasbt/LLMs-from-scratch๐Ÿ”ฅ active

Implement a ChatGPT-like LLM in PyTorch from scratch, step by step

๐Ÿค– Language Models๐Ÿง  Natural Language Processing
Jupyter NotebookOtherupdated Sep 22, 2026
GitHub โ†—โ˜… 105.5Kโ‘‚ 16.2K
Tokenization & Preprocessing

explosion/spaCy

๐Ÿ’ซ Industrial-strength Natural Language Processing (NLP) in Python

๐Ÿท๏ธ Named Entity Recognition๐Ÿ—‚๏ธ Text Classification๐Ÿง  Natural Language Processing
PythonMITupdated Aug 24, 2026
GitHub โ†—โ˜… 33.9Kโ‘‚ 4.7K

Data from GitHub ยท snapshot Sep 24, 2026