๐Ÿ† #2,826 overall#121 of 148 in Tokenization & Preprocessing

kyegomez /MambaByte

Implementation of MambaByte in "MambaByte: Token-free Selective State Space Model" in Pytorch and Zeta

$ git clone https://github.com/kyegomez/MambaByte.git
GitHub social preview for kyegomez/MambaByte
Stars
129
129
Forks
9
9
Language
Python
License
MIT
Created
Jan 26, 2024
2.7 years old
Last push
Aug 28, 2026

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

explosion/spaCy

๐Ÿ’ซ Industrial-strength Natural Language Processing (NLP) in Python

๐Ÿท๏ธ Named Entity Recognition๐Ÿ—‚๏ธ Text Classification๐Ÿง  Natural Language Processing
PythonMITupdated Aug 24, 2026
GitHub โ†—โ˜… 33.9Kโ‘‚ 4.7K
Tokenization & Preprocessing

analyticalrohit/llms-from-scratch

Build a ChatGPT like LLM from scratch in PyTorch, explained step by step.

๐Ÿค– Language Models
Jupyter NotebookMITupdated Jun 8, 2026
GitHub โ†—โ˜… 383โ‘‚ 87
Tokenization & Preprocessing

jshuadvd/LongRoPE

Implementation of the LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens Paper

๐Ÿค– Language Models๐Ÿง  Natural Language Processing
Pythonno licenseupdated Jul 20, 2024
GitHub โ†—โ˜… 154โ‘‚ 13
Tokenization & Preprocessing

rasbt/LLMs-from-scratch๐Ÿ”ฅ active

Implement a ChatGPT-like LLM in PyTorch from scratch, step by step

๐Ÿค– Language Models๐Ÿง  Natural Language Processing
Jupyter NotebookOtherupdated Sep 22, 2026
GitHub โ†—โ˜… 105.5Kโ‘‚ 16.2K
Tokenization & Preprocessing

nlp-uoregon/trankit

Trankit is a Light-Weight Transformer-based Python Toolkit for Multilingual Natural Language Processing

๐Ÿค– Language Models๐Ÿง  Natural Language Processing
PythonApache-2.0updated Jul 22, 2025
GitHub โ†—โ˜… 799โ‘‚ 106
Tokenization & Preprocessing

microsoft/Tokenizer

Typescript and .NET implementation of BPE tokenizer for OpenAI LLMs.

C#MITupdated Sep 10, 2026
GitHub โ†—โ˜… 213โ‘‚ 37

Data from GitHub ยท snapshot Sep 24, 2026