Tokenization & Preprocessing
explosion/spaCy
๐ซ Industrial-strength Natural Language Processing (NLP) in Python
๐ท๏ธ Named Entity Recognition๐๏ธ Text Classification๐ง Natural Language Processing
PythonMITupdated Aug 24, 2026
Implementation of MambaByte in "MambaByte: Token-free Selective State Space Model" in Pytorch and Zeta
$ git clone https://github.com/kyegomez/MambaByte.git๐ซ Industrial-strength Natural Language Processing (NLP) in Python
Build a ChatGPT like LLM from scratch in PyTorch, explained step by step.
Implementation of the LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens Paper
Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
Trankit is a Light-Weight Transformer-based Python Toolkit for Multilingual Natural Language Processing
Typescript and .NET implementation of BPE tokenizer for OpenAI LLMs.
Data from GitHub ยท snapshot Sep 24, 2026