Tokenization & Preprocessing
marcelroed/gigatoken
Language model tokenization at GB/s
๐ง Natural Language Processing
RustMITupdated Sep 2, 2026
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
$ git clone https://github.com/adbar/trafilatura.gitLanguage model tokenization at GB/s
Text2Text Language Modeling Toolkit
Simple multilingual lemmatizer for Python, especially useful for speed and efficiency
Language AI Engineering Lab, a place where you can deeply understand and build modern Language AI systems, from fundamentals to production.
Token Cost Parity: Multilingual LLM Efficiency Analysis 2026
Text preprocessing, representation and visualization from zero to hero.
Data from GitHub ยท snapshot Sep 24, 2026