Tokenization & Preprocessing
wangfenjin/simple
支持中文和拼音的 SQLite fts5 全文搜索扩展 | A SQLite3 fts5 tokenizer which supports Chinese and PinYin
C++Otherupdated May 17, 2026
Tokenizes Chinese texts into words.
$ git clone https://github.com/yishn/chinese-tokenizer.git支持中文和拼音的 SQLite fts5 全文搜索扩展 | A SQLite3 fts5 tokenizer which supports Chinese and PinYin
数据标注是一款专门对文本数据进行处理和标注的工具,通过简化快捷的文本标注流程和动态的算法反馈,支持用户快速标注关键词并能通过算法持续减少人工标注的成本和时间。数据标注的过程先由人工标注构建基础,再由自动标注反哺人工标注,最后由人工标注进行纠偏,从而大幅度提高标注的精准度和高效性。数据标注需要依赖开源的数字底座进行人员岗位管控。
phoneme tokenizer and grapheme-to-phoneme model for 8k languages
A small library for converting tokenized PHP source code into XML (and potentially other formats)
Language model tokenization at GB/s
A non-validating SQL parser module for Python
Data from GitHub · snapshot Sep 24, 2026