🏆 #3,079 overall#144 of 148 in Tokenization & Preprocessing

yishn /chinese-tokenizer

Tokenizes Chinese texts into words.

$ git clone https://github.com/yishn/chinese-tokenizer.git
GitHub social preview for yishn/chinese-tokenizer
Stars
102
102
Forks
25
25
Language
JavaScript
License
MIT
Created
Sep 14, 2016
10.0 years old
Last push
Dec 21, 2022

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

wangfenjin/simple

支持中文和拼音的 SQLite fts5 全文搜索扩展 | A SQLite3 fts5 tokenizer which supports Chinese and PinYin

C++Otherupdated May 17, 2026
GitHub ↗★ 866⑂ 113
Tokenization & Preprocessing

risesoft-y9/Data-Labeling

数据标注是一款专门对文本数据进行处理和标注的工具,通过简化快捷的文本标注流程和动态的算法反馈,支持用户快速标注关键词并能通过算法持续减少人工标注的成本和时间。数据标注的过程先由人工标注构建基础,再由自动标注反哺人工标注,最后由人工标注进行纠偏,从而大幅度提高标注的精准度和高效性。数据标注需要依赖开源的数字底座进行人员岗位管控。

JavaGPL-3.0updated Jun 23, 2025
GitHub ↗★ 700⑂ 106
Tokenization & Preprocessing

xinjli/transphone

phoneme tokenizer and grapheme-to-phoneme model for 8k languages

PythonMITupdated Jun 9, 2023
GitHub ↗★ 176⑂ 19
Tokenization & Preprocessing

theseer/tokenizer

A small library for converting tokenized PHP source code into XML (and potentially other formats)

PHPOtherupdated Feb 3, 2026
GitHub ↗★ 5.2K⑂ 25
Tokenization & Preprocessing

marcelroed/gigatoken

Language model tokenization at GB/s

🧠 Natural Language Processing
RustMITupdated Sep 2, 2026
GitHub ↗★ 4.1K⑂ 220
Tokenization & Preprocessing

andialbrecht/sqlparse

A non-validating SQL parser module for Python

PythonBSD-3-Clauseupdated Aug 13, 2026
GitHub ↗★ 4K⑂ 755

Data from GitHub · snapshot Sep 24, 2026