🏆 #2,930 overall#131 of 148 in Tokenization & Preprocessing

taishan1994 /sentencepiece_chinese_bpe

使用sentencepiece中BPE训练中文词表,并在transformers中进行使用。

$ git clone https://github.com/taishan1994/sentencepiece_chinese_bpe.git
GitHub social preview for taishan1994/sentencepiece_chinese_bpe
Stars
118
118
Forks
17
17
Language
Python
License
None
Created
Jun 24, 2023
3.3 years old
Last push
Jun 24, 2023

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

OpenNMT/Tokenizer

Fast and customizable text tokenization library with BPE and SentencePiece support

🌍 Machine Translation🧠 Natural Language Processing
C++MITupdated Jan 10, 2026
GitHub ↗★ 340⑂ 84
Tokenization & Preprocessing

toon-format/toon

🎒 Token-Oriented Object Notation (TOON) – compact, human-readable serialization of JSON data for LLM prompts. TypeScript SDK, CLI, benchmarks.

TypeScriptMITupdated Sep 3, 2026
GitHub ↗★ 25.4K⑂ 1.1K
Tokenization & Preprocessing

marcelroed/gigatoken

Language model tokenization at GB/s

🧠 Natural Language Processing
RustMITupdated Sep 2, 2026
GitHub ↗★ 4.1K⑂ 220
Tokenization & Preprocessing

AgentOps-AI/tokencost

Easy token price estimates for 400+ LLMs. TokenOps.

🤖 Language Models
PythonMITupdated Sep 5, 2025
GitHub ↗★ 2K⑂ 106
Tokenization & Preprocessing

cbaziotis/ekphrasis

Ekphrasis is a text processing tool, geared towards text from social networks, such as Twitter or Facebook. Ekphrasis performs tokenization, word normalization, word segmentation (for splitting hashtags) and spell correction, using word statistics from 2 big corpora (english Wikipedia, twitter - 330mil english tweets).

🧠 Natural Language Processing
PythonMITupdated Jun 2, 2025
GitHub ↗★ 673⑂ 92

Data from GitHub · snapshot Sep 24, 2026