πŸ† #2,759 overall#119 of 148 in Tokenization & Preprocessing

Kensuke-Mitsuzawa /JapaneseTokenizers

aim to use JapaneseTokenizer as easy as possible

$ git clone https://github.com/Kensuke-Mitsuzawa/JapaneseTokenizers.git
GitHub social preview for Kensuke-Mitsuzawa/JapaneseTokenizers
Stars
138
138
Forks
21
21
Language
Python
License
MIT
Created
Sep 1, 2015
11.1 years old
Last push
Mar 25, 2019

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

polm/fugashi

A Cython MeCab wrapper for fast, pythonic Japanese tokenization and morphological analysis.

🧠 Natural Language Processing
C++MITupdated Oct 24, 2025
GitHub β†—β˜… 535β‘‚ 40
Tokenization & Preprocessing

ku-nlp/jumanpp

Juman++ (a Morphological Analyzer Toolkit)

🧠 Natural Language Processing
C++Apache-2.0updated Apr 17, 2026
GitHub β†—β˜… 413β‘‚ 47
Tokenization & Preprocessing

marcelroed/gigatoken

Language model tokenization at GB/s

🧠 Natural Language Processing
RustMITupdated Sep 2, 2026
GitHub β†—β˜… 4.1Kβ‘‚ 220
Tokenization & Preprocessing

natasha/natasha

Solves basic Russian NLP tasks, API for lower level Natasha projects

🏷️ Named Entity Recognition🧠 Natural Language Processing
PythonMITupdated Apr 13, 2026
GitHub β†—β˜… 1.4Kβ‘‚ 121
Tokenization & Preprocessing

lovit/soynlp

ν•œκ΅­μ–΄ μžμ—°μ–΄μ²˜λ¦¬λ₯Ό μœ„ν•œ 파이썬 λΌμ΄λΈŒλŸ¬λ¦¬μž…λ‹ˆλ‹€. 단어 μΆ”μΆœ/ ν† ν¬λ‚˜μ΄μ € / ν’ˆμ‚¬νŒλ³„/ μ „μ²˜λ¦¬μ˜ κΈ°λŠ₯을 μ œκ³΅ν•©λ‹ˆλ‹€.

🧠 Natural Language Processing
PythonOtherupdated Mar 10, 2026
GitHub β†—β˜… 993β‘‚ 183
Tokenization & Preprocessing

cbaziotis/ekphrasis

Ekphrasis is a text processing tool, geared towards text from social networks, such as Twitter or Facebook. Ekphrasis performs tokenization, word normalization, word segmentation (for splitting hashtags) and spell correction, using word statistics from 2 big corpora (english Wikipedia, twitter - 330mil english tweets).

🧠 Natural Language Processing
PythonMITupdated Jun 2, 2025
GitHub β†—β˜… 673β‘‚ 92

Data from GitHub Β· snapshot Sep 24, 2026