🏆 #1,038 overall#19 of 148 in Tokenization & Preprocessing

lovit /soynlp

한국어 자연어처리를 위한 파이썬 라이브러리입니다. 단어 추출/ 토크나이저 / 품사판별/ 전처리의 기능을 제공합니다.

$ git clone https://github.com/lovit/soynlp.git
GitHub social preview for lovit/soynlp
Stars
993
993
Forks
183
183
Language
Python
License
Other
Created
May 13, 2017
9.4 years old
Last push
Mar 10, 2026

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

ongjin/garu🔥 active

초경량 한국어 형태소 분석기 — 1MB 모델로 브라우저에서 실행(WASM). Ultra-lightweight Korean morphological analyzer in the browser. npm: garu-ko

🧠 Natural Language Processing
PythonMITupdated Sep 21, 2026
GitHub ↗★ 177⑂ 9
Tokenization & Preprocessing

marcelroed/gigatoken

Language model tokenization at GB/s

🧠 Natural Language Processing
RustMITupdated Sep 2, 2026
GitHub ↗★ 4.1K⑂ 220
Tokenization & Preprocessing

natasha/natasha

Solves basic Russian NLP tasks, API for lower level Natasha projects

🏷️ Named Entity Recognition🧠 Natural Language Processing
PythonMITupdated Apr 13, 2026
GitHub ↗★ 1.4K⑂ 121
Tokenization & Preprocessing

cbaziotis/ekphrasis

Ekphrasis is a text processing tool, geared towards text from social networks, such as Twitter or Facebook. Ekphrasis performs tokenization, word normalization, word segmentation (for splitting hashtags) and spell correction, using word statistics from 2 big corpora (english Wikipedia, twitter - 330mil english tweets).

🧠 Natural Language Processing
PythonMITupdated Jun 2, 2025
GitHub ↗★ 673⑂ 92
Tokenization & Preprocessing

open-korean-text/open-korean-text

Open Korean Text Processor - An Open-source Korean Text Processor

🧠 Natural Language Processing
ScalaApache-2.0updated Mar 12, 2024
GitHub ↗★ 669⑂ 95
Tokenization & Preprocessing

polm/fugashi

A Cython MeCab wrapper for fast, pythonic Japanese tokenization and morphological analysis.

🧠 Natural Language Processing
C++MITupdated Oct 24, 2025
GitHub ↗★ 535⑂ 40

Data from GitHub · snapshot Sep 24, 2026