๐Ÿ† #1,264 overall#27 of 148 in Tokenization & Preprocessing

BLKSerene /Wordless

An Integrated Corpus Tool With Multilingual Support for the Study of Language, Literature, and Translation

$ git clone https://github.com/BLKSerene/Wordless.git
GitHub social preview for BLKSerene/Wordless
Stars
758
758
Forks
99
99
Language
Python
License
GPL-3.0
Created
Aug 23, 2018
8.1 years old
Last push
Oct 30, 2025

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

bitextor/bitextor

Bitextor generates translation memories from multilingual websites

๐ŸŒ Machine Translation
PythonGPL-3.0updated Nov 11, 2024
GitHub โ†—โ˜… 300โ‘‚ 41
Tokenization & Preprocessing

adbar/simplemma๐Ÿ”ฅ active

Simple multilingual lemmatizer for Python, especially useful for speed and efficiency

๐Ÿง  Natural Language Processing
PythonMITupdated Sep 21, 2026
GitHub โ†—โ˜… 219โ‘‚ 17
Tokenization & Preprocessing

Dadmatech/DadmaTools

DadmaTools is a Persian NLP tools developed by Dadmatech Co.

๐Ÿท๏ธ Named Entity Recognition๐Ÿง  Natural Language Processing
PythonApache-2.0updated Aug 14, 2025
GitHub โ†—โ˜… 214โ‘‚ 45
Tokenization & Preprocessing

adbar/trafilatura๐Ÿ”ฅ active

Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML

๐Ÿง  Natural Language Processing๐Ÿ’Ž Text Mining
PythonApache-2.0updated Sep 21, 2026
GitHub โ†—โ˜… 6.9Kโ‘‚ 436
Tokenization & Preprocessing

roshan-research/hazm

Persian NLP Toolkit

๐Ÿง  Natural Language Processing
PythonMITupdated Apr 1, 2026
GitHub โ†—โ˜… 1.4Kโ‘‚ 206
Tokenization & Preprocessing

adobe/NLP-Cube

Natural Language Processing Pipeline - Sentence Splitting, Tokenization, Lemmatization, Part-of-speech Tagging and Dependency Parsing

โ›๏ธ Information Extraction๐ŸŒ Machine Translation
HTMLApache-2.0updated Nov 3, 2024
GitHub โ†—โ˜… 564โ‘‚ 92

Data from GitHub ยท snapshot Sep 24, 2026