๐Ÿ† #1,223 overall#26 of 148 in Tokenization & Preprocessing

nlp-uoregon /trankit

Trankit is a Light-Weight Transformer-based Python Toolkit for Multilingual Natural Language Processing

$ git clone https://github.com/nlp-uoregon/trankit.git
GitHub social preview for nlp-uoregon/trankit
Stars
799
799
Forks
106
106
Language
Python
License
Apache-2.0
Created
Jan 8, 2021
5.7 years old
Last push
Jul 22, 2025

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

jshuadvd/LongRoPE

Implementation of the LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens Paper

๐Ÿค– Language Models๐Ÿง  Natural Language Processing
Pythonno licenseupdated Jul 20, 2024
GitHub โ†—โ˜… 154โ‘‚ 13
Tokenization & Preprocessing

explosion/spaCy

๐Ÿ’ซ Industrial-strength Natural Language Processing (NLP) in Python

๐Ÿท๏ธ Named Entity Recognition๐Ÿ—‚๏ธ Text Classification๐Ÿง  Natural Language Processing
PythonMITupdated Aug 24, 2026
GitHub โ†—โ˜… 33.9Kโ‘‚ 4.7K
Tokenization & Preprocessing

explosion/spacy-streamlit

๐Ÿ‘‘ spaCy building blocks and visualizers for Streamlit apps

๐Ÿท๏ธ Named Entity Recognition๐Ÿ—‚๏ธ Text Classification๐Ÿง  Natural Language Processing
PythonMITupdated Jul 29, 2024
GitHub โ†—โ˜… 863โ‘‚ 113
Tokenization & Preprocessing

CogComp/cogcomp-nlp

CogComp's Natural Language Processing Libraries and Demos: Modules include lemmatizer, ner, pos, prep-srl, quantifier, question type, relation-extraction, similarity, temporal normalizer, tokenizer, transliteration, verb-sense, and more.

๐Ÿท๏ธ Named Entity Recognitionโ›๏ธ Information Extraction๐Ÿง  Natural Language Processing
JavaOtherupdated Jul 7, 2023
GitHub โ†—โ˜… 479โ‘‚ 143
Tokenization & Preprocessing

janlukasschroeder/nlp-cheat-sheet-python

NLP Cheat Sheet, Python, spacy, LexNPL, NLTK, tokenization, stemming, sentence detection, named entity recognition

๐Ÿท๏ธ Named Entity Recognition๐Ÿง  Natural Language Processing
Jupyter Notebookno licenseupdated Feb 11, 2023
GitHub โ†—โ˜… 260โ‘‚ 77
Tokenization & Preprocessing

mit-ccc/TweebankNLP

[LREC 2022] An off-the-shelf pre-trained Tweet NLP Toolkit (NER, tokenization, lemmatization, POS tagging, dependency parsing) + Tweebank-NER dataset

๐Ÿท๏ธ Named Entity Recognition๐Ÿง  Natural Language Processing
PythonApache-2.0updated Jan 24, 2024
GitHub โ†—โ˜… 106โ‘‚ 10

Data from GitHub ยท snapshot Sep 24, 2026