๐Ÿ† #675 overall#283 of 1,075 in Natural Language Processing

nltk /nltk_data

NLTK Data

$ git clone https://github.com/nltk/nltk_data.git
GitHub social preview for nltk/nltk_data
Stars
1.8K
1,832
Forks
1.1K
1,089
Language
Python
License
Apache-2.0
Created
May 10, 2012
14.4 years old
Last push
Jul 1, 2026

Categories

GitHub topics

More in Natural Language Processing

Natural Language Processing

nltk/nltk๐Ÿ”ฅ active

NLTK Source

PythonApache-2.0updated Sep 23, 2026
GitHub โ†—โ˜… 14.7Kโ‘‚ 3K
Natural Language Processing

sloria/TextBlob๐Ÿ”ฅ active

Simple, Pythonic, text processing--Sentiment analysis, part-of-speech tagging, noun phrase extraction, translation, and more.

PythonMITupdated Sep 22, 2026
GitHub โ†—โ˜… 9.5Kโ‘‚ 1.2K
Natural Language Processing

juand-r/entity-recognition-datasets

A collection of corpora for named entity recognition (NER) and entity recognition tasks. These annotated datasets cover a variety of languages, domains and entity types.

๐Ÿท๏ธ Named Entity Recognition
PythonMITupdated Jul 2, 2026
GitHub โ†—โ˜… 1.6Kโ‘‚ 245
Natural Language Processing

proycon/pynlpl

PyNLPl, pronounced as 'pineapple', is a Python library for Natural Language Processing. It contains various modules useful for common, and less common, NLP tasks. PyNLPl can be used for basic tasks such as the extraction of n-grams and frequency lists, and to build simple language model. There are also more complex data types and algorithms. Moreover, there are parsers for file formats common in NLP (e.g. FoLiA/Giza/Moses/ARPA/Timbl/CQL). There are also clients to interface with various NLP specific servers. PyNLPl most notably features a very extensive library for working with FoLiA XML (Format for Linguistic Annotation).

PythonGPL-3.0updated Sep 14, 2023
GitHub โ†—โ˜… 476โ‘‚ 66

Data from GitHub ยท snapshot Sep 24, 2026