Text Mining
deanmalmgren/textract
extract text from any document. no muss. no fuss.
๐ง Natural Language Processing
HTMLMITupdated Sep 1, 2026
Reworked https://www.readability.com/ parsing library (now https://mercury.postlight.com/ is living alternative)
$ git clone https://github.com/bookieio/breadability.gitextract text from any document. no muss. no fuss.
Library to scrape and clean web pages to create massive datasets.
Python package for Korean natural language processing.
Python implementation of the Rapid Automatic Keyword Extraction algorithm using NLTK.
Fake News Detection in Python
Data from GitHub ยท snapshot Sep 24, 2026