Natural Language Processing
deanmalmgren/textract
extract text from any document. no muss. no fuss.
π Text Mining
HTMLMITupdated Sep 1, 2026
Term frequencyβinverse document frequency for Chinese novel/documents implemented in python.
$ git clone https://github.com/Jasonnor/tf-idf-python.gitextract text from any document. no muss. no fuss.
Library to scrape and clean web pages to create massive datasets.
Starter code to solve real world text data problems. Includes: Gensim Word2Vec, phrase embeddings, Text Classification with Logistic Regression, word count with pyspark, simple text preprocessing, pre-trained embeddings and more.
Various Algorithms for Short Text Mining
Automatically extract chemical information from scientific documents
Jupyter notebooks for our O'Reilly book "Blueprints for Text Analysis Using Python"
Data from GitHub Β· snapshot Sep 24, 2026