Text Mining
chiphuyen/lazynlp
Library to scrape and clean web pages to create massive datasets.
๐ค Language Models๐ง Natural Language Processing
Pythonno licenseupdated Nov 11, 2020
Multiple and Large PDF Documents Text Extraction.
$ git clone https://github.com/ahmedkhemiri95/PDFs-TextExtract.gitLibrary to scrape and clean web pages to create massive datasets.
a curated list of R tutorials for Data Science, NLP and Machine Learning
cntext is a Python library for social science text analysis, offering word frequency, sentiment, word embeddings, and semantic projection to measure constructs like attitudes and psychological states from Chinese text.
๐ฃ๏ธ Tool to generate adversarial text examples and test machine learning models against them
Labelling platform for text using weak supervision.
Materials for GWU DNSC 6279 and DNSC 6290.
Data from GitHub ยท snapshot Sep 24, 2026