πŸ† #289 overall#124 of 1,075 in Natural Language Processing

deanmalmgren /textract

extract text from any document. no muss. no fuss.

$ git clone https://github.com/deanmalmgren/textract.git
GitHub social preview for deanmalmgren/textract
Stars
4.7K
4,722
Forks
720
720
Language
HTML
License
MIT
Created
Jul 3, 2014
12.2 years old
Last push
Sep 1, 2026

Categories

GitHub topics

More in Natural Language Processing

Natural Language Processing

Jasonnor/tf-idf-python

Term frequency–inverse document frequency for Chinese novel/documents implemented in python.

πŸ’Ž Text Mining
PythonMITupdated Sep 15, 2018
GitHub β†—β˜… 104β‘‚ 35
Natural Language Processing

chiphuyen/lazynlp

Library to scrape and clean web pages to create massive datasets.

πŸ€– Language ModelsπŸ’Ž Text Mining
Pythonno licenseupdated Nov 11, 2020
GitHub β†—β˜… 2.3Kβ‘‚ 325
Natural Language Processing

mcs07/ChemDataExtractor

Automatically extract chemical information from scientific documents

⛏️ Information ExtractionπŸ’Ž Text Mining
PythonMITupdated Jul 27, 2023
GitHub β†—β˜… 366β‘‚ 125
Natural Language Processing

TiesdeKok/Python_NLP_Tutorial

This repository provides everything to get started with Python for Text Mining / Natural Language Processing (NLP)

πŸ’Ž Text Mining
Jupyter Notebookno licenseupdated Jun 5, 2020
GitHub β†—β˜… 129β‘‚ 66

Data from GitHub Β· snapshot Sep 24, 2026