๐Ÿ† #2,806 overall#105 of 131 in Text Mining

ahmedkhemiri95 /PDFs-TextExtract

Multiple and Large PDF Documents Text Extraction.

$ git clone https://github.com/ahmedkhemiri95/PDFs-TextExtract.git
GitHub social preview for ahmedkhemiri95/PDFs-TextExtract
Stars
132
132
Forks
65
65
Language
Python
License
Apache-2.0
Created
May 7, 2020
6.4 years old
Last push
Feb 10, 2025

Categories

GitHub topics

More in Text Mining

Text Mining

chiphuyen/lazynlp

Library to scrape and clean web pages to create massive datasets.

๐Ÿค– Language Models๐Ÿง  Natural Language Processing
Pythonno licenseupdated Nov 11, 2020
GitHub โ†—โ˜… 2.3Kโ‘‚ 325
Text Mining

hiDaDeng/cntext

cntext is a Python library for social science text analysis, offering word frequency, sentiment, word embeddings, and semantic projection to measure constructs like attitudes and psychological states from Chinese text.

๐Ÿ’– Sentiment Analysis๐Ÿงฌ Embeddings๐Ÿง  Natural Language Processing
PythonMITupdated May 3, 2026
GitHub โ†—โ˜… 465โ‘‚ 43
Text Mining

airbnb/artificial-adversary

๐Ÿ—ฃ๏ธ Tool to generate adversarial text examples and test machine learning models against them

๐Ÿ—‚๏ธ Text Classification
PythonMITupdated Jan 7, 2022
GitHub โ†—โ˜… 402โ‘‚ 56
Text Mining

dataqa/nlp-labelling

Labelling platform for text using weak supervision.

๐Ÿท๏ธ Named Entity Recognition๐Ÿ—‚๏ธ Text Classification๐Ÿง  Natural Language Processing
JavaScriptGPL-3.0updated Jun 24, 2022
GitHub โ†—โ˜… 260โ‘‚ 18

Data from GitHub ยท snapshot Sep 24, 2026