🏆 #1,104 overall#455 of 1,075 in Natural Language Processing

lionsoul2014 /jcseg

Jcseg is a light weight NLP framework developed with Java. Provide CJK and English segmentation based on MMSEG algorithm, With also keywords extraction, key sentence extraction, summary extraction implemented based on TEXTRANK algorithm. Jcseg had a build-in http server and search modules for lucene,solr,elasticsearch,opensearch

$ git clone https://github.com/lionsoul2014/jcseg.git
GitHub social preview for lionsoul2014/jcseg
Stars
920
920
Forks
210
210
Language
Java
License
Apache-2.0
Created
Mar 31, 2014
12.5 years old
Last push
Sep 18, 2023

Categories

GitHub topics

More in Natural Language Processing

Natural Language Processing

brightmart/nlp_chinese_corpus

大规模中文自然语言处理语料 Large Scale Chinese Corpus for NLP

🗂️ Text Classification❓ Question Answering🧬 Embeddings🤖 Language Models
OtherMITupdated Feb 6, 2026
GitHub ↗★ 9.9K⑂ 1.6K
Natural Language Processing

NLPchina/ansj_seg

ansj分词.ict的真正java实现.分词效果速度都超过开源版的ict. 中文分词,人名识别,词性标注,用户自定义词典

JavaApache-2.0updated Nov 19, 2023
GitHub ↗★ 6.5K⑂ 2.3K
Natural Language Processing

houbb/sensitive-word

👮‍♂️The sensitive word tool for java.(敏感词/违禁词/违法词/脏词。基于 DFA 算法实现的高性能 java 敏感词过滤工具框架。内置支持单词标签分类分级。请勿发布涉及政治、广告、营销、翻墙、违反国家法律法规等内容。高性能敏感词检测过滤组件,附带繁体简体互换,支持全角半角互换,汉字转拼音,模糊搜索等功能。)

JavaApache-2.0updated Mar 23, 2026
GitHub ↗★ 6.1K⑂ 806
Natural Language Processing

HIT-SCIR/ltp

Language Technology Platform

Pythonno licenseupdated Mar 11, 2026
GitHub ↗★ 5.3K⑂ 1.1K
Natural Language Processing

esbatmop/MNBVC

MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集。对标chatGPT训练的40T数据。MNBVC数据集不但包括主流文化,也包括各个小众文化甚至火星文的数据。MNBVC数据集包括新闻、作文、小说、书籍、杂志、论文、台词、帖子、wiki、古诗、歌词、商品介绍、笑话、糗事、聊天记录等一切形式的纯文本中文数据。

OtherMITupdated Sep 13, 2026
GitHub ↗★ 4.3K⑂ 296
Natural Language Processing

ownthink/Jiagu

Jiagu深度学习自然语言处理工具 知识图谱关系抽取 中文分词 词性标注 命名实体识别 情感分析 新词发现 关键词 文本摘要 文本聚类

🏷️ Named Entity Recognition
PythonMITupdated May 7, 2022
GitHub ↗★ 3.4K⑂ 607

Data from GitHub · snapshot Sep 24, 2026