๐Ÿ† #3,099 overall#388 of 388 in Search & Retrieval

JuliaText /WordTokenizers.jl

High performance tokenizers for natural language processing and other related tasks

$ git clone https://github.com/JuliaText/WordTokenizers.jl.git
GitHub social preview for JuliaText/WordTokenizers.jl
Stars
100
100
Forks
25
25
Language
Julia
License
Other
Created
Apr 10, 2018
8.5 years old
Last push
Dec 30, 2021

Categories

GitHub topics

More in Search & Retrieval

Search & Retrieval

rth/vtext

Simple NLP in Rust with Python bindings

โœ‚๏ธ Tokenization & Preprocessing๐Ÿง  Natural Language Processing
RustApache-2.0updated Jul 6, 2023
GitHub โ†—โ˜… 153โ‘‚ 10
Search & Retrieval

piskvorky/gensim

Topic Modelling for Humans

๐Ÿซง Topic Modeling๐Ÿงฌ Embeddings๐Ÿง  Natural Language Processing
PythonLGPL-2.1updated Nov 1, 2025
GitHub โ†—โ˜… 16.5Kโ‘‚ 4.4K
Search & Retrieval

LSYS/LexicalRichness

:smile_cat: :speech_balloon: A module to compute textual lexical richness (aka lexical diversity).

๐Ÿง  Natural Language Processing
PythonMITupdated May 2, 2026
GitHub โ†—โ˜… 113โ‘‚ 22
Search & Retrieval

artitw/text2text

Text2Text Language Modeling Toolkit

โœ‚๏ธ Tokenization & Preprocessing๐Ÿง  Natural Language Processing
PythonOtherupdated Jan 14, 2025
GitHub โ†—โ˜… 304โ‘‚ 40
Search & Retrieval

neuml/txtai๐Ÿ”ฅ active

๐Ÿ’ก All-in-one AI framework for semantic search, LLM orchestration and language model workflows

๐Ÿงฌ Embeddings๐Ÿค– Language Models๐Ÿง  Natural Language Processing
PythonApache-2.0updated Sep 23, 2026
GitHub โ†—โ˜… 13Kโ‘‚ 899
Search & Retrieval

beir-cellar/beir

A Heterogeneous Benchmark for Information Retrieval. Easy to use, evaluate your models across 15+ diverse IR datasets.

๐Ÿง  Natural Language Processing
PythonApache-2.0updated Oct 16, 2025
GitHub โ†—โ˜… 2.3Kโ‘‚ 251

Data from GitHub ยท snapshot Sep 24, 2026