ashvardanian/StringZilla๐ฅ active
Up to 100x faster strings for C, C++, CUDA, Python, Rust, Swift, JS, & Go, leveraging NEON, AVX2, AVX-512, SVE, GPGPU, & SWAR to accelerate search, hashing, sorting, edit distances, sketches, and memory ops ๐ฆ
Provides a common interface to many IR ranking datasets.
$ git clone https://github.com/allenai/ir_datasets.gitUp to 100x faster strings for C, C++, CUDA, Python, Rust, Swift, JS, & Go, leveraging NEON, AVX2, AVX-512, SVE, GPGPU, & SWAR to accelerate search, hashing, sorting, edit distances, sketches, and memory ops ๐ฆ
A Heterogeneous Benchmark for Information Retrieval. Easy to use, evaluate your models across 15+ diverse IR datasets.
Links to conference/journal publications in automated fact-checking (resources for the TACL22/EMNLP23 paper).
Dataset and benchmark for RAG on company internal documents.
HDLTex: Hierarchical Deep Learning for Text Classification
A large-scale multilingual dataset for Information Retrieval. Thorough human-annotations across 18 diverse languages.
Data from GitHub ยท snapshot Sep 24, 2026