beir-cellar/beir
A Heterogeneous Benchmark for Information Retrieval. Easy to use, evaluate your models across 15+ diverse IR datasets.
A large-scale multilingual dataset for Information Retrieval. Thorough human-annotations across 18 diverse languages.
$ git clone https://github.com/project-miracl/miracl.gitA Heterogeneous Benchmark for Information Retrieval. Easy to use, evaluate your models across 15+ diverse IR datasets.
Dataset and benchmark for RAG on company internal documents.
Up to 100x faster strings for C, C++, CUDA, Python, Rust, Swift, JS, & Go, leveraging NEON, AVX2, AVX-512, SVE, GPGPU, & SWAR to accelerate search, hashing, sorting, edit distances, sketches, and memory ops π¦
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corporaβwith retrieval and evaluation tooling included.
Links to conference/journal publications in automated fact-checking (resources for the TACL22/EMNLP23 paper).
Data from GitHub Β· snapshot Sep 24, 2026