RWKV/rwkv.cpp
INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model
Run Mixtral-8x7B models in Colab or consumer desktops
$ git clone https://github.com/dvmazur/mixtral-offloading.gitINT4/INT5/INT8 and FP16 inference on CPU for RWKV language model
Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization
From scratch implementation of a sparse mixture of experts language model inspired by Andrej Karpathy's makemore :)
High-performance In-browser LLM Inference Engine
RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable). We are at RWKV-7 "Goose". So it's combining the best of RNN and transformer - great performance, linear time, constant space (no kv-cache), fast training, infinite ctx_len, and free sentence embedding.
Data from GitHub ยท snapshot Sep 24, 2026