Language Models
RahulSChand/gpu_poor
Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization
JavaScriptno licenseupdated Dec 3, 2024
LLaMA 2 implemented from scratch in PyTorch
$ git clone https://github.com/hkproj/pytorch-llama.gitCalculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization
On-device LLM Inference Powered by X-Bit Quantization
RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable). We are at RWKV-7 "Goose". So it's combining the best of RNN and transformer - great performance, linear time, constant space (no kv-cache), fast training, infinite ctx_len, and free sentence embedding.
Code for loralib, an implementation of "LoRA: Low-Rank Adaptation of Large Language Models"
A PyTorch-based Speech Toolkit
Data from GitHub ยท snapshot Sep 24, 2026