Language Models
dvmazur/mixtral-offloading
Run Mixtral-8x7B models in Colab or consumer desktops
PythonMITupdated Apr 8, 2024
Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization
$ git clone https://github.com/RahulSChand/gpu_poor.gitRun Mixtral-8x7B models in Colab or consumer desktops
INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model
A PyTorch-based Speech Toolkit
ไธญๆLLaMA-2 & Alpaca-2ๅคงๆจกๅไบๆ้กน็ฎ + 64K่ถ ้ฟไธไธๆๆจกๅ (Chinese LLaMA-2 & Alpaca-2 LLMs with 64K long context models)
Eagle: Frontier Vision-Language Models with Data-Centric Strategies
Data from GitHub ยท snapshot Sep 24, 2026