Natural Language Processing
🏆 #836 overall#367 of 1,075 in Natural Language Processing
datawhalechina /diy-llm
Covers pre-training data, Tokenizer, Transformer, MoE,distributed training, Scaling Laws, inference & alignment .6 progressive code assignments for full-stack LLM learning | 涵盖预训练数据、分词器、Transformer、MoE、分布式训练、缩放定律、推理与对齐,6 项渐进代码作业,掌握 LLM 全栈知识
$ git clone https://github.com/datawhalechina/diy-llm.gitStars
1.4K
1,414
Forks
146
146
Language
Jupyter Notebook
License
None
Created
Nov 24, 2025
0.8 years old
Last push
Sep 10, 2026
Categories
More in Natural Language Processing
Natural Language Processing
hiyouga/EasyR1🔥 active
EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
PythonApache-2.0updated Sep 19, 2026
Natural Language Processing
google/langextract🔥 active
A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.
🤖 Language Models
PythonApache-2.0updated Sep 21, 2026
Natural Language Processing
ymcui/Chinese-LLaMA-Alpaca
中文LLaMA&Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA & Alpaca LLMs)
🤖 Language Models
PythonApache-2.0updated Apr 19, 2026
Natural Language Processing
graykode/nlp-tutorial
Natural Language Processing Tutorial for Deep Learning Researchers
Jupyter NotebookMITupdated Feb 21, 2024
Natural Language Processing
botpress/botpress🔥 active
The open-source hub to build & deploy GPT/LLM Agents ⚡️
TypeScriptMITupdated Sep 23, 2026
Data from GitHub · snapshot Sep 24, 2026