Building an LLM from scratch
不依赖任何高层黑盒库,纯使用 Python 与 PyTorch 从底层字节流和矩阵运算出发,端到端实现 BPE 分词、现代 Transformer 架构、内存映射二进制数据流水线、混合精度预训练引擎、自回归 KV-Cache 推理、SFT 指令微调与 LoRA 适配器,并在消费级显卡上完整闭环。
SEQUENCE
Series sequence
12 guides
- 01
Building a 0.04B (36M) LLM from Scratch: End-to-End Full-Stack Overview
Beginner - 02
0.04B Model Architecture Design, Mathematical Modeling, and Intuition
Beginner - 03
Building a Byte-Level BPE Tokenizer from Scratch
Intermediate - 04
Implementing Modern Transformer Architecture from Scratch
Advanced - 05
First-Principles Mathematical Derivations for Modern Transformers
Intermediate - 06
Industrial Data Sanitization and Zero-Copy Binary Pipelines with mmap
Intermediate - 07
Engineering Pretraining Engines and Optimization on Consumer GPUs
Intermediate - 08
Autoregressive Streaming Inference and KV-Cache Memory Optimization
Intermediate - 09
0.04B End-to-End Hands-on Practice and Telemetry Diagnostics
Intermediate - 10
Supervised Fine-Tuning (SFT) and Instruction Alignment
Advanced - 11
Parameter-Efficient Fine-Tuning with Handcrafted LoRA on Modern LLMs
Advanced - 12
Full Lifecycle Review: From Pretraining Convergence to SFT and LoRA
Intermediate