Part 5:LLM 基础:以 GPT-2 为例
| Title | Author | Date |
|---|---|---|
| 11.1 语言模型在预测什么:Next-Token Prediction | 2026-06-18 | |
| 11.2 MiniGPT:从 Causal GPT Block 到语言模型 | 2026-06-20 | |
| 11.3 Tokenizer:字符、BPE 与词表 | 2026-06-18 | |
| 11.4 Embedding、LM Head 与 Weight Tying | 2026-06-18 | |
| 11.5 在 TinyStories 上训练 MiniGPT | 2026-06-22 | |
| 11.6 从训练到生成:Temperature、Top-k、Top-p | 2026-06-18 | |
| 11.7 GPT-2:从 MiniGPT 到预训练语言模型 | 2026-06-23 | |
| 11.8 Chapter 11 练习题:从零实现 GPT-2 | 2026-09-28 | |
| 12.1 LLM 训练显存账本:参数、激活与状态 | 2026-08-15 | |
| 12.10 大模型 Checkpoint:模型、优化器与分布式状态如何恢复 | 2026-09-17 | |
| 12.11 Chapter 12 练习题:LLM 训练工程 | 2026-09-28 | |
| 12.2 FLOPs、Memory 与 Arithmetic Intensity:为什么不是只看计算量 | 2026-09-16 | |
| 12.3 Profiling:怎么知道代码慢在哪里 | 2026-09-10 | |
| 12.4 Mixed Precision:FP32、FP16、BF16 与 Loss Scaling | 2026-09-10 | |
| 12.5 Gradient Accumulation:显存不够时如何模拟大 Batch | 2026-09-08 | |
| 12.6 Activation Checkpointing:用重计算换显存 | 2026-09-13 | |
| 12.7 现代 Attention API 与 Hugging Face Kernels | 2026-09-12 | |
| 12.8 Triton 入门:什么时候需要自己写 Kernel | 2026-09-11 | |
| 12.9 分布式训练入门:DDP、ZeRO 与 FSDP 的直觉 | 2026-09-14 |
No matching items