Series Path · 5 parts
Inference Optimization
Make models faster at inference — quantization, edge deployment, compilation, and efficient attention.
~60 min total·5 articles
Start from Part 1Curriculum — 5 parts
- 01
Complete Guide to Quantization and Pruning
7 min read - 02
PyTorch Model Deployment on Edge Devices - Complete Code Guide
9 min read - 03
PyTorch 2.x Compilation Pipeline: From FX to Hardware
11 min read - 04
FlashAttention for Image Models
21 min read - 05
SGLang: Comprehensive Guide to Structured Generation Language
12 min read