Series Path · 3 parts
Mixture of Experts
Sparse MoE from theory through GShard to Switch Transformer.
~32 min total·3 articles
Start from Part 1Curriculum — 3 parts
- 01
Mixture of Experts: A Deep Overview
Mixture of Experts (MoE) represents a fundamental paradigm shift in machine learning architecture design, offering a scalable approach to building models that…
9 min read - 02
GShard: Scaling Giant Neural Networks with Conditional Computation
GShard represents a pivotal advancement in neural network scaling, introduced by Google Research in 2020. This innovative approach addresses one of the most…
13 min read - 03
Switch Transformer: Scaling Neural Networks with Sparsity
The Switch Transformer represents a groundbreaking advancement in neural network architecture, introduced by Google Research in 2021. This innovative model…
10 min read