MoEs provide two knobs for scaling: model size (total params) + FLOPs-per-token (via active params).
What’s the right scaling strategy? And how does it depend on the pretraining budget?
Our work introduces sparsity-aware scaling laws for MoE LMs to tackle these questions!
🧵👇
🚨 One question that has always intrigued me is the role of different ways to increase a model's capacity: parameters, parallelizable compute, or sequential compute?
We explored this through the lens of MoEs:



