An always-on expert that every token passes through, alongside whatever the router picks.
Introduced inJan 2024DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
DeepSeek-style MoE adds one or more shared experts that process every token unconditionally, in parallel with the routed top-k experts. The shared expert captures common knowledge every token needs, freeing the routed experts to specialize harder. Which measurably improves quality at the same compute. Most 2025-generation MoE models (DeepSeek, GLM, Qwen-MoE, Nemotron) include one.
Share of new models that have included a shared expert over time.
Open any of these on hfviewer to find this block in the interactive architecture graph.
Browse all 345 models with this in the catalog →
hfviewer renders the full architecture of 3,500+ Hugging Face models as interactive graphs. Hover any block to see what it does, with this model’s real numbers.
Browse all model graphs →