MoE Analysis, which experts fire, and how sure is the router?

We took the MT-Bench prompts (first turn only), ran inference on each of them with two Mixture-of-Experts models, and captured the statistics of every MoE layer: for each token, which experts the router picked, and how confident that choice was.

Model
Colour by
Layer

Loading data/manifest.json…