Language models & other things.

Mainly fiddling with small and large language models.

Notes & experiments

MoE Analysis: which experts fire, and how sure is the router?

The MT-Bench prompts answered by Qwen3.5 and Qwen3.6 (35B-A3B), with every MoE routing decision captured: which of the 256 experts fired for each token in each of the 40 layers, how peaked the router's choice was, and how much probability sat just outside the top-8.