Tag
This paper investigates how trace-driven evaluation can mislead assessments of MoE expert caching, identifying replay semantics, workload contamination, and operating regimes as confounding axes that can reverse policy rankings. After correcting these issues, it shows that a large offline-optimal gap overstates the gains actually recovered by lightweight causal caching mechanisms.