MoEXBench tests how compression methods interact in mixture-of-experts language models
A new arXiv benchmark evaluates expert pruning, weight quantization and KV-cache compression together across 10 mixture-of-experts language models, finding that combined deployment trade-offs cannot be inferred reliably from isolated tests.