MMLU · Massive Multitask Language Understanding
A broad knowledge benchmark used as the quality yardstick when judging whether a lower-precision training or quantization recipe has degraded a model.
Current numbers
≤1%NVFP4 inference accuracy drop vs FP8 on DeepSeek-R1 (MMLU-Pro 84% vs 85%)
~29%estimated MMLU contamination across public web corpora; clean-mirror retests drop scores high-single to low-double digits