The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
GuideGlossaryMMLU

MMLU · Massive Multitask Language Understanding

A broad knowledge benchmark used as the quality yardstick when judging whether a lower-precision training or quantization recipe has degraded a model.

Current numbers

≤1%NVFP4 inference accuracy drop vs FP8 on DeepSeek-R1 (MMLU-Pro 84% vs 85%)as of 2025 · register ↗
~29%estimated MMLU contamination across public web corpora; clean-mirror retests drop scores high-single to low-double digitsas of 2025 · register ↗

← All terms