MLPerf benchmark results are generally characterized by NVIDIA winning. This new MLPerf Training 5.1 benchmark was especially so, given that only NVIDIA and AMD GPUs made it into the submission pool, with around 79% of the configurations sporting NVIDIA accelerators. The results are very difficult to compare because for the 91 submissions, each configuration averaged to submit on roughly 2 of 7 benchmarks, making the data very sparse. Further, the systems submitted are often clusters of varying sizes and accelerators, making it even more challenging to compare results. Instead of trying to compare GPUs and accelerators from a sparse set of results, we decided to take a look at the configurations.
Folks can submit whatever configurations they want, but generally, they use these submissions as a marketing tool for configurations they offer. Given the current state of GPU supply, they also tend to come from pools of systems that are either what is available in a cloud or are regularly evaluated and shipped to customers. We know this because it appears that 2-3 of the configurations tested in this round of MLPerf we reviewed on STH over the last few months. These were likely the same systems used.
Given that these are either deployed systems in clusters or they are systems often from a customer demo pool, I wanted to take a look at the CPUs that are being used.
The NVIDIA Grace CPU is a well-known Arm Neoverse-V2 part with 72 cores. It is certainly not the most modern CPU out there as we were looking at Grace Hopper in 2024, before the latest generations of Intel Xeon and AMD EPYC CPUs were launched, let alone really Blackwell and Blackwell Ultra becoming available.
In this latest round, NVIDIA CPUs are credited as being used in 22 of the 91 submissions. We will get into why that might actually be 23 of 91, but in either case, that gives NVIDIA roughly 25% share of the configurations. That makes a lot of sense. The NVIDIA Grace Blackwell GB200 and GB300 platforms have only the company’s Grace CPU as an option currently.
That is what makes one submission strange. Historically, MLPerf uses a peer review process to catch anything that looks off. There may be one submission with an Intel Xeon submission that is actually a NVIDIA Grace CPU, or the GPU being reported is likely incorrect. Here is the Lambda GB300 submission that lists being a NVIDIA GB300 platform, or a Grace plus Blackwell Ultra, and yet is listed as using Intel Xeon Platinum 8570. The core count is listed as 72, which would align with a Grace Blackwell, but the Platinum 8570 is a 56-core part per Intel.
Another line looks strange in the configuration JSON:
“host_memory_configuration”: “32x 64GB HMCG94AGBRA179N”
That would be thirty-two 64GB DDR5 RDIMMs, for 2TB of host memory. Still, the NVIDIA GB300 uses LPDDR5x, not Hynix DDR5-5600 ECC RDIMMs.
It seems either the accelerator being listed as the GB300 was incorrect, or the CPU being reported should have been “Neoverse V2” Grace instead of Intel Xeon. Given that when we sort the submissions for 18-node results, the Lambda versus reference NVIDIA GB300 submissions are relatively close, and the core counts are listed as 72, it seems likely that this was probably a Grace system, and the configuration information was incorrect.
For our AMD versus Intel analysis, we are going to remove this entry because the hardware information reported on the official MLPerf GitHub should not exist in the same system. Still, we should note that if this is actually an NVIDIA Grace CPU, then just over 25% of the submissions will have been made on Arm CPUs. Our sense is that the relevant teams will update it once they see this (and I will send them a heads-up), but we will remove it from our AMD EPYC versus Intel Xeon analysis in the meantime.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.