• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

NVIDIA Wins Every MLPerf Training v5.1 Benchmark

erek

8=D
2FA
Joined
Dec 19, 2005
Messages
17,774
"New Benchmarks, New Records
NVIDIA also set performance records on the two new benchmarks added this round: Llama 3.1 8B and FLUX.1.

Llama 3.1 8B—a compact yet highly capable LLM—replaced the long-running BERT-large model, adding a modern, smaller LLM to the benchmark suite. NVIDIA submitted results with up to 512 Blackwell Ultra GPUs, setting the bar at 5.2 minutes to train.

In addition, FLUX.1—a state-of-the-art image generation model—replaced Stable Diffusion v2, with only the NVIDIA platform submitting results on the benchmark. NVIDIA submitted results using 1,152 Blackwell GPUs, setting a record time to train of 12.5 minutes.

NVIDIA continued to hold records on the existing graph neural network, object detection and recommender system tests.

A Broad and Deep Partner Ecosystem
The NVIDIA ecosystem participated extensively this round, with compelling submissions from 15 organizations including ASUSTeK, Dell Technologies, Giga Computing, Hewlett Packard Enterprise, Krai, Lambda, Lenovo, Nebius, Quanta Cloud Technology, Supermicro, University of Florida, Verda (formerly DataCrunch) and Wiwynn.

NVIDIA is innovating at a one-year rhythm, driving significant and rapid performance increases across pretraining, post-training and inference—paving the way to new levels of intelligence and accelerating AI adoption."

1762987306431.png

Source: https://www.techpowerup.com/342860/nvidia-wins-every-mlperf-training-v5-1-benchmark
 

"AMD Unveils First MLPerf 5.1 Training Results on Instinct MI350 Series GPUs

Press Release by Nomad76 Today, 10:28 Discuss (0 Comments)
The latest MLPerf 5.1 Training results mark an important milestone—the first MLPerf training submission using AMD Instinct MI350 Series GPUs. These new benchmarks demonstrate breakthrough generational performance gains and highlight broad ecosystem participation across some of today's most demanding AI training workloads.

Building on the strong MLPerf 5.1 Inference results released earlier this year, this submission represents the first time the AMD Instinct MI350 Series—including both the MI355X and MI350X GPUs—have been publicly benchmarked for AI training. The results show clear progress in scalability, efficiency, and compute performance, underscoring how the AMD Instinct MI350 Series is accelerating the development of next-generation generative AI models."
 
Didnt knew 4/8 bits was now used so much for the training phase....

10 minutes to "train" 405B Llama 3.1 ... must be a single pass and a whole training would invovle reinforcement phase and what not but still....

https://public.tableau.com/app/prof...ns-Training_16993769118290/MLCommons-Training

Ability to scale for this is really something, 80 NVL72 (5120) do not take half the time as a 40NVL72 system but take 53% of the time, almost perfect scaling.

On llama70b going from 8 to 16 MI325x or B200, ~62% of the time (using standard CPU instead of Nvidia platform).

There is way more Intel Xeon platform showing up than expected.

Comparing with H100 generation benchmark, seem a general rules that Blackwell+passage of time software wise 2x+ hopper.
 
  • Like
Reactions: erek
like this
Back
Top