MLPerf Inference v6.1: Supermicro NVIDIA HGX™ B300 8-GPU Systems Deliver LLM Inference Performance for the Era of Agentic AI

42_2026-09_MLPERF-Blog_R01_1200x628

Overview

Supermicro has made a number of submissions to the MLPerf Inference 6.1 benchmark program, demonstrating that the NVIDIA HGX B300 system with eight NVIDIA Blackwell Ultra GPUs connected with NVIDIA NVLink for LLM inference. MLPerf Inference 6.1 includes four LLM benchmarks.

These LLMs are strong proxies for AI workloads in the data center, supporting chatbots, image queries, reasoning, summarization, and interactions with many AI agents.

Test Systems

Supermicro produced these results using the following systems, each configured with 8x NVIDIA Blackwell Ultra GPUs connected with fifth-generation NVIDIA NVLink and features an aggregate of 2.1TB of HBM3e memory:

  • Supermicro SYS-822GS-NB3RT with Intel® Xeon® 6 CPUs
  • Supermicro AS-8126GS-NB3RT with AMD EPYC™ 9005 Series CPUs

Benchmark Results

Benchmark Tokens/second (rounded) Model

Total Parameters

Active Parameters

Reasoning

Server: 67,000

Offline: 70,000 
 DeepSeek-R1  671B  ~37B
Vision-language

Server: 112 queries/sec

Offline: 129 samples/sec 
 Qwen3-VL-235B-A22B  235B  22B
 LLM reasoning/coding, RAG query generation 

Server: 107,000

Offline: 109,000 
 GPT-OSS-120B  117B  5.1B
Question Answering

Server: 110,000

Offline: 111,000 
 Llama 2 70B  70B  70B

 

MLPerf Inference v6.1 data center benchmarks highlight the importance of serving LLMs to AI agents running everywhere. Supermicro tested Mixture-of-Experts (MoE) models, including DeepSeek, Qwen, and GPT-OSS, where the full set of experts is loaded into GPU memory but only a subset of experts is activated for each token, as well as the Llama dense model, where every parameter is activated for every token. Together, these benchmarks provide a representative mix of the model types used for LLM serving in the data center today. Across all of them, Supermicro’s NVIDIA HGX B300 server demonstrated very high inference performance.

Comparing NVIDIA Blackwell Ultra Platforms

MLPerf Inference v6.1 published results for a wide range of GPU systems. Most of these results are based on 8x GPU configurations, so it’s useful to compare three of these systems:

  • NVIDIA GB300 NVL72, using 8x Blackwell Ultra GPUs
  • NVIDIA HGX B300 8-GPU System
  • PCIe GPU System with 8x NVIDIA RTX PROTM 6000 Blackwell Server Edition GPUs

MLPerf Training v2 bar chart sept 2026
The chart shows relative performance for the Llama 2 70B benchmark, normalized to tokens per second, for this specific 8-GPU inference configuration. The NVIDIA GB300 NVL72 shows an edge here because its rack-scale design runs each Blackwell Ultra GPU at a higher power envelope, giving it greater per-GPU compute throughput — one data point for one model at this scale. GB300 NVL72's full-rack architecture is built for the industry's largest and most demanding workloads, including training and inference for frontier-class models with trillions of parameters, where its scale offers a clear advantage, at the cost of a full 72-GPU rack requiring over 132kW of power and liquid cooling.

Built for Flexible, Scalable Deployment

Many existing data centers have only air cooling with constrained power budgets. NVIDIA HGX B300 8-GPU systems and NVIDIA RTX PRO Servers provide flexibility and scalability, especially for inference-focused AI infrastructure.

Supermicro offers the full spectrum of NVIDIA Blackwell and NVIDIA Blackwell Ultra optimized systems, NVIDIA-Certified systems, combined with NVIDIA Grace CPU, AMD EPYC, or Intel Xeon CPUs, supported with air-cooling or liquid-cooling.

Deploy AI Infrastructure Faster with Supermicro DCBBS

Turning strong benchmark results like these into production AI infrastructure is where many enterprises get stuck, weighing GPU, CPU, networking, power, and cooling choices against their own workloads and timelines. Supermicro’s Data Center Building Block Solutions® (DCBBS) address that challenge directly, pairing validated, workload-optimized building blocks with the deployment experience needed to bring them online quickly. With deep expertise across NVIDIA Blackwell and Blackwell Ultra platforms, air- and liquid-cooled system designs, and a broad range of CPU and GPU configurations, Supermicro helps enterprises identify the right AI infrastructure for their needs and deploy it in the fastest way possible.

Looking ahead, NVIDIA's MLPerf Inference v6.1 submission gave the industry an early look at the next-generation Vera Rubin platform — a substantial leap in scale and capability for AI infrastructure. As NVIDIA Vera Rubin platform becomes available across the ecosystem, Supermicro's DCBBS approach is designed to evolve alongside it, helping enterprises navigate this next tier of AI infrastructure and bring it into production.

Learn more: Supermicro AI Server Solutions