Home / Servers / NVIDIA GB200 NVL72 – Grace Blackwell Rack-Scale AI Supercomputer

NVIDIA GB200 NVL72 – Grace Blackwell Rack-Scale AI Supercomputer

The NVIDIA GB200 NVL72 is a rack-scale accelerated computing platform built on the NVIDIA Grace Blackwell architecture, combining 36 NVIDIA Grace CPUs and 72 NVIDIA Blackwell GPUs into a single high-performance NVLink domain. Designed for trillion-parameter AI models, generative AI, large-scale inference, model training, and HPC, the GB200 NVL72 delivers massive GPU memory capacity, extremely high memory bandwidth, and ultra-high-speed GPU-to-GPU communication.

The system provides up to 13.4TB of HBM3e GPU memory, 576TB/s aggregate GPU memory bandwidth, and 130TB/s of NVLink bandwidth, creating a tightly coupled 72-GPU computing domain optimized for large-scale AI workloads.

Category:

Description

The NVIDIA GB200 NVL72 is a next-generation rack-scale AI computing platform based on the NVIDIA Grace Blackwell architecture. Rather than functioning as an individual graphics card or conventional server, it integrates CPUs, GPUs, NVLink switches, networking, power delivery, and liquid cooling into a unified rack-scale infrastructure designed for extreme AI and HPC workloads.

At the heart of the system are 36 NVIDIA Grace CPUs and 72 NVIDIA Blackwell GPUs. Each GB200 Grace Blackwell Superchip combines one Grace CPU with two Blackwell GPUs using the high-bandwidth NVLink-C2C interconnect. NVIDIA specifies up to 3.6TB/s NVLink-C2C bandwidth per GB200 Superchip in the current GB200 NVL72 specifications.

Across the complete NVL72 platform, the 72 Blackwell GPUs provide up to 13.4TB of HBM3e GPU memory and up to 576TB/s of aggregate GPU memory bandwidth. This enormous memory subsystem is designed to keep very large AI models and datasets close to the GPU compute resources while minimizing data movement bottlenecks.

The system’s fifth-generation NVIDIA NVLink fabric provides up to 130TB/s of aggregate NVLink bandwidth, allowing all 72 GPUs to participate in a single large NVLink domain. This architecture is particularly important for trillion-parameter models where conventional PCIe or smaller GPU clusters can introduce significant communication overhead.

The GB200 NVL72 also incorporates Blackwell’s second-generation Transformer Engine, enabling advanced low-precision AI computing including FP4, alongside FP8, FP6, BF16, FP16, TF32, FP32, and FP64 processing. NVIDIA specifies up to 1,440 PFLOPS of sparse NVFP4 Tensor Core performance for the complete NVL72 configuration.

Unlike conventional air-cooled GPU servers, the GB200 NVL72 uses a liquid-cooled rack-scale design. NVIDIA’s reference architecture includes compute trays, NVLink switch trays, power shelves, a bus bar, liquid-cooling manifolds, and an NVLink passive copper cable backplane.


3. Complete Technical Specification

Specification NVIDIA GB200 NVL72
Manufacturer NVIDIA
Product Family NVIDIA Grace Blackwell
Model GB200 NVL72
System Type Rack-scale AI / HPC computing platform
CPU Architecture NVIDIA Grace
CPU Count 36
CPU Architecture Arm Neoverse V2
Total CPU Cores 2,592
GPU Architecture NVIDIA Blackwell
GPU Count 72
GPU Configuration 36 GB200 Grace Blackwell Superchips
Blackwell GPUs per Superchip 2
GPU Memory Up to 13.4TB HBM3e
GPU Memory Bandwidth Up to 576TB/s aggregate
CPU Memory Up to 17TB LPDDR5X across the rack
CPU Memory Bandwidth Up to 14TB/s aggregate
NVLink Generation 5th Generation NVIDIA NVLink
Aggregate NVLink Bandwidth Up to 130TB/s
NVLink-C2C Up to 3.6TB/s per GB200 Superchip
NVFP4 Tensor Performance Up to 1,440 PFLOPS sparse
FP8 / FP6 Tensor Performance Up to 720 PFLOPS sparse
INT8 Tensor Performance Up to 720 POPS sparse
FP16 / BF16 Tensor Performance Up to 360 PFLOPS sparse
TF32 Tensor Performance Up to 180 PFLOPS sparse
FP32 Performance Up to 5,760 TFLOPS
FP64 / FP64 Tensor Up to 2,880 TFLOPS
Transformer Engine 2nd Generation
GPU Interconnect NVIDIA NVLink + NVLink Switch System
Networking Qualified high-speed data-center networking, including InfiniBand/Ethernet architectures
Cooling Liquid cooled
Rack Architecture NVIDIA NVL72 rack-scale architecture
Compute Trays 18 × 1RU trays in NVIDIA’s current reference configuration
NVLink Switch Trays 9 × 1RU switch trays
GPU Domain 72-GPU unified NVLink domain
Primary Workloads Generative AI, LLM training, LLM inference, HPC, data analytics
Deployment Class Enterprise / Data Center / AI Factory

NVIDIA’s published specifications list the complete GB200 NVL72 at 2,592 Arm Neoverse V2 CPU cores, 13.4TB HBM3e, 576TB/s GPU memory bandwidth, 130TB/s NVLink bandwidth, and up to 17TB LPDDR5X CPU memory.


4. Applications

The GB200 NVL72 is designed for extremely large-scale computing environments, including:

  • Trillion-parameter LLM training
  • Large Language Model inference
  • Generative AI
  • AI model pre-training
  • AI model fine-tuning
  • Mixture-of-Experts (MoE) models
  • Reasoning and agentic AI workloads
  • High-Performance Computing (HPC)
  • Scientific computing
  • Large-scale data analytics
  • Recommendation systems
  • Natural language processing
  • Computer vision
  • Drug discovery and computational chemistry
  • Digital twins and simulation
  • Climate and weather modeling
  • Enterprise AI factories
  • Cloud AI infrastructure
  • Hyperscale AI clusters
  • Research and supercomputing

NVIDIA specifically positions the GB200 NVL72 for real-time trillion-parameter inference and large-scale training, with the 72-GPU NVLink domain acting as a single massive accelerator fabric.


5. Compatibility

The GB200 NVL72 is a complete rack-scale computing platform, not a standalone GPU or conventional server component.

Platform Compatibility

Designed for:

  • NVIDIA GB200 NVL72 reference architecture
  • NVIDIA DGX GB200 infrastructure
  • Qualified NVIDIA GB200 OEM systems
  • Enterprise AI data centers
  • Hyperscale AI infrastructure
  • NVIDIA-certified AI factories
  • Large-scale HPC environments
  • NVLink-based multi-GPU deployments

The NVIDIA reference NVL72 rack architecture uses 18 × 1RU compute trays, 9 × 1RU NVLink switch trays, management/top-of-rack networking, power shelves, and liquid-cooling infrastructure.

Important Compatibility Note

The GB200 NVL72 cannot be installed into a standard 19-inch GPU server as a conventional graphics accelerator. It requires the complete rack-scale architecture, including compatible compute trays, Grace Blackwell modules, NVLink switches, power infrastructure, liquid cooling, firmware, networking, and system management.


6. Key Benefits

  • 72 NVIDIA Blackwell GPUs in one unified NVLink domain
  • 36 NVIDIA Grace CPUs
  • 2,592 Arm Neoverse V2 CPU cores
  • Up to 13.4TB HBM3e GPU memory
  • Up to 576TB/s aggregate GPU memory bandwidth
  • Up to 130TB/s NVLink bandwidth
  • Second-generation Transformer Engine
  • Advanced FP4 / FP6 / FP8 AI acceleration
  • Up to 1,440 PFLOPS sparse NVFP4 performance
  • Designed for trillion-parameter AI models
  • Extremely high GPU-to-GPU communication bandwidth
  • Liquid-cooled rack-scale architecture
  • Optimized for large-scale LLM inference and training
  • Supports AI and HPC workloads in the same infrastructure
  • High-density enterprise AI computing
  • Designed for NVIDIA AI factory architectures
  • Scalable to multi-rack deployments

7. SEO Title

NVIDIA GB200 NVL72 Grace Blackwell AI Supercomputer – 72 Blackwell GPUs, 13.4TB HBM3e