NVIDIA GB200 NVL72 – Grace Blackwell Rack-Scale AI Supercomputer
The NVIDIA GB200 NVL72 is a rack-scale accelerated computing platform built on the NVIDIA Grace Blackwell architecture, combining 36 NVIDIA Grace CPUs and 72 NVIDIA Blackwell GPUs into a single high-performance NVLink domain. Designed for trillion-parameter AI models, generative AI, large-scale inference, model training, and HPC, the GB200 NVL72 delivers massive GPU memory capacity, extremely high memory bandwidth, and ultra-high-speed GPU-to-GPU communication.
The system provides up to 13.4TB of HBM3e GPU memory, 576TB/s aggregate GPU memory bandwidth, and 130TB/s of NVLink bandwidth, creating a tightly coupled 72-GPU computing domain optimized for large-scale AI workloads.
Description
The NVIDIA GB200 NVL72 is a next-generation rack-scale AI computing platform based on the NVIDIA Grace Blackwell architecture. Rather than functioning as an individual graphics card or conventional server, it integrates CPUs, GPUs, NVLink switches, networking, power delivery, and liquid cooling into a unified rack-scale infrastructure designed for extreme AI and HPC workloads.
At the heart of the system are 36 NVIDIA Grace CPUs and 72 NVIDIA Blackwell GPUs. Each GB200 Grace Blackwell Superchip combines one Grace CPU with two Blackwell GPUs using the high-bandwidth NVLink-C2C interconnect. NVIDIA specifies up to 3.6TB/s NVLink-C2C bandwidth per GB200 Superchip in the current GB200 NVL72 specifications.
Across the complete NVL72 platform, the 72 Blackwell GPUs provide up to 13.4TB of HBM3e GPU memory and up to 576TB/s of aggregate GPU memory bandwidth. This enormous memory subsystem is designed to keep very large AI models and datasets close to the GPU compute resources while minimizing data movement bottlenecks.
The system’s fifth-generation NVIDIA NVLink fabric provides up to 130TB/s of aggregate NVLink bandwidth, allowing all 72 GPUs to participate in a single large NVLink domain. This architecture is particularly important for trillion-parameter models where conventional PCIe or smaller GPU clusters can introduce significant communication overhead.
The GB200 NVL72 also incorporates Blackwell’s second-generation Transformer Engine, enabling advanced low-precision AI computing including FP4, alongside FP8, FP6, BF16, FP16, TF32, FP32, and FP64 processing. NVIDIA specifies up to 1,440 PFLOPS of sparse NVFP4 Tensor Core performance for the complete NVL72 configuration.
Unlike conventional air-cooled GPU servers, the GB200 NVL72 uses a liquid-cooled rack-scale design. NVIDIA’s reference architecture includes compute trays, NVLink switch trays, power shelves, a bus bar, liquid-cooling manifolds, and an NVLink passive copper cable backplane.
3. Complete Technical Specification
| Specification | NVIDIA GB200 NVL72 |
|---|---|
| Manufacturer | NVIDIA |
| Product Family | NVIDIA Grace Blackwell |
| Model | GB200 NVL72 |
| System Type | Rack-scale AI / HPC computing platform |
| CPU Architecture | NVIDIA Grace |
| CPU Count | 36 |
| CPU Architecture | Arm Neoverse V2 |
| Total CPU Cores | 2,592 |
| GPU Architecture | NVIDIA Blackwell |
| GPU Count | 72 |
| GPU Configuration | 36 GB200 Grace Blackwell Superchips |
| Blackwell GPUs per Superchip | 2 |
| GPU Memory | Up to 13.4TB HBM3e |
| GPU Memory Bandwidth | Up to 576TB/s aggregate |
| CPU Memory | Up to 17TB LPDDR5X across the rack |
| CPU Memory Bandwidth | Up to 14TB/s aggregate |
| NVLink Generation | 5th Generation NVIDIA NVLink |
| Aggregate NVLink Bandwidth | Up to 130TB/s |
| NVLink-C2C | Up to 3.6TB/s per GB200 Superchip |
| NVFP4 Tensor Performance | Up to 1,440 PFLOPS sparse |
| FP8 / FP6 Tensor Performance | Up to 720 PFLOPS sparse |
| INT8 Tensor Performance | Up to 720 POPS sparse |
| FP16 / BF16 Tensor Performance | Up to 360 PFLOPS sparse |
| TF32 Tensor Performance | Up to 180 PFLOPS sparse |
| FP32 Performance | Up to 5,760 TFLOPS |
| FP64 / FP64 Tensor | Up to 2,880 TFLOPS |
| Transformer Engine | 2nd Generation |
| GPU Interconnect | NVIDIA NVLink + NVLink Switch System |
| Networking | Qualified high-speed data-center networking, including InfiniBand/Ethernet architectures |
| Cooling | Liquid cooled |
| Rack Architecture | NVIDIA NVL72 rack-scale architecture |
| Compute Trays | 18 × 1RU trays in NVIDIA’s current reference configuration |
| NVLink Switch Trays | 9 × 1RU switch trays |
| GPU Domain | 72-GPU unified NVLink domain |
| Primary Workloads | Generative AI, LLM training, LLM inference, HPC, data analytics |
| Deployment Class | Enterprise / Data Center / AI Factory |
NVIDIA’s published specifications list the complete GB200 NVL72 at 2,592 Arm Neoverse V2 CPU cores, 13.4TB HBM3e, 576TB/s GPU memory bandwidth, 130TB/s NVLink bandwidth, and up to 17TB LPDDR5X CPU memory.
4. Applications
The GB200 NVL72 is designed for extremely large-scale computing environments, including:
- Trillion-parameter LLM training
- Large Language Model inference
- Generative AI
- AI model pre-training
- AI model fine-tuning
- Mixture-of-Experts (MoE) models
- Reasoning and agentic AI workloads
- High-Performance Computing (HPC)
- Scientific computing
- Large-scale data analytics
- Recommendation systems
- Natural language processing
- Computer vision
- Drug discovery and computational chemistry
- Digital twins and simulation
- Climate and weather modeling
- Enterprise AI factories
- Cloud AI infrastructure
- Hyperscale AI clusters
- Research and supercomputing
NVIDIA specifically positions the GB200 NVL72 for real-time trillion-parameter inference and large-scale training, with the 72-GPU NVLink domain acting as a single massive accelerator fabric.
5. Compatibility
The GB200 NVL72 is a complete rack-scale computing platform, not a standalone GPU or conventional server component.
Platform Compatibility
Designed for:
- NVIDIA GB200 NVL72 reference architecture
- NVIDIA DGX GB200 infrastructure
- Qualified NVIDIA GB200 OEM systems
- Enterprise AI data centers
- Hyperscale AI infrastructure
- NVIDIA-certified AI factories
- Large-scale HPC environments
- NVLink-based multi-GPU deployments
The NVIDIA reference NVL72 rack architecture uses 18 × 1RU compute trays, 9 × 1RU NVLink switch trays, management/top-of-rack networking, power shelves, and liquid-cooling infrastructure.
Important Compatibility Note
The GB200 NVL72 cannot be installed into a standard 19-inch GPU server as a conventional graphics accelerator. It requires the complete rack-scale architecture, including compatible compute trays, Grace Blackwell modules, NVLink switches, power infrastructure, liquid cooling, firmware, networking, and system management.
6. Key Benefits
- 72 NVIDIA Blackwell GPUs in one unified NVLink domain
- 36 NVIDIA Grace CPUs
- 2,592 Arm Neoverse V2 CPU cores
- Up to 13.4TB HBM3e GPU memory
- Up to 576TB/s aggregate GPU memory bandwidth
- Up to 130TB/s NVLink bandwidth
- Second-generation Transformer Engine
- Advanced FP4 / FP6 / FP8 AI acceleration
- Up to 1,440 PFLOPS sparse NVFP4 performance
- Designed for trillion-parameter AI models
- Extremely high GPU-to-GPU communication bandwidth
- Liquid-cooled rack-scale architecture
- Optimized for large-scale LLM inference and training
- Supports AI and HPC workloads in the same infrastructure
- High-density enterprise AI computing
- Designed for NVIDIA AI factory architectures
- Scalable to multi-rack deployments
7. SEO Title
NVIDIA GB200 NVL72 Grace Blackwell AI Supercomputer – 72 Blackwell GPUs, 13.4TB HBM3e



