End-to-end NVIDIA reference architectures: the anatomy of a DGX system, DGX SuperPOD, rack-scale NVLink domains with GB200 NVL72, InfiniBand networking and BlueField DPUs, GPUDirect storage. Then real-world operations: Base Command, NVIDIA AI Enterprise, Slurm and Kubernetes on DGX, MIG, distributed training and resilience, DCGM, sovereignty — all the way to the DGX-Ready Data Center program and DGX Cloud.
Running the cluster: Base Command, Slurm, Kubernetes + GPU Operator, and managing multi-user GPU queues.