Insights

Infrastructure9 min read

Building a hosting-grade AI data platform

Power per rack, cooling choices, server specifications, storage and the project stages from capacity plan to commissioning. What hosting providers should decide before the first purchase order.

Cold aisle of a data center with amber light at the far end

A GPU hosting platform is not a rack of servers with a fast network. It is a power and cooling design that determines which servers you can buy, a procurement plan that determines when they arrive, and a commissioning process that determines when they start earning. This is the framework we use with hosting providers and data center operators, in the order the decisions actually have to be made.

Stage 1: capacity plan

Start from the product you will sell: dedicated GPU servers, GPU-hours through an orchestrator, inference endpoints, or managed clusters. Each maps to a different server type and utilization model. Then fix three numbers: GPUs in the first phase, target phase two within 12 months, and the power envelope your facility can commit to per rack. The last one drives everything else.

Stage 2: power and cooling

Server classPower per nodeNodes per 40 kW rackCooling
4U, 8× PCIe GPUs (RTX PRO 6000 / L40S / H200 NVL)5-7 kW5-6Air, hot-aisle containment
HGX H200 8-GPU (SXM)10-11 kW3Air with rear-door heat exchanger, or DLC
HGX B200 8-GPU (SXM)14-15 kW2 (air) / 4-5 in 80 kW DLC rackDLC recommended
GB200 / GB300 NVL72 rack~120-140 kW per rack1 rack = 72 GPUsDLC mandatory
  • Air cooling: practical to ~30-40 kW per rack with hot-aisle containment and adequate CRAH capacity. Fine for PCIe-GPU servers and small HGX H200 deployments.
  • Rear-door heat exchangers (RDHx): water-cooled doors that capture most of the rack exhaust. Extend air-cooled halls to ~50-70 kW per rack without touching the servers. Good bridge technology for HGX nodes in existing facilities.
  • Direct liquid cooling (DLC): cold plates on GPUs and CPUs, a CDU per row or per rack, facility water at 30-45 °C. Required for NVL72 and the sensible choice for B200/B300 at density. Needs manifolds, leak detection and a water-quality regime.

Specify PDUs at the real load: 8× 600 W GPUs plus CPUs, memory, NICs and fans push a 4U server to 6-7 kW at peak, which on 230 V three-phase means dual 32 A feeds per rack at minimum for a five-node rack. Ask for switched, metered PDUs so that per-outlet power can be reported to tenants and used to detect failing PSUs.

Stage 3: server specification

Supermicro, Dell and HPE all ship qualified GPU servers; pick based on support model, delivery time and the density you need rather than on marginal spec differences.

PCIe inference node
Supermicro SYS-421GE/521GE-class 4U or Dell PowerEdge XE7745 / R760xa, 8× dual-slot GPUs, 2× EPYC 9005 or Xeon 6, 1-1.5 TB DDR5, 2× 400G NIC, 4-8× NVMe.
HGX training node
Supermicro SYS-821GE / Dell XE9680 / HPE Cray XD670, HGX H200 or B200, 2 TB-3 TB DDR5, 8× ConnectX-7/8, 8× NVMe, dual BlueField or 200G storage NICs.
Storage
NVMe-based parallel or scale-out file system (WEKA, VAST, Lustre, or Ceph for lower tiers), 200-400G connected; plan 1-2 GB/s per GPU for training datasets, far less for inference.
Management
Out-of-band BMC network, 25/100G in-band, a jump host per pod, Slurm or Kubernetes with GPU operator.

Stage 4: procurement and allocation

PCIe GPUs, memory, NICs and optics are generally in stock and ship within days from Dubai or Hong Kong. HGX and NVL systems are allocated by NVIDIA and the OEMs, with lead times that move between 8 and 20 weeks depending on generation and quarter. A workable plan orders the PCIe-based phase for immediate revenue while the HGX phase is in allocation, and orders long-lead optics and switches together with the servers rather than after.

Stage 5: staging, burn-in and commissioning

  • Staging: rack, cable and firmware-level every node in a staging area or in the target row before handover to the platform team. Serial numbers, MAC addresses and rack positions recorded in the inventory system.
  • Burn-in: 48-72 hours of GPU stress (DCGM diagnostics level 3-4, NCCL tests, memory tests) at full power. Infant-mortality failures are found here, not in a customer’s job.
  • Fabric validation: per-link bandwidth and error counters, all-reduce at pod scale, storage throughput per node.
  • Commissioning: power and thermal logs at load for a full day, alerting configured, spares on the shelf, documentation handed over.

Spares and operations

Keep on site: 2-3% of GPUs, one PSU per server model, 5% of transceivers and cables, a set of DIMMs and NVMe drives per node type. Set up RMA paths with the manufacturer before day one; NVIDIA data center GPUs carry a 3-year warranty and OEM servers 3-5 years, but the advance-replacement terms differ by region and reseller.

Where Nodeforge fits

We work with hosting providers at three points. Design: capacity plan, power and cooling envelope, server BOM, fabric and storage architecture. Supply: direct manufacturer supply of GPUs, servers, memory, networking and optics with consolidated delivery to the site. Commissioning: staging, burn-in and fabric validation with our engineers on site or with the operator’s team. Providers can engage us for one stage or all three; either way the specification stays the operator’s.

  • Data center
  • Hosting
  • Cooling
  • Supermicro
  • Dell
  • HPE

Next step

Need help choosing the hardware?

Send the model and the workload. An engineer replies with a BOM, availability in Dubai and Hong Kong, and lead times.

Related

Keep reading