Choosing GPUs for LLM inference vs training
Inference is a memory problem, training is a bandwidth-and-interconnect problem. How to size VRAM, read precision figures and decide between PCIe cards and NVLink systems.
Insights
Comparisons and infrastructure notes from the engineers who spec, source and commission AI hardware. No press releases.
Inference is a memory problem, training is a bandwidth-and-interconnect problem. How to size VRAM, read precision figures and decide between PCIe cards and NVLink systems.
Three PCIe cards, three price bands. A side-by-side of memory, bandwidth, power and cooling, and the workloads where each is the right answer rather than the expensive one.
How many channels to populate, how much system RAM per GB of VRAM, and which speed bins and ranks to specify so the memory keeps up with the GPUs.
NDR InfiniBand or 400/800G RoCE Ethernet, rail-optimized fabrics, and the optics and cables that connect them. When Ethernet is enough, and when it is not.
Power per rack, cooling choices, server specifications, storage and the project stages from capacity plan to commissioning. What hosting providers should decide before the first purchase order.
DGX Spark, RTX PRO Blackwell workstation cards and GeForce RTX 5090/5080 for development and local inference. What each runs comfortably, and where the desk stops being enough.
Memory, bandwidth, TFLOPS by precision, TDP, form factor, PCIe generation, MIG, NVLink, licences and warranty. A field guide to the numbers on a data center GPU datasheet and the ones that are missing.