High-performance computing cluster for LLM training and working with Big Data
Project Description
- Scope of activity: AI development, deep learning and high performance computing (HPC).
- Software used: Frameworks PyTorch, TensorFlow, Docker containers for neural network computing, data processing libraries.
Key requirement: Uncompromising GPU processing power and the fastest storage subsystem. An important condition was high bandwidth between graphics cards for parallel computing.
Upgrade configuration
- Graphics card:4 x NVIDIA H200 (141 GB HBM3e, 16,896 CUDA cores)
- CPU:2 x AMD EPYC 9374F (32 cores, up to 4.1 GHz, 256 MB L3 cache)
- Motherboard:ASUS ESC8000A-E13P 4U (dual socket, supports up to 8 GPUs)
- RAM:24 x 64GB Samsung ECC DDR5 (1536GB)
- Drives:2 x 960GB/4x3.84TB Samsung PM9A3
- Cooling:HYPERPC CUSTOM
- Power supply:3000W ASUS PRO WS [80+ Platinum]
- Case:CORSAIR h200 White
Selection process and decision
Standard server solutions were not suitable for implementing such a task. We chose the ASUS ESC8000A-E13P 4U platform, which is the standard for GPU-oriented systems.
- Graphics power: The choice fell on the latest NVIDIA H200 accelerators with a memory capacity of 141 GB each. This is the “gold standard” for working with neural networks in 2026. To combine them into a single computing ecosystem, we used NVLink bridges, which provide data exchange at speeds not available on the conventional PCIe bus.
- Computing center: Two AMD EPYC 9374F processors provide 64 high-frequency cores, which is critical for data preparation (preprocessing) before it is sent to the GPU.
- Memory and storage: We equipped the system with 1.5 TB of RAM and advanced Samsung PM9D3a NVMe drives with PCIe 5.0 support. A read speed of 12,000 MB/s ensures that there are no bottlenecks when working with huge datasets.
Network interface: To integrate the server into the customer's existing infrastructure, a dual-port Mellanox ConnectX-6 adapter is installed, supporting speeds of up to 25 Gbit/s.
Result
Our company's engineers assembled and tested one of the most powerful servers in its class.
- Performance: The use of NVLink bridges and the H200 architecture has reduced the training time for specific customer models by several times compared to the previous generation of systems.
- Reliability: The server has passed 48-hour stress testing. Thanks to the ASUS server platform with redundant cooling and power supplies, the system is ready for operation 24/7/365.
Scalability: The configuration leaves the possibility of adding 4 more NVIDIA H200 accelerators without replacing the basic components of the system.
Did you like the upgrade?
Want to order a similar computer or create your own unique project? Contact us for a consultation - we will discuss every detail of your future PC.