High-performance computing cluster for LLM training and working with Big Data

Project Description

The client approached us with a request to create a powerful server solution capable of handling extreme workloads in the field of artificial intelligence. The main focus is training large language models (LLM) and complex data analytics.
  • Scope of activity: AI development, deep learning and high performance computing (HPC).
  • Software used: Frameworks PyTorch, TensorFlow, Docker containers for neural network computing, data processing libraries.

Key requirement: Uncompromising GPU processing power and the fastest storage subsystem. An important condition was high bandwidth between graphics cards for parallel computing.

Upgrade configuration

  • Graphics card:
    4 x NVIDIA H200 (141 GB HBM3e, 16,896 CUDA cores)
  • CPU:
    2 x AMD EPYC 9374F (32 cores, up to 4.1 GHz, 256 MB L3 cache)
  • Motherboard:
    ASUS ESC8000A-E13P 4U (dual socket, supports up to 8 GPUs)
  • RAM:
    24 x 64GB Samsung ECC DDR5 (1536GB)
  • Drives:
    2 x 960GB/4x3.84TB Samsung PM9A3
  • Cooling:
    HYPERPC CUSTOM
  • Power supply:
    3000W ASUS PRO WS [80+ Platinum]
  • Case:
    CORSAIR h200 White

Selection process and decision

Standard server solutions were not suitable for implementing such a task. We chose the ASUS ESC8000A-E13P 4U platform, which is the standard for GPU-oriented systems.

Configuration logic:
  • Graphics power: The choice fell on the latest NVIDIA H200 accelerators with a memory capacity of 141 GB each. This is the “gold standard” for working with neural networks in 2026. To combine them into a single computing ecosystem, we used NVLink bridges, which provide data exchange at speeds not available on the conventional PCIe bus.
  • Computing center: Two AMD EPYC 9374F processors provide 64 high-frequency cores, which is critical for data preparation (preprocessing) before it is sent to the GPU.
  • Memory and storage: We equipped the system with 1.5 TB of RAM and advanced Samsung PM9D3a NVMe drives with PCIe 5.0 support. A read speed of 12,000 MB/s ensures that there are no bottlenecks when working with huge datasets.

Network interface: To integrate the server into the customer's existing infrastructure, a dual-port Mellanox ConnectX-6 adapter is installed, supporting speeds of up to 25 Gbit/s.

Result

Our company's engineers assembled and tested one of the most powerful servers in its class.

  • Performance: The use of NVLink bridges and the H200 architecture has reduced the training time for specific customer models by several times compared to the previous generation of systems.
  • Reliability: The server has passed 48-hour stress testing. Thanks to the ASUS server platform with redundant cooling and power supplies, the system is ready for operation 24/7/365.

Scalability: The configuration leaves the possibility of adding 4 more NVIDIA H200 accelerators without replacing the basic components of the system.

Did you like the upgrade?

Want to order a similar computer or create your own unique project? Contact us for a consultation - we will discuss every detail of your future PC.

Contact the manager Other projects

Copyright ©2026 HYPERPC


main version