High-performance computing cluster for LLM training and working with Big Data

Project Description

The client approached us with a request to create a powerful server solution capable of handling extreme workloads in the field of artificial intelligence. The main focus is training large language models (LLM) and complex data analytics.
  • Scope of activity: AI development, deep learning and high performance computing (HPC).
  • Software used: Frameworks PyTorch, TensorFlow, Docker containers for neural network computing, data processing libraries.

Key requirement: Uncompromising GPU processing power and the fastest storage subsystem. An important condition was high bandwidth between graphics cards for parallel computing.

Upgrade configuration

  • Graphics card:
    4 x NVIDIA H200 (141 GB HBM3e, 16,896 CUDA cores)
  • CPU:
    2 x AMD EPYC 9374F (32 cores, up to 4.1 GHz, 256 MB L3 cache)
  • Motherboard:
    ASUS ESC8000A-E13P 4U (dual socket, supports up to 8 GPUs)
  • RAM:
    24 x 64GB Samsung ECC DDR5 (1536GB)
  • Drives:
    2 x 960GB/4x3.84TB Samsung PM9A3
  • Cooling:
    HYPERPC CUSTOM
  • Power supply:
    3000W ASUS PRO WS [80+ Platinum]
  • Case:
    CORSAIR h200 White

Selection process and decision

Standard server solutions were not suitable for implementing such a task. We chose the ASUS ESC8000A-E13P 4U platform, which is the standard for GPU-oriented systems.

Configuration logic:
  • Graphics power: The choice fell on the latest NVIDIA H200 accelerators with a memory capacity of 141 GB each. This is the “gold standard” for working with neural networks in 2026. To combine them into a single computing ecosystem, we used NVLink bridges, which provide data exchange at speeds not available on the conventional PCIe bus.
  • Computing center: Two AMD EPYC 9374F processors provide 64 high-frequency cores, which is critical for data preparation (preprocessing) before it is sent to the GPU.
  • Memory and storage: We equipped the system with 1.5 TB of RAM and advanced Samsung PM9D3a NVMe drives with PCIe 5.0 support. A read speed of 12,000 MB/s ensures that there are no bottlenecks when working with huge datasets.

Network interface: To integrate the server into the customer's existing infrastructure, a dual-port Mellanox ConnectX-6 adapter is installed, supporting speeds of up to 25 Gbit/s.

Result

Our company's engineers assembled and tested one of the most powerful servers in its class.

  • Performance: The use of NVLink bridges and the H200 architecture has reduced the training time for specific customer models by several times compared to the previous generation of systems.
  • Reliability: The server has passed 48-hour stress testing. Thanks to the ASUS server platform with redundant cooling and power supplies, the system is ready for operation 24/7/365.

Scalability: The configuration leaves the possibility of adding 4 more NVIDIA H200 accelerators without replacing the basic components of the system.

Did you like the upgrade?

Want to order a similar computer or create your own unique project? Contact us for a consultation - we will discuss every detail of your future PC.

Contact Us
Contact Us
Every HYPERPC computer is the result of 15 years of experience and expertise. Our experts know exactly what a gaming PC, workstation, or server should be like.
To get started, we just need to talk. Tell us about your tasks, timelines, and budget, and we will offer the best solution.
Call us or request a callback:
Message us:
Send an email:
sales@hyperpc.ae
Need to quickly know the cost?
Working hours: Daily from 10 AM to 7 PM.
English
+971 4 526 3600
Dayly from AM 10:00 to 7:00 PM