
DGMY has put a new large-model training cluster into service. The hall is built for long training runs: GPU servers stay on the job, and the data they need is staged before the run starts.
Networks are split by job
Training traffic, data traffic, storage, the business network and out-of-band management each have their own path. A heavy checkpoint no longer has to share a link with the systems operators use to watch the cluster.
Hot data stays close to the GPUs
Data used in the current run is cached on the GPU servers. Older data is kept as objects. That split keeps the expensive processors busy and leaves the storage system for everything that does not have to be read this hour.
The same design is described in our AIGC large model training solution.
The latest news, articles, and resources, sent to your inbox weekly