NVLS all-reduce · 4 GPUs
A = ⊕aᵢ
B = ⊕bᵢ
C = ⊕cᵢ
D = ⊕dᵢ
⊕
= selected NCCL reduction
UC
= per-GPU backing replica
MC
= one multicast address over 4 replicas
3rd-gen NVSwitch · NVLink 4+
Previous
Next
Input
Step 0 of 4 · One tensor per GPU
One logical chunk group; NCCL pipelines and overlaps these phases across chunks and channels. Registered buffers can bypass staging copies.