Description
Provides collective communication primitives for ROCm GPU workloads. It helps distributed machine learning and high-performance computing applications coordinate data across AMD GPUs.
Developers and researchers use it as an HPC library. Performance and correctness depend on GPU drivers, topology, and cluster configuration, so test with the target hardware.