mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-27 21:46:57 +02:00
Opt-in path that replaces host-memory staging in the tensor-parallel AllReduce with direct peer reads of another GPU's VRAM over PCIe P2P: each device's partial lives in an exportable device buffer (DMA_BUF), peers import it, and the existing comm reads host_buf[k][i] -- now peer VRAM -- unchanged. Targets the O(n^2)-host-bandwidth scaling collapse; pairs with the O(n) ring. Ordering still uses native/CPU-proxy semaphores (proxy auto-selects on RADV), so this is the "D2D data + proxy semaphores" combo. Enables VK_KHR_external_memory_fd + VK_EXT_external_memory_dma_buf. DMA_BUF is the cross-device handle type (OPAQUE_FD memory is spec-locked to one physical device); this matches the amdgpu PCIe-P2P dma-buf mechanism RADV/ROCm use. Status: the fast path is AMD-targeted and UNVALIDATED -- NVIDIA's Vulkan driver rejects cross-device fd import (vkGetMemoryFdPropertiesKHR -> memoryTypeBits=0; confirmed against the Vulkan spec's same-deviceUUID rule and NVIDIA's own statements), and NVIDIA has no fd-import or device-group P2P for unlinked GPUs. So on NVIDIA it logs once and gracefully falls back to host staging -- verified byte-identical and at host speed (pp2048 692 vs 695) on 4x A16, not the slow butterfly. An AMD multi-GPU rig is needed to validate the actual P2P fast path (test recipe accompanies this work, incl. how to prove real VRAM P2P vs a silent amdgpu GTT fallback). Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ApKCQ32VLqUW4Kus6tUvBL