Files
llama.cpp/ggml
Piotr WilkinandClaude Opus 4.8 aa52642881 vulkan: prototype device-to-device peer buffers for -sm tensor (GGML_VK_COMM_D2D)
Opt-in path that replaces host-memory staging in the tensor-parallel AllReduce
with direct peer reads of another GPU's VRAM over PCIe P2P: each device's
partial lives in an exportable device buffer (DMA_BUF), peers import it, and the
existing comm reads host_buf[k][i] -- now peer VRAM -- unchanged. Targets the
O(n^2)-host-bandwidth scaling collapse; pairs with the O(n) ring. Ordering still
uses native/CPU-proxy semaphores (proxy auto-selects on RADV), so this is the
"D2D data + proxy semaphores" combo.

Enables VK_KHR_external_memory_fd + VK_EXT_external_memory_dma_buf. DMA_BUF is
the cross-device handle type (OPAQUE_FD memory is spec-locked to one physical
device); this matches the amdgpu PCIe-P2P dma-buf mechanism RADV/ROCm use.

Status: the fast path is AMD-targeted and UNVALIDATED -- NVIDIA's Vulkan driver
rejects cross-device fd import (vkGetMemoryFdPropertiesKHR -> memoryTypeBits=0;
confirmed against the Vulkan spec's same-deviceUUID rule and NVIDIA's own
statements), and NVIDIA has no fd-import or device-group P2P for unlinked GPUs.
So on NVIDIA it logs once and gracefully falls back to host staging -- verified
byte-identical and at host speed (pp2048 692 vs 695) on 4x A16, not the slow
butterfly. An AMD multi-GPU rig is needed to validate the actual P2P fast path
(test recipe accompanies this work, incl. how to prove real VRAM P2P vs a silent
amdgpu GTT fallback).

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ApKCQ32VLqUW4Kus6tUvBL
2026-06-28 15:58:08 +02:00
..
2024-07-13 18:12:39 +02:00