mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-27 21:46:57 +02:00
Three related changes to the Vulkan -sm tensor (tensor-parallel) all-reduce, developed and measured on 2x Radeon AI PRO R9700 (RDNA4/GFX1201, Mesa 26.1.2): 1. Shrink the comm staging buffers on the prefill->decode transition. ensure() only ever grew the host/tmp buffers, to the peak prefill micro-batch (~10 MB at n_ubatch=512), then reused them for the tiny (~20 KB) decode all-reduces. On RADV the imported external host memory is made visible across devices on every timeline-semaphore signal, at a cost proportional to the resident buffer size, so an oversized leftover cap stalled every decode step (cross-device wait ~770 us vs ~30 us), collapsing decode from ~30 to ~2.4 t/s and staying stuck for the whole session (the AMD "multi-turn crawl"). ensure() now also shrinks when the request is much smaller than cap, with a one-time semaphore wait so the realloc is safe against the previous async all-reduce. Decode after a large prefill: 2.4 -> ~31 t/s, flat across prefill sizes. 2. Decide proxy vs native cross-device sync before creating the progress timelines. They were created as exportable up front, which aborted init on devices that cannot export timeline semaphores (e.g. llvmpipe, RADV on older Mesa) instead of falling back to the portable CPU proxy. Now create exportable timelines only for the native path and plain ones for the proxy path. 3. Add ggml_backend_vk_comm_allreduce_tree: an opt-in recursive halving/doubling all-reduce for power-of-two device counts (2*log2(n) cross-device steps vs the ring's 2*(n-1); bandwidth-optimal). Enabled via GGML_VK_COMM_TREE; the ring stays the default and handles non-power-of-two counts. The step schedule is built as in a reference simulation verified for n=2..8; validated at n=2 (native and forced proxy) to match the ring's greedy output byte-for-byte. Assisted-by: Claude Opus 4.8 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01SazLJJfgpjt9Kq7JXKKnuw