Files
llama.cpp/ggml
Piotr Wilkin 13ac901567 vulkan: fix -sm tensor decode crawl, comm init robustness, add opt-in tree all-reduce
Three related changes to the Vulkan -sm tensor (tensor-parallel) all-reduce,
developed and measured on 2x Radeon AI PRO R9700 (RDNA4/GFX1201, Mesa 26.1.2):

1. Shrink the comm staging buffers on the prefill->decode transition.
   ensure() only ever grew the host/tmp buffers, to the peak prefill micro-batch
   (~10 MB at n_ubatch=512), then reused them for the tiny (~20 KB) decode
   all-reduces. On RADV the imported external host memory is made visible across
   devices on every timeline-semaphore signal, at a cost proportional to the
   resident buffer size, so an oversized leftover cap stalled every decode step
   (cross-device wait ~770 us vs ~30 us), collapsing decode from ~30 to ~2.4 t/s
   and staying stuck for the whole session (the AMD "multi-turn crawl"). ensure()
   now also shrinks when the request is much smaller than cap, with a one-time
   semaphore wait so the realloc is safe against the previous async all-reduce.
   Decode after a large prefill: 2.4 -> ~31 t/s, flat across prefill sizes.

2. Decide proxy vs native cross-device sync before creating the progress
   timelines. They were created as exportable up front, which aborted init on
   devices that cannot export timeline semaphores (e.g. llvmpipe, RADV on older
   Mesa) instead of falling back to the portable CPU proxy. Now create exportable
   timelines only for the native path and plain ones for the proxy path.

3. Add ggml_backend_vk_comm_allreduce_tree: an opt-in recursive halving/doubling
   all-reduce for power-of-two device counts (2*log2(n) cross-device steps vs the
   ring's 2*(n-1); bandwidth-optimal). Enabled via GGML_VK_COMM_TREE; the ring
   stays the default and handles non-power-of-two counts. The step schedule is
   built as in a reference simulation verified for n=2..8; validated at n=2
   (native and forced proxy) to match the ring's greedy output byte-for-byte.

Assisted-by: Claude Opus 4.8 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01SazLJJfgpjt9Kq7JXKKnuw
2026-06-29 11:49:48 +01:00
..
2024-07-13 18:12:39 +02:00