mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-27 21:46:57 +02:00
The ring transported fp32 chunks; the pipeline already halves host/peer traffic by staging F16. Bring the ring to parity: keep tensors[i] as the fp32 accumulator (matching the pipeline's precision) but cast each chunk to F16 for transport. The cast is folded into the recv step (the just-reduced chunk is the next step's send), so it costs only one extra pre-cast prog value rather than a doubled scheme; per-step up16 slots avoid a send-buffer WAR. GGML_VK_COMM_FP32 still forces fp32. Verified byte-identical greedy output vs fp32 ring / pipeline / butterfly on 4x A16, and no regression (ring-f16 ~= ring-fp32 ~= pipeline on A16 and 4090). The bandwidth win only shows when comm is exposed (comm-bound hosts); on this NVIDIA box the ring's transfer/compute overlap hides the comm, so F16 is neutral here -- same comm-hidden reason the other comm micro-opts are neutral on NVIDIA. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ApKCQ32VLqUW4Kus6tUvBL