Files
llama.cpp/ggml
Piotr WilkinandClaude Opus 4.8 831bc206da vulkan: F16 staging for the ring AllReduce
The ring transported fp32 chunks; the pipeline already halves host/peer traffic
by staging F16. Bring the ring to parity: keep tensors[i] as the fp32
accumulator (matching the pipeline's precision) but cast each chunk to F16 for
transport. The cast is folded into the recv step (the just-reduced chunk is the
next step's send), so it costs only one extra pre-cast prog value rather than a
doubled scheme; per-step up16 slots avoid a send-buffer WAR. GGML_VK_COMM_FP32
still forces fp32.

Verified byte-identical greedy output vs fp32 ring / pipeline / butterfly on
4x A16, and no regression (ring-f16 ~= ring-fp32 ~= pipeline on A16 and 4090).
The bandwidth win only shows when comm is exposed (comm-bound hosts); on this
NVIDIA box the ring's transfer/compute overlap hides the comm, so F16 is neutral
here -- same comm-hidden reason the other comm micro-opts are neutral on NVIDIA.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ApKCQ32VLqUW4Kus6tUvBL
2026-06-28 16:56:58 +02:00
..
2024-07-13 18:12:39 +02:00