mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-27 21:46:57 +02:00
With the ring as the default there is no reason to keep the slower paths: - Remove ggml_backend_vk_comm_allreduce_pipeline (the O(n^2) all-to-all) and GGML_VK_COMM_PIPELINE. The ring is now the unconditional large-tensor path; the comm->ring flag and pipe_round are gone, and pipeline_ok (the "has two queues" gate the ring needs) is renamed ring_ok. - Remove fp32 staging and GGML_VK_COMM_FP32. The ring always stages F16 (fp32 accumulator preserved); its fp32 branch and use_f16 are removed. Net ~-260 lines. Verified byte-identical greedy output (ring / proxy) and clean build on 4x A16. The decode single-shot and the meta-backend butterfly fallback are untouched. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ApKCQ32VLqUW4Kus6tUvBL