Commit Graph
10718 Commits
Author SHA1 Message Date
Ruben Ortlam ab94ffdd30 use 4-byte loads where possible 2026-08-29 11:51:15 +02:00
Ruben Ortlam 5037f027aa use shmem arrays for LUTs 2026-08-29 11:51:15 +02:00
Ruben Ortlam ba8402e9f6 remove elem row/col fast path, invalid for RDNA4 2026-08-29 11:51:15 +02:00
Ruben Ortlam dc08fa9101 support iq4_nl and mxfp4 2026-08-29 11:51:15 +02:00
Ruben Ortlam 3cd76432d7 fix mul_mat_id bug 2026-08-29 11:51:15 +02:00
Ruben Ortlam d8ed37209a fix segfault 2026-08-29 11:51:15 +02:00
Ruben Ortlam fa37db93ce enable mul_mat_id support 2026-08-29 11:51:15 +02:00
Ruben Ortlam ed294e8724 restructure mmq cm1 functions 2026-08-29 11:51:15 +02:00
Ruben Ortlam d2a5ac69c9 add q4_1, q5_0, q5_1 support 2026-08-29 11:51:15 +02:00
Ruben Ortlam eb46b8cf0a move quant-specific prefetch function out of main file 2026-08-29 11:51:15 +02:00
Ruben Ortlam 1825270eb4 fix compilation 2026-08-29 11:51:15 +02:00
Ruben Ortlam 42a6269bda use BK_STEP 4 2026-08-29 11:51:15 +02:00
Ruben Ortlam ab6e84bbdd only force subgroup size 32 on AMD RDNA 2026-08-29 11:51:15 +02:00
Ruben Ortlam 57dd3f8bd3 skip computation for inactive tiles 2026-08-29 11:51:15 +02:00
Ruben Ortlam 2b9c53381e restructure for vgpr use 2026-08-29 11:51:15 +02:00
Ruben Ortlam ef5c62992f Revert "increase large tile size"
This reverts commit 7fabc25c5e0c9047a7a25b23bba5cad017de23ec.
2026-08-29 11:51:15 +02:00
Ruben Ortlam a9e0e89c0b increase large tile size 2026-08-29 11:51:15 +02:00
Ruben Ortlam 57b4276784 use wave32 2026-08-29 11:51:15 +02:00
Ruben Ortlam 15b3ce0de3 Revert "revert load reordering and scale pre-loading"
This reverts commit fbaefe0eaaccbb6ab9d733ad84580428f18db410.
2026-08-29 11:51:15 +02:00
Ruben Ortlam 9f069f42e8 clean up 2026-08-29 11:51:15 +02:00
Ruben Ortlam 23356762ca workgroup scheduling for cache proximity 2026-08-29 11:51:15 +02:00
Ruben Ortlam 31b9bc683e revert load reordering and scale pre-loading 2026-08-29 11:51:15 +02:00
Ruben Ortlam ad015d6a44 add faster RDNA int->float conversion 2026-08-29 11:51:15 +02:00
Ruben Ortlam 95d14c99ea use float for scales 2026-08-29 11:51:15 +02:00
Ruben Ortlam 14a6080dd5 coopmat load first, then wmma 2026-08-29 11:51:15 +02:00
Ruben Ortlam b453be38a3 preload scales 2026-08-29 11:51:15 +02:00
Ruben OrtlamandPiotr Wilkin 5553b89912 double buffering
Co-authored-by: Piotr Wilkin (ilintar) <[email protected]>
2026-08-29 11:51:13 +02:00
Ruben OrtlamandPiotr Wilkin 2a878d9ea9 use larger workgroups
Co-authored-by: Piotr Wilkin (ilintar) <[email protected]>
2026-08-29 11:51:10 +02:00
Ruben OrtlamandPiotr Wilkin b121ef3b33 add BK_STEP to shader, default to 2
Co-authored-by: Piotr Wilkin (ilintar) <[email protected]>
2026-08-29 11:51:08 +02:00
Ruben Ortlam 513f43600c add q8_0 support 2026-08-29 11:51:08 +02:00
Ruben OrtlamandPiotr Wilkin 162560885a probe and directly access coopmat values instead of going through shmem
Co-authored-by: Piotr Wilkin (ilintar) <[email protected]>
2026-08-29 11:51:04 +02:00
Ruben Ortlam 5f034d6086 use scalar sums 2026-08-29 11:30:17 +02:00
Ruben Ortlam e086b9ca05 apply scales inline 2026-08-29 11:30:17 +02:00
Ruben Ortlam 13f6d8e389 vulkan: add int8 coopmat quantized matmul shader 2026-08-29 11:30:17 +02:00
Nick Farrell cc83d7b482 sycl: make --fit respect --fit-target better (#27629)
improve the --fit algorithm to take into account the actual peak
required VRAM for a given context size on a SYCL backend.

This includes both properly accounting for how much VRAM is required
when the allocated context is fully used (which makes the reported
context drop below what it did before, but stop it OOMing) as well
as preventing some overly-conservative calculations which meant too much
VRAM was being reserved.

Tested on a Arc b70 with unsloth's qwen3.8 (Q4_K_XL), able to get 262144 context,
fully usable, with q8_0 KV and MTP and 4k ubatch size using --fit-target 1
b10684
2026-08-29 05:00:09 -04:00
Jeff Bolz c9ca51c1f6 vulkan: combine duplicated fastdiv functions, rename the one optimizing small divs (#27526)
* vulkan: combine duplicated fastdiv functions, rename the one optimizing small divs

* remove one more fastdiv
b10683
2026-08-29 10:59:48 +03:00
Jhen-Jie Hong 5ea1b124e7 metal : add fa-vec tunings for M1 Max (#27932) b10682 2026-08-29 15:12:23 +08:00
Jeff Bolz 77f132cb1d vulkan: Change mul_mat_id to pad K rather than N (#27925)
The N padding is needed for mul_mat, but not mul_mat_id. For mul_mat_id,
we indirect the row index through a shared memory lookup table which avoids
any OOB row coordinate. But that callback doesn't bounds check K, so we
actually need K padding instead.
b10681
2026-08-29 10:09:24 +03:00
kurquharandKristopher Urquhart d7bd3bfcad snapdragon: python SDK setup (Windows) (#27903)
* port setup-build.ps1 to setup_sdk.py, to facilitate installation of Hexagon and OpenCL SDKs on Windows

* rename setup_sdk.py -> setup-sdk.py

* flake8 fix: print() -> logger.info()

---------

Co-authored-by: Kristopher Urquhart <[email protected]>
b10680
2026-08-28 14:01:59 -07:00
Xuan-Son Nguyen 50f068ffff bench: add --tensor-read-lazy (#27881)
* bench: add --tensor-read-lazy

* rm the alias

* rename to LLAMA_LAZY_MODE_*
b10679
2026-08-28 20:51:05 +02:00
Xuan-Son Nguyen 6fe7498016 model: qwen4exp: reduce number of graph splits (#27880) b10678 2026-08-28 19:24:46 +02:00
Eric A StaleeandJeff Bolz b387ddfd84 vulkan: fix missing view-alias dependencies in ggml_vk_graph_optimize (#27812)
* vulkan: fix missing view-alias dependencies in ggml_vk_graph_optimize

is_src_of doesn't treat two views of one tensor as dependent, so the optimizer reorders nodes across aliased reads and writes. 

Result: silently wrong tokens under greedy decoding, different output on every server start, and invalid speculative-decoding acceptance, with nothing logged.

Hits Qwen3.8's recurrent state (and any model with view-aliased state) on AMD and NVIDIA Vulkan.  CUDA is clean. 

Compare view_src bases on both sides.

Fixes #27805

* vulkan: don't treat view/no-op nodes as aliasing dependencies

Nodes whose op is NONE, RESHAPE, TRANSPOSE, VIEW or PERMUTE execute nothing, so aliasing through them is not a real dependency. The previous base comparison matched them anyway, which only costs the optimizer reordering freedom.

Co-authored-by: Jeff Bolz <[email protected]>

* vulkan: make the lambda parameter const and capture is_empty in is_src_of

Code will not compile without these changes.  
is_src_of has an empty capture list, so is_empty was not visible inside it, and is_empty took a non-const pointer, while is_src_of receives const ones. Other call sites pass non-const pointers, which still convert as usual.

---------

Co-authored-by: Jeff Bolz <[email protected]>
b10677
2026-08-28 19:12:33 +02:00
Tekin ErtekinandGeorgi Gerganov a43c3986b4 ggml : fix conv_transpose_2d for multiple batches (#26132)
* ggml : fix conv_transpose_2d for multiple batches

ggml_compute_forward_conv_transpose_2d_impl only computed the first
batch (ne[3] of the destination); every batch after the first was left
as zero. Both the src1 permutation and the main compute loop now iterate
over the batch dimension, and the work buffer size in ggml_graph_plan is
scaled by the src1 batch count so the extra permuted batches fit. A
multi-batch test case is added to test-backend-ops.

Fixes ggml-org/ggml#1448

* metal : fix conv_transpose_2d for multiple batches

The kernel only computed batch 0 of the input (src1->ne[3]); every
output batch after the first was left as zero, so multi-batch
conv_transpose_2d results diverged from the CPU reference.

The grid now covers all batches (OW x OH x OC x N), the kernel decodes
the batch from the grid z coordinate and offsets both the input and
destination indices accordingly. nb3 is passed in the kernel args.

Assisted-by: pi:llama.cpp/Qwen3.8-27B

---------

Co-authored-by: Georgi Gerganov <[email protected]>
b10676
2026-08-28 20:09:08 +03:00
ravel7524andJeff Bolz 90c26fcd4b Vulkan: add hoisting support for row IDs and expert count in shaders (#26686)
* vulkan: add hoisting support for row IDs and expert count in shaders

* use hoisted row ids in coopmat2

* vulkan: address review feedback on count_experts
- use vk_op_count_experts_push_constants instead of a raw uint vector
- apply the fastdiv trick to the ne00 div/mod in count_experts
- compute the per-expert offsets with subgroupExclusiveAdd when the
  device supports it, keeping the serial path as fallback
- document the data_d layout and the hoisted_row_id_words bound
- drop a leftover debug print in ggml_vk_matmul_id

* vulkan: use init_pushconst_fastdiv for count_experts push constants

* vulkan: refine comments for row ID hoisting and data layout in count_experts shader

* Whitespace

---------

Co-authored-by: Jeff Bolz <[email protected]>
b10675
2026-08-28 16:52:49 +02:00
Georgi Gerganov 8663224818 context : disable non-fused GDN and LID ops (#27877) 2026-08-28 16:34:26 +03:00
StrongtutandStrongtut f5e85d43a0 metal : add fa-vec tunings for M4 (#27875)
This adds fa_vec_tuned_table records for Apple M4 to ggml-metal-tuning.cpp.

Includes F16, Q4_0, Q4_1, Q5_0, Q5_1, and Q8_0. (M4, 10 GPU Cores)

Co-authored-by: Strongtut <[email protected]>
b10673
2026-08-28 15:37:37 +03:00
511f9c1379 OpenVINO: Update OV to 2026.3.1, whisper.cpp support, Qwen3.5 on NPU, and new ops (#27843)
* OpenVINO Backend: Fuse IM2COL + MatMul convolution into OpenVINO convolution

* ci:ggml-ov: Skip recurrent state rollback tests

* ci:ggml-ov: Skip recurrent state rollback tests

* Update OPENVINO.md

* ggml-openvino : add env-var gated op support debugging

* Fix ggml_rope_set_offset case

* OpenVINO backend: Support Whisper.cpp

* Fix code style

* openvino : enable qwen35 on NPU

Static shapes:
- get_graph_input_shape() left the s_copy / s_copy-leaf inputs dynamic
  ([1,1,1,-1]) even in static mode, which propagated a dynamic slot dim through
  GET_ROWS into the conv/GDN state, the state reshapes and the GDN output.
- With -np 1 the s_copy defrag remainder gathers zero rows; short-circuit that
  CPY to the untouched cache instead of emitting a degenerate Slice/Concat, and
  skip binding its zero-byte ggml tensor as an output (the dynamic path already
  did the latter, the static path wrote the full cache over a 0-byte buffer).

Token-count independence:
- In static mode the compiled model's token count is the prefill chunk size or
  1, not the captured cgraph's. Offsets derived from the captured count were
  therefore wrong. Anchor the GDN state slice at the end of the packed
  [attn | state] output and drop the rs_src_begin runtime inputs, and make
  VIEWs over the GDN output / conv_input pass through so the consumer does the
  slicing.
- CONT could not identify its token axis when the graph was captured with a
  single token (every trailing dim has the same stride and size 1) and baked
  the captured shape into the prefill model.

Chunked prefill:
- The last chunk is padded with fabricated tokens. Attention masks them, but
  the recurrent path folded them into cache_r/cache_s permanently. Add a
  chunk_valid_len runtime input, use it to zero g and beta for padded steps
  (making the recurrence an exact identity) and to end the conv snapshot window
  at the last valid token, and disable the recurrent-cache reset after the
  first chunk so earlier chunks are not wiped.
- get_is_prefill() and the chunk loop bound read inp_pos->ne[0] directly, but
  IMROPE stacks 4 position planes, so every decode step was run through the
  padded prefill model and the loop ran extra out-of-bounds chunks.

cache_rs_reset_idx/len now stay runtime Parameters in static mode, since
can_reuse_statically() does not invalidate the cached model on ComputeParams
changes. Add GGML_OPENVINO_FORCE_STATIC to exercise the static path on CPU.

* Update to OpenVINO 2026.3.1

* ggml-openvino: forward NPU compilation mode parameters

Add GGML_OPENVINO_NPU_COMPILE_CONFIG to the backend's cached environment so callers can configure the NPU compiler without using the generic property escape hatch.

When the value is non-empty, pass it to OpenVINO as NPU_COMPILATION_MODE_PARAMS. This enables settings such as optimization-level=3 for NPU compilation while preserving the existing behavior when the variable is unset and leaving CPU and GPU configuration unchanged.

Document the variable, its NPU-only scope, and the optimization-level=3 example in the OpenVINO backend runtime configuration table.

* ggml-openvino : support RELU, POOL_2D, QUICK_GEGLU, and ROLL ops

* reorder op table

* exclude GPU/NPU failing POOL_2D case

* move op type detection to compute_op_case

* Relax rope supported cases

* Fix pool case

* Update openvino doc, gpu driver in ov docker

* openvino: remove unused static remote context branch

* openvino: parallelize static model build

* Apply editorconfig

---------

Co-authored-by: Mostafa Faheem <[email protected]>
Co-authored-by: Ravi Panchumarthy <[email protected]>
Co-authored-by: zhaixuejun1993 <[email protected]>
b10672
2026-08-28 14:42:07 +03:00
Xuan-Son Nguyen b19cbe925b convert: prevent ndarray conversion in LazyChunkedTensor (#27869) 2026-08-28 11:46:30 +02:00
Ozymandias_EBON d077b4c214 sycl: use TILE for quantized KV decode on BMG (#26689)
Route quantized KV decode to TILE on Xe2 (BMG) only, keep VEC on other archs until validated there.
b10670
2026-08-28 11:58:58 +03:00
Titaniumtown be876204aa sycl: bind the f16 KV cache in place for the oneDNN SDPA path (#27468)
Measured at a live KV length of 34816 (32768 depth plus one 2048 ubatch),
on Qwen3.8 27B Q4_K_S:

  per tensor         4 * 34816 * 256 * 2 B  =  71.3 MB
  staged per call    K and V, so 2x         = 142.6 MB
  traffic per call   read once, write once  = 285.2 MB
  traffic per ubatch 285.2 MB * 16 calls    =   4.56 GB

One ubatch is one ggml_cgraph submission (llama_context::process_ubatch ->
graph_compute), so that 4.56 GB is the cost of a single 2048-token prefill
chunk, and it scales with the live KV length: the first ubatch of the same run,
at seq = 2048, moves 0.27 GB.

Reproduce the two measured inputs with:

  GGML_SCHED_DEBUG=2 llama-bench -m MODEL -p 8 -n 0 -r 1 -ngl 0 \
      -fa on -ctk f16 -ctv f16 -v > nd.txt 2>&1
  grep -E 'n_layer|n_head_kv|n_embd_head_k' nd.txt
  awk '/node #  0 /{g++} g==1 && /\(FLASH_ATTN\)/{n++} END{print n+0}' nd.txt
b10669
2026-08-28 11:53:31 +03:00