Ruben Ortlam
6421fbe532
clean up
2026-08-29 11:51:15 +02:00
Ruben Ortlam
c30a8740f2
improve offset application
2026-08-29 11:51:15 +02:00
Ruben Ortlam
0eab519531
add RDNA4 architecture, use for hardcoded coopmat elem thread access, set BK_STEP back to 4
2026-08-29 11:51:15 +02:00
Ruben Ortlam
f252045d89
undo uint8_t, gate to RDNA3/4
2026-08-29 11:51:15 +02:00
Ruben Ortlam
9ef5996da4
merge shmem arrays
2026-08-29 11:51:15 +02:00
Ruben Ortlam
2ba1d225e1
dedup b scales
2026-08-29 11:51:15 +02:00
Ruben Ortlam
eeeacedd4e
improvements
2026-08-29 11:51:15 +02:00
Ruben Ortlam
092661df98
improve performance
2026-08-29 11:51:15 +02:00
Ruben Ortlam
3985f64670
improve performance
2026-08-29 11:51:15 +02:00
Ruben Ortlam
22f588c10d
fix l warptile
2026-08-29 11:51:15 +02:00
Ruben Ortlam
4695f108bb
add q3_k, q4_k, q5_k, q6_k and nvfp4 support
2026-08-29 11:51:15 +02:00
Ruben Ortlam
ab94ffdd30
use 4-byte loads where possible
2026-08-29 11:51:15 +02:00
Ruben Ortlam
5037f027aa
use shmem arrays for LUTs
2026-08-29 11:51:15 +02:00
Ruben Ortlam
ba8402e9f6
remove elem row/col fast path, invalid for RDNA4
2026-08-29 11:51:15 +02:00
Ruben Ortlam
dc08fa9101
support iq4_nl and mxfp4
2026-08-29 11:51:15 +02:00
Ruben Ortlam
3cd76432d7
fix mul_mat_id bug
2026-08-29 11:51:15 +02:00
Ruben Ortlam
d8ed37209a
fix segfault
2026-08-29 11:51:15 +02:00
Ruben Ortlam
fa37db93ce
enable mul_mat_id support
2026-08-29 11:51:15 +02:00
Ruben Ortlam
ed294e8724
restructure mmq cm1 functions
2026-08-29 11:51:15 +02:00
Ruben Ortlam
d2a5ac69c9
add q4_1, q5_0, q5_1 support
2026-08-29 11:51:15 +02:00
Ruben Ortlam
eb46b8cf0a
move quant-specific prefetch function out of main file
2026-08-29 11:51:15 +02:00
Ruben Ortlam
1825270eb4
fix compilation
2026-08-29 11:51:15 +02:00
Ruben Ortlam
42a6269bda
use BK_STEP 4
2026-08-29 11:51:15 +02:00
Ruben Ortlam
ab6e84bbdd
only force subgroup size 32 on AMD RDNA
2026-08-29 11:51:15 +02:00
Ruben Ortlam
57dd3f8bd3
skip computation for inactive tiles
2026-08-29 11:51:15 +02:00
Ruben Ortlam
2b9c53381e
restructure for vgpr use
2026-08-29 11:51:15 +02:00
Ruben Ortlam
ef5c62992f
Revert "increase large tile size"
...
This reverts commit 7fabc25c5e0c9047a7a25b23bba5cad017de23ec.
2026-08-29 11:51:15 +02:00
Ruben Ortlam
a9e0e89c0b
increase large tile size
2026-08-29 11:51:15 +02:00
Ruben Ortlam
57b4276784
use wave32
2026-08-29 11:51:15 +02:00
Ruben Ortlam
15b3ce0de3
Revert "revert load reordering and scale pre-loading"
...
This reverts commit fbaefe0eaaccbb6ab9d733ad84580428f18db410.
2026-08-29 11:51:15 +02:00
Ruben Ortlam
9f069f42e8
clean up
2026-08-29 11:51:15 +02:00
Ruben Ortlam
23356762ca
workgroup scheduling for cache proximity
2026-08-29 11:51:15 +02:00
Ruben Ortlam
31b9bc683e
revert load reordering and scale pre-loading
2026-08-29 11:51:15 +02:00
Ruben Ortlam
ad015d6a44
add faster RDNA int->float conversion
2026-08-29 11:51:15 +02:00
Ruben Ortlam
95d14c99ea
use float for scales
2026-08-29 11:51:15 +02:00
Ruben Ortlam
14a6080dd5
coopmat load first, then wmma
2026-08-29 11:51:15 +02:00
Ruben Ortlam
b453be38a3
preload scales
2026-08-29 11:51:15 +02:00
Ruben Ortlam and Piotr Wilkin
5553b89912
double buffering
...
Co-authored-by: Piotr Wilkin (ilintar) <[email protected] >
2026-08-29 11:51:13 +02:00
Ruben Ortlam and Piotr Wilkin
2a878d9ea9
use larger workgroups
...
Co-authored-by: Piotr Wilkin (ilintar) <[email protected] >
2026-08-29 11:51:10 +02:00
Ruben Ortlam and Piotr Wilkin
b121ef3b33
add BK_STEP to shader, default to 2
...
Co-authored-by: Piotr Wilkin (ilintar) <[email protected] >
2026-08-29 11:51:08 +02:00
Ruben Ortlam
513f43600c
add q8_0 support
2026-08-29 11:51:08 +02:00
Ruben Ortlam and Piotr Wilkin
162560885a
probe and directly access coopmat values instead of going through shmem
...
Co-authored-by: Piotr Wilkin (ilintar) <[email protected] >
2026-08-29 11:51:04 +02:00
Ruben Ortlam
5f034d6086
use scalar sums
2026-08-29 11:30:17 +02:00
Ruben Ortlam
e086b9ca05
apply scales inline
2026-08-29 11:30:17 +02:00
Ruben Ortlam
13f6d8e389
vulkan: add int8 coopmat quantized matmul shader
2026-08-29 11:30:17 +02:00
Nick Farrell
cc83d7b482
sycl: make --fit respect --fit-target better ( #27629 )
...
improve the --fit algorithm to take into account the actual peak
required VRAM for a given context size on a SYCL backend.
This includes both properly accounting for how much VRAM is required
when the allocated context is fully used (which makes the reported
context drop below what it did before, but stop it OOMing) as well
as preventing some overly-conservative calculations which meant too much
VRAM was being reserved.
Tested on a Arc b70 with unsloth's qwen3.8 (Q4_K_XL), able to get 262144 context,
fully usable, with q8_0 KV and MTP and 4k ubatch size using --fit-target 1
b10684
2026-08-29 05:00:09 -04:00
Jeff Bolz
c9ca51c1f6
vulkan: combine duplicated fastdiv functions, rename the one optimizing small divs ( #27526 )
...
* vulkan: combine duplicated fastdiv functions, rename the one optimizing small divs
* remove one more fastdiv
b10683
2026-08-29 10:59:48 +03:00
Jhen-Jie Hong
5ea1b124e7
metal : add fa-vec tunings for M1 Max ( #27932 )
b10682
2026-08-29 15:12:23 +08:00
Jeff Bolz
77f132cb1d
vulkan: Change mul_mat_id to pad K rather than N ( #27925 )
...
The N padding is needed for mul_mat, but not mul_mat_id. For mul_mat_id,
we indirect the row index through a shared memory lookup table which avoids
any OOB row coordinate. But that callback doesn't bounds check K, so we
actually need K padding instead.
b10681
2026-08-29 10:09:24 +03:00
kurquhar and Kristopher Urquhart
d7bd3bfcad
snapdragon: python SDK setup (Windows) ( #27903 )
...
* port setup-build.ps1 to setup_sdk.py, to facilitate installation of Hexagon and OpenCL SDKs on Windows
* rename setup_sdk.py -> setup-sdk.py
* flake8 fix: print() -> logger.info()
---------
Co-authored-by: Kristopher Urquhart <[email protected] >
b10680
2026-08-28 14:01:59 -07:00