Alberto Cabrera Pérez
afc0e89698
sycl: refactor quantization to q8_1 ( #14815 )
...
* sycl: quantization to q8_1 refactor
* Refactored src1 copy logic in op_mul_mat
2025-07-28 11:05:53 +01:00
Alberto Cabrera Pérez
cb4a63aad6
sycl: fixed semantics of block offset calculation ( #14814 )
2025-07-24 11:09:57 +01:00
Alberto Cabrera Pérez
725f23f1f3
sycl : backend documentation review ( #13544 )
...
* sycl: reviewing and updating docs
* Updates Runtime error codes
* Improves OOM troubleshooting entry
* Added a llama 3 sample
* Updated supported models
* Updated releases table
2025-05-19 14:38:20 +01:00
Alberto Cabrera Pérez
f71f40a284
ci : upgraded oneAPI version in SYCL workflows and dockerfile ( #13532 )
2025-05-19 11:46:09 +01:00
Alberto Cabrera Pérez and romain.biessy
17512a94d6
sycl : implementation of reordered Q4_0 MMVQ for Intel GPUs ( #12858 )
...
* sycl : Implemented reorder Q4_0 mmvq
Signed-off-by: Alberto Cabrera <[email protected] >
* sycl : Fixed mmvq being called when reorder is disabled
* sycl : Improved comments in the quants header
Signed-off-by: Alberto Cabrera <[email protected] >
* Use static_assert
* safe_div -> ceil_div
* Clarify qi comment
* change the reorder tensor from init to execute OP
* dbg
* Undo changes to test-backend-ops
* Refactor changes on top of q4_0 reorder fix
* Missing Reverts
* Refactored opt_for_reorder logic to simplify code path
* Explicit inlining and unroll
* Renamed mul_mat_algo enum for consistency
---------
Signed-off-by: Alberto Cabrera <[email protected] >
Co-authored-by: romain.biessy <[email protected] >
2025-05-09 16:34:08 +01:00
Alberto Cabrera Pérez
8733e0cf6e
sycl: addressing non-contiguous src1 mul_mats (nc and batched) ( #13343 )
...
* sycl: fixed non-contiguous src1 mul_mats (nc and batched)
* Fixed wrong static_cast inside kernel
2025-05-08 10:08:01 +01:00
Alberto Cabrera Pérez
5a63980117
llama-bench: fixed size of fields to correctly map to values ( #13183 )
2025-04-29 17:24:36 +02:00
Alberto Cabrera Pérez
363f8c5d67
sycl : variable sg_size support for mmvq kernels ( #12336 )
2025-03-12 09:57:32 +00:00
Alberto Cabrera Pérez
0f77aae560
sycl : offload of get_rows set to 0 ( #10432 )
2024-11-29 20:38:45 +08:00
Alberto Cabrera Pérez
266b8519ee
sycl : Reroute permuted mul_mats through oneMKL ( #10408 )
...
This PR fixes the failing MUL_MAT tests for the sycl backend.
2024-11-29 09:49:43 +00:00
Alberto Cabrera Pérez
557924f222
sycl: Revert MUL_MAT_OP support changes ( #10385 )
2024-11-19 08:50:04 +08:00
Alberto Cabrera Pérez
2e82ffa4af
sycl : Fixes to broken builds and test-backend-ops ( #10257 )
...
* Fixes broken build for the SYCL CUDA backend caused by non-explicit gemm call in outprod (merged in with RWKV6 in
Optimize RWKV6 Operator Naming and Implement Multi-core CPU/ SYCL Acceleration #10133 )
* Marks permuted MUL_MAT as unsupported to be able to run test-backend-ops
* Fixes asserts in norm to fix debug builds.
2024-11-13 09:40:57 +00:00
Alberto Cabrera Pérez
f536f4c439
[SYCL] Initial cmake support of SYCL for AMD GPUs ( #9658 )
...
sycl: initial cmake support of SYCL for AMD GPUs
2024-10-02 13:57:18 +01:00
Alberto Cabrera Pérez
51b6038636
sycl : update support conditions ( #9394 )
...
* sycl : update support condition to im2col
Signed-off-by: Alberto Cabrera <[email protected] >
* Added TODO to remind supporting FP32 im2col
---------
Signed-off-by: Alberto Cabrera <[email protected] >
2024-09-11 08:53:42 +08:00
Alberto Cabrera Pérez
5b0b8d8cfb
sycl : Reenabled mmvq path for the SYCL Nvidia Backend ( #8372 )
...
* SYCL : Reenabled mmvq path for the SYCL Nvidia Backend
* Reduced verbosity of comment
2024-07-09 22:03:15 +08:00
Alberto Cabrera Pérez
a130eccef4
labeler : updated sycl to match docs and code refactor ( #8373 )
2024-07-08 22:35:17 +02:00