Default Branch

f114f91f9e · tests : initialize the L2_NORM batch array (#28553) · Updated 2026-09-07 19:54:13 +02:00

Branches

f648ca2cee · llama : add llama_sampling API + move grammar in libllama · Updated 2024-09-03 09:31:54 +02:00    Superminaren

7190
1

40fa68cb46 · readme : add API change notice · Updated 2024-09-02 17:32:24 +02:00    Superminaren

7199
3

a95225cdfd · metal : another fix for the fa kernel · Updated 2024-08-26 14:08:38 +02:00    Superminaren

7223
1

aa931d0375 · metal : fix fa kernel · Updated 2024-08-26 12:09:50 +02:00    Superminaren

7223
1

6494509801 · backup · Updated 2024-08-26 10:58:54 +02:00    Superminaren

7233
2

ccb45186d0 · docs : remove references · Updated 2024-08-26 08:52:17 +02:00    Superminaren

7227
2

8062650343 · llama : fix simple splits when the batch contains embeddings · Updated 2024-08-21 21:09:03 +02:00    Superminaren

7238
19

9127800d83 · wip · Updated 2024-08-17 01:51:06 +02:00    Superminaren

7271
2

62d7b6c87f · cuda : re-add q4_0 · Updated 2024-08-14 12:37:03 +02:00    Superminaren

7267
3

93ec58b932 · server : fix typo in comment · Updated 2024-08-14 04:12:26 +02:00    Superminaren

7269
4

faaac59d16 · llama : support NUL bytes in tokens · Updated 2024-08-12 03:00:03 +02:00    Superminaren

7280
1

73bc9350cd · gguf-py : Numpy dequantization for grid-based i-quants · Updated 2024-08-10 05:47:31 +02:00    Superminaren

7300
2

9329953a61 · llama : avoid double tensor copy when saving session to buffer · Updated 2024-08-07 22:03:34 +02:00    Superminaren

7308
2

cad8abb49b · add tool to allow plotting tensor allocation maps within buffers · Updated 2024-08-06 22:09:51 +02:00    Superminaren

7317
1

6e299132e7 · clip : style changes · Updated 2024-08-06 10:44:29 +02:00    Superminaren

7641
56

bddcc5f985 · llama : better replace_all · Updated 2024-08-04 12:42:08 +02:00    Superminaren

7342
1

229c35cb59 · gguf-py : remove LlamaFileTypeMap · Updated 2024-08-04 03:22:37 +02:00    Superminaren

7345
5

eab4a88210 · Using dp4a ptx intrinsics for an improved Mul8MAT perf [By Alcpz] · Updated 2024-07-29 17:52:29 +02:00    Superminaren

7363
1

9cddd9aeec · llama : cast seq_id in comparison with unsigned n_seq_max · Updated 2024-07-27 21:50:23 +02:00    Superminaren

7401
7

9aeb0e1f75 · sycl add conv support · Updated 2024-07-25 14:15:02 +02:00    Superminaren

7390
1