Default Branch

d230ddd763 · llama: fix whole source code rebuilt on each new commit (#28278) · Updated 2026-09-03 23:53:04 +02:00

Branches

29fe516913 · wip · Updated 2023-10-31 17:36:37 +01:00    Superminaren

9347
1

dab42893c9 · scripts : working curl pipe · Updated 2023-10-31 16:03:56 +01:00    Superminaren

9347
3

7923b70cb8 · llama : add llm_build_inp_embd helper · Updated 2023-10-31 15:43:08 +01:00    Superminaren

9352
37

4b3cb98d46 · ggml-impl : move extern "C" to start of file · Updated 2023-10-30 18:05:58 +01:00    Superminaren

9348
7
lto

bc28aaa8c2 · make : use -lfto=auto to avoid warnings and maintain perf · Updated 2023-10-30 15:00:53 +01:00    Superminaren

9348
5

15267192c0 · llama : refactor tensor offloading as callback · Updated 2023-10-29 12:04:36 +01:00    Superminaren

9352
15

8a86b95e87 · quantize : --pure option for disabling k-quant mixtures · Updated 2023-10-28 22:37:03 +02:00    Superminaren

9353
3

de7e0912b6 · convert : ignore tokens if their IDs are within [0, vocab_size) · Updated 2023-10-28 14:01:36 +02:00    Superminaren

9356
1

bbfc62ac2f · sampling : temp == 0.0 -> no probs, temp < 0.0 -> probs · Updated 2023-10-28 13:04:57 +02:00    Superminaren

9364
3

cd3e20fb50 · cuda : fix multi-gpu with tensor cores · Updated 2023-10-27 22:11:50 +02:00    Superminaren

9363
3

49af767fad · build : add compile option to force use of MMQ kernels · Updated 2023-10-27 12:21:04 +02:00    Superminaren

9365
7

d798a17c34 · cuda : add TODO for calling cublas from kernel + using mem pool · Updated 2023-10-24 15:33:24 +02:00    Superminaren

9379
10

6966474928 · cuda : play with faster Q4_0 dequantization · Updated 2023-10-24 09:29:40 +02:00    Superminaren

9379
8

b9bb4cbe86 · Separate bug and enhancement template + no default title · Updated 2023-10-23 17:59:11 +02:00    Superminaren

9379
1

c0f4d54870 · server : add comment about changing slot_state to bool · Updated 2023-10-22 21:24:39 +02:00    Superminaren

9385
72

cb79f8a2d8 · llama : add SKIP_KQ_KQV option · Updated 2023-10-22 08:58:29 +02:00    Superminaren

9385
3

56ba00b923 · sampling : hide prev behind API and apply #3661 · Updated 2023-10-20 17:53:27 +02:00    Superminaren

9388
6

ad2727d091 · Merge branch 'master' into speculative-tree · Updated 2023-10-18 09:50:58 +02:00    Superminaren

9399
18

932589c0ef · Honor -ngl option for Cuda offloading in llava · Updated 2023-10-14 02:12:10 +02:00    Superminaren

9413
1

5261aee8d8 · sampling : one sequence per sampling context · Updated 2023-10-12 19:36:44 +02:00    Superminaren

9416
1