Default Branch

d230ddd763 · llama: fix whole source code rebuilt on each new commit (#28278) · Updated 2026-09-03 23:53:04 +02:00

Branches

7216af5c09 · ggml : fix 32-bit ARM compat (cont) · Updated 2024-01-09 09:33:16 +01:00    Superminaren

8999
2

d57cb9c294 · passkey : add readme · Updated 2024-01-08 10:13:44 +01:00    Superminaren

9009
7

7cfde78190 · llama : remove redundant GQA check · Updated 2024-01-06 15:04:20 +01:00    Superminaren

9017
1

9f51f3e695 · metal : opt mul_mm_id · Updated 2024-01-02 19:50:18 +01:00    Superminaren

9043
17

4cc78d3873 · ggml : force F32 precision for ggml_mul_mat · Updated 2024-01-02 16:54:56 +01:00    Superminaren

9042
1

b5af7ad84f · llama : refactor quantization to avoid <mutex> header · Updated 2024-01-02 14:56:57 +01:00    Superminaren

9045
1

120a1a5515 · llama : auto download HF models if URL provided · Updated 2024-01-02 12:29:06 +01:00    Superminaren

9046
1

f64e4f04e7 · ggml : testing GPU FP precision via quantized CPY · Updated 2023-12-30 18:11:40 +01:00    Superminaren

9064
1

f32f30bc57 · test · Updated 2023-12-26 16:52:42 +01:00    Superminaren

9094
1

ab1b75166f · Merge branch 'master' into gg/ggml_scale · Updated 2023-12-21 21:35:11 +01:00    Superminaren

9117
4

7c87353e61 · common : remove incorrect --model-draft default · Updated 2023-12-21 18:17:12 +01:00    Superminaren

9125
1

a40f6110f0 · ggml : force F32 precision for ggml_mul_mat · Updated 2023-12-19 15:34:59 +01:00    Superminaren

9132
1

3c734f4941 · plamo : testing · Updated 2023-12-18 16:06:05 +01:00    Superminaren

9137
13

a462159c43 · cuda : ggml_cuda_op_mul_mat_cublas support F32 precision · Updated 2023-12-18 13:24:29 +01:00    Superminaren

9137
16

1b05817112 · decode : fix logits_valid for old API · Updated 2023-12-18 00:49:21 +01:00    Superminaren

9138
1

865066621b · llama.swiftui : improve bench · Updated 2023-12-17 18:37:22 +01:00    Superminaren

9152
12

f86b9d152c · lookup : minor · Updated 2023-12-17 16:25:28 +01:00    Superminaren

9150
9

d2f1e0dacc · Merge branch 'cuda-cublas-opts' into gg/phi-2 · Updated 2023-12-17 07:41:46 +01:00    Superminaren

9148
17

b0547d2196 · gguf-py : fail fast on nonsensical special token IDs · Updated 2023-12-16 00:06:42 +01:00    Superminaren

9150
1

c8554b80be · Merge branch 'master' of https://github.com/ggerganov/llama.cpp into ceb/fix-cuda-warning-flags · Updated 2023-12-13 18:06:01 +01:00    Superminaren

9162
12