Logo
Explore Help
Sign In
Superminaren/llama.cpp
Watch 1
Star 0
Fork 0
mirror of https://github.com/ggml-org/llama.cpp.git synced 2026-09-04 10:47:38 +02:00
Code Issues Packages Projects Releases Wiki Activity
Files
b6978
llama.cpp/tools
T
History
Georgi Gerganov 7956bb4d7f bench : cache the llama_context state at computed depth (#16944)
* bench : cache llama_context state at depth

* cont : handle failures to restore the old state

* cont : print information when the state is being reused
2025-11-07 21:23:11 +02:00
..
batched-bench
scripts : add script to bench models (#16894)
2025-11-02 00:15:31 +02:00
cvector-generator
…
export-lora
…
gguf-split
ci : use smaller model (#16168)
2025-09-22 09:11:39 +03:00
imatrix
Manually link -lbsd to resolve flock symbol on AIX (#16610)
2025-10-23 19:37:31 +08:00
llama-bench
bench : cache the llama_context state at computed depth (#16944)
2025-11-07 21:23:11 +02:00
main
llama-cli: prevent spurious assistant token (#16202)
2025-09-29 10:03:12 +03:00
mtmd
hparams : add n_embd_inp() to support extended embed (#16928)
2025-11-07 19:27:58 +01:00
perplexity
perplexity : show more kl-divergence data (#16321)
2025-09-29 09:30:45 +03:00
quantize
ci : use smaller model (#16168)
2025-09-22 09:11:39 +03:00
rpc
rpc : report actual free memory (#16616)
2025-10-17 18:02:52 +03:00
run
Manually link -lbsd to resolve flock symbol on AIX (#16610)
2025-10-23 19:37:31 +08:00
server
kv-cache : pad the cache size to 256 for performance (#17046)
2025-11-07 20:03:25 +02:00
tokenize
…
tts
model : Apertus model implementation (#15852)
2025-10-02 20:43:22 +03:00
CMakeLists.txt
…
Powered by Gitea Version: 1.27.2 Page: 2026ms Template: 21ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API