Files
llama.cpp/src
JamePeng cc231cb0da dflash: pass missing NVFP4 scales to attention operations (#28000)
- DFlash2 NVFP4 draft models produced almost no accepted speculative
tokens because the Q, K, V, and output projection scales were not
passed to the corresponding graph operations.
2026-08-30 11:34:39 +03:00
..
2026-08-21 19:52:34 +02:00
2026-08-21 19:52:34 +02:00
2026-06-29 16:58:51 +08:00