The collapsed \p{S} class was missing '~', which split " ~" into
separate pre-tokens and prevented the Ġ~ BPE merge used by DeepSeek V4.
This caused re-tokenized prompts to diverge from sampled tokens and
broke KV cache reuse.
Assisted-by: Codex
* common/chat: update DeepSeek V4 templates
Align the DeepSeek V4 templates with the official encoders while keeping parser behavior out of this change.
- Default drop_thinking for DeepSeek V4 history so prior thinking is omitted unless preserve_reasoning is requested or tools are present.
- Add structured output response-format instructions to the V4 templates and pass the schema into template rendering.
- Add a separate Flash 0731 template for the updated high and max reasoning effort mapping.
- Cover reasoning effort, drop_thinking, structured output prompts, preserved reasoning, continuations, and empty tool arguments in template rendering tests.
Official references:
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/encoding/encoding_dsv4.pyhttps://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py
Assisted-by: Codex
* Fix deepseek v4 0731 template selection
* remove unneeded lower normalization
* Fix DSML parser to consume the tool call separator
* address aldehir requests
* address aldehir comment
- Add a supports_mtp_export capability to ModelBase so architectures can opt
into --mtp and --no-mtp without extending a central class allowlist.
- Enable the capability for the existing Qwen3.5/3.6 and Step3.5/3.7
implementations, and for HY V3, whose converter already supports
filtering the appended MTP layers.
Before this commit, --cache-ram was not a hard limit:
- The cache always kept at least one entry, even if that entry exceeded the
RAM/token limits.
- Old entries were only evicted for the RAM/token limits after saving the new
one, which could cause the cache to temporarily exceed the RAM/token limits
even if individual entries were below the limit.
Now, ensure that the RAM limit is strict with these changes:
- Skip saving state to cache if by itself it exceeds the RAM limit.
- Evict old entries as necessary to make the new entry fit.
Additionally, token-limit cleanup may now evict the last remaining cache entry
instead of always preserving one.