Conversation
The K2-Horizon regex had no arm in unicode_regex_split_custom and fell through to the
general std::regex fallback, which fails two ways.
On MSVC std::regex rejects \p{...}, so no K2-Horizon GGUF loads on Windows at all:
llama-quantize, llama-imatrix and llama-perplexity all abort with
regex_error(error_escape) before a token is produced.
Where the fallback does compile it is still wrong. unicode_regex_split collapses each
codepoint to a single byte naming its Unicode category before matching, and U+200C/U+200D
are category Control, which has no entry in k_ucat_cpt, so both become the 0xD0 fallback
byte. The literal and alternatives in K2's regex can then never match and
every ZWNJ or ZWJ ends a letter run.
The splitter is the existing llama3 one with a single rule widened, since K2's regex
differs from llama3's only in that a letter run also takes marks, ZWNJ and ZWJ.
tests/test-unicode.cpp gains a case for this: it fails before the change with
[Amy] [ZWNJ khaham] and passes after with the run intact.
…ed-tests tests: expand K2 Horizon unicode splitter coverage
Assisted-by: Codex
…nizer unicode : add the K2-Horizon pre-tokenizer splitter
Assisted-by: Codex
Includes the K2 Horizon implementation from ifm-ai/llama.cpp with converter, tensor-parallel and model save/reload fixes. Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
|
Hi @bitalov, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
Assisted-by: Codex
Constrain final JSON after reasoning, accept flexible JSON tool envelopes, enforce XML dialects, and handle repeated or alternate thinking markers. Load YaRN beta metadata instead of retaining the default values. Add schema, streaming, continuation, and model reload regressions. Validate CUDA and CPU builds and 0.9B, 4B, and MoVA conversation/tool round trips. Assisted-by: Codex
| // like jinja2, an all-digit attribute is an index into a sequence item, | ||
| // e.g. rejectattr('0', 'equalto', '$ref') on the (key, value) pairs of dict|items |
| if not part_names and not self.is_mistral_format: | ||
| part_names = ModelBase.get_model_part_names(self.dir_model, "pytorch_model", ".safetensors") |
There was a problem hiding this comment.
Which model is this for?
Either way, potential incoming fix here #29650 (comment)
|
|
||
| def set_vocab(self): | ||
| super().set_vocab() | ||
|
|
||
| # the 0.9B repo keeps an older chat_template.jinja next to the served chat_template_generation.jinja | ||
| tmpl_file = self.dir_model / "chat_template_generation.jinja" | ||
| if tmpl_file.is_file(): | ||
| self.gguf_writer.remove_key(gguf.Keys.Tokenizer.CHAT_TEMPLATE) | ||
| self.gguf_writer.add_chat_template(tmpl_file.read_text(encoding="utf-8")) | ||
| logger.info(f"gguf: using {tmpl_file.name} as the chat template") |
There was a problem hiding this comment.
| def set_vocab(self): | |
| super().set_vocab() | |
| # the 0.9B repo keeps an older chat_template.jinja next to the served chat_template_generation.jinja | |
| tmpl_file = self.dir_model / "chat_template_generation.jinja" | |
| if tmpl_file.is_file(): | |
| self.gguf_writer.remove_key(gguf.Keys.Tokenizer.CHAT_TEMPLATE) | |
| self.gguf_writer.add_chat_template(tmpl_file.read_text(encoding="utf-8")) | |
| logger.info(f"gguf: using {tmpl_file.name} as the chat template") |
Even if it does, chat_template.jinja is the one transformers picks, so you better make sure that's the right one.
|
|
||
| gating_funcs = {"sigmoid": gguf.ExpertGatingFuncType.SIGMOID, "softmax": gguf.ExpertGatingFuncType.SOFTMAX} | ||
| router_func = hparams.get("router_score_func") | ||
| if router_func not in gating_funcs: | ||
| raise ValueError(f"Unsupported router_score_func: {router_func!r}") |
There was a problem hiding this comment.
| gating_funcs = {"sigmoid": gguf.ExpertGatingFuncType.SIGMOID, "softmax": gguf.ExpertGatingFuncType.SOFTMAX} | |
| router_func = hparams.get("router_score_func") | |
| if router_func not in gating_funcs: | |
| raise ValueError(f"Unsupported router_score_func: {router_func!r}") |
Add router_score_func here:
Line 1532 in 757e62e
| self.gguf_writer.add_expert_shared_feed_forward_length(n_ff_exp * n_shared) | ||
| if (router_scale := hparams.get("router_scaling_factor")) is not None: | ||
| self.gguf_writer.add_expert_weights_scale(float(router_scale)) | ||
| self.gguf_writer.add_expert_gating_func(gating_funcs[router_func]) |
There was a problem hiding this comment.
| self.gguf_writer.add_expert_gating_func(gating_funcs[router_func]) |
| # K2 Horizon. 2 hashes because various sets of tokens depending on size | ||
| {"name": "k2-horizon", "tokt": TOKENIZER_TYPE.BPE, "repo": "https://huggingface.co/IFM/K2-Horizon-0.9B", "chkhsh": "1f9825a388f700a6b591722f17d470cbbcf10973ece35d2fd14239a14110ae1a"}, | ||
| {"name": "k2-horizon", "tokt": TOKENIZER_TYPE.BPE, "repo": "https://huggingface.co/IFM/K2-Horizon-36B", "chkhsh": "a9af07a84191f55098b248ae6f3dfe9e32d3190bebe8eafd91c1ddec9bc3449f"}, |
There was a problem hiding this comment.
Put one in the regular models list.
This PR adds support for K2 Horizon, both the dense and MoVA variants. It covers HF-to-GGUF conversion, tokenizer support, model loading, and the inference graph. Everything is built on existing GGML operators, so there are no new kernels.
Along the way, a few supporting changes were needed:
Numeric indexing in
selectattrandrejectattr. The K2 chat template processes tool schemas withdict | items | rejectattr('0', 'equalto', '$ref'), which needs to reach the first element of each key/value pair. The existing filters couldn't do that, so I added support for it.A dedicated K2 chat parser. The auto-parser figures out reasoning and tool-call formats by rendering sample conversations. Some of those samples include assistant turns without
reasoning_content, and the K2 template rejects them, so discovery fails. The dedicated parser handles K2's reasoning and tool calls, its end-of-turn marker, and the case wheretool_choice=requiredjumps straight to a tool call.Template capability detection in
common/jinja/caps.cpp. The same missing field can make capability probes fail, which then wrongly reports that tools, object arguments, or parallel calls aren't supported. Now, if a probe fails, it retries once with an emptyreasoning_contentadded to any assistant turns that lack it. This does touch shared capability detection, but when I compared results across 75 templates, only the K2 templates changed.Multi-GPU support with
-sm tensor. Value-expert weights follow the same split rule as the value weights. Q/K normalization weights are loaded per head, and the MoVA expert sum uses 3D views so the tensor-parallel backend can keep track of the split. None of this requires a change to the GGUF format.MoVA save/reload support. The saver now writes the value-expert metadata needed to load saved models back in.
AI usage disclosure: YES - Codex reviewed the changes, made targeted fixes, and ran additional validation.