Finding
ADR-189 (proposed in #554) names ettin-reranker-17m-v1 and -32m-v1 (Apache-2.0) as the lead gate candidates. The pinned llama.cpp (ec2b787) cannot run them correctly as rerankers yet:
- In
src/llama-graph.cpp, the LLAMA_POOLING_TYPE_RANK branch hard-codes mean pooling for LLM_ARCH_MODERN_BERT, to match gte-reranker-modernbert-base.
- Ettin rerankers use CLS pooling, then Dense, GELU, LayerNorm and Dense.
- Their head lives in separate Sentence Transformers module folders, which the converter does not pick up.
Proposal
- Patch the rank path to select pooling from GGUF metadata rather than from the architecture.
- Repack the Ettin head in
convert_hf_to_gguf.py so the classifier tensors land where the rank path expects them.
- Offer the change upstream, and carry it on the submodule until it lands.
- Verify score parity against the Sentence Transformers reference to within 1e-3 (ADR-189 eligibility).
If the patch cannot be carried, ADR-189 falls back to the baselines (ms-marco-MiniLM-L6-v2, jina-reranker-v1-tiny-en), which need no change.
Finding
ADR-189 (proposed in #554) names
ettin-reranker-17m-v1and-32m-v1(Apache-2.0) as the lead gate candidates. The pinned llama.cpp (ec2b787) cannot run them correctly as rerankers yet:src/llama-graph.cpp, theLLAMA_POOLING_TYPE_RANKbranch hard-codes mean pooling forLLM_ARCH_MODERN_BERT, to matchgte-reranker-modernbert-base.Proposal
convert_hf_to_gguf.pyso the classifier tensors land where the rank path expects them.If the patch cannot be carried, ADR-189 falls back to the baselines (
ms-marco-MiniLM-L6-v2,jina-reranker-v1-tiny-en), which need no change.