Skip to content

way-embed: verify scaled LoRA adapters on encoder rerankers in llama.cpp #557

Description

@aaronsb

Finding

ADR-190 (proposed in #554) stores each user's personal delta as a LoRA adapter on the reranker. It applies the adapter at a scale α: θ = (1−α)·θ_base + α·θ_delta, with α rising on reheat and annealing back. That relies on llama.cpp applying a LoRA adapter at a runtime scale to an encoder model (BERT or ModernBERT) running with rank pooling. This is unverified on the pinned commit (ec2b787).

Proposal

  1. Train a toy LoRA adapter on one of the ADR-189 gate candidates.
  2. Load it with a scale in llama.cpp on the rank path, and check that scores move with α and match a merged reference at α = 0 and α = 1.
  3. If scaled adapters do not work on this path, measure the fallback ADR-190 names: merging the adapter and re-exporting the GGUF. Record how long that takes at 17–32M parameters, so the α anneal can run as a re-export.

Record the outcome in ADR-190 section 3, under "Personal delta".

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:waysWays CLI, matching, steering layereffort:smallOne sitting

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions