Repository navigation
Conversation
Contributor
Author
|
Sorry, agent went crazy - I won't ask for a review of 9k lines of code. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Decoder embedding checkpoints need their declared bidirectional attention mode to be honored. This adds that support to Llama, Mistral, and Qwen3, and corrects Gemma3Text bidirectional window handling, while retaining causal defaults.
Depends on #474. This draft targets upstream
main, so its current diff includes the prerequisite RoPE commit. The embedding-only changes are in commit 11d6863. Merge the prerequisite first, then rebase this branch before merging.Validation: 83 targeted tests passed, including existing Llama, Mistral, Qwen3, Gemma3Text, and Phi3 tests. Formatting, compilation with
--warnings-as-errors, unused-dependency checks, and diff whitespace checks pass. All 12 fixture checkpoints were regenerated using PyTorch 2.7.1 and Transformers commit3693f8d26311305e914735a6373fb03468d6aaa0; their configurations, weights, and Python reference outputs reproduce exactly. The generator is included undertest/fixtures/embedding_models/.The full test suite and opt-in real EmbeddingGemma integration were not run. This supports the transformer backbone, not the complete trained EmbeddingGemma projection pipeline. The dependency reference above links a PR and does not close an issue.
AI disclosure: this contribution was developed with substantial LLM/coding-agent assistance, including implementation, tests, and review. Validation above was executed by the coding agent.