Repository navigation
Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
Contributor
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Comment |
kaix-nv
added this pull request to stack #2563
October 5, 2026 06:11
kaix-nv
removed this pull request from stack #2563
October 5, 2026 06:15
kaix-nv
added this pull request to stack #2658
October 5, 2026 06:15
This was referenced Oct 5, 2026
kaix-nv
force-pushed
the
kaix/linear-attention-qat-example
branch
from
October 6, 2026 17:27
2829885 to
3837d28
Compare
kaix-nv
force-pushed
the
kaix/linear-attention-qat-example
branch
from
October 6, 2026 19:55
3837d28 to
b7401f7
Compare
kaix-nv
force-pushed
the
kaix/linear-attention-qat-example
branch
from
October 6, 2026 21:26
b7401f7 to
0357915
Compare
Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
Update the linear-attention example for the flat execution config, native state-only training, and checkpoint migration. Remove the retired CPU reference script while retaining its historical results and source link. Record the focused CPU, native GPU, Megatron checkpoint, and config migration validation. Documentation pre-commit hooks passed. Signed-off-by: Kai Xu <kaix@nvidia.com>
Document the QDQ scheduling and arithmetic mismatch, the native prefix/handoff/suffix solution, gradient flow, and the limits of current validation. Align configuration guidance with the current state QAT API. Signed-off-by: Kai Xu <kaix@nvidia.com>
Remove raw experiment JSON from the example and shorten the usage and state-alignment documents. Retain the mismatch derivation, supported policies, concise numerical results, and validation limits. Preserve development evidence in an untracked local archive. Training code and minimal QAT/QAD tests are unchanged. Validation: scoped pre-commit, documentation links and anchors, Python snippet syntax, archive integrity, and git diff --check. Signed-off-by: Kai Xu <kaix@nvidia.com>
kaix-nv
force-pushed
the
kaix/linear-attention-qat-example
branch
from
October 8, 2026 05:35
0d898ac to
7717ea4
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## kaix/linear-attention-decode-first #2657 +/- ##
===================================================================
Coverage 78.62% 78.62%
===================================================================
Files 650 650
Lines 71262 71262
===================================================================
Hits 56029 56029
Misses 15233 15233
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Linear-attention series — 7 PRs (1 merged, 6 open)
mainThe six open PRs form native GitHub stack #2658 in the order shown; #2497 is retained as the merged foundation in this seven-PR series. #2497 is merged, so #2519 targets
main. #2657 contains the training example split from #2519; #2562 now targets #2657. #2657 is rebased onto the latest #2519. Rebase the remaining descendants after their immediate parent merges.#2541 applies TensorQuantizer before native vLLM prefill/decode calls. Serving-time prefill-GEMM quantization remains deferred until an optimized fused kernel is available. #2506 and #2509 are superseded and closed.
What does this PR do?
Type of change: New example.
Add a Megatron Bridge example for recurrent-state QAT and QAD. The training entry point selects a native chunked prefix and recurrent suffix, applies loss to the suffix, and keeps the phase context active through backward. Bridge owns optimization, distributed scheduling, and checkpoints. A frozen unquantized teacher enables QAD.
The example contains the training script, dependency files, launcher, two minimal training tests, and concise usage/alignment documentation. Raw experiment records and development links are kept outside the release PR. State formats, execution policies, and kernels are supplied by the parent PR.
Usage
Add
--teacher-model /path/to/unquantized-modelfor QAD. The ordinary INT8 recipe uses public vLLM kernels. INT8 + Hadamard and replay require a compatible native ReplaySSM fork. A KDA trainer additionally requires a compatible Bridge provider. TP/PP/EP options follow Bridge; context parallelism stays at one, and each local pipeline chunk must contain linear attention.Testing
After rebasing onto the latest #2519:
pytest tests/examples/megatron_bridge/test_linear_attention.py: 2 passed in 99.85 seconds on one RTX A6000, using public vLLM 0.15.1 kernels and matching Megatron Bridge/Core dependencies. Shared setup took 84.22 seconds; QAT and QAD calls each took 4.07 seconds. The tests check student updates and a frozen unquantized QAD teacher.git diff --checkpassed. Relative documentation links/anchors and Python snippet syntax were checked.Previous workflow checks exercised the real GDN Bridge training entry point with tiny random weights and mock data, including Hadamard token mode and replay QAT/QAD. KDA checks exercised Megatron layers rather than a Bridge KDA trainer. These do not establish pretrained-model quality recovery, full serving-engine equivalence, or performance.
Before your PR is "Ready for review"
CONTRIBUTING.md?: The optional public vLLM dependency and its license are listed separately. INT8/Hadamard and replay require the documented native fork.Additional Information
Step 3/7. Merge #2519 first, then this example, followed by #2562. This example trains state quantization; prefill GEMM quantization and approximate inverse remain follow-up work. Rebased onto #2519 at
94c5ead61c; the review diff contains only this example (8 files, 821 added lines).