Skip to content

[1/7] GDN state/W QAT foundation - #2497

Merged
kaix-nv merged 7 commits into
mainfrom
kaix/linear-attention-qat-m1
Oct 2, 2026
Merged

kaix-nv merged 7 commits into
mainfrom
kaix/linear-attention-qat-m1

Conversation

@kaix-nv

@kaix-nv kaix-nv commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Linear-attention series — 7 PRs (1 merged, 6 open)

Order PR Depends on
1/7 #2497 GDN state/W QAT foundation main
2/7 #2519 Torch GDN/KDA decode QAT + INT8 #2497
3/7 #2657 Megatron Bridge linear attention QAT/QAD example #2519
4/7 #2562 Fused Triton GDN/KDA decode QAT #2657
5/7 #2541 vLLM GDN/KDA state-only fake quantization #2562
6/7 #2503 GDN/KDA prefill GEMM quantization #2541
7/7 #2507 Experimental GDN/KDA approximate inverse #2503

The six open PRs form native GitHub stack #2658 in the order shown; #2497 is retained as the merged foundation in this seven-PR series. #2497 is merged, so #2519 targets main. #2657 contains the training example split from #2519; #2562 now targets #2657. Rebase each remaining descendant after its immediate parent merges.

#2541 applies TensorQuantizer before native vLLM prefill/decode calls. Serving-time prefill-GEMM quantization remains deferred until an optimized fused kernel is available. #2506 and #2509 are superseded and closed.

What does this PR do?

Type of change: new feature

GatedDeltaNet training keeps recurrent states inside a chunked kernel, so projection quantizers cannot emulate rounding at state boundaries. This PR adds dynamic per-tile FP8 E4M3 fake QDQ to the recurrent state and independent dynamic FP8 fake QDQ to WY-transformed W activations, with identity straight-through gradients for QAT/QAD.

Both sites use the standard quant_cfg interface and start disabled. State QDQ uses 64-token chunks and recomputes amax at each boundary over each full-key by 64-value-column tile, independently per sequence and head. Each tile has its own scalar scale (amax / 448, with a zero guard); fp8_scalar_qdq applies that supplied scale rather than choosing tensor-wide grouping. W grouping is applied by TensorQuantizer. Quantizer settings use normal ModelOpt checkpoint state. There is no QuantizeConfig.linear_attention field in this PR; #2519 introduces execution policies for decode and ReplaySSM, and later PRs extend them for prefill and approximate inverse. Configurations or checkpoints from earlier experimental drafts that use those execution policies require #2519; those selecting Triton decode also require #2562.

The Megatron adapter supports the direct-forward and older split-forward call layouts, restores the original kernel when disabled, and removes temporary quantizer attributes on export. Independent recurrent/chunk numerical references live under tests/_test_utils/torch/quantization/; shared runtime capability checks live in linear_attention/utils.py.

The fused path requires fla-core==0.5.1 and chunk size 64. State FP8 emulation requires SM89 or newer. The Hopper path has additional dtype/TileLang restrictions enforced before launch. This PR simulates numerical error; it does not add compressed state storage or faster inference.

Usage

import modelopt.torch.quantization as mtq

model = mtq.quantize(model, {
    "quant_cfg": [
        {"quantizer_name": "*", "enable": False},
        {"quantizer_name": "*gdn_state_quantizer",
         "cfg": {"num_bits": (4, 3), "type": "dynamic", "axis": (0, 1)}},
        {"quantizer_name": "*gdn_w_quantizer",
         "cfg": {"num_bits": (4, 3), "type": "dynamic", "axis": (0, 1, 2)}},
    ],
    "algorithm": None,
})
# Continue with the framework's normal forward/backward/optimizer steps.

Dynamic scales require no calibration.

Testing

The focused GPU suite contains four cases: three BF16 numerical forward/backward checks (disabled, W QDQ, and state+W QDQ) using one shared shape, plus one single-rank, one-layer Megatron QAT/checkpoint test. The Megatron test checks quantizer enable/disable behavior, checkpoint restore, gradients, and an optimizer update; it enables state QDQ when the GPU supports native FP8 conversion. Compilation runs in setup fixtures, and functional calls retain the normal 120-second timeout. There are no dtype, layout, tile-width, or parallelism sweeps.

The pinned FLA/TileLang/TVM-FFI dependencies live in the dev-fla optional extra, installed by both GPU nox sessions.

Validation of the consolidated changes on RTX A6000 (SM86), Python 3.12.8, Torch 2.9.1+cu128, Triton 3.5.1, fla-core 0.5.1, TileLang 0.1.8, Megatron Core 0.19.2, and Transformer Engine 2.16.0:

  • Cold and warm focused runs: 3 passed, 1 hardware skip each. The state+W numerical case requires SM89+; the local Megatron test exercised W QDQ.
  • Fresh Triton/TileLang cache: 363.09s total, including setup and teardown. Kernel setup took 66.38s + 44.46s; Megatron setup, including shared extension setup and worker startup, took 245.26s. Functional calls totaled about 2.56s.
  • Same cache, new pytest process: 38.20s total, with about 2.41s in functional calls.
  • Pre-commit checks passed for the four changed files. Dependency-group wiring and installed pinned versions were checked.
PYTHONPATH=. python -m pytest -q \
  tests/gpu/torch/kernels/quantization/linear_attention/test_fla_chunk_gated_delta_rule.py \
  tests/gpu_megatron/torch/quantization/plugins/test_megatron_gated_delta_net.py \
  --durations=0

These timings describe local test setup and execution, not inference performance. Native FP8 state QDQ and Hopper still require suitable GPU/CI runs. This minimal suite does not qualify tensor/context/pipeline parallelism, checkpoint resharding, or model-quality recovery. Mamba compilation coverage is tracked separately in #2572.

Before your PR is "Ready for review"

Contributor and security guidance reviewed. Commits are signed and signed off.

  • Is this change backward compatible?: ✅ Disabled-by-default quantizers, standard-recipe exclusions, and legacy-checkpoint coverage; enabled experimental configurations have explicit capability restrictions.
  • If you copied code from any other sources or added a new PIP dependency, did you follow guidance in CONTRIBUTING.md: ❌ Internal third-party approval tracking still needs confirmation. Upstream attribution, MIT/Apache headers, LICENSE notice, and license-hook exclusions are included. FLA/TileLang and TVM-FFI license files were reviewed.
  • Did you write any new necessary tests?: ✅ Numerical, gradient, conversion/checkpoint, and real framework tests.
  • Did you update Changelog?: ✅ Experimental quantization feature entry.
  • Did you get Claude approval on this PR?: ❌ Bot feedback addressed or discussed; renewed approval pending.

Additional Information

Related: #2455. This is the first integration slice and does not assume #2455 has merged. Later milestones will extend the numerical boundaries after choosing their approximation contracts.

Summary by CodeRabbit

  • New Features
    • Added experimental dynamic FP8 fake quantization for GatedDeltaNet recurrent states and WY activations during training.
    • Added PTQ configuration options for state and WY activation quantization. State quantization requires an SM89-or-newer GPU; the fused path requires fla-core==0.5.1 and a chunk size of 64.
  • Bug Fixes
    • Improved quantizer configuration validation and restoration for linear-attention models.

Signed-off-by: Kai Xu <kaix@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Sep 22, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@coderabbitai

coderabbitai Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: NVIDIA/Model-Optimizer/.coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 8ca47d70-f31a-47f9-8259-8476e024cd8b

📥 Commits

Reviewing files that changed from the base of the PR and between bfbb13e and 95766de.

📒 Files selected for processing (2)
  • tests/gpu/torch/kernels/quantization/linear_attention/test_fla_chunk_gated_delta_rule.py
  • tests/gpu_megatron/torch/quantization/plugins/test_megatron_gated_delta_net.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.


📝 Walkthrough

Walkthrough

This change adds experimental dynamic FP8 fake quantization for GatedDeltaNet recurrent states and WY activations. It adds chunked Triton kernels, quantizer and Megatron integration, PTQ configurations, and reference and GPU tests. The documented fused path uses fla-core==0.5.1 and chunk size 64.

Changes

GatedDeltaNet Quantization

Layer / File(s) Summary
Chunked state kernels
modelopt/torch/kernels/quantization/linear_attention/*, .pre-commit-config.yaml, docs/source/_templates/autosummary/module.rst, pyproject.toml, noxfile.py, LICENSE
Adds forward and backward Triton kernels with optional dynamic FP8 state QDQ, state-tile selection, and launch wrappers. Supporting changes update lint and autosummary exclusions, license-hook exclusions, dependencies for GPU test sessions, and copyright attribution.
Fused chunk GatedDeltaNet API
modelopt/torch/kernels/quantization/linear_attention/fla_chunk_gated_delta_rule.py, tests/gpu/torch/kernels/quantization/linear_attention/*, tests/_test_utils/torch/quantization/linear_attention_reference.py, tests/unit/torch/quantization/test_linear_attention_reference.py, CHANGELOG.rst, .github/workflows/gpu_tests.yml
Adds the chunked forward and backward API with optional WY quantization and state QDQ. Reference and GPU tests cover outputs, gradients, packed sequences, state layouts, and supported options. The changelog describes the experimental feature, and the GPU test timeout increases to 75 minutes.
Quantizer configuration and model lifecycle
modelopt/torch/quantization/linear_attention/*, modelopt/torch/quantization/conversion.py, modelopt/torch/quantization/model_quant.py, modelopt/torch/quantization/plugins/*, modelopt_recipes/configs/ptq/units/*, tests/unit/torch/quantization/plugins/test_gated_delta_net.py
Adds state and WY quantizers, validates their configurations during conversion and restoration, and adds PTQ recipes and tests for quantizer behavior and saved state.
Megatron GatedDeltaNet wiring
modelopt/torch/quantization/plugins/megatron.py, tests/gpu_megatron/torch/quantization/plugins/test_megatron_gated_delta_net.py, noxfile.py
Registers quantized Megatron GatedDeltaNet support when the upstream class is available. Tests cover distributed execution, checkpoint restore, and context-parallel rejection.

Priority: ⬇️ Low

Estimated code review effort: 4 (Complex) | ~60 minutes

Change: Feature

Merge Risk: ⚪ Minimal · up to 95766

The change reorganizes test setup so compilation is warmed before functional tests. No actionable merge-blocking risk was found.

🚥 Pre-merge checks | ✅ 5 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 22.11% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 95 functions across 18 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (5 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Security Anti-Patterns ✅ Passed PASS. The changed modelopt and examples Python files add no torch.load(..., weights_only=False), numpy.load(..., allow_pickle=True), hardcoded trust_remote_code=True, external-input eval()…
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly identifies the main change: adding the foundation for GatedDeltaNet state and W QAT. The [1/6] prefix is acceptable for a series of related pull requests.
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-10-02 20:34 UTC

@codecov

codecov Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 27.72397% with 597 lines in your changes missing coverage. Please review.
✅ Project coverage is 78.34%. Comparing base (f094f89) to head (f6a334e).
⚠️ Report is 1 commits behind head on main.

Files with missing lines Patch % Lines
...quantization/linear_attention/fla_chunk_delta_h.py 8.00% 540 Missing ⚠️
...ion/linear_attention/fla_chunk_gated_delta_rule.py 63.38% 52 Missing ⚠️
modelopt/torch/quantization/plugins/megatron.py 84.84% 5 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main    #2497      +/-   ##
==========================================
+ Coverage   69.49%   78.34%   +8.85%     
==========================================
  Files         614      618       +4     
  Lines       68631    69455     +824     
==========================================
+ Hits        47694    54414    +6720     
+ Misses      20937    15041    -5896     
Flag Coverage Δ
examples-diffusers 21.08% <3.51%> (-0.26%) ⬇️
examples-gpt-oss 13.44% <2.78%> (-0.17%) ⬇️
examples-hf_ptq 22.90% <3.26%> (-0.29%) ⬇️
examples-llm_distill 13.50% <2.78%> (-0.18%) ⬇️
examples-llm_eval 17.42% <3.26%> (-0.21%) ⬇️
examples-llm_qat 17.56% <3.51%> (-0.22%) ⬇️
examples-llm_sparsity 15.85% <2.78%> (-0.22%) ⬇️
examples-megatron_bridge 26.71% <7.99%> (+0.03%) ⬆️
examples-specdec_bench 13.21% <2.78%> (-0.17%) ⬇️
examples-speculative_decoding 17.76% <3.26%> (-0.25%) ⬇️
examples-torch_onnx 21.63% <3.26%> (-0.30%) ⬇️
examples-torch_trt 15.24% <3.26%> (-0.18%) ⬇️
examples-vllm_serve 13.68% <2.78%> (-0.18%) ⬇️
gpu 58.20% <26.39%> (+36.63%) ⬆️
regression 15.10% <2.78%> (-0.21%) ⬇️
unit 58.71% <7.26%> (-0.64%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@kaix-nv
kaix-nv added this pull request to stack #2510 September 22, 2026 21:27
@kaix-nv
kaix-nv removed this pull request from stack #2510 September 23, 2026 01:07
@kaix-nv
kaix-nv added this pull request to stack #2520 September 23, 2026 01:07
@kaix-nv
kaix-nv removed this pull request from stack #2520 September 23, 2026 01:08
@kaix-nv kaix-nv changed the title Add GDN state/W QAT emulation and linear-attention references [1/4] GDN state/W QAT foundation Sep 23, 2026
@kaix-nv
kaix-nv added this pull request to stack #2521 September 23, 2026 01:22
@kaix-nv kaix-nv changed the title [1/4] GDN state/W QAT foundation [1/5] GDN state/W QAT foundation Sep 24, 2026
@kaix-nv
kaix-nv removed this pull request from stack #2521 September 24, 2026 06:17
@kaix-nv
kaix-nv added this pull request to stack #2542 September 24, 2026 06:18
@kaix-nv
kaix-nv removed this pull request from stack #2542 September 24, 2026 06:31
@kaix-nv
kaix-nv added this pull request to stack #2543 September 24, 2026 06:31
@kaix-nv
kaix-nv force-pushed the kaix/linear-attention-qat-m1 branch 2 times, most recently from bff7ddf to 6686c9b Compare September 25, 2026 01:54
@kaix-nv
kaix-nv marked this pull request as ready for review September 25, 2026 02:05
@kaix-nv
kaix-nv requested review from a team as code owners September 25, 2026 02:05
@kaix-nv
kaix-nv requested a review from h-guo18 September 25, 2026 02:05
chunks.append((n, qc, kc, gc, decay, u))
all_w.append(w.transpose(0, 1))
w = torch.cat(all_w).reshape(q.shape)
if w_quantizer is not None:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why are we quantizing W?

@ifed-ucsd

Copy link
Copy Markdown

After we QAT a model with this PR, how do we evaluate it? Do you have a corresponding vLLM patch?

@kaix-nv
kaix-nv requested a review from ifed-ucsd September 28, 2026 23:40
@kaix-nv

kaix-nv commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

#2541 vLLM GDN/KDA state-only fake quantization

Yes, see this PR.

Warm standalone and Megatron forward/backward kernels before functional tests. Remove the 300-second overrides so test calls retain the default 120-second limit and report execution separately from compilation.

Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
from modelopt.torch.kernels.quantization.common.fp8_quant import fp8_scalar_qdq

# ``STATE_QDQ`` modes of the forward state kernel.
STATE_QDQ_OFF = 0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this how we are controlling the behavior of the QAT, through these global variables? Is there a way to make this more programatic?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These constants identify kernel mode. The QAT config is controlled per module through ModelOpt’s quant_cfg.


# ``STATE_QDQ`` modes of the forward state kernel.
STATE_QDQ_OFF = 0
STATE_QDQ_FP8_DYNAMIC = 1 # FP8 E4M3, one dynamic scale per program tile ([K, BV] of one head)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we also add support for integer state?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added dynamic INT8 state fake quantization by moving the basic implementation and tests from PR #2519 into this PR.

@kaix-nv
kaix-nv requested review from a team, kevalmorabia97 and shengliangxu September 30, 2026 17:10
Comment thread noxfile.py Outdated
Comment thread noxfile.py Outdated
Centralize the pinned FLA test dependencies in the dev-fla optional extra used by both GPU nox sessions. Replace the exhaustive compilation and numerical matrices with three shared-shape BF16 checks and one small Megatron QAT/checkpoint test; compile only the selected paths in setup fixtures.

Consolidates the dependency, extra-naming, and minimal-test follow-ups without adding Triton/TileLang CI cache plumbing. On RTX A6000, cold and warm focused runs each passed 3 tests with 1 SM89+ hardware skip; the warm run took 38.20s total and 2.41s in functional calls. Dependency wiring and pre-commit checks passed.

Signed-off-by: Kai Xu <kaix@nvidia.com>
@kaix-nv
kaix-nv force-pushed the kaix/linear-attention-qat-m1 branch from 7236936 to 7b43e6e Compare October 2, 2026 04:59
Comment thread .github/workflows/gpu_tests.yml Outdated
- example: gpu
timeout: 60
# Includes dependency builds and cold compilation of the FLA forward/backward tests.
timeout: 75

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@kaix-nv gpu tests now only take 40mins after your simplifications. Can we revert this change and leave it as 60mins now?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in d4895f8353: restored the general GPU job timeout to 60 minutes and removed the comment added with the increase. Workflow YAML validation and pre-commit checks passed.

Signed-off-by: Kai Xu <kaix@nvidia.com>
Preserve both GDN and upstream indexer entries in the changelog and PTQ unit table. Keep the reviewed GDN implementation, minimal GPU tests, and 60-minute GPU job timeout.

Validation: 36 focused GDN/reference unit tests passed; pre-commit and conflict checks passed.
Signed-off-by: Kai Xu <kaix@nvidia.com>
@kaix-nv
kaix-nv merged commit e4884ba into main Oct 2, 2026
55 checks passed
@kaix-nv
kaix-nv deleted the kaix/linear-attention-qat-m1 branch October 2, 2026 20:34
kevalmorabia97 added a commit that referenced this pull request Oct 2, 2026
_QuantGatedDeltaNet (#2497) gets the class-level get/set_extra_state overrides that torch needs to
route its quantizer state through _extra_state; GatedDeltaNet defines none, so without them the
state was dropped, as test_registered_megatron_quant_modules_checkpoint_quantizer_state flags.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
@kaix-nv
kaix-nv restored the kaix/linear-attention-qat-m1 branch October 5, 2026 06:10
@kaix-nv
kaix-nv deleted the kaix/linear-attention-qat-m1 branch October 5, 2026 06:11
@kaix-nv
kaix-nv restored the kaix/linear-attention-qat-m1 branch October 5, 2026 06:12
@kaix-nv
kaix-nv deleted the kaix/linear-attention-qat-m1 branch October 5, 2026 06:15
@kaix-nv kaix-nv changed the title [1/6] GDN state/W QAT foundation [1/7] GDN state/W QAT foundation Oct 5, 2026

@cjluo-nv cjluo-nv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bot review (bedrock-claude-opus-5-5) — DM the bot to share feedback.

Nudge: the code-level concerns from the last review are fixed, but the PR is still 2000 core-logic lines against a 500-line budget. It also vendors MIT-licensed FLA code, which needs a human licensing sign-off.

Needs action:

  • ✂️ Split this PR into stacked [x/N] PRs. Each PR must build and pass CI on its own and carry its own tests. Suggested order:

    • [1/4]: plugins/gated_delta_net.py, linear_attention/utils.py, the conversion.py/model_quant.py hooks, the recipe units, and the unit and reference tests.
    • [2/4]: fla_chunk_delta_h.py.
    • [3/4]: fla_chunk_gated_delta_rule.py with its GPU test.
    • [4/4]: the Megatron adapter in plugins/megatron.py.

    The kernel directory (~1723 lines) is still over budget on its own. Consider landing an unmodified upstream copy first, then the [ModelOpt] edits. Renumber the 7-PR stack to match.

  • Confirm third-party approval with a human. That covers the vendored fla-core kernels, the LICENSE copyright entry, the Apache-2.0 AND MIT SPDX headers and the new dev-fla pins (fla-core, tilelang, apache-tvm-ffi). The PR body says approval is still pending.

  • Confirm where INT8 state support lives. A reply says it was moved into this PR, but the diff only implements FP8, and test_validate_state_quantizer_rejects_unsupported rejects INT8.

  • Fix the PR body: the title says [1/7] and the URL is #2497, but the body says #2497 is already merged.

No action needed:

  • ✔️ Resolved since the last review:
    • Kernel detection now checks identity against FLA's function and unwraps functools.partial.
    • Leftover kwargs now raise TypeError.
    • The CI timeout, workflow and Mamba-timeout edits were reverted.
    • The dependencies moved to the dev-fla extra.
    • Kernel compilation moved into fixtures.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants