Skip to content

fix(cuda-macros,codegen): keep all-generic library .oxart via #72 anchor - #1372

Open
dundysm wants to merge 1 commit into
NVIDIA:mainfrom
dundysm:fix/1365-all-generic-cuda-module-anchor
Open

dundysm wants to merge 1 commit into
NVIDIA:mainfrom
dundysm:fix/1365-all-generic-cuda-module-anchor

Conversation

@dundysm

@dundysm dundysm commented Oct 1, 2026 •

Copy link
Copy Markdown

Summary

Fixes #1365.

Depends on #1446. The backend stub calls oxide_artifacts::build_host_strong_anchor_object_for_target(_with_legacy_anchor), which #1446 adds in oxide-artifacts 0.2.2. The backend consumes oxide-artifacts from crates.io, so this PR builds once 0.2.2 is published and #1446 moves the requirement and relocks (the bump, publish, relock order from check-oxide-artifacts-parity.sh). After #1446 merges I will rebase this onto it and relock the new example with scripts/sync-example-locks.sh so it resolves 0.2.2 too.

An all-generic #[cuda_module] in a library that monomorphizes locally still embeds .oxart with an artifact anchor, but the macro skipped the #72 rlib keep-alive under the assumption that generics never produce an artifact in the defining crate. Nothing referenced the anchor, so the linker dropped the archive member and load_all_ptx_bundles_merged returned NoModules.

Adding one unused concrete kernel "fixed" it only because concrete kernels restored the keep-alive. This PR restores the handshake for generics without reverting #222 merge loading.

Root cause

Shape Library .oxart? Macro anchor ref? Result
Concrete in lib (#72) Yes Yes Works
Generics mono only in binary (#222) Usually no No (was skipped) Works via merge
Generics mono in library, thin binary (#1365) Yes No (bug) NoModules

Approach

Preserves the intentional design layers:

  1. ModuleNotFound after I move a #[cuda_module] into a separate crate #72 archive handshake: keep referencing the package (or v2) artifact anchor from load_named.
  2. Panic DriverError(500, "named symbol not found") when calling generic kernels from a separate crate #222 merge load: generic modules still call load_all_ptx_bundles_merged (no named-only regression).
  3. No false undefined symbols: when the defining crate has merge markers but kernel_count == 0, the backend emits a strong anchor-only .oxlink stub (no fake empty PTX payload).

Macro (cuda_module_artifact_anchor_references)

  • Remove the early return that skipped when every kernel is generic.
  • Emit the same cfg-guarded black_box anchor refs for generic kernels via effective_cfg_attrs.
  • Keep owner-filter / missing CARGO_PKG_* skips.

Backend (rustc-codegen-cuda)

Tests / example

Orthogonal to #1367 / #1166 (finalizer expected-kernel inventory).

Test plan

  • cargo test -p cuda-macros (118 unit tests, including the two generic anchor tests)
  • cargo clippy -p cuda-macros --all-targets -- -D warnings
  • cargo fmt --check for the SIMT workspace, the backend, and the new example
  • rustc-codegen-cuda: cargo check, cargo clippy --all-targets -- -D warnings, and cargo test against the feat(oxide-artifacts): add strong anchor-only host objects, bump to 0.2.2 #1446 crate via a local [patch.crates-io]
  • scripts/sync-example-locks.sh --check and every cuda-oxide/scripts/check-*.sh, including the cargo-deny example license policy
  • cargo oxide run generic_mono_in_lib and --verify-bundles
  • cargo oxide run cuda_module_in_lib and cross_crate_embedded

Reviewers

cc @nihalpasham @xavierforge (prior #72 / from-anchor / embedding owners; CODEOWNERS currently declares no path owners in-tree)

@copy-pr-bot

copy-pr-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

… anchor

All-generic #[cuda_module] skipped the rlib artifact-anchor keep-alive
because the macro assumed "no artifact here". When the library itself
monomorphizes, codegen still embeds .oxart with an anchor; nothing
referenced it, so the linker dropped the archive member and
load_all_ptx_bundles_merged returned NoModules (issue NVIDIA#1365).

Emit the same cfg-guarded black_box anchor refs for generic kernels.
When owner-selected with PTX-merge markers but kernel_count == 0, emit
a strong anchor-only .oxlink stub so the NVIDIA#222 (mono only in the binary)
shape still links. Keep merge loading; do not invent empty PTX.

The anchor reference tokens are built by a small helper that takes the
resolved symbol, so the generic keep-alive is unit tested without
cargo's per-crate environment. Adds the generic_mono_in_lib example
under cuda-oxide/crates/rustc-codegen-cuda/examples.

The stub uses build_host_strong_anchor_object_for_target(_with_legacy_anchor)
from oxide-artifacts 0.2.2 (PR 1446), so this builds once that release
is published and the requirement is relocked.

Signed-off-by: Dundy Pasupuleti <dundysm@gmail.com>
@dundysm
dundysm force-pushed the fix/1365-all-generic-cuda-module-anchor branch from fd345bc to 2da9429 Compare October 8, 2026 03:04
@dundysm
dundysm requested a review from nihalpasham as a code owner October 8, 2026 03:04
@dundysm

dundysm commented Oct 8, 2026

Copy link
Copy Markdown
Author

split the oxide-artifacts part out into #1446 (new strong anchor-only writers, bump to 0.2.2). the backend reads oxide-artifacts from crates.io, so this one now builds once 0.2.2 is published and relocked; i'll rebase it on top after that lands. also squashed here: the example moved under cuda-oxide/crates/rustc-codegen-cuda/examples, rustfmt, and the generic anchor tests now go through a small token builder so they pass under plain cargo test.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

all-generic #[cuda_module] in a library gives "load module: NoModules"

1 participant