Skip to content

all-generic #[cuda_module] in a library gives "load module: NoModules" #1365

Description

@martinjrobins

Description

If a library's #[cuda_module] contains only generic kernels, its PTX bundle is not linked into
downstream binaries. This happens even when the library itself monomorphizes the kernels. Adding
one non-generic kernel to the module, never launched, fixes it.

Minimal reproducer

Two crates A and B, crate A defines a cuda_module like so:

use cuda_core::simt::LaunchConfig;
use cuda_core::{CudaContext, DeviceBuffer};
use cuda_device::{cuda_module, kernel, thread, DisjointSlice};

#[cuda_module]
pub mod kernels {
    use super::*;

    #[kernel]
    pub fn fill<T: Copy>(mut out: DisjointSlice<T>, value: T) {
        if let Some(x) = out.get_mut(thread::index_1d()) {
            *x = value;
        }
    }
}

/// Non-generic entry point: `fill::<f64>` is monomorphized in this crate.
pub fn fill_f64(n: usize, value: f64) -> Vec<f64> {
    let ctx = CudaContext::new(0).expect("no CUDA device");
    let stream = ctx.default_stream();
    let module = unsafe { kernels::load(&ctx) }.expect("load module");
    let mut out = DeviceBuffer::<f64>::zeroed(&stream, n).expect("alloc");
    unsafe {
        module.fill::<f64>(&stream, LaunchConfig::for_num_elems(n as u32), &mut out, value)
    }
    .expect("launch fill::<f64>");
    out.to_host_vec(&stream).expect("copy back")
}

crate B then calls fill_f64 using A as a dependency:

fn main() {
    let out = kernels::fill_f64(4, 1.5);
    assert_eq!(out, [1.5; 4]);
    println!("ok");
}

Expected behavior

should print "ok".

Actual behavior

thread 'main' (1604530) panicked at kernels/src/lib.rs:27:49:
load module: NoModules
stack backtrace:
   0: __rustc::rust_begin_unwind
             at /rustc/e457a7b0d326d67b4322ef0d11bd715cfaeda48f/library/std/src/panicking.rs:676:5
   1: core::panicking::panic_fmt
             at /rustc/e457a7b0d326d67b4322ef0d11bd715cfaeda48f/library/core/src/panicking.rs:80:14
   2: core::result::unwrap_failed
             at /rustc/e457a7b0d326d67b4322ef0d11bd715cfaeda48f/library/core/src/result.rs:1870:5
   3: <core::result::Result<kernels::kernels::LoadedModule, cuda_host::embedded::EmbeddedModuleError>>::expect
             at /home/mrobins/.rustup/toolchains/nightly-2026-08-28-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/core/src/result.rs:1184:23
   4: kernels::fill_f64
             at ./kernels/src/lib.rs:27:49
   5: app::main
             at ./app/src/main.rs:2:15
   6: <fn() as core::ops::function::FnOnce<()>>::call_once
             at /home/mrobins/.rustup/toolchains/nightly-2026-08-28-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/core/src/ops/function.rs:250:5

Environment

cargo oxide doctor
❯ cargo-oxide doctor
cargo-oxide environment check
==============================

Rust nightly toolchain... ✓ rustc 1.100.0-nightly (e457a7b0d 2026-08-27)
rust-toolchain.toml... ✓ channel nightly-2026-08-28
Pinned toolchain active... ✓ nightly-2026-08-28-x86_64-unknown-linux-gnu (overridden by '/home/mrobins/git/tmp/cuda-oxide-generic-anchor-repro/rust-toolchain.toml')
Required rustup components... ✓ rust-src, rustc-dev, rust-analyzer, clippy, rustfmt, llvm-tools
Codegen backend... ✓ /home/mrobins/.cargo/cuda-oxide/librustc_codegen_cuda.so
Project config (.cargo/cuda-oxide.toml)... - not present (using defaults)
Shared cache (external projects)... ✓ /home/mrobins/.cargo/cuda-oxide/librustc_codegen_cuda.so
Backend source (cuda-oxide commit)... ✓ cache built from ec4aa47979, matching this project's dependency
CUDA headers (cuda.h)... ✓ /usr/local/cuda/include/cuda.h
CUDA toolkit (nvcc)... ✓ Cuda compilation tools, release 12.4, V12.4.131 (/usr/bin/nvcc)
libNVVM (libnvvm.so)... ✓ libNVVM 2.0
nvJitLink (libnvJitLink.so)... ✓ nvJitLink 13.3
libdevice (libdevice.10.bc)... ✓ /usr/local/cuda/nvvm/libdevice/libdevice.10.bc
llc (LLVM)... ✓ LLVM version 23.1.0-rust-1.100.0-nightly (/home/mrobins/.rustup/toolchains/nightly-2026-08-28-x86_64-unknown-linux-gnu/lib/rustlib/x86_64-unknown-linux-gnu/bin/llc)
clang / libclang resource dir... ✓ /usr/lib/llvm-21/lib/clang/21
NVIDIA driver / GPU... ✓ NVIDIA A40 (compute capability 8.6, driver 580.173.02)
cuda-gdb (optional)... ✓ NVIDIA (R) cuda-gdb 12.4 (/usr/bin/cuda-gdb)
compute-sanitizer (optional)... ✓ Version 2024.1.1.0 (build 34043715) (public-release) (/usr/bin/compute-sanitizer)

✅ Environment looks good!

Activity

  1. added
    bugSomething isn't working
    TBDNot yet triaged; scope or disposition undecided
    on Sep 29, 2026
  2. dundysm commented on Oct 1, 2026

    @dundysm

    taking this (@dundysm).

    root cause: all-generic #[cuda_module] skips the #72 rlib artifact-anchor keep-alive, but when the library itself monomorphizes, codegen still embeds .oxart with an anchor; nothing references it, so the linker drops the archive member and load_all_ptx_bundles_merged returns NoModules.

    fix plan: emit the same cfg-guarded anchor keep-alive for generic kernels, and when owner-selected with merge markers but kernel_count == 0, emit an anchor-only host stub so the #222 (mono only in the binary) shape still links. keep merge loading; do not depend on #1367.

  3. dundysm commented on Oct 1, 2026

    @dundysm

    opened #1372 with the #72 keep-alive extension for all-generic modules (and a strong anchor-only stub when there is no local mono). details and test plan are in the pr body.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    TBDNot yet triaged; scope or disposition undecidedbugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions