Skip to content

fix(mir-importer): normalize pointer address spaces in tuple construction - #1454

Open
uurl wants to merge 1 commit into
NVIDIA:mainfrom
uurl:fix/as7-tuple-normalization
Open

uurl wants to merge 1 commit into
NVIDIA:mainfrom
uurl:fix/as7-tuple-normalization

Conversation

@uurl

@uurl uurl commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Summary

Fix MIR tuple construction when a pointer operand uses a CUDA address space that differs from the address space expected by the tuple's declared element type.

Previously, constructing a tuple containing a cluster-shared (AS7) pointer could fail MIR verification because the tuple expected a generic (AS0) pointer.

Fixes #1453.

Root Cause

The MIR importer already normalizes pointer representations when constructing structs, enums, and arrays. However, AggregateKind::Tuple was missing the corresponding normalization step.

As a result, MirConstructTupleOp could receive operands whose address spaces did not match the declared tuple element types.

The failure was reproduced with:

MirConstructTupleOp operand 0 type mismatch.

Expected:
mir.ptr <builtin.integer ui32,mutable:false,addrspace:0,kind:RawConst>

Actual:
mir.ptr <builtin.integer ui32,mutable:false,addrspace:7,kind:RawConst>

Changes

  • Normalize tuple operands against their declared element types before constructing MirConstructTupleOp.
  • Reuse the existing cast_to_expected_pointer_type_if_needed helper.
  • Add a unit test verifying AS7-to-AS0 normalization while preserving MirPointerKind::RawConst.
  • Extend the cluster example with a tuple containing a cluster-shared pointer passed through a device function.
  • Extend verify-code-shape.sh to verify the tuple round-trip call in LLVM IR.

The implementation follows the existing pointer normalization approach used for other MIR aggregates.

Regression Coverage

The reproduction uses cluster::map_shared_rank to obtain an AS7 pointer and constructs a tuple containing that pointer:

#[inline(never)]
#[device]
fn bounce_cluster_pair(
    pair: (*const u32, u64),
) -> (*const u32, u64) {
    pair
}

// Inside the existing cluster kernel:
let pair = bounce_cluster_pair((neighbor_ptr, neighbor_rank as u64));
let bounced = unsafe { *pair.0 };

Before the fix, this triggers a MIR verification failure.

After the fix, the importer inserts the required pointer cast before tuple construction, and compilation succeeds.

The regression verifies the device-function call in LLVM IR. Later PTX optimization may eliminate the call, so the test does not require it to remain in the final PTX.

Validation

All checks passed:

  • cargo test -p mir-importer: 1164 passed, 0 failed.
  • cargo clippy -p mir-importer --all-targets -- -D warnings: PASS.
  • cargo fmt --all -- --check: PASS.
  • git diff --check: PASS.
  • cluster compile-only smoketest (sm_90): PASS.
  • cluster/verify-code-shape.sh: PASS.
  • cargo oxide run cluster: PASS on an NVIDIA GeForce RTX 5060 Laptop GPU.
  • CUDA_OXIDE_NO_OPT=1 cargo oxide run cluster: PASS.

The DSMEM ring exchange test produced the expected values for all four cluster blocks.

The final compile-only smoketest also passed after removing the experimental instrumentation used during the investigation.

Scope

This PR addresses tuple operand normalization in the MIR importer.

It does not modify mir-lower or the aggregate-memory cast fallback investigated in #883.

@uurl
uurl requested a review from nihalpasham as a code owner October 8, 2026 08:52
@copy-pr-bot

copy-pr-bot Bot commented Oct 8, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

Signed-off-by: Raul Estrada <raulestradaa@gmail.com>
@uurl
uurl force-pushed the fix/as7-tuple-normalization branch from ff2b278 to 95133c9 Compare October 8, 2026 08:55

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG]: MIR tuple construction fails for cluster-shared AS7 pointer operands

1 participant