Skip to content

Add grid-stride + ROCm cap to split optimizer update kernel - #6086

Open
q10 wants to merge 1 commit into
pytorch:mainfrom
q10:export-D113351687
Open

Add grid-stride + ROCm cap to split optimizer update kernel#6086
q10 wants to merge 1 commit into
pytorch:mainfrom
q10:export-D113351687

Conversation

@q10

@q10 q10 commented Jul 28, 2026

Copy link
Copy Markdown
Contributor

Summary:
The standalone split optimizer path launches split_{optimizer}_update_kernel
with grid = div_round_up(grad_dev_indices.numel(), kMaxThreads / kThreadGroupSize)
and block = dim3(kThreadGroupSize, kMaxThreads / kThreadGroupSize, 1), so total
threads ~= num_unique_indices * kThreadGroupSize exceeds the HIP 2^32
threads-per-launch limit on ROCm for large index counts.

Cap the launch with utils::cuda::cap_grid_dim_x(..., OverflowOnly) and add a
ROCm grid-stride loop over run_id to the kernel (single guard converted to the
loop bound; no internal early-returns). No-op on CUDA.

Reviewed By: henrylhtsang

Differential Revision: D113351687

Summary:
The standalone split optimizer path launches split_{optimizer}_update_kernel
with grid = div_round_up(grad_dev_indices.numel(), kMaxThreads / kThreadGroupSize)
and block = dim3(kThreadGroupSize, kMaxThreads / kThreadGroupSize, 1), so total
threads ~= num_unique_indices * kThreadGroupSize exceeds the HIP 2^32
threads-per-launch limit on ROCm for large index counts.

Cap the launch with utils::cuda::cap_grid_dim_x(..., OverflowOnly) and add a
ROCm grid-stride loop over run_id to the kernel (single guard converted to the
loop bound; no internal early-returns). No-op on CUDA.

Reviewed By: henrylhtsang

Differential Revision: D113351687
@meta-codesync

meta-codesync Bot commented Jul 28, 2026

Copy link
Copy Markdown
Contributor

@q10 has exported this pull request. If you are a Meta employee, you can view the originating Diff in D113351687.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant