Repository navigation
Conversation
…ry query cuMemGetInfo had no public wrapper: the shim's mem_get_info sits in a pub(crate) module, so callers reached for the raw binding. Add the method on CudaContext next to the other device queries; it binds the context first. The leak test used the raw call and now uses the method. Signed-off-by: midagedev <midagedev@gmail.com>
midagedev
requested review from
elibol,
kkraus14,
nihalpasham and
rparolin
as code owners
October 8, 2026 02:38
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Recreated from NVlabs/cutile-rs#309 after the move. Closes #1438.
CudaContext::mem_info()returns(free, total)device memory bytes. There was no public way to ask this. The shim hasmem_get_info, but it is in apub(crate)module, so users callcuda_bindings::cuMemGetInfo_v2directly. Our engine does this to check VRAM before loading a model.Changes
CudaContextnext tocompute_capability. It binds the context first, like the other device queries.vram_returns_to_baseline_after_buffer_cyclesnow uses the method instead of the raw call.totalequalscuDeviceTotalMemandfree <= total.[Unreleased]incutile-rs/CHANGELOG.md.I replayed the commit with the "Transplanting a cutile-rs pull request" steps in
AGENTS.md. Git put the CHANGELOG line in the 0.4.0 section, so I moved it to[Unreleased].Testing
On this branch, CUDA 13:
cargo test -p cuda-core --test simt_device_buffer_leakson an RTX 3060 (sm_86): 4 passed.cargo fmt --checkandcargo clippy -p cuda-core --all-targets -- -D warnings: clean.just -f cuda-oxide/Justfile check. The change is additive and touches onlycuda-core.Checklist
git commit -s)I used an AI assistant to draft the patch and this text.