Avoid aggregate graph access descriptor initialization - #2688
Merged
Conversation
Contributor
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
Contributor
Author
|
/ok to test |
This comment has been minimized.
This comment has been minimized.
rwgk
marked this pull request as ready for review
August 23, 2026 22:19
juenglin
approved these changes
Aug 24, 2026
Contributor
Author
Additional rationale for merging this PRThe need for this change appears to arise from a Cython code-generation issue exposed by cybind's representation of CUDA's anonymous nested struct:
The field-by-field initialization in this PR is therefore a narrowly scoped workaround that preserves the intended descriptor contents and behavior. |
2 tasks
This comment has been minimized.
This comment has been minimized.
1 similar comment
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Follow-up to #2683.
PR #2683 centralizes
CUmemLocationconstruction, but the graph allocation path still embeds the returned location in aCUmemAccessDescaggregate constructor. With the CUDA 13.4 generated declaration, that nested aggregate causes Cython to emit an invalid conversion for the addedlocalizedunion arm.Declare the access descriptor separately and assign its
locationandflagsfields before adding it to the descriptor vector. This keeps the helper introduced by #2683 and preserves the existing runtime behavior.Please see ctk-next PR 539 for background on how this escaped attention while validating #2683.
Validation
pre-commit run --all-filesChecklist