Skip to content

cuda.core: make Device methods use their bound context - #2750

Open
Andy-Jost wants to merge 1 commit into
NVIDIA:mainfrom
Andy-Jost:ajost/issue-2311-device-context
Open

cuda.core: make Device methods use their bound context#2750
Andy-Jost wants to merge 1 commit into
NVIDIA:mainfrom
Andy-Jost:ajost/issue-2311-device-context

Conversation

@Andy-Jost

@Andy-Jost Andy-Jost commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Description

closes #2311

Device methods that create resources or synchronize (create_stream, create_event, create_opaque_array, create_mipmapped_array, create_texture_object, create_surface_object, sync) previously operated on whatever context was current on the calling thread, so dev1.create_stream() could silently create a stream on device 0 while labeling it device_id == 1.

These methods now run against the Device's bound context and restore the caller's current context afterward, including when no context is current. Details:

  • Stream, event, array, mipmapped-array, texture, and surface creation take the target context explicitly; the C++ handle layer switches context around the driver call and undoes the creation if the caller's context cannot be restored.
  • Device.sync() synchronizes the bound context. Device.set_current() returns the previously current context.
  • Texture and surface creation reject resources that belong to a different device or context.
  • Context-sensitive cleanup (default-stream frees, texture/surface destroys) runs in the owning context and warns instead of raising; handle-based destroys are left alone since they resolve their own context.
  • _SynchronousMemoryResource binds its context at construction.
  • Release notes for 1.2.0 and the interoperability docs describe the new behavior.

Checklist

  • New or existing tests cover these changes.
  • The documentation is up to date with these changes.

🤖 Generated with Claude Code

@copy-pr-bot

copy-pr-bot Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@github-actions github-actions Bot added the cuda.core Everything related to the cuda.core module label Sep 1, 2026
@Andy-Jost

Copy link
Copy Markdown
Contributor Author

/ok to test

@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown

@Andy-Jost
Andy-Jost force-pushed the ajost/issue-2311-device-context branch 3 times, most recently from a09bebe to 838f850 Compare September 1, 2026 22:12
Run context-sensitive Device operations against the Device's bound context while preserving caller state. Centralize context-aware cleanup and synchronous allocation handling so resource lifetimes remain correct.
@Andy-Jost
Andy-Jost force-pushed the ajost/issue-2311-device-context branch from 838f850 to 83a3d46 Compare September 1, 2026 22:26
@Andy-Jost

Copy link
Copy Markdown
Contributor Author

/ok to test

Comment on lines +282 to +311
// Run a creation operation and undo it if context restoration fails.
// Context-independent undo always runs. Context-sensitive undo runs only
// after verifying that the target context remains current; otherwise the
// resource leaks rather than risking cleanup in the wrong context.
template <typename Fn, typename Undo>
CUresult invoke_in_context_or_undo(const ContextHandle& h_context, Fn&& operation,
Undo&& undo, bool undo_requires_target_context) noexcept {
ASSERT_NOTHROW_INVOCABLE(Fn&&);
ASSERT_NOTHROW_INVOCABLE(Undo&&);
CUcontext previous = nullptr;
int changed = 0;
CUresult status = enter_context(h_context, &previous, &changed);
if (status != CUDA_SUCCESS) {
return status;
}
status = std::invoke(std::forward<Fn>(operation));
CUresult composite = exit_context(previous, changed, status);
if (status == CUDA_SUCCESS && composite != CUDA_SUCCESS) {
bool undo_ok = true;
if (undo_requires_target_context) {
CUcontext current = nullptr;
undo_ok = p_cuCtxGetCurrent(&current) == CUDA_SUCCESS
&& current == as_cu(h_context);
}
if (undo_ok) {
std::invoke(std::forward<Undo>(undo));
}
status_ = p_cuCtxSetCurrent(target);
changed_ = status_ == CUDA_SUCCESS;
}
return composite;
}

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I played a lot of defense here, and it is easy to see how complex this gets as errors cascade. I partly question the value of this kind of code. Would it be reasonable to call std::abort instead when a key invariant cannot be maintained? If context restoration fails, then whether or not the undo succeeds, the wrong context may remain current, and the program is in big trouble either way.

event_registry.unregister_handle(b->resource);
GILReleaseGuard gil;
p_cuEventDestroy(b->resource);
pw_cuEventDestroy(b->resource);

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Replaced deallocation functions (p_*) in destructors with wrapped versions that issue warnings on failure (pw_*).

@Andy-Jost Andy-Jost self-assigned this Sep 2, 2026
@Andy-Jost Andy-Jost added the bug Something isn't working label Sep 2, 2026
@Andy-Jost Andy-Jost added this to the cuda.core next milestone Sep 2, 2026
@Andy-Jost
Andy-Jost marked this pull request as ready for review September 2, 2026 00:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working cuda.core Everything related to the cuda.core module

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG]: Device methods operate on the current context, not the device they're called on

1 participant