Add automatic workspace management feature. - #159
Merged
Conversation
Signed-off-by: romerojosh <joshr@nvidia.com>
Collaborator
Author
|
/build |
|
🚀 Build workflow triggered! View run |
|
❌ Build workflow failed! View run |
Collaborator
Author
|
/build |
|
🚀 Build workflow triggered! View run |
|
✅ Build workflow passed! View run |
Signed-off-by: romerojosh <joshr@nvidia.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR adds automatically managed workspaces to cuDecomp to simplify application usage. Instead of querying workspace requirements and managing caller-owned allocations, users can opt in to automatic workspace management by passing CUDECOMP_WORKSPACE_AUTO (or a null pointer in C/C++) as the workspace argument to the transpose and halo communication APIs.
Automatically managed workspaces are owned by the cuDecomp handle, reused across its grid descriptors, and grown as needed. Operations using automatic workspace management on the same handle are serialized on a library-owned CUDA stream, while preserving ordering with the caller-provided stream. Users must make the same automatic-versus-explicit workspace choice on all participating ranks.
Automatically managed workspace memory is not exposed for reuse or aliasing by other components, such as FFT libraries. Automatic workspace management is also unavailable during external CUDA Graph capture. Applications requiring concurrent operations, workspace aliasing, or external graph capture should continue using caller-managed workspaces.