[llvm] [offload] Use pinned memory for KLE (PR #213767)
Robert Imschweiler via llvm-commits
llvm-commits at lists.llvm.org
Tue Aug 25 01:05:49 PDT 2026
ro-i wrote:
> ```
> // Otherwise, use two-step copy with an intermediate pinned host buffer.
> AMDGPUMemoryManagerTy &PinnedMemoryManager =
> HostDevice.getPinnedMemoryManager();
> if (auto Err = PinnedMemoryManager.allocate(Size, &PinnedPtr))
> return Err;
>
> if (auto Err = getStream(AsyncInfoWrapper, Stream))
> return Err;
>
> return Stream->pushMemoryCopyH2DAsync(TgtPtr, HstPtr, PinnedPtr, Size,
> PinnedMemoryManager);
> ```
>
> Didn't we already end up in this case if we just to dataSubmit?
Sorry, forgot to send my reply. We *do* want to use dataSubmit. But for dataSubmit to use the fast path, aka
https://github.com/llvm/llvm-project/blob/05f70782cc14770f0250de0468ab1f9196b4391a/offload/plugins-nextgen/amdgpu/src/rtl.cpp#L2849-L2855
we need to have the KLE already registered as pinned memory, which is what the dataAlloc in this patch does. Otherwise, we don't get the fast pushPinnedMemoryCopyAsync, but only the slower pushMemoryCopyH2DAsync.
https://github.com/llvm/llvm-project/pull/213767
More information about the llvm-commits
mailing list