[llvm] [offload] Use pinned memory for KLE (PR #213767)

Robert Imschweiler via llvm-commits llvm-commits at lists.llvm.org
Tue Aug 25 01:05:49 PDT 2026


ro-i wrote:

> ```
>      // Otherwise, use two-step copy with an intermediate pinned host buffer.
>     AMDGPUMemoryManagerTy &PinnedMemoryManager =
>         HostDevice.getPinnedMemoryManager();
>     if (auto Err = PinnedMemoryManager.allocate(Size, &PinnedPtr))
>       return Err;
> 
>     if (auto Err = getStream(AsyncInfoWrapper, Stream))
>       return Err;
> 
>     return Stream->pushMemoryCopyH2DAsync(TgtPtr, HstPtr, PinnedPtr, Size,
>                                           PinnedMemoryManager);
> ```
> 
> Didn't we already end up in this case if we just to dataSubmit?

Sorry, forgot to send my reply. We *do* want to use dataSubmit. But for dataSubmit to use the fast path, aka

https://github.com/llvm/llvm-project/blob/05f70782cc14770f0250de0468ab1f9196b4391a/offload/plugins-nextgen/amdgpu/src/rtl.cpp#L2849-L2855

we need to have the KLE already registered as pinned memory, which is what the dataAlloc in this patch does. Otherwise, we don't get the fast pushPinnedMemoryCopyAsync, but only the slower pushMemoryCopyH2DAsync.

https://github.com/llvm/llvm-project/pull/213767


More information about the llvm-commits mailing list