[all-commits] [llvm/llvm-project] 4a74a4: [flang][cuda] Flatten memref descriptors in GPU ke...
Zhen Wang via All-commits
all-commits at lists.llvm.org
Thu Apr 23 10:06:02 PDT 2026
Branch: refs/heads/main
Home: https://github.com/llvm/llvm-project
Commit: 4a74a4346c34a39836ebfe1814498043af8be4ca
https://github.com/llvm/llvm-project/commit/4a74a4346c34a39836ebfe1814498043af8be4ca
Author: Zhen Wang <zhenw at nvidia.com>
Date: 2026-04-23 (Thu, 23 Apr 2026)
Changed paths:
M flang/lib/Optimizer/Transforms/CUDA/CUFGPUToLLVMConversion.cpp
M flang/test/Fir/CUDA/cuda-gpu-launch-func.mlir
Log Message:
-----------
[flang][cuda] Flatten memref descriptors in GPU kernel argument packing (#193651)
createKernelArgArray in CUFGPUToLLVMConversion packed one pointer per
operand, treating a memref argument as a single value. But gpu-to-nvvm
expands each memref into 3 + 2*rank scalar parameters, so the host-side
kernelParams did not match the device kernel signature and
cudaLaunchKernel failed with cudaErrorInvalidValue.
Host-side packing now unpacks each memref descriptor into its scalar
fields (allocatedPtr, alignedPtr, offset, sizes..., strides...),
matching the NVVM ABI. Non-memref operands pass through unchanged.
To unsubscribe from these emails, change your notification settings at https://github.com/llvm/llvm-project/settings/notifications
More information about the All-commits
mailing list