[all-commits] [llvm/llvm-project] 4a74a4: [flang][cuda] Flatten memref descriptors in GPU ke...

Zhen Wang via All-commits all-commits at lists.llvm.org
Thu Apr 23 10:06:02 PDT 2026


  Branch: refs/heads/main
  Home:   https://github.com/llvm/llvm-project
  Commit: 4a74a4346c34a39836ebfe1814498043af8be4ca
      https://github.com/llvm/llvm-project/commit/4a74a4346c34a39836ebfe1814498043af8be4ca
  Author: Zhen Wang <zhenw at nvidia.com>
  Date:   2026-04-23 (Thu, 23 Apr 2026)

  Changed paths:
    M flang/lib/Optimizer/Transforms/CUDA/CUFGPUToLLVMConversion.cpp
    M flang/test/Fir/CUDA/cuda-gpu-launch-func.mlir

  Log Message:
  -----------
  [flang][cuda] Flatten memref descriptors in GPU kernel argument packing (#193651)

createKernelArgArray in CUFGPUToLLVMConversion packed one pointer per
operand, treating a memref argument as a single value. But gpu-to-nvvm
expands each memref into 3 + 2*rank scalar parameters, so the host-side
kernelParams did not match the device kernel signature and
cudaLaunchKernel failed with cudaErrorInvalidValue.

Host-side packing now unpacks each memref descriptor into its scalar
fields (allocatedPtr, alignedPtr, offset, sizes..., strides...),
matching the NVVM ABI. Non-memref operands pass through unchanged.



To unsubscribe from these emails, change your notification settings at https://github.com/llvm/llvm-project/settings/notifications


More information about the All-commits mailing list