[flang-commits] [flang] [flang][cuda] Preserve array lower bounds in implicit device-to-host transfer (PR #227824)
Siddhanth gupta via flang-commits
flang-commits at lists.llvm.org
Thu Oct 1 00:04:41 PDT 2026
Siddhanthguptaa wrote:
Thanks for the review @clementval.
I've updated the implementation based on both pieces of feedback:
**1. "This is probably better in BoxValue as a helper"**
Added ir::updateRuntimeLBounds(exv, lbounds) to BoxValue.h/BoxValue.cpp, modeled directly on the existing ir::substBase pattern. It handles all four ExtendedValue array representations:
- ir::ArrayBoxValue - replaces lbounds, keeps addr/extents/sourceBox
- ir::CharArrayBoxValue - replaces lbounds, keeps addr/len/extents
- ir::BoxValue - replaces lbounds, keeps addr/explicitParams/explicitExtents
- ir::MutableBoxValue - converts to BoxValue with explicit lbounds (see below)
- scalar/unboxed passthrough
**2. "You are also missing MutableBoxValue"**
Device allocatable arrays (integer, device, allocatable :: a(:)) lower-bound through ir::MutableBoxValue, whose lower bounds live in the runtime descriptor rather than in the static representation. Fixed by:
- Promoting the file-static getNonDefaultLowerBounds() in HLFIRTools.cpp to a public hlfir::getNonDefaultLowerBounds(). This function correctly handles the mutable-box case by dereferencing the descriptor and emitting ir.box_dims to read bounds at runtime.
- In genCUDAImplicitDataTransfer, calling hlfir::getNonDefaultLowerBounds(entity) *before* createTempFromMold(), then passing the result to ir::updateRuntimeLBounds(). This replaces the previous 80-line multi-arm match block.
A regression test case for the device allocatable path was added as Test 6 in cuda-implicit-transfer-lbounds.cuf, verifying that ir.shape_shift with the runtime lower bound appears on the rebound declaration and no data_attr = #cuf.cuda<device> is present.
https://github.com/llvm/llvm-project/pull/227824
More information about the flang-commits
mailing list