[flang-commits] [flang] [flang][cuda] Preserve array lower bounds in implicit device-to-host transfer (PR #227824)

Siddhanth gupta via flang-commits flang-commits at lists.llvm.org
Thu Oct 1 00:04:41 PDT 2026


Siddhanthguptaa wrote:

Thanks for the review @clementval.

I've updated the implementation based on both pieces of feedback:

**1. "This is probably better in BoxValue as a helper"**

Added ir::updateRuntimeLBounds(exv, lbounds) to BoxValue.h/BoxValue.cpp, modeled directly on the existing ir::substBase pattern. It handles all four ExtendedValue array representations:
- ir::ArrayBoxValue - replaces lbounds, keeps addr/extents/sourceBox
- ir::CharArrayBoxValue - replaces lbounds, keeps addr/len/extents
- ir::BoxValue - replaces lbounds, keeps addr/explicitParams/explicitExtents
- ir::MutableBoxValue - converts to BoxValue with explicit lbounds (see below)
- scalar/unboxed passthrough

**2. "You are also missing MutableBoxValue"**

Device allocatable arrays (integer, device, allocatable :: a(:)) lower-bound through ir::MutableBoxValue, whose lower bounds live in the runtime descriptor rather than in the static representation. Fixed by:

- Promoting the file-static getNonDefaultLowerBounds() in HLFIRTools.cpp to a public hlfir::getNonDefaultLowerBounds(). This function correctly handles the mutable-box case by dereferencing the descriptor and emitting ir.box_dims to read bounds at runtime.
- In genCUDAImplicitDataTransfer, calling hlfir::getNonDefaultLowerBounds(entity) *before* createTempFromMold(), then passing the result to ir::updateRuntimeLBounds(). This replaces the previous 80-line multi-arm match block.

A regression test case for the device allocatable path was added as Test 6 in cuda-implicit-transfer-lbounds.cuf, verifying that ir.shape_shift with the runtime lower bound appears on the rebound declaration and no data_attr = #cuf.cuda<device> is present.

https://github.com/llvm/llvm-project/pull/227824


More information about the flang-commits mailing list