[flang-commits] [flang] [flang] Fix TRANSFER into derived type with tail padding zeroing pad bytes (PR #223814)
via flang-commits
flang-commits at lists.llvm.org
Wed Sep 16 20:37:05 PDT 2026
================
@@ -1552,15 +1472,24 @@ void fir::factory::genRecordAssignment(fir::FirOpBuilder &builder,
return;
}
- // Otherwise, the derived type has compile time constant size and for which
- // the component by component assignment can be replaced by a memory copy.
- // Since we do not know the size of the derived type in lowering, do a
- // component by component assignment. Note that a single fir.load/fir.store
- // could be used on "small" record types, but as the type size grows, this
- // leads to issues in LLVM (long compile times, long IR files, and even
- // asserts at some point). Since there is no good size boundary, just always
- // use component by component assignment here.
- genComponentByComponentAssignment(builder, loc, lhs, rhs, isTemporaryLHS);
+ // Otherwise, the derived type has compile time constant size, no
+ // allocatable components, and no user-defined assignment. The size of the
+ // type is not known at this point in lowering, but fir.copy defers the size
+ // computation to codegen where the LLVM data layout is available. That allows
+ // it to copy the full allocated storage including any ABI tail-padding bytes,
+ // which a field-by-field copy would silently skip. Preserving those bytes
+ // matters for SEQUENCE types whose storage is reinterpreted via EQUIVALENCE
+ // or TRANSFER.
+ mlir::Value fromAddr = fir::getBase(rhs);
+ mlir::Value toAddr = fir::getBase(lhs);
+ // Ensure we have raw ref<RecordType> pointers for fir.copy.
+ auto refTy = builder.getRefType(recTy);
+ if (fromAddr.getType() != refTy)
+ fromAddr = builder.createConvert(loc, refTy, fromAddr);
+ if (toAddr.getType() != refTy)
+ toAddr = builder.createConvert(loc, refTy, toAddr);
+ // disjoint == true at this point (guaranteed by the condition above).
+ fir::CopyOp::create(builder, loc, fromAddr, toAddr, /*noOverlap=*/true);
----------------
MattPD wrote:
In a kernel, `type(dim3) :: idx; idx = threadIdx` now lowers to `fir.copy %threadidx to %idx` instead of three `fir.coordinate_of` and `fir.load` pairs, and `fir-opt --cuf-predefined-var-to-gpu` crashes on the new form. `processDeclareOp` at `CUFPredefinedVarToGPU.cpp:71-86` casts every use of the predefined variable's `fir.declare` to `fir.coordinate_of` without a null check, and `processCoordinateOp` calls `getFieldIndices()` on the null result. The parent's pass crashes the same way on `call unpack3(threadIdx, out)`, a `fir.call` use, so the missing check predates this patch. The patch adds whole-record assignment to the affected inputs. The upstream driver doesn't run `cuf-predefined-var-to-gpu`. `fir-opt` and downstream pipelines can run it. If @clementval agrees, the fix likely belongs in the pass: `processDeclareOp` would need to skip or handle uses other than `fir.coordinate_of`.
https://github.com/llvm/llvm-project/pull/223814
More information about the flang-commits
mailing list