[flang-commits] [flang] [flang] Fix TRANSFER into derived type with tail padding zeroing pad bytes (PR #223814)

via flang-commits flang-commits at lists.llvm.org
Wed Sep 16 20:37:05 PDT 2026


================
@@ -1552,15 +1472,24 @@ void fir::factory::genRecordAssignment(fir::FirOpBuilder &builder,
     return;
   }
 
-  // Otherwise, the derived type has compile time constant size and for which
-  // the component by component assignment can be replaced by a memory copy.
-  // Since we do not know the size of the derived type in lowering, do a
-  // component by component assignment. Note that a single fir.load/fir.store
-  // could be used on "small" record types, but as the type size grows, this
-  // leads to issues in LLVM (long compile times, long IR files, and even
-  // asserts at some point). Since there is no good size boundary, just always
-  // use component by component assignment here.
-  genComponentByComponentAssignment(builder, loc, lhs, rhs, isTemporaryLHS);
+  // Otherwise, the derived type has compile time constant size, no
+  // allocatable components, and no user-defined assignment. The size of the
+  // type is not known at this point in lowering, but fir.copy defers the size
+  // computation to codegen where the LLVM data layout is available. That allows
+  // it to copy the full allocated storage including any ABI tail-padding bytes,
+  // which a field-by-field copy would silently skip. Preserving those bytes
+  // matters for SEQUENCE types whose storage is reinterpreted via EQUIVALENCE
+  // or TRANSFER.
+  mlir::Value fromAddr = fir::getBase(rhs);
+  mlir::Value toAddr = fir::getBase(lhs);
+  // Ensure we have raw ref<RecordType> pointers for fir.copy.
+  auto refTy = builder.getRefType(recTy);
+  if (fromAddr.getType() != refTy)
+    fromAddr = builder.createConvert(loc, refTy, fromAddr);
+  if (toAddr.getType() != refTy)
+    toAddr = builder.createConvert(loc, refTy, toAddr);
+  // disjoint == true at this point (guaranteed by the condition above).
+  fir::CopyOp::create(builder, loc, fromAddr, toAddr, /*noOverlap=*/true);
----------------
MattPD wrote:

In a kernel, `type(dim3) :: idx; idx = threadIdx` now lowers to `fir.copy %threadidx to %idx` instead of three `fir.coordinate_of` and `fir.load` pairs, and `fir-opt --cuf-predefined-var-to-gpu` crashes on the new form. `processDeclareOp` at `CUFPredefinedVarToGPU.cpp:71-86` casts every use of the predefined variable's `fir.declare` to `fir.coordinate_of` without a null check, and `processCoordinateOp` calls `getFieldIndices()` on the null result. The parent's pass crashes the same way on `call unpack3(threadIdx, out)`, a `fir.call` use, so the missing check predates this patch. The patch adds whole-record assignment to the affected inputs. The upstream driver doesn't run `cuf-predefined-var-to-gpu`. `fir-opt` and downstream pipelines can run it. If @clementval agrees, the fix likely belongs in the pass: `processDeclareOp` would need to skip or handle uses other than `fir.coordinate_of`.

https://github.com/llvm/llvm-project/pull/223814


More information about the flang-commits mailing list