[llvm] [AMDGPU] Price scalar integer to fp casts by source width and sign (PR #225343)

Krzysztof Drewniak via llvm-commits llvm-commits at lists.llvm.org
Fri Oct 2 16:56:57 PDT 2026


================
@@ -1098,22 +1098,34 @@ InstructionCost GCNTTIImpl::getCastInstrCost(unsigned Opcode, Type *Dst,
   };
 
   if (IsIntToFP) {
-    const unsigned ExtOps = UsesInt64 && SrcBits < 64 ? (IsSigned ? 2 : 1) : 0;
+    // A scalar load of 24, 40, 48 or 56 bits is split and its high part load
+    // extends the source. A constant or invariant load aligned to 4 bytes may
+    // be widened instead and then still needs the extension.
+    const auto *Load =
+        I && Src->isIntegerTy() && I->getOperand(0)->getType() == Src
+            ? dyn_cast<LoadInst>(I->getOperand(0))
+            : nullptr;
+    const bool LoadMayWiden =
+        Load && Load->getAlign() >= Align(4) &&
+        (AMDGPU::isConstantAddressSpace(Load->getPointerAddressSpace()) ||
----------------
krzysz00 wrote:

... I think we can do a better check for this. For one thing, !noclobber works. For another, I think you're wanting "extended global address space", and for a third, we have - or should grow - a utility for "this'll be promotable to an s_*"

https://github.com/llvm/llvm-project/pull/225343


More information about the llvm-commits mailing list