[llvm] [AMDGPU] Price scalar integer to fp casts by source width and sign (PR #225343)
Krzysztof Drewniak via llvm-commits
llvm-commits at lists.llvm.org
Fri Oct 2 16:56:57 PDT 2026
================
@@ -1098,22 +1098,34 @@ InstructionCost GCNTTIImpl::getCastInstrCost(unsigned Opcode, Type *Dst,
};
if (IsIntToFP) {
- const unsigned ExtOps = UsesInt64 && SrcBits < 64 ? (IsSigned ? 2 : 1) : 0;
+ // A scalar load of 24, 40, 48 or 56 bits is split and its high part load
+ // extends the source. A constant or invariant load aligned to 4 bytes may
+ // be widened instead and then still needs the extension.
+ const auto *Load =
+ I && Src->isIntegerTy() && I->getOperand(0)->getType() == Src
+ ? dyn_cast<LoadInst>(I->getOperand(0))
+ : nullptr;
+ const bool LoadMayWiden =
+ Load && Load->getAlign() >= Align(4) &&
+ (AMDGPU::isConstantAddressSpace(Load->getPointerAddressSpace()) ||
----------------
krzysz00 wrote:
... I think we can do a better check for this. For one thing, !noclobber works. For another, I think you're wanting "extended global address space", and for a third, we have - or should grow - a utility for "this'll be promotable to an s_*"
https://github.com/llvm/llvm-project/pull/225343
More information about the llvm-commits
mailing list