[llvm] [AMDGPU] Skip fold_pow constant-exponent shortcuts for powr on possibly-negative base (PR #200579)
Arseniy Obolenskiy via llvm-commits
llvm-commits at lists.llvm.org
Thu Jul 9 04:57:57 PDT 2026
================
@@ -898,6 +898,16 @@ bool AMDGPULibCalls::fold_pow(FPMathOperator *FPOp, IRBuilder<> &B,
// 0x1111111 means that we don't do anything for this call.
int ci_opr1 = (CINT ? (int)CINT->getSExtValue() : 0x1111111);
+ // OpenCL powr(x<0, y) = NaN, but the folds below would turn it into a
+ // finite number. Skip them unless NaNs are ignored or the base is known
+ // non-negative.
+ if ((CF || CINT) && !FPOp->hasNoNaNs() &&
+ (FInfo.getId() == AMDGPULibFunc::EI_POWR ||
+ FInfo.getId() == AMDGPULibFunc::EI_POWR_FAST) &&
+ !cannotBeOrderedLessThanZero(
+ opr0, SQ.getWithInstruction(cast<Instruction>(FPOp))))
+ return false;
----------------
aobolensk wrote:
Actually, both are valid and produce the same result, but the diff is much more verbose. Applied.
https://github.com/llvm/llvm-project/pull/200579
More information about the llvm-commits
mailing list