[llvm] [SLP] Vectorize select-addressed loads as masked-load blends (PR #210455)
Ryan Buchner via llvm-commits
llvm-commits at lists.llvm.org
Mon Jul 20 20:07:09 PDT 2026
================
@@ -6995,6 +7002,29 @@ isMaskedLoadCompress(ArrayRef<Value *> VL, ArrayRef<Value *> PointerOps,
CompressMask, LoadVecTy);
}
+/// Returns the cost of a BlendedLoadVectorize node loading \p VecTy: two
+/// masked loads (one per candidate base) blended by a select, plus building
+/// the blend mask from the scalar conditions and negating it for the
+/// false-lane load.
+static InstructionCost getBlendedLoadCost(const TargetTransformInfo &TTI,
+ Type *VecTy, Align Alignment,
+ unsigned AddressSpace,
+ TTI::TargetCostKind CostKind) {
+ Type *CmpTy = CmpInst::makeCmpResultType(VecTy);
+ InstructionCost MaskBuildCost = getScalarizationOverhead(
+ TTI, CmpTy->getScalarType(), cast<VectorType>(CmpTy),
+ APInt::getAllOnes(getNumElements(VecTy)), /*Insert=*/true,
+ /*Extract=*/false, CostKind);
+ return 2 * TTI.getMemIntrinsicInstrCost(
+ MemIntrinsicCostAttributes(Intrinsic::masked_load, VecTy,
+ Alignment, AddressSpace),
+ CostKind) +
+ MaskBuildCost +
+ TTI.getArithmeticInstrCost(Instruction::Xor, CmpTy, CostKind) +
+ TTI.getCmpSelInstrCost(Instruction::Select, VecTy, CmpTy,
+ CmpInst::BAD_ICMP_PREDICATE, CostKind);
----------------
bababuck wrote:
I'm not sure if the issue lies here or on the RISCV backend, but this seems to be costed incorrectly for RISCV. I don't really see a way for the backend to solve this issue though because currently it is only told about the sequence one instruction at a time. That said, I don't think this blocking since we will be erring on the side of not-vectorizing which matches the behavior before the patch.
The final RISCV sequence will be something along the lines of:
```
vsetivli zero, 16, e32, m4, ta, ma
vmsne.vi v0, v12, 0 -> comparison
vle32.v v8, (a1), v0.t -> masked load
vmseq.vi v0, v12, 0 -> negate
vsetvli zero, zero, e32, m4, ta, mu -> I think this shouldn't be here, am debugging
vle32.v v8, (a0), v0.t -> second masked load
ret
```
the second mask load also is able to act as the select instruction in this case, so the cost for `TTI.getCmpSelInstrCost()` is unneeded (for RISCV).
https://github.com/llvm/llvm-project/pull/210455
More information about the llvm-commits
mailing list