[llvm] [SLP] Vectorize select-addressed loads as masked-load blends (PR #210455)

Ryan Buchner via llvm-commits llvm-commits at lists.llvm.org
Mon Jul 20 20:07:09 PDT 2026


================
@@ -6995,6 +7002,29 @@ isMaskedLoadCompress(ArrayRef<Value *> VL, ArrayRef<Value *> PointerOps,
                               CompressMask, LoadVecTy);
 }
 
+/// Returns the cost of a BlendedLoadVectorize node loading \p VecTy: two
+/// masked loads (one per candidate base) blended by a select, plus building
+/// the blend mask from the scalar conditions and negating it for the
+/// false-lane load.
+static InstructionCost getBlendedLoadCost(const TargetTransformInfo &TTI,
+                                          Type *VecTy, Align Alignment,
+                                          unsigned AddressSpace,
+                                          TTI::TargetCostKind CostKind) {
+  Type *CmpTy = CmpInst::makeCmpResultType(VecTy);
+  InstructionCost MaskBuildCost = getScalarizationOverhead(
+      TTI, CmpTy->getScalarType(), cast<VectorType>(CmpTy),
+      APInt::getAllOnes(getNumElements(VecTy)), /*Insert=*/true,
+      /*Extract=*/false, CostKind);
+  return 2 * TTI.getMemIntrinsicInstrCost(
+                 MemIntrinsicCostAttributes(Intrinsic::masked_load, VecTy,
+                                            Alignment, AddressSpace),
+                 CostKind) +
+         MaskBuildCost +
+         TTI.getArithmeticInstrCost(Instruction::Xor, CmpTy, CostKind) +
+         TTI.getCmpSelInstrCost(Instruction::Select, VecTy, CmpTy,
+                                CmpInst::BAD_ICMP_PREDICATE, CostKind);
----------------
bababuck wrote:

I'm not sure if the issue lies here or on the RISCV backend, but this seems to be costed incorrectly for RISCV. I don't really see a way for the backend to solve this issue though because currently it is only told about the sequence one instruction at a time. That said, I don't think this blocking since we will be erring on the side of not-vectorizing which matches the behavior before the patch.

The final RISCV sequence will be something along the lines of:
```
        vsetivli        zero, 16, e32, m4, ta, ma
        vmsne.vi        v0, v12, 0                        -> comparison
        vle32.v v8, (a1), v0.t                              -> masked load
        vmseq.vi        v0, v12, 0                        -> negate
        vsetvli zero, zero, e32, m4, ta, mu        -> I think this shouldn't be here, am debugging
        vle32.v v8, (a0), v0.t                               -> second masked load
        ret
```
the second mask load also is able to act as the select instruction in this case, so the cost for `TTI.getCmpSelInstrCost()` is unneeded (for RISCV).

https://github.com/llvm/llvm-project/pull/210455


More information about the llvm-commits mailing list