[llvm] [RISCV][SLP] Use common alignment when checking constant-stride loads (PR #222520)
via llvm-commits
llvm-commits at lists.llvm.org
Wed Sep 9 23:16:52 PDT 2026
llvmorg-github-actions[bot] wrote:
<!--LLVM PR SUMMARY COMMENT-->
@llvm/pr-subscribers-backend-risc-v
Author: Kiva (imkiva)
<details>
<summary>Changes</summary>
Fix SLP's constant-stride load legality check, which can widen naturally aligned 32-bit loads into misaligned 64-bit RVV accesses and cause `SIGBUS`.
In `BoUpSLP::canVectorizeLoads`, the check uses the first sorted load's alignment, which may be stronger than that of later widened groups. Pass `CommonAlignment` to `analyzeConstantStrideCandidate` instead, matching the existing cost model and code generation. This is conservative: it fixes the legality check without deriving a more precise alignment for each widened group's starting address, so it may reject some otherwise legal widening.
### Reproducer
The following reproduces the failure on the tested [SpacemiT K1 / Banana Pi F3](https://docs.banana-pi.org/en/BPI-F3/BananaPi_BPI-F3) running RV64 Linux, compiled with `-O3 -march=rv64gcv -mabi=lp64d -mstrict-align`:
```c
struct Row {
unsigned x, y, padding[5];
};
_Static_assert(sizeof(unsigned) == 4, "requires 32-bit unsigned");
_Static_assert(sizeof(struct Row) == 28, "requires a 28-byte row");
__attribute__((noinline))
void gather_fields(const struct Row *p, unsigned *restrict out) {
p = __builtin_assume_aligned(p, 16);
out[0] = p[0].x + 3;
out[1] = p[0].y + 3;
out[2] = p[1].x + 3;
out[3] = p[1].y + 3;
}
int main(void) {
_Alignas(16) struct Row rows[2] = {
{5, 9, {0}}, {11, 17, {0}}};
unsigned out[4];
gather_fields(rows, out);
return out[0] != 8 || out[1] != 12 ||
out[2] != 14 || out[3] != 20;
}
```
The alignment assumption is satisfied, and the four loads at byte offsets `0`, `4`, `28`, and `32` are naturally aligned and in bounds. SLP widens the adjacent pairs into:
```llvm
%v = call <2 x i64> @<!-- -->llvm.experimental.vp.strided.load.v2i64.p0.i64(
ptr align 4 %p, i64 28, <2 x i1> splat (i1 true), i32 2)
```
On RV64 this lowers to:
```asm
li a2, 28
vsetivli zero, 2, e64, m1, ta, ma
vlse64.v v8, (a0), a2
```
whose second active lane accesses `base + 28`, which is not eight-byte aligned. The legality check incorrectly uses alignment 16, although the emitted intrinsic already carries `align 4`.
On the tested machine, the example raises `SIGBUS` before the fix and returns zero afterward. The fixed compiler uses 32-bit vector loads and retains vectorized arithmetic.
---
Full diff: https://github.com/llvm/llvm-project/pull/222520.diff
2 Files Affected:
- (modified) llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp (+4-5)
- (added) llvm/test/Transforms/SLPVectorizer/RISCV/strided-load-common-alignment.ll (+63)
``````````diff
diff --git a/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp b/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
index 108b8a52d54bc..0f54e773b8486 100644
--- a/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+++ b/llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
@@ -6796,11 +6796,10 @@ BoUpSLP::LoadsState BoUpSLP::canVectorizeLoads(
cast<Instruction>(V), UserIgnoreList);
}))
return LoadsState::CompressVectorize;
- Align Alignment =
- cast<LoadInst>(Order.empty() ? VL.front() : VL[Order.front()])
- ->getAlign();
- if (analyzeConstantStrideCandidate(PointerOps, ScalarTy, Alignment, Order,
- Diff, Ptr0, SPtrInfo))
+ // Widened strided loads must be legal for every group, not just the first
+ // pointer, which may have a stronger alignment than the remaining loads.
+ if (analyzeConstantStrideCandidate(PointerOps, ScalarTy, CommonAlignment,
+ Order, Diff, Ptr0, SPtrInfo))
return LoadsState::StridedVectorize;
}
if (!IsMaskedGatherLegal())
diff --git a/llvm/test/Transforms/SLPVectorizer/RISCV/strided-load-common-alignment.ll b/llvm/test/Transforms/SLPVectorizer/RISCV/strided-load-common-alignment.ll
new file mode 100644
index 0000000000000..4ec15de2532fd
--- /dev/null
+++ b/llvm/test/Transforms/SLPVectorizer/RISCV/strided-load-common-alignment.ll
@@ -0,0 +1,63 @@
+; RUN: opt -S -passes=slp-vectorizer -mtriple=riscv64 -mattr=+v,-unaligned-vector-mem < %s | FileCheck %s --check-prefixes=CHECK,RV64
+; RUN: opt -S -passes=slp-vectorizer -mtriple=riscv32 -mattr=+v,-unaligned-vector-mem < %s | FileCheck %s --check-prefixes=CHECK,RV32
+; RUN: opt -S -passes=slp-vectorizer -mtriple=riscv64 -mattr=+v,+unaligned-vector-mem < %s | FileCheck %s --check-prefixes=UNALIGNED,RV64
+;
+; The first load's alignment does not describe every group. Widening each
+; adjacent pair to i64 would leave the group at byte offset 28 misaligned.
+
+define void @gather_fields(ptr align 16 %base, ptr noalias %out) {
+; CHECK-LABEL: define void @gather_fields(
+; CHECK-NOT: call <2 x i64> @llvm.experimental.vp.strided.load
+; CHECK: ret void
+; UNALIGNED-LABEL: define void @gather_fields(
+; UNALIGNED: call <2 x i64> @llvm.experimental.vp.strided.load.v2i64.p0.i64(ptr align 4 %base, i64 28, <2 x i1> splat (i1 true), i32 2)
+; UNALIGNED: ret void
+ %p1 = getelementptr inbounds i32, ptr %base, i64 1
+ %p2 = getelementptr inbounds i32, ptr %base, i64 7
+ %p3 = getelementptr inbounds i32, ptr %base, i64 8
+ %a = load i32, ptr %base, align 16
+ %b = load i32, ptr %p1, align 4
+ %c = load i32, ptr %p2, align 4
+ %d = load i32, ptr %p3, align 16
+ %v0 = add i32 %a, 3
+ %v1 = add i32 %b, 3
+ %v2 = add i32 %c, 3
+ %v3 = add i32 %d, 3
+ %q1 = getelementptr inbounds i32, ptr %out, i64 1
+ %q2 = getelementptr inbounds i32, ptr %out, i64 2
+ %q3 = getelementptr inbounds i32, ptr %out, i64 3
+ store i32 %v0, ptr %out, align 4
+ store i32 %v1, ptr %q1, align 4
+ store i32 %v2, ptr %q2, align 4
+ store i32 %v3, ptr %q3, align 4
+ ret void
+}
+
+; Naturally aligned strided loads remain legal.
+define void @aligned_values(ptr %base, ptr noalias %out) {
+; RV64-LABEL: define void @aligned_values(
+; RV64: call <4 x i64> @llvm.experimental.vp.strided.load.v4i64.p0.i64(ptr align 8 %base, i64 64, <4 x i1> splat (i1 true), i32 4)
+; RV64: ret void
+; RV32-LABEL: define void @aligned_values(
+; RV32: call <4 x i64> @llvm.experimental.vp.strided.load.v4i64.p0.i32(ptr align 8 %base, i32 64, <4 x i1> splat (i1 true), i32 4)
+; RV32: ret void
+ %p1 = getelementptr inbounds i64, ptr %base, i64 8
+ %p2 = getelementptr inbounds i64, ptr %base, i64 16
+ %p3 = getelementptr inbounds i64, ptr %base, i64 24
+ %a = load i64, ptr %base, align 16
+ %b = load i64, ptr %p1, align 8
+ %c = load i64, ptr %p2, align 8
+ %d = load i64, ptr %p3, align 8
+ %v0 = add i64 %a, 3
+ %v1 = add i64 %b, 3
+ %v2 = add i64 %c, 3
+ %v3 = add i64 %d, 3
+ %q1 = getelementptr inbounds i64, ptr %out, i64 1
+ %q2 = getelementptr inbounds i64, ptr %out, i64 2
+ %q3 = getelementptr inbounds i64, ptr %out, i64 3
+ store i64 %v0, ptr %out, align 8
+ store i64 %v1, ptr %q1, align 8
+ store i64 %v2, ptr %q2, align 8
+ store i64 %v3, ptr %q3, align 8
+ ret void
+}
``````````
</details>
https://github.com/llvm/llvm-project/pull/222520
More information about the llvm-commits
mailing list