[llvm] [VectorCombine] Consolidate fragmented loads from shuffle chains into wide loads (PR #177571)

Yunbo Ni via llvm-commits llvm-commits at lists.llvm.org
Thu Mar 5 23:02:24 PST 2026


cardigan1008 wrote:

Hi @ParkHanbum , there is a related miscompilation case:

```llvm
define <4 x i32> @test_widen_insufficient_align(ptr %x, ptr %y) {
  %x0 = load <2 x i32>, ptr %x, align 1
  %xa = getelementptr i8, ptr %x, i64 8
  %x1 = load <1 x i32>, ptr %xa, align 1
  %y0 = load <2 x i32>, ptr %y, align 1
  %ya = getelementptr i8, ptr %y, i64 8
  %y1 = load <2 x i32>, ptr %ya, align 1
  %vx = shufflevector <2 x i32> %x0, <2 x i32> poison, <4 x i32> <i32 0, i32 1, i32 poison, i32 poison>
  %v1 = shufflevector <1 x i32> %x1, <1 x i32> poison, <4 x i32> <i32 0, i32 poison, i32 poison, i32 poison>
  %vx_comb = shufflevector <4 x i32> %vx, <4 x i32> %v1, <4 x i32> <i32 0, i32 1, i32 4, i32 poison>
  %vy = shufflevector <2 x i32> %y0, <2 x i32> %y1, <4 x i32> <i32 0, i32 1, i32 2, i32 3>
  %res = shufflevector <4 x i32> %vx_comb, <4 x i32> %vy, <4 x i32> <i32 1, i32 2, i32 5, i32 7>
  ret <4 x i32> %res
}

define <4 x i32> @test() {
  %x = alloca [3 x i32], align 16
  %y = alloca [4 x i32], align 16
  store i32 1, ptr %x
  %x_1 = getelementptr i32, ptr %x, i64 1
  store i32 2, ptr %x_1
  %x_2 = getelementptr i32, ptr %x, i64 2
  store i32 3, ptr %x_2
  store <4 x i32> <i32 5, i32 6, i32 7, i32 8>, ptr %y
  %r = call <4 x i32> @test_widen_insufficient_align(ptr %x, ptr %y)
  ret <4 x i32> %r
}

define <4 x i32> @main(i32 %argc, ptr %argv) {
entry:
  %r = call <4 x i32> @test()
  ret <4 x i32> %r
}
```

With opt built on this patch, it's transformed into:

```llvm
define <4 x i32> @test_widen_insufficient_align(ptr %x, ptr %y) {
  %x0 = load <2 x i32>, ptr %x, align 1
  %xa = getelementptr i8, ptr %x, i64 8
  %x1 = load <1 x i32>, ptr %xa, align 1
  %y0 = load <2 x i32>, ptr %y, align 1
  %ya = getelementptr i8, ptr %y, i64 8
  %y1 = load <2 x i32>, ptr %ya, align 1
  %vx = shufflevector <2 x i32> %x0, <2 x i32> poison, <4 x i32> <i32 0, i32 1, i32 poison, i32 poison>
  %v1 = shufflevector <1 x i32> %x1, <1 x i32> poison, <4 x i32> <i32 0, i32 poison, i32 poison, i32 poison>
  %vx_comb = shufflevector <4 x i32> %vx, <4 x i32> %v1, <4 x i32> <i32 0, i32 1, i32 4, i32 poison>
  %vy = shufflevector <2 x i32> %y0, <2 x i32> %y1, <4 x i32> <i32 0, i32 1, i32 2, i32 3>
  %res = shufflevector <4 x i32> %vx_comb, <4 x i32> %vy, <4 x i32> <i32 1, i32 2, i32 5, i32 7>
  ret <4 x i32> %res
}

define <4 x i32> @test() {
  %x = alloca [3 x i32], align 16
  %y = alloca [4 x i32], align 16
  store i32 1, ptr %x
  %x_1 = getelementptr i32, ptr %x, i64 1
  store i32 2, ptr %x_1
  %x_2 = getelementptr i32, ptr %x, i64 2
  store i32 3, ptr %x_2
  store <4 x i32> <i32 5, i32 6, i32 7, i32 8>, ptr %y
  %r = call <4 x i32> @test_widen_insufficient_align(ptr %x, ptr %y)
  ret <4 x i32> %r
}

define <4 x i32> @main(i32 %argc, ptr %argv) {
entry:
  %r = call <4 x i32> @test()
  ret <4 x i32> %r
}
```

Ran llubi on the transformed case,  we got:

```sh
UB triggered: Out of bound mem op, bound = 12, access range = [8, 16)
Exited with immediate UB.
Stacktrace:
    %x1 = load <2 x i32>, ptr %xa, align 1 at @test_widen_insufficient_align
    %r = call <4 x i32> @test_widen_insufficient_align(ptr %x, ptr %y) at @test
    %r = call <4 x i32> @test() at @main
```

> This is a review assisted with a self-built agent. The reproducer was validated manually. Please let me know if anything is wrong.

**Bug Triggering Analysis:**
The provided test case allocates a `[3 x i32]` array for `%x`. The function loads `<2 x i32>` from `%x` and `<1 x i32>` from `%x + 8`. The optimization `foldShuffleOfFragmentedLoads` groups these loads because they share the same base pointer `%x`. Since the loads have different types (`<2 x i32>` and `<1 x i32>`), it decides to widen the narrower load to the larger type (`<2 x i32>`). It uses `canWidenLoad` to check if widening is safe, but `canWidenLoad` only checks if the scalar type and vector size are valid for the target, it does not check if the widened load is fully within a dereferenceable region. As a result, the `<1 x i32>` load at offset 8 is widened to a `<2 x i32>` load, which reads 8 bytes starting from offset 8. However, the allocation is only 12 bytes (`[3 x i32]`), so reading 8 bytes from offset 8 reads past the end of the allocation, causing an out-of-bounds memory access (UB).

**Fix Weakness Analysis:**
The weakness in the fix is that it relies on `canWidenLoad` to determine if it's safe to widen a load to a larger vector type. However, `canWidenLoad` does not guarantee that the extra bytes read by the widened load are within a dereferenceable object. The fix should explicitly check if the widened load is safe to execute unconditionally, for example by using `isSafeToLoadUnconditionally` with the new widened type and alignment, or by ensuring that the original loads already cover the entire range of the widened load. Without this check, the optimization can introduce out-of-bounds memory accesses, leading to undefined behavior or crashes.

https://github.com/llvm/llvm-project/pull/177571


More information about the llvm-commits mailing list