[llvm] [VPlan] Extend cse to eliminate redundant widened loads (PR #212543)
via llvm-commits
llvm-commits at lists.llvm.org
Thu Jul 30 00:20:05 PDT 2026
github-actions[bot] wrote:
<!--PREMERGE ADVISOR COMMENT: Windows-->
# :window: Windows x64 Test Results
* 139187 tests passed
* 3607 tests skipped
* 5 tests failed
## Failed Tests
(click on a test name to see its output)
### LLVM
<details>
<summary>LLVM.Transforms/LoopVectorize/AArch64/partial-reduce-fdot-product.ll</summary>
```
Exit Code: 1
Command Output (stdout):
--
# RUN: at line 2
c:\_work\llvm-project\llvm-project\build\bin\opt.exe -passes=loop-vectorize -enable-epilogue-vectorization=false -S < C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\partial-reduce-fdot-product.ll | c:\_work\llvm-project\llvm-project\build\bin\filecheck.exe C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\partial-reduce-fdot-product.ll
# executed command: 'c:\_work\llvm-project\llvm-project\build\bin\opt.exe' -passes=loop-vectorize -enable-epilogue-vectorization=false -S
# note: command had no output on stdout or stderr
# executed command: 'c:\_work\llvm-project\llvm-project\build\bin\filecheck.exe' 'C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\partial-reduce-fdot-product.ll'
# .---command stderr------------
# | C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\partial-reduce-fdot-product.ll:1038:15: error: CHECK-NEXT: expected string not found in input
# | ; CHECK-NEXT: [[WIDE_LOAD7:%.*]] = load <vscale x 8 x half>, ptr [[GEP_A]], align 1
# | ^
# | <stdin>:964:167: note: scanning from here
# | %partial.reduce = call reassoc contract <vscale x 4 x float> @llvm.vector.partial.reduce.fadd.nxv4f32.nxv8f32(<vscale x 4 x float> %vec.phi, <vscale x 8 x float> %7)
# | ^
# | <stdin>:964:167: note: with "GEP_A" equal to "%3"
# | %partial.reduce = call reassoc contract <vscale x 4 x float> @llvm.vector.partial.reduce.fadd.nxv4f32.nxv8f32(<vscale x 4 x float> %vec.phi, <vscale x 8 x float> %7)
# | ^
# | <stdin>:964:167: note: pattern attempts to capture variables: "WIDE_LOAD7"
# | %partial.reduce = call reassoc contract <vscale x 4 x float> @llvm.vector.partial.reduce.fadd.nxv4f32.nxv8f32(<vscale x 4 x float> %vec.phi, <vscale x 8 x float> %7)
# | ^
# | <stdin>:986:5: note: possible intended match here
# | %load.a = load half, ptr %gep.a, align 1
# | ^
# |
# | Input file: <stdin>
# | Check file: C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\partial-reduce-fdot-product.ll
# |
# | -dump-input=help explains the following input dump.
# |
# | Input was:
# | <<<<<<
# | .
# | .
# | .
# | 959: %4 = getelementptr half, ptr %b, i64 %index
# | 960: %wide.load1 = load <vscale x 8 x half>, ptr %4, align 1
# | 961: %5 = fpext <vscale x 8 x half> %wide.load1 to <vscale x 8 x float>
# | 962: %6 = fpext <vscale x 8 x half> %wide.load to <vscale x 8 x float>
# | 963: %7 = fmul <vscale x 8 x float> %5, %6
# | 964: %partial.reduce = call reassoc contract <vscale x 4 x float> @llvm.vector.partial.reduce.fadd.nxv4f32.nxv8f32(<vscale x 4 x float> %vec.phi, <vscale x 8 x float> %7)
# | next:1038'0 { search range start (exclusive)
# | next:1038'1 error: no match found in search range
# | next:1038'2 with "GEP_A" equal to "%3"
# | next:1038'3 pattern attempts to capture variables: "WIDE_LOAD7"
# | 965: %8 = fmul <vscale x 8 x float> %6, %5
# | 966: %9 = fneg <vscale x 8 x float> %8
# | 967: %partial.reduce2 = call reassoc contract <vscale x 4 x float> @llvm.vector.partial.reduce.fadd.nxv4f32.nxv8f32(<vscale x 4 x float> %partial.reduce, <vscale x 8 x float> %9)
# | 968: %index.next = add nuw i64 %index, %2
# | 969: %10 = icmp eq i64 %index.next, %n.vec
# | .
# | .
# | .
# | 981:
# | 982: for.body: ; preds = %scalar.ph, %for.body
# | 983: %iv = phi i64 [ %bc.resume.val, %scalar.ph ], [ %iv.next, %for.body ]
# | 984: %accum = phi float [ %bc.merge.rdx, %scalar.ph ], [ %sub, %for.body ]
# | 985: %gep.a = getelementptr half, ptr %a, i64 %iv
# | 986: %load.a = load half, ptr %gep.a, align 1
# | next:1038'4 ? possible intended match
# | 987: %ext.a = fpext half %load.a to float
# | 988: %gep.b = getelementptr half, ptr %b, i64 %iv
# | 989: %load.b = load half, ptr %gep.b, align 1
# | 990: %ext.b = fpext half %load.b to float
# | 991: %mul = fmul float %ext.b, %ext.a
# | .
# | .
# | .
# | 1002: for.exit: ; preds = %middle.block, %for.body
# | 1003: %sub.lcssa = phi float [ %sub, %for.body ], [ %11, %middle.block ]
# | 1004: ret float %sub.lcssa
# | 1005: }
# | 1006:
# | 1007: define float @reduce_fsub_fadd_chain_without_mul(ptr %a, ptr noalias %b) #0 {
# | next:1038'5 } search range end (exclusive)
# | 1008: entry:
# | 1009: %0 = call i64 @llvm.vscale.i64()
# | 1010: %1 = shl nuw i64 %0, 4
# | 1011: %min.iters.check = icmp ult i64 1025, %1
# | 1012: br i1 %min.iters.check, label %scalar.ph, label %vector.ph
# | .
# | .
# | .
# | >>>>>>
# `-----------------------------
# error: command failed with exit status: 1
--
```
</details>
<details>
<summary>LLVM.Transforms/LoopVectorize/AArch64/partial-reduce-sub-epilogue-vec.ll</summary>
```
Exit Code: 1
Command Output (stdout):
--
# RUN: at line 2
c:\_work\llvm-project\llvm-project\build\bin\opt.exe -passes=loop-vectorize -force-vector-interleave=1 -enable-epilogue-vectorization=true -epilogue-vectorization-force-VF=4 -vectorizer-maximize-bandwidth -S < C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\partial-reduce-sub-epilogue-vec.ll | c:\_work\llvm-project\llvm-project\build\bin\filecheck.exe C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\partial-reduce-sub-epilogue-vec.ll --check-prefix=CHECK-EPI
# executed command: 'c:\_work\llvm-project\llvm-project\build\bin\opt.exe' -passes=loop-vectorize -force-vector-interleave=1 -enable-epilogue-vectorization=true -epilogue-vectorization-force-VF=4 -vectorizer-maximize-bandwidth -S
# note: command had no output on stdout or stderr
# executed command: 'c:\_work\llvm-project\llvm-project\build\bin\filecheck.exe' 'C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\partial-reduce-sub-epilogue-vec.ll' --check-prefix=CHECK-EPI
# .---command stderr------------
# | C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\partial-reduce-sub-epilogue-vec.ll:27:19: error: CHECK-EPI-NEXT: expected string not found in input
# | ; CHECK-EPI-NEXT: [[WIDE_LOAD1:%.*]] = load <vscale x 16 x i8>, ptr [[TMP4]], align 4
# | ^
# | <stdin>:27:55: note: scanning from here
# | %wide.load = load <vscale x 16 x i8>, ptr %3, align 4
# | ^
# | <stdin>:27:55: note: with "TMP4" equal to "%3"
# | %wide.load = load <vscale x 16 x i8>, ptr %3, align 4
# | ^
# | <stdin>:27:55: note: pattern attempts to capture variables: "WIDE_LOAD1"
# | %wide.load = load <vscale x 16 x i8>, ptr %3, align 4
# | ^
# | <stdin>:55:8: note: possible intended match here
# | %wide.load3 = load <4 x i8>, ptr %10, align 4
# | ^
# | C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\partial-reduce-sub-epilogue-vec.ll:214:101: error: undefined variable: PROF3
# | ; CHECK-EPI-NEXT: br i1 false, label %[[VEC_EPILOG_VECTOR_BODY]], label %[[VEC_EPILOG_PH]], !prof [[PROF3]]
# | ^
# | <stdin>:123:23: note: with "VEC_EPILOG_VECTOR_BODY" equal to "vec.epilog.scalar.ph"
# | vec.epilog.iter.check: ; preds = %middle.block
# | ^
# | <stdin>:123:23: note: with "VEC_EPILOG_PH" equal to "vec.epilog.ph"
# | vec.epilog.iter.check: ; preds = %middle.block
# | ^
# | <stdin>:149:2: note: possible intended match here
# | br i1 false, label %exit, label %vec.epilog.scalar.ph
# | ^
# |
# | Input file: <stdin>
# | Check file: C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\partial-reduce-sub-epilogue-vec.ll
# |
# | -dump-input=help explains the following input dump.
# |
# | Input was:
# | <<<<<<
# | .
# | .
# | .
# | 22:
# | 23: vector.body: ; preds = %vector.body, %vector.ph
# | 24: %index = phi i32 [ 0, %vector.ph ], [ %index.next, %vector.body ]
# | 25: %vec.phi = phi <vscale x 4 x i32> [ zeroinitializer, %vector.ph ], [ %partial.reduce, %vector.body ]
# | 26: %3 = getelementptr i8, ptr %src1, i32 %index
# | 27: %wide.load = load <vscale x 16 x i8>, ptr %3, align 4
# | next:27'0 { search range start (exclusive)
# | next:27'1 error: no match found in search range
# | next:27'2 with "TMP4" equal to "%3"
# | next:27'3 pattern attempts to capture variables: "WIDE_LOAD1"
# | 28: %4 = sext <vscale x 16 x i8> %wide.load to <vscale x 16 x i32>
# | 29: %5 = mul <vscale x 16 x i32> %4, %4
# | 30: %partial.reduce = call <vscale x 4 x i32> @llvm.vector.partial.reduce.add.nxv4i32.nxv16i32(<vscale x 4 x i32> %vec.phi, <vscale x 16 x i32> %5)
# | 31: %index.next = add nuw i32 %index, %2
# | 32: %6 = icmp eq i32 %index.next, %n.vec
# | .
# | .
# | .
# | 50:
# | 51: vec.epilog.vector.body: ; preds = %vec.epilog.vector.body, %vec.epilog.ph
# | 52: %index1 = phi i32 [ %vec.epilog.resume.val, %vec.epilog.ph ], [ %index.next4, %vec.epilog.vector.body ]
# | 53: %vec.phi2 = phi <4 x i32> [ %9, %vec.epilog.ph ], [ %13, %vec.epilog.vector.body ]
# | 54: %10 = getelementptr i8, ptr %src1, i32 %index1
# | 55: %wide.load3 = load <4 x i8>, ptr %10, align 4
# | next:27'4 ? possible intended match
# | 56: %11 = sext <4 x i8> %wide.load3 to <4 x i32>
# | 57: %12 = mul <4 x i32> %11, %11
# | 58: %13 = sub <4 x i32> %vec.phi2, %12
# | 59: %index.next4 = add nuw i32 %index1, 4
# | 60: %14 = icmp eq i32 %index.next4, 36
# | .
# | .
# | .
# | 88: %sub.lcssa = phi i32 [ %sub, %loop ], [ %8, %middle.block ], [ %15, %vec.epilog.middle.block ]
# | 89: ret i32 %sub.lcssa
# | 90: }
# | 91:
# | 92: ; Function Attrs: vscale_range(1,16)
# | 93: define float @fsub_reduction(float %startval, ptr %src1, ptr %src2) #1 {
# | next:27'5 } search range end (exclusive)
# | 94: iter.check:
# | 95: br i1 false, label %vec.epilog.scalar.ph, label %vector.main.loop.iter.check
# | 96:
# | 97: vector.main.loop.iter.check: ; preds = %iter.check
# | 98: br i1 false, label %vec.epilog.ph, label %vector.ph
# | .
# | .
# | .
# | 118: middle.block: ; preds = %vector.body
# | 119: %6 = call reassoc contract float @llvm.vector.reduce.fadd.v8f32(float -0.000000e+00, <8 x float> %partial.reduce)
# | 120: %7 = fsub float %startval, %6
# | 121: br i1 false, label %exit, label %vec.epilog.iter.check
# | 122:
# | 123: vec.epilog.iter.check: ; preds = %middle.block
# | next:214'0 { search range start (exclusive)
# | next:214'1 error: match failed for invalid pattern
# | next:214'2 undefined variable: PROF3
# | next:214'3 with "VEC_EPILOG_VECTOR_BODY" equal to "vec.epilog.scalar.ph"
# | next:214'4 with "VEC_EPILOG_PH" equal to "vec.epilog.ph"
# | 124: br i1 false, label %vec.epilog.scalar.ph, label %vec.epilog.ph, !prof !3
# | 125:
# | 126: vec.epilog.ph: ; preds = %vector.main.loop.iter.check, %vec.epilog.iter.check
# | 127: %vec.epilog.resume.val = phi i32 [ 32, %vec.epilog.iter.check ], [ 0, %vector.main.loop.iter.check ]
# | 128: %bc.merge.rdx = phi float [ %7, %vec.epilog.iter.check ], [ %startval, %vector.main.loop.iter.check ]
# | .
# | .
# | .
# | 144: br i1 %13, label %vec.epilog.middle.block, label %vec.epilog.vector.body, !llvm.loop !7
# | 145:
# | 146: vec.epilog.middle.block: ; preds = %vec.epilog.vector.body
# | 147: %14 = call reassoc contract float @llvm.vector.reduce.fadd.v2f32(float -0.000000e+00, <2 x float> %partial.reduce6)
# | 148: %15 = fsub float %bc.merge.rdx, %14
# | 149: br i1 false, label %exit, label %vec.epilog.scalar.ph
# | next:214'5 ? possible intended match
# | 150:
# | 151: vec.epilog.scalar.ph: ; preds = %iter.check, %vec.epilog.iter.check, %vec.epilog.middle.block
# | 152: %bc.resume.val = phi i32 [ 40, %vec.epilog.middle.block ], [ 32, %vec.epilog.iter.check ], [ 0, %iter.check ]
# | 153: %bc.merge.rdx8 = phi float [ %15, %vec.epilog.middle.block ], [ %7, %vec.epilog.iter.check ], [ %startval, %iter.check ]
# | 154: br label %loop
# | .
# | .
# | .
# | 172: %sub.lcssa = phi float [ %sub, %loop ], [ %7, %middle.block ], [ %15, %vec.epilog.middle.block ]
# | 173: ret float %sub.lcssa
# | 174: }
# | 175:
# | 176: ; Function Attrs: vscale_range(1,16)
# | 177: define float @fsub_reduction_nsz(ptr %a, ptr %b, ptr %c, i64 %n) #1 {
# | next:214'6 } search range end (exclusive)
# | 178: iter.check:
# | 179: br i1 false, label %vec.epilog.scalar.ph, label %vector.main.loop.iter.check
# | 180:
# | 181: vector.main.loop.iter.check: ; preds = %iter.check
# | 182: br i1 false, label %vec.epilog.ph, label %vector.ph
# | .
# | .
# | .
# | >>>>>>
# `-----------------------------
# error: command failed with exit status: 1
--
```
</details>
<details>
<summary>LLVM.Transforms/LoopVectorize/AArch64/transform-narrow-interleave-to-widen-memory-with-wide-ops-chained.ll</summary>
```
Exit Code: 1
Command Output (stdout):
--
# RUN: at line 2
c:\_work\llvm-project\llvm-project\build\bin\opt.exe -p loop-vectorize -force-vector-width=2 -force-vector-interleave=1 -S C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory-with-wide-ops-chained.ll | c:\_work\llvm-project\llvm-project\build\bin\filecheck.exe --check-prefixes=VF2 C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory-with-wide-ops-chained.ll
# executed command: 'c:\_work\llvm-project\llvm-project\build\bin\opt.exe' -p loop-vectorize -force-vector-width=2 -force-vector-interleave=1 -S 'C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory-with-wide-ops-chained.ll'
# note: command had no output on stdout or stderr
# executed command: 'c:\_work\llvm-project\llvm-project\build\bin\filecheck.exe' --check-prefixes=VF2 'C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory-with-wide-ops-chained.ll'
# .---command stderr------------
# | C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory-with-wide-ops-chained.ll:669:13: error: VF2-NEXT: expected string not found in input
# | ; VF2-NEXT: [[WIDE_LOAD1:%.*]] = load <2 x i64>, ptr [[TMP3]], align 4
# | ^
# | <stdin>:395:46: note: scanning from here
# | %wide.load = load <2 x i64>, ptr %3, align 4
# | ^
# | <stdin>:395:46: note: with "TMP3" equal to "%3"
# | %wide.load = load <2 x i64>, ptr %3, align 4
# | ^
# | <stdin>:395:46: note: pattern attempts to capture variables: "WIDE_LOAD1"
# | %wide.load = load <2 x i64>, ptr %3, align 4
# | ^
# | <stdin>:399:2: note: possible intended match here
# | store <2 x i64> %5, ptr %6, align 4
# | ^
# |
# | Input file: <stdin>
# | Check file: C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory-with-wide-ops-chained.ll
# |
# | -dump-input=help explains the following input dump.
# |
# | Input was:
# | <<<<<<
# | .
# | .
# | .
# | 390:
# | 391: vector.body: ; preds = %vector.body, %vector.ph
# | 392: %index = phi i64 [ 0, %vector.ph ], [ %index.next, %vector.body ]
# | 393: %2 = shl nuw nsw i64 %index, 1
# | 394: %3 = getelementptr inbounds i64, ptr %src, i64 %2
# | 395: %wide.load = load <2 x i64>, ptr %3, align 4
# | next:669'0 { search range start (exclusive)
# | next:669'1 error: no match found in search range
# | next:669'2 with "TMP3" equal to "%3"
# | next:669'3 pattern attempts to capture variables: "WIDE_LOAD1"
# | 396: %4 = xor <2 x i64> %wide.load, splat (i64 -1)
# | 397: %5 = select <2 x i1> %1, <2 x i64> %wide.load, <2 x i64> %4
# | 398: %6 = getelementptr inbounds i64, ptr %dst, i64 %2
# | 399: store <2 x i64> %5, ptr %6, align 4
# | next:669'4 ? possible intended match
# | 400: %index.next = add nuw i64 %index, 1
# | 401: %7 = icmp eq i64 %index.next, 4
# | 402: br i1 %7, label %middle.block, label %vector.body, !llvm.loop !13
# | 403:
# | 404: middle.block: ; preds = %vector.body
# | .
# | .
# | .
# | 419: !8 = distinct !{!8, !1, !2}
# | 420: !9 = distinct !{!9, !1, !2}
# | 421: !10 = distinct !{!10, !1, !2}
# | 422: !11 = distinct !{!11, !1, !2}
# | 423: !12 = distinct !{!12, !1, !2}
# | 424: !13 = distinct !{!13, !1, !2}
# | next:669'5 } search range end (exclusive)
# | >>>>>>
# `-----------------------------
# error: command failed with exit status: 1
--
```
</details>
<details>
<summary>LLVM.Transforms/LoopVectorize/AArch64/transform-narrow-interleave-to-widen-memory-with-wide-ops.ll</summary>
```
Exit Code: 1
Command Output (stdout):
--
# RUN: at line 2
c:\_work\llvm-project\llvm-project\build\bin\opt.exe -p loop-vectorize -force-vector-width=2 -force-vector-interleave=1 -S C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory-with-wide-ops.ll | c:\_work\llvm-project\llvm-project\build\bin\filecheck.exe --check-prefixes=VF2 C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory-with-wide-ops.ll
# executed command: 'c:\_work\llvm-project\llvm-project\build\bin\opt.exe' -p loop-vectorize -force-vector-width=2 -force-vector-interleave=1 -S 'C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory-with-wide-ops.ll'
# note: command had no output on stdout or stderr
# executed command: 'c:\_work\llvm-project\llvm-project\build\bin\filecheck.exe' --check-prefixes=VF2 'C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory-with-wide-ops.ll'
# .---command stderr------------
# | C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory-with-wide-ops.ll:579:13: error: VF2-NEXT: expected string not found in input
# | ; VF2-NEXT: [[WIDE_LOAD1:%.*]] = load <2 x i64>, ptr [[TMP1]], align 8
# | ^
# | <stdin>:231:52: note: scanning from here
# | %2 = getelementptr inbounds i64, ptr %data, i64 %1
# | ^
# | <stdin>:231:52: note: with "TMP1" equal to "%0"
# | %2 = getelementptr inbounds i64, ptr %data, i64 %1
# | ^
# | <stdin>:231:52: note: pattern attempts to capture variables: "WIDE_LOAD1"
# | %2 = getelementptr inbounds i64, ptr %data, i64 %1
# | ^
# | <stdin>:238:10: note: possible intended match here
# | %wide.load1 = load <2 x i64>, ptr %6, align 8
# | ^
# | C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory-with-wide-ops.ll:1525:13: error: VF2-NEXT: expected string not found in input
# | ; VF2-NEXT: [[WIDE_LOAD1:%.*]] = load <2 x double>, ptr [[TMP1]], align 8
# | ^
# | <stdin>:655:49: note: scanning from here
# | %wide.load = load <2 x double>, ptr %1, align 8
# | ^
# | <stdin>:655:49: note: with "TMP1" equal to "%1"
# | %wide.load = load <2 x double>, ptr %1, align 8
# | ^
# | <stdin>:655:49: note: pattern attempts to capture variables: "WIDE_LOAD1"
# | %wide.load = load <2 x double>, ptr %1, align 8
# | ^
# | <stdin>:658:2: note: possible intended match here
# | store <2 x double> %3, ptr %1, align 8
# | ^
# |
# | Input file: <stdin>
# | Check file: C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory-with-wide-ops.ll
# |
# | -dump-input=help explains the following input dump.
# |
# | Input was:
# | <<<<<<
# | .
# | .
# | .
# | 226: vector.body: ; preds = %vector.body, %vector.ph
# | 227: %index = phi i64 [ 0, %vector.ph ], [ %index.next, %vector.body ]
# | 228: %0 = getelementptr inbounds i64, ptr %src.0, i64 %index
# | 229: %wide.load = load <2 x i64>, ptr %0, align 8
# | 230: %1 = shl nsw i64 %index, 1
# | 231: %2 = getelementptr inbounds i64, ptr %data, i64 %1
# | next:579'0 { search range start (exclusive)
# | next:579'1 error: no match found in search range
# | next:579'2 with "TMP1" equal to "%0"
# | next:579'3 pattern attempts to capture variables: "WIDE_LOAD1"
# | 232: %3 = mul <2 x i64> %wide.load, %wide.load
# | 233: %4 = or disjoint i64 %1, 1
# | 234: %5 = getelementptr inbounds i64, ptr %data, i64 %4
# | 235: %wide.vec = load <4 x i64>, ptr %5, align 8
# | 236: %strided.vec = shufflevector <4 x i64> %wide.vec, <4 x i64> poison, <2 x i32> <i32 0, i32 2>
# | 237: %6 = getelementptr inbounds i64, ptr %src.1, i64 %index
# | 238: %wide.load1 = load <2 x i64>, ptr %6, align 8
# | next:579'4 ? possible intended match
# | 239: %7 = mul <2 x i64> %wide.load1, %strided.vec
# | 240: %8 = shufflevector <2 x i64> %3, <2 x i64> %7, <4 x i32> <i32 0, i32 1, i32 2, i32 3>
# | 241: %interleaved.vec = shufflevector <4 x i64> %8, <4 x i64> poison, <4 x i32> <i32 0, i32 2, i32 1, i32 3>
# | 242: store <4 x i64> %interleaved.vec, ptr %2, align 8
# | 243: %index.next = add nuw i64 %index, 2
# | .
# | .
# | .
# | 272:
# | 273: exit: ; preds = %loop
# | 274: ret void
# | 275: }
# | 276:
# | 277: define void @test_3xi64(ptr noalias %data, ptr noalias %factor) {
# | next:579'5 } search range end (exclusive)
# | 278: entry:
# | 279: br label %vector.ph
# | 280:
# | 281: vector.ph: ; preds = %entry
# | 282: br label %vector.body
# | .
# | .
# | .
# | 650:
# | 651: vector.body: ; preds = %vector.body, %vector.ph
# | 652: %index = phi i64 [ 0, %vector.ph ], [ %index.next, %vector.body ]
# | 653: %0 = shl nsw i64 %index, 1
# | 654: %1 = getelementptr inbounds double, ptr %data, i64 %0
# | 655: %wide.load = load <2 x double>, ptr %1, align 8
# | next:1525'0 { search range start (exclusive)
# | next:1525'1 error: no match found in search range
# | next:1525'2 with "TMP1" equal to "%1"
# | next:1525'3 pattern attempts to capture variables: "WIDE_LOAD1"
# | 656: %2 = fcmp ogt <2 x double> %wide.load, zeroinitializer
# | 657: %3 = select <2 x i1> %2, <2 x double> zeroinitializer, <2 x double> %wide.load
# | 658: store <2 x double> %3, ptr %1, align 8
# | next:1525'4 ? possible intended match
# | 659: %index.next = add nuw i64 %index, 1
# | 660: %4 = icmp eq i64 %index.next, 100
# | 661: br i1 %4, label %middle.block, label %vector.body, !llvm.loop !23
# | 662:
# | 663: middle.block: ; preds = %vector.body
# | 664: br label %exit
# | 665:
# | 666: exit: ; preds = %middle.block
# | 667: ret void
# | 668: }
# | 669:
# | 670: define void @narrowed_fneg_intersects_fmf(ptr noalias %data) {
# | next:1525'5 } search range end (exclusive)
# | 671: entry:
# | 672: br label %vector.ph
# | 673:
# | 674: vector.ph: ; preds = %entry
# | 675: br label %vector.body
# | .
# | .
# | .
# | >>>>>>
# `-----------------------------
# error: command failed with exit status: 1
--
```
</details>
<details>
<summary>LLVM.Transforms/LoopVectorize/AArch64/transform-narrow-interleave-to-widen-memory.ll</summary>
```
Exit Code: 1
Command Output (stdout):
--
# RUN: at line 2
c:\_work\llvm-project\llvm-project\build\bin\opt.exe -p loop-vectorize -force-vector-width=2 -force-vector-interleave=1 -S C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory.ll | c:\_work\llvm-project\llvm-project\build\bin\filecheck.exe --check-prefixes=VF2 C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory.ll
# executed command: 'c:\_work\llvm-project\llvm-project\build\bin\opt.exe' -p loop-vectorize -force-vector-width=2 -force-vector-interleave=1 -S 'C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory.ll'
# note: command had no output on stdout or stderr
# executed command: 'c:\_work\llvm-project\llvm-project\build\bin\filecheck.exe' --check-prefixes=VF2 'C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory.ll'
# .---command stderr------------
# | C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory.ll:616:13: error: VF2-NEXT: expected string not found in input
# | ; VF2-NEXT: [[WIDE_LOAD1:%.*]] = load <2 x double>, ptr [[TMP0]], align 8
# | ^
# | <stdin>:249:49: note: scanning from here
# | %wide.load = load <2 x double>, ptr %0, align 8
# | ^
# | <stdin>:249:49: note: with "TMP0" equal to "%0"
# | %wide.load = load <2 x double>, ptr %0, align 8
# | ^
# | <stdin>:249:49: note: pattern attempts to capture variables: "WIDE_LOAD1"
# | %wide.load = load <2 x double>, ptr %0, align 8
# | ^
# | <stdin>:251:21: note: possible intended match here
# | store <2 x double> %wide.load, ptr %1, align 8
# | ^
# |
# | Input file: <stdin>
# | Check file: C:\_work\llvm-project\llvm-project\llvm\test\Transforms\LoopVectorize\AArch64\transform-narrow-interleave-to-widen-memory.ll
# |
# | -dump-input=help explains the following input dump.
# |
# | Input was:
# | <<<<<<
# | .
# | .
# | .
# | 244: br label %vector.body
# | 245:
# | 246: vector.body: ; preds = %vector.body, %vector.ph
# | 247: %index = phi i64 [ 0, %vector.ph ], [ %index.next, %vector.body ]
# | 248: %0 = getelementptr { double, double }, ptr %A, i64 %index
# | 249: %wide.load = load <2 x double>, ptr %0, align 8
# | next:616'0 { search range start (exclusive)
# | next:616'1 error: no match found in search range
# | next:616'2 with "TMP0" equal to "%0"
# | next:616'3 pattern attempts to capture variables: "WIDE_LOAD1"
# | 250: %1 = getelementptr { double, double }, ptr %B, i64 %index
# | 251: store <2 x double> %wide.load, ptr %1, align 8
# | next:616'4 ? possible intended match
# | 252: %2 = getelementptr { double, double }, ptr %C, i64 %index
# | 253: store <2 x double> %wide.load, ptr %2, align 8
# | 254: %index.next = add nuw i64 %index, 1
# | 255: %3 = icmp eq i64 %index.next, 1000
# | 256: br i1 %3, label %middle.block, label %vector.body, !llvm.loop !11
# | .
# | .
# | .
# | 271: !6 = distinct !{!6, !1, !2}
# | 272: !7 = distinct !{!7, !1, !2}
# | 273: !8 = distinct !{!8, !1, !2}
# | 274: !9 = distinct !{!9, !1, !2}
# | 275: !10 = distinct !{!10, !1, !2}
# | 276: !11 = distinct !{!11, !1, !2}
# | next:616'5 } search range end (exclusive)
# | >>>>>>
# `-----------------------------
# error: command failed with exit status: 1
--
```
</details>
If these failures are unrelated to your changes (for example tests are broken or flaky at HEAD), please open an issue at https://github.com/llvm/llvm-project/issues and add the `infrastructure` label.
https://github.com/llvm/llvm-project/pull/212543
More information about the llvm-commits
mailing list