[llvm] [VPlan][NFC] Simplify reverse access pattern detection in EVL vectorization (PR #199510)
Mel Chen via llvm-commits
llvm-commits at lists.llvm.org
Thu Jun 4 07:18:13 PDT 2026
Mel-Chen wrote:
> I tested this PR out on llvm-test-suite, it looks like it causes the VF/LMUL of some loops to be reduced:
>
> ```
> --- build.rva23u64-O3-a//MultiSource/Benchmarks/7zip/CMakeFiles/7zip-benchmark.dir/CPP/Common/UT
> FConvert.s 2026-05-25 19:21:20
> +++ build.rva23u64-O3-b/MultiSource/Benchmarks/7zip/CMakeFiles/7zip-benchmark.dir/CPP/Common/UTF
> Convert.s 2026-05-25 19:15:25
> @@ -533,6 +533,8 @@
> lui a2, 1034754
> .Lpcrel_hi1:
> auipc a5, %pcrel_hi(_ZL11kUtf8Limits)
> + vsetvli a3, zero, e32, m4, ta, ma
> + vid.v v8
> li t2, -1
> li s9, 6
> li a3, 63
> @@ -612,39 +614,38 @@
> add a0, t5, a1
> sh1add a5, a1, a1
> lbu s8, -1(a0)
> - vsetvli a0, zero, e32, m8, ta, ma
> - vmv.v.x v16, s1
> + vsetvli a0, zero, e32, m4, ta, ma
> + vmv.v.x v12, s1
> slli a5, a5, 1
> srlw a0, s1, a5
> add a0, a0, s8
> add a5, t6, s4
> addi s1, s4, 1
> - vmv.v.x v24, a1
> + vmv.v.x v16, a1
> sb a0, 0(a5)
> add s4, s1, a1
> - vid.v v8
> - vmacc.vx v24, t2, v8
> + vmacc.vx v16, t2, v8
> add s1, s1, t6
> ```
>
> Are we somehow failing to fold some splices back into vp.reverse in VPlan? The splices will also be folded by InstCombine, but missing it in VPlan could be affecting either the register pressure estimation or the cost model.
That's interesting! I'll check the test suite results for the new implementation tomorrow. Thanks for your report.
https://github.com/llvm/llvm-project/pull/199510
More information about the llvm-commits
mailing list