[llvm] [VPlan][NFC] Simplify reverse access pattern detection in EVL vectorization (PR #199510)

Mel Chen via llvm-commits llvm-commits at lists.llvm.org
Thu Jun 4 07:18:13 PDT 2026


Mel-Chen wrote:

> I tested this PR out on llvm-test-suite, it looks like it causes the VF/LMUL of some loops to be reduced:
> 
> ```
> --- build.rva23u64-O3-a//MultiSource/Benchmarks/7zip/CMakeFiles/7zip-benchmark.dir/CPP/Common/UT
> FConvert.s      2026-05-25 19:21:20
> +++ build.rva23u64-O3-b/MultiSource/Benchmarks/7zip/CMakeFiles/7zip-benchmark.dir/CPP/Common/UTF
> Convert.s       2026-05-25 19:15:25
> @@ -533,6 +533,8 @@
>         lui     a2, 1034754
>  .Lpcrel_hi1:
>         auipc   a5, %pcrel_hi(_ZL11kUtf8Limits)
> +       vsetvli a3, zero, e32, m4, ta, ma
> +       vid.v   v8
>         li      t2, -1
>         li      s9, 6
>         li      a3, 63
> @@ -612,39 +614,38 @@
>         add     a0, t5, a1
>         sh1add  a5, a1, a1
>         lbu     s8, -1(a0)
> -       vsetvli a0, zero, e32, m8, ta, ma
> -       vmv.v.x v16, s1
> +       vsetvli a0, zero, e32, m4, ta, ma
> +       vmv.v.x v12, s1
>         slli    a5, a5, 1
>         srlw    a0, s1, a5
>         add     a0, a0, s8
>         add     a5, t6, s4
>         addi    s1, s4, 1
> -       vmv.v.x v24, a1
> +       vmv.v.x v16, a1
>         sb      a0, 0(a5)
>         add     s4, s1, a1
> -       vid.v   v8
> -       vmacc.vx        v24, t2, v8
> +       vmacc.vx        v16, t2, v8
>         add     s1, s1, t6
> ```
> 
> Are we somehow failing to fold some splices back into vp.reverse in VPlan? The splices will also be folded by InstCombine, but missing it in VPlan could be affecting either the register pressure estimation or the cost model.

That's interesting! I'll check the test suite results for the new implementation tomorrow. Thanks for your report.

https://github.com/llvm/llvm-project/pull/199510


More information about the llvm-commits mailing list