[llvm] [LoopInterchange] Extract statically bounded outer epilogues (PR #224196)

via llvm-commits llvm-commits at lists.llvm.org
Wed Sep 16 23:16:07 PDT 2026


MattPD wrote:

Thanks! With the runtime-versioning (to be submitted as a follow-up):

| CPU microarchitecture | ISA/vector context | Control (s) | Transformed (s) | Reference (s) | Reduction |
|---|---|---:|---:|---:|---:|
| AMD Zen 4 | AVX-512 | 17.79 | 8.64 | 8.60 | 51% |
| AMD Zen 5 | AVX-512 | 12.78 | 6.47 | 6.48 | 49% |
| Fujitsu A64FX | SVE 512 | 43.84 | 24.84 | 24.84 | 43% |
| Arm Neoverse N1 | Neon | 30.90 | 18.45 | 18.50 | 40% |
| AMD Zen 2 | AVX2 | 24.16 | 17.58 | 18.24 | 27% |
| Intel Sierra Forest | AVX2 (E-core) | 20.15 | 15.44 | 15.37 | 23% |
| Intel Cascade Lake | AVX-512 | 27.75 | 22.86 | 22.91 | 18% |
| Intel Granite Rapids | AVX-512/AMX | 14.00 | 12.40 | 12.38 | 11% |

Mean 171.swim run time over 10 runs, comparable only within a row. Control: flang -O3 -fassociative-math -fno-signed-zeros with a target-specific -march or -mcpu. Transformed: the same plus -floop-interchange, with the outer-epilogue fission and runtime-versioning options enabled. Reference: a source-interchanged swim.f with the control flags.


https://github.com/llvm/llvm-project/pull/224196


More information about the llvm-commits mailing list