[llvm] [LAA] Switch -stencil-runtime-check-merge to auto by default (PR #228373)

Igor Kirillov via llvm-commits llvm-commits at lists.llvm.org
Fri Oct 2 03:37:34 PDT 2026


igogo-x86 wrote:

@fhahn, Yes, I updated it - couldn't reference non-existent PRs at the moment. The full story. The code in `709.cactus` has around 20–30 read-only objects, heavy calculations, and 6–11 output objects for the results. Each input object can be accessed in up to 223 places per iteration, but these accesses share the same i, with different stencil offsets. Currently, LAA creates a separate checking group for each of these accesses. With this patch, eligible accesses to the same object are merged into one group, reducing the number of runtime checks required:

| Kernel   | Merging off | Merging enabled |
|----------|------------:|----------------:|
| SplitBy1 |       2,676 |             180 |
| SplitBy2 |      29,304 |             330 |
| SplitBy3 |      12,459 |             147 |

These are runtime overlap checks between groups. This patch, together with raising the runtime-check threshold, can give the 709.cactus benchmark around +39% score on Neoverse V2. It's very likely good for any other CPUs/architectures that support vectorisation.

`SplitBy1` and `SplitBy3` take approximately 50% of benchmark's time and `SplitBy2` takes the rest. We propose to raise the threshold by 50% at first and raising it to 330+ may be discussed later (I see how this number may feel too big).

Stencil merging can also in theory help cactus from SPEC2017 but it requires extra work in loop unswitching.

https://github.com/llvm/llvm-project/pull/228373


More information about the llvm-commits mailing list