[llvm] [LoopFlatten] Invalidate SCEV on each widening (PR #211819)

Arda Serdar Pektezol via llvm-commits llvm-commits at lists.llvm.org
Wed Sep 9 13:58:13 PDT 2026


pektezol wrote:

@nikic I didn't come up with the original solution, so I had less command on the subject than I should have. I have now checked it more, debugged with AI assistance, and come up with the following.

The issue had an assertion in LoopVectorize which fails `BackedgeTakenCount == PSE.getBackedgeTakenCount()`, and forcing recomputation of SCEV in between LoopFlatten and LoopVectorize didn't reproduce the issue.

1. LoopFlatten first checks SCEV for backedge count for the loops in the IR. The inner loop has an i32 counter which gets cached by SCEV.
2. LoopFlatten widens the counter to i64.
3. LoopFlatten bails the flattening for the reproducer. However, the widening stays at i64. If it didn't bail and finished as normal, SCEV would have been invalidated and there would be no issues. Now we have widened i64 counters from LoopFlatten and stale SCEV cache with i32 counters.
4. Vectorizer begins and transforms the preceding loop. Invalidating that loop's exit values also invalidates the count for loop header. Inner loop's i32 count is still cached.
5. During the inner loop analysis from the Vectorizer, PSE gets i32 backedge count from SCEV and caches it as well.
6. On Vectorizer's memory access check, LoopAccessInfo is requested for the inner loop. During that, SCEV needs to know the range information for the induction exprs from loop header. To calculate that range, loop headers max backedge count is requested. Because this was invalidated at the start of the Vectorizer, it is recomputated.
7. After recomputation, SCEV invalidates dependant expressions which include the inner loop's backedge count. Right now, the SCEV cache for the inner loop is gone and it is not replaced with the widened count, and PSE is still keeping the i32 count cached.
8. When PSE  later requests the symbolic max backedge count, SCEV recomputes from the widened IR and returns an i64 expression. This is compared with the PSE cached count, which is what fails the assertion based on i32 expr != i64 expr.

I did update the PR description after your comment, but I didn't realize more detail would have been better for this. I apologize.

https://github.com/llvm/llvm-project/pull/211819


More information about the llvm-commits mailing list