[all-commits] [llvm/llvm-project] a209ff: [InstCombine] Limit canonicalization of extractele...
azwolski via All-commits
all-commits at lists.llvm.org
Thu Jan 8 13:43:35 PST 2026
Branch: refs/heads/main
Home: https://github.com/llvm/llvm-project
Commit: a209ff855bfaafe7930dd35a1731220e65aebaf9
https://github.com/llvm/llvm-project/commit/a209ff855bfaafe7930dd35a1731220e65aebaf9
Author: azwolski <antoni.zwolski at intel.com>
Date: 2026-01-08 (Thu, 08 Jan 2026)
Changed paths:
M llvm/lib/Transforms/InstCombine/InstCombineVectorOps.cpp
M llvm/test/Transforms/InstCombine/vec_extract_var_elt.ll
Log Message:
-----------
[InstCombine] Limit canonicalization of extractelement(cast) to constant index or same basic block (#166227)
The current canonicalization of extractelement(cast) requires that the
CastInst has only one use. However, when that use occurs inside a loop,
it still satisfies this condition, even though the cast is effectively
used multiple times, once per iteration, rather than truly being used
once.
```cpp
} else if (auto *CI = dyn_cast<CastInst>(I)) {
// Canonicalize extractelement(cast) -> cast(extractelement).
// Bitcasts can change the number of vector elements, and they cost
// nothing.
if (CI->hasOneUse() && (CI->getOpcode() != Instruction::BitCast)){
```
Before
```llvm
%34 = fptosi <4 x float> %33 to <4 x i32>
;/loop{
%40 = extractelement <4 x i32> %34, i32 %36
```
After
```llvm
;/loop{
%37 = extractelement <4 x float> %30, i32 %32
%38 = fptosi float %37 to i32
```
After canonicalization, for this particular example, it no longer uses a single instruction to cast the entire vector at once, but instead performs the cast for every element separately, which is less performant.
Ideally, we would like to check if the cast instruction **has one use and that this use is not called inside a loop**. However, InstCombine/InstCombineVectorOps.cpp does not provide utilities like `LoopInfo` to check that. It might be possible to approximate this by analyzing basic block successors or by building a dominance tree, but that may be a costly solution.
A solution to prevent this optimization could be to check if the index is an immediate value and if the use is inside the same basic block as the cast instruction:
```cpp
if (CI->hasOneUse() && (CI->getOpcode() != Instruction::BitCast)) {
Instruction *U = cast<Instruction>(*CI->user_begin());
if (U->getParent() == CI->getParent() || isa<ConstantInt>(Index)){
```
Fixes: https://github.com/llvm/llvm-project/issues/165793
To unsubscribe from these emails, change your notification settings at https://github.com/llvm/llvm-project/settings/notifications
More information about the All-commits
mailing list