[all-commits] [llvm/llvm-project] a209ff: [InstCombine] Limit canonicalization of extractele...

azwolski via All-commits all-commits at lists.llvm.org
Thu Jan 8 13:43:35 PST 2026


  Branch: refs/heads/main
  Home:   https://github.com/llvm/llvm-project
  Commit: a209ff855bfaafe7930dd35a1731220e65aebaf9
      https://github.com/llvm/llvm-project/commit/a209ff855bfaafe7930dd35a1731220e65aebaf9
  Author: azwolski <antoni.zwolski at intel.com>
  Date:   2026-01-08 (Thu, 08 Jan 2026)

  Changed paths:
    M llvm/lib/Transforms/InstCombine/InstCombineVectorOps.cpp
    M llvm/test/Transforms/InstCombine/vec_extract_var_elt.ll

  Log Message:
  -----------
  [InstCombine] Limit canonicalization of extractelement(cast) to constant index or same basic block (#166227)

The current canonicalization of extractelement(cast) requires that the
CastInst has only one use. However, when that use occurs inside a loop,
it still satisfies this condition, even though the cast is effectively
used multiple times, once per iteration, rather than truly being used
once.

```cpp
} else if (auto *CI = dyn_cast<CastInst>(I)) {
  // Canonicalize extractelement(cast) -> cast(extractelement).
  // Bitcasts can change the number of vector elements, and they cost
  // nothing.
  if (CI->hasOneUse() && (CI->getOpcode() != Instruction::BitCast)){
 ```
Before
```llvm
%34 = fptosi <4 x float> %33 to <4 x i32>
;/loop{
%40 = extractelement <4 x i32> %34, i32 %36
```
 
  After
 ```llvm
 ;/loop{
 %37 = extractelement <4 x float> %30, i32 %32
 %38 = fptosi float %37 to i32
  ```

After canonicalization, for this particular example, it no longer uses a single instruction to cast the entire vector at once, but instead performs the cast for every element separately, which is less performant.

Ideally, we would like to check if the cast instruction **has one use and that this use is not called inside a loop**. However, InstCombine/InstCombineVectorOps.cpp does not provide utilities like `LoopInfo` to check that. It might be possible to approximate this by analyzing basic block successors or by building a dominance tree, but that may be a costly solution.

A solution to prevent this optimization could be to check if the index is an immediate value and if the use is inside the same basic block as the cast instruction:
```cpp
if (CI->hasOneUse() && (CI->getOpcode() != Instruction::BitCast)) {
    Instruction *U = cast<Instruction>(*CI->user_begin());
    if (U->getParent() == CI->getParent() || isa<ConstantInt>(Index)){
```

Fixes: https://github.com/llvm/llvm-project/issues/165793



To unsubscribe from these emails, change your notification settings at https://github.com/llvm/llvm-project/settings/notifications


More information about the All-commits mailing list