[llvm] X86: Defend against regression from SimplifyDemandedVectorElts load support (PR #213611)
Matt Arsenault via llvm-commits
llvm-commits at lists.llvm.org
Mon Aug 3 00:17:57 PDT 2026
https://github.com/arsenm created https://github.com/llvm/llvm-project/pull/213611
It doesn't appear possible to test this independently.
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
>From 195b799bc7715cb7ad1a8ec70187a81f7e92428d Mon Sep 17 00:00:00 2001
From: Matt Arsenault <Matthew.Arsenault at amd.com>
Date: Thu, 23 Jul 2026 08:25:14 +0200
Subject: [PATCH] X86: Defend against regression from
SimplifyDemandedVectorElts load support
It doesn't appear possible to test this independently.
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
---
llvm/lib/Target/X86/X86ISelLowering.cpp | 20 ++++++++++++--------
1 file changed, 12 insertions(+), 8 deletions(-)
diff --git a/llvm/lib/Target/X86/X86ISelLowering.cpp b/llvm/lib/Target/X86/X86ISelLowering.cpp
index 44c847565cb69..2cc2a54e4ea76 100644
--- a/llvm/lib/Target/X86/X86ISelLowering.cpp
+++ b/llvm/lib/Target/X86/X86ISelLowering.cpp
@@ -3476,15 +3476,19 @@ bool X86TargetLowering::shouldReduceLoadWidth(
if (const auto *GA = dyn_cast<GlobalAddressSDNode>(BasePtr.getOperand(0)))
return GA->getTargetFlags() != X86II::MO_GOTTPOFF;
- // If this is an (1) AVX vector load with (2) multiple uses and (3) all of
- // those uses are extracted directly into a store, then the extract + store
- // can be store-folded, or (4) any use will be used by legal full width
- // instruction. Then, it's probably not worth splitting the load.
+ // If this is a (1) 128-bit or wider vector load, and any use will be used by
+ // a legal full width instruction, then the load can typically be memory
+ // folded into that instruction, so it's probably not worth splitting the
+ // load. Additionally, for (2) AVX vector loads with (3) multiple uses where
+ // (4) all of those uses are extracted directly into a store, the extract +
+ // store can be store-folded, so it's again not worth splitting.
EVT VT = Load->getValueType(0);
- if ((VT.is256BitVector() || VT.is512BitVector()) &&
- !SDValue(Load, 0).hasOneUse()) {
+ if (VT.is128BitVector() || VT.is256BitVector() || VT.is512BitVector()) {
bool FullWidthUse = false;
- bool AllExtractStores = true;
+ // The extract + store folding only helps the AVX split case, which requires
+ // multiple uses of the load.
+ bool AllExtractStores = (VT.is256BitVector() || VT.is512BitVector()) &&
+ !SDValue(Load, 0).hasOneUse();
for (SDUse &Use : Load->uses()) {
// Skip uses of the chain value. Result 0 of the node is the load value.
if (Use.getResNo() != 0)
@@ -3493,7 +3497,7 @@ bool X86TargetLowering::shouldReduceLoadWidth(
const SDNode *User = PeekThroughOneUserBitcasts(Use.getUser());
// If this use is an extract + store, it's probably not worth splitting.
- if (User->getOpcode() == ISD::EXTRACT_SUBVECTOR &&
+ if (AllExtractStores && User->getOpcode() == ISD::EXTRACT_SUBVECTOR &&
all_of(User->uses(), [&](const SDUse &U) {
const SDNode *Inner = PeekThroughOneUserBitcasts(U.getUser());
return Inner->getOpcode() == ISD::STORE;
More information about the llvm-commits
mailing list