[llvm] [AMDGPU] Add waterfall intrinsics (PR #192409)

via llvm-commits llvm-commits at lists.llvm.org
Fri Apr 17 09:09:33 PDT 2026


https://github.com/gretay-amd updated https://github.com/llvm/llvm-project/pull/192409

>From a78a8ac234f391e29ef937a73fe7f6f21a56eeeb Mon Sep 17 00:00:00 2001
From: Greta Yorsh <Greta.Yorsh at amd.com>
Date: Thu, 2 Apr 2026 11:18:22 +0100
Subject: [PATCH 1/5] Add waterfall intrinsics

Co-authored-by: David Stuttard <David.Stuttard at amd.com>
Co-authored-by: Carl Ritson <Carl.Ritson at amd.com>
Co-authored-by: Jay Foad <Jay.Foad at amd.com>
Co-authored-by: Piotr Sobczak <piotr.sobczak at amd.com>
Co-authored-by: Joseph Nash <Joseph.Nash at amd.com>
Co-authored-by: Sebastian Neubauer <Sebastian.Neubauer at amd.com>
Assisted-by: Claude
---
 llvm/docs/AMDGPUUsage.rst                     |    53 +
 llvm/include/llvm/IR/IntrinsicsAMDGPU.td      |    67 +
 llvm/lib/Target/AMDGPU/AMDGPU.h               |     4 +
 .../Target/AMDGPU/AMDGPUAtomicOptimizer.cpp   |    89 +-
 .../Target/AMDGPU/AMDGPUCodeGenPrepare.cpp    |    70 +-
 .../AMDGPU/AMDGPUInstCombineIntrinsic.cpp     |    26 +
 llvm/lib/Target/AMDGPU/AMDGPUPassRegistry.def |     1 +
 .../Target/AMDGPU/AMDGPURegBankCombiner.cpp   |    54 +
 .../Target/AMDGPU/AMDGPURegisterBankInfo.cpp  |    50 +
 .../Target/AMDGPU/AMDGPUSearchableTables.td   |     4 +
 .../lib/Target/AMDGPU/AMDGPUTargetMachine.cpp |     4 +
 llvm/lib/Target/AMDGPU/CMakeLists.txt         |     1 +
 llvm/lib/Target/AMDGPU/SIISelLowering.cpp     |    21 +
 llvm/lib/Target/AMDGPU/SIInsertWaterfall.cpp  |   896 ++
 llvm/lib/Target/AMDGPU/SIInsertWaterfall.h    |    22 +
 llvm/lib/Target/AMDGPU/SIInstructions.td      |   107 +
 .../AMDGPU/MIR/llvm.amdgcn.waterfall.mir      |    51 +
 .../AMDGPU/llvm.amdgcn.waterfall.ll           |    29 +
 .../AMDGPU/amdgcn.waterfall.atomic.opt.ll     |   406 +
 llvm/test/CodeGen/AMDGPU/llc-pipeline-npm.ll  |     3 +
 llvm/test/CodeGen/AMDGPU/llc-pipeline.ll      |    16 +
 .../CodeGen/AMDGPU/llvm.amdgcn.waterfall.ll   | 12656 ++++++++++++++++
 .../AMDGPU/si-insert-waterfall-licm.ll        |    43 +
 .../CodeGen/AMDGPU/si-insert-waterfall.mir    |    54 +
 .../InstCombine/AMDGPU/waterfall.ll           |    73 +
 25 files changed, 14795 insertions(+), 5 deletions(-)
 create mode 100644 llvm/lib/Target/AMDGPU/SIInsertWaterfall.cpp
 create mode 100644 llvm/lib/Target/AMDGPU/SIInsertWaterfall.h
 create mode 100644 llvm/test/Analysis/UniformityAnalysis/AMDGPU/MIR/llvm.amdgcn.waterfall.mir
 create mode 100644 llvm/test/Analysis/UniformityAnalysis/AMDGPU/llvm.amdgcn.waterfall.ll
 create mode 100644 llvm/test/CodeGen/AMDGPU/amdgcn.waterfall.atomic.opt.ll
 create mode 100644 llvm/test/CodeGen/AMDGPU/llvm.amdgcn.waterfall.ll
 create mode 100644 llvm/test/CodeGen/AMDGPU/si-insert-waterfall-licm.ll
 create mode 100644 llvm/test/CodeGen/AMDGPU/si-insert-waterfall.mir
 create mode 100644 llvm/test/Transforms/InstCombine/AMDGPU/waterfall.ll

diff --git a/llvm/docs/AMDGPUUsage.rst b/llvm/docs/AMDGPUUsage.rst
index fb440bde68b8f..96ad9c0b755f6 100644
--- a/llvm/docs/AMDGPUUsage.rst
+++ b/llvm/docs/AMDGPUUsage.rst
@@ -1776,6 +1776,59 @@ The AMDGPU backend implements the following LLVM IR intrinsics.
                                                    * :ref:`Synchronization Scope<amdgpu-intrinsics-syncscope-metadata-operand>`.
                                                      Note that the scope used must ensure that the L2 cache will be hit.
 
+  llvm.amdgcn.waterfall.begin                      Marks the beginning of a waterfall region of code.
+
+                                                   The compiler generates a waterfall loop around the region.
+                                                   A waterfall loop handles the case where an operation that requires
+                                                   a uniform operand (e.g., held in an SGPR) is applied to a divergent operand
+                                                   (held in a VGPR, with values varying per lane).
+                                                   Each iteration of the waterfall loop activates a subset of lanes that
+                                                   share the same value of the VGPR (the value in the first active lane).
+                                                   The operation is then executed using that value as the uniform operand.
+                                                   If the VGPR is already uniform, the waterfall loop executes only once.
+                                                   The worst case for a waterfall loop is one iteration per lane (all lanes
+                                                   have different values of the VGPR), but it is not common in practice.
+
+                                                   The intrinsic takes a previous token
+                                                   (``i32``; use a null/zero value if this is the first ``waterfall.begin`` in
+                                                   a waterfall group) and a VGPR.
+
+                                                   The intrinsic returns a new token that must be threaded through the
+                                                   corresponding ``waterfall.readfirstlane`` and ``waterfall.end`` or
+                                                   ``waterfall.last_use`` intrinsics, forming a waterfall group of intrinsics
+                                                   that together define a waterfall region.
+
+                                                   All intrinsics in a waterfall group must reside in the same basic block.
+
+                                                   Multiple ``waterfall.begin`` intrinsics can be chained by
+                                                   passing the token of the preceding ``waterfall.begin`` as the first argument.
+                                                   This allows a front-end to create one waterfall loop for
+                                                   multiple non-uniform values.
+                                                   Later compiler passes may remove values determined as uniform.
+                                                   The final token is used for other waterfall intrinsics in the same group.
+
+  llvm.amdgcn.waterfall.readfirstlane              Reads the first active lane's value of the VGPR and returns it as
+                                                   an SGPR for use within a waterfall region.
+
+                                                   Takes the ``i32`` token from the final ``waterfall.begin`` in the waterfall
+                                                   group and the VGPR. Returns the uniform (SGPR) result.
+
+                                                   If the VGPR is determined to be uniform at compile time, this intrinsic
+                                                   is optimized away (the input VGPR value is used directly).
+
+  llvm.amdgcn.waterfall.end                        Marks the end of a waterfall region.
+                                                   Takes the ``i32`` token from the final ``waterfall.begin``
+
+  llvm.amdgcn.waterfall.last_use                   Variant of ``waterfall.end`` for values whose last use is in a
+                                                   non-defining operation such as a store. Marks that the use of the value
+                                                   constitutes the end of the waterfall region.
+
+  llvm.amdgcn.waterfall.last_use_vgpr              Variant of ``waterfall.last_use`` for values that remain in a VGPR.
+
+  llvm.amdgcn.waterfall.loop_end                   Inserted later by the compiler to be used with ```waterfall.last_use*``
+                                                   to mark the loop-end point for special
+                                                   handling such as SCC clobber tracking.
+
   ==============================================   ==========================================================
 
 .. TODO::
diff --git a/llvm/include/llvm/IR/IntrinsicsAMDGPU.td b/llvm/include/llvm/IR/IntrinsicsAMDGPU.td
index 9528fb2b446bc..e4aee0ad0310f 100644
--- a/llvm/include/llvm/IR/IntrinsicsAMDGPU.td
+++ b/llvm/include/llvm/IR/IntrinsicsAMDGPU.td
@@ -2693,6 +2693,73 @@ def int_amdgcn_cs_chain:
             ],
             [IntrConvergent, IntrNoReturn, ImmArg<ArgIndex<4>>]>;
 
+// Waterfall intrinsics are used to mark a region as requiring waterfall loops to
+// activate and deactivate lanes in a loop (and sometimes transform VGPR values
+// into SGPR values).
+// Best case (uniform index in VGPR) the loop will execute once, worst case (all
+// values in the index are different) the loop will execute wave-size times.
+// The waterfall.begin intrinsic returns a new token that must be threaded through the
+// corresponding waterfall.readfirstlane and waterfall.end or
+// waterfall.last_use intrinsics, forming a waterfall group of intrinsics that together
+// define a waterfall region.
+// All intrinsics in a waterfall group must reside within the same basic block.
+// Where the intrinsic specifies "llvm_any_ty", the intrinsic will accept any of
+// the following types:
+// i16, v2i16, v4i16, i32, v2i32, v4i32, v8i32, f16, v2f16, v4f16, f32, v2f32,
+// v4f32, v8f32
+// TODO: extend the support to allow spanning multiple blocks
+def int_amdgcn_waterfall_begin :
+    Intrinsic<[llvm_i32_ty],  // Returns sgpr token
+              [llvm_i32_ty,   // Previous begin token / null if first
+               llvm_any_ty],  // VGPR index for waterfall
+              [IntrConvergent]>;
+
+// Often the index from begin will be used as the readfirstlane value from
+// waterfall code - however, this isn't always the case. The readfirstlane style
+// intrinsic is required to be used, but often is a no-op if the index is the
+// used directly.
+def int_amdgcn_waterfall_readfirstlane :
+    Intrinsic<[llvm_any_ty],     // Returns any value
+              [llvm_i32_ty,      // Token from final begin
+               llvm_any_ty],     // VGPR value to be used in readfirstlane
+              [IntrConvergent]>;
+
+// Effectively a no-op, but this intrinsic is used to indicate the end of the
+// region for waterfall by tagging a value as the end of the section.
+// Waterfall will happen for all values dependent on the begin until the end
+// intrinsic is reached,
+def int_amdgcn_waterfall_end :
+    Intrinsic<[llvm_any_ty],      // Returns any value
+              [llvm_i32_ty,       // Token from final begin
+               LLVMMatchType<0>], // Same as return value
+              [IntrConvergent]>;
+
+// Effectively a no-op, no code results directly from this intrinsic. This is
+// used to indicate that the use of the defined value constitutes the end of the
+// waterfall propagation. This is roughly equivalent to the end intrinsic, but
+// can be used to tag values that are used in non-defining operations such as a
+// store.
+def int_amdgcn_waterfall_last_use :
+    Intrinsic<[llvm_any_ty],      // Returns any value
+              [llvm_i32_ty,       // Token from begin
+               LLVMMatchType<0>], // Same as return value
+              [IntrConvergent]>;
+
+// Special variant where the last-use uses a VGPR that needs to be uniform.
+def int_amdgcn_waterfall_last_use_vgpr :
+    Intrinsic<[llvm_any_ty],      // Returns any value
+              [llvm_i32_ty,       // Token from begin
+               LLVMMatchType<0>], // Same as return value
+              []>;
+
+// Special waterfall loop end intrinsic to be used with last_use
+// This is inserted late, before codegen, mainly to mark the loop end
+// for special handling such as scc clobbers.
+def int_amdgcn_waterfall_loop_end :
+    Intrinsic<[],
+              [llvm_i32_ty],      // Token from begin
+              [IntrConvergent]>;
+
 // Run a function with all the lanes enabled. Only direct calls are allowed. The
 // first argument is the callee, which must have the `amdgpu_gfx_whole_wave`
 // calling convention and must not be variadic. The remaining arguments to the
diff --git a/llvm/lib/Target/AMDGPU/AMDGPU.h b/llvm/lib/Target/AMDGPU/AMDGPU.h
index 878f374110159..4edc25a391669 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPU.h
+++ b/llvm/lib/Target/AMDGPU/AMDGPU.h
@@ -51,6 +51,7 @@ FunctionPass *createSIFixSGPRCopiesLegacyPass();
 FunctionPass *createLowerWWMCopiesPass();
 FunctionPass *createSIMemoryLegalizerPass();
 FunctionPass *createSIInsertWaitcntsPass();
+FunctionPass *createSIInsertWaterfallPass();
 FunctionPass *createSIPreAllocateWWMRegsLegacyPass();
 FunctionPass *createSIFormMemoryClausesLegacyPass();
 
@@ -232,6 +233,9 @@ extern char &SILateBranchLoweringPassID;
 void initializeSIOptimizeExecMaskingLegacyPass(PassRegistry &);
 extern char &SIOptimizeExecMaskingLegacyID;
 
+void initializeSIInsertWaterfallPass(PassRegistry &);
+extern char &SIInsertWaterfallID;
+
 void initializeSIPreAllocateWWMRegsLegacyPass(PassRegistry &);
 extern char &SIPreAllocateWWMRegsLegacyID;
 
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUAtomicOptimizer.cpp b/llvm/lib/Target/AMDGPU/AMDGPUAtomicOptimizer.cpp
index b4d51522e28af..499ce33cfe149 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUAtomicOptimizer.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUAtomicOptimizer.cpp
@@ -28,6 +28,7 @@
 #include "llvm/CodeGen/TargetPassConfig.h"
 #include "llvm/IR/IRBuilder.h"
 #include "llvm/IR/InstVisitor.h"
+#include "llvm/IR/IntrinsicInst.h"
 #include "llvm/IR/IntrinsicsAMDGPU.h"
 #include "llvm/InitializePasses.h"
 #include "llvm/Target/TargetMachine.h"
@@ -89,6 +90,8 @@ class AMDGPUAtomicOptimizerImpl
   void optimizeAtomic(Instruction &I, AtomicRMWInst::BinOp Op, unsigned ValIdx,
                       bool ValDivergent) const;
 
+  void processBB(BasicBlock &BB);
+
 public:
   AMDGPUAtomicOptimizerImpl() = delete;
 
@@ -157,7 +160,8 @@ bool AMDGPUAtomicOptimizerImpl::run() {
   if (ST.isSingleLaneExecution(F))
     return false;
 
-  visit(F);
+  for (auto &BB : F)
+    processBB(BB);
   if (ToReplace.empty())
     return false;
 
@@ -181,6 +185,89 @@ static bool isLegalCrossLaneType(Type *Ty) {
   }
 }
 
+static Instruction *findLastInWaterfall(Instruction &Begin) {
+  // Given Begin as the first begin for a waterfall group, look through the
+  // instructions tagged as part of the waterfall and return the last use.
+  // If the group is malformed, then return nullptr.
+
+  Instruction *FinalBegin = &Begin;
+
+  // Drill through any begin intrinsics
+  do {
+    if (FinalBegin->hasOneUse()) {
+      User *U = *FinalBegin->user_begin();
+      IntrinsicInst *Intrin = dyn_cast<IntrinsicInst>(U);
+      if (Intrin &&
+          Intrin->getIntrinsicID() == Intrinsic::amdgcn_waterfall_begin) {
+        FinalBegin = Intrin;
+        continue;
+      }
+    }
+  } while (false);
+
+  Instruction *Last = FinalBegin;
+
+  for (auto Use : FinalBegin->users()) {
+    if (auto *Intrin = dyn_cast<IntrinsicInst>(Use)) {
+      switch (Intrin->getIntrinsicID()) {
+      default: {
+        // Unexpected intrinsic
+        return nullptr;
+      }
+      case Intrinsic::amdgcn_waterfall_begin:
+        // Badly formed WF group - should already have discovered the last begin
+        // intrinsic before entering this loop.
+        return nullptr;
+      case Intrinsic::amdgcn_waterfall_end:
+        if (Last->comesBefore(Intrin))
+          Last = Intrin;
+        break;
+      case Intrinsic::amdgcn_waterfall_last_use:
+      case Intrinsic::amdgcn_waterfall_last_use_vgpr: {
+        // Find the actual last use
+        for (auto &LastUse : Intrin->uses()) {
+          auto *LUI = static_cast<Instruction *>(LastUse.getUser());
+          if (LUI->getParent() != Intrin->getParent()) {
+            // LUI has to be in same BB as waterfall intrinsics.
+            return nullptr;
+          }
+          if (Last->comesBefore(LUI))
+            Last = LUI;
+        }
+        break;
+      }
+      case Intrinsic::amdgcn_waterfall_readfirstlane:
+        // Always in the middle of a group - so doesn't have any effect on group
+        // detection
+        break;
+      }
+    }
+  }
+
+  return Last;
+}
+
+void AMDGPUAtomicOptimizerImpl::processBB(BasicBlock &BB) {
+  // Visit all instructions in a BB unless they're inside a waterfall loop
+  for (auto I = BB.begin(); I != BB.end(); ++I) {
+    if (auto *Intrin = dyn_cast<IntrinsicInst>(&*I)) {
+      if (Intrin->getIntrinsicID() == Intrinsic::amdgcn_waterfall_begin) {
+        auto *Last = findLastInWaterfall(*I);
+        if (!Last) {
+          // Malformed waterfall group - assume that all instructions in this
+          // BB are inside a waterfall group
+          return;
+        }
+        I = Last->getIterator();
+        continue;
+      }
+    }
+    // Safe to perform transformations on this instruction
+    // which is not inside a waterfall group
+    visit(*I);
+  }
+}
+
 void AMDGPUAtomicOptimizerImpl::visitAtomicRMWInst(AtomicRMWInst &I) {
   // Early exit for unhandled address space atomic instructions.
   switch (I.getPointerAddressSpace()) {
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUCodeGenPrepare.cpp b/llvm/lib/Target/AMDGPU/AMDGPUCodeGenPrepare.cpp
index d049df810c476..2c175f49c273c 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUCodeGenPrepare.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUCodeGenPrepare.cpp
@@ -18,6 +18,7 @@
 #include "llvm/ADT/SetVector.h"
 #include "llvm/Analysis/AssumptionCache.h"
 #include "llvm/Analysis/ConstantFolding.h"
+#include "llvm/Analysis/LoopInfo.h"
 #include "llvm/Analysis/TargetLibraryInfo.h"
 #include "llvm/Analysis/TargetTransformInfo.h"
 #include "llvm/Analysis/UniformityAnalysis.h"
@@ -102,6 +103,7 @@ class AMDGPUCodeGenPrepareImpl
   const AMDGPUTargetMachine &TM;
   const TargetLibraryInfo *TLI;
   const UniformityInfo &UA;
+  const LoopInfo &LI;
   const DataLayout &DL;
   SimplifyQuery SQ;
   const bool HasFP32DenormalFlush;
@@ -112,11 +114,14 @@ class AMDGPUCodeGenPrepareImpl
 
   DenseMap<const PHINode *, bool> BreakPhiNodesCache;
 
+  IntrinsicInst *CurrentWaterfall = nullptr;
+
   AMDGPUCodeGenPrepareImpl(Function &F, const AMDGPUTargetMachine &TM,
                            const TargetLibraryInfo *TLI, AssumptionCache *AC,
-                           const DominatorTree *DT, const UniformityInfo &UA)
+                           const DominatorTree *DT, const UniformityInfo &UA,
+                           const LoopInfo &LI)
       : F(F), ST(TM.getSubtarget<GCNSubtarget>(F)), TM(TM), TLI(TLI), UA(UA),
-        DL(F.getDataLayout()), SQ(DL, TLI, DT, AC),
+        LI(LI), DL(F.getDataLayout()), SQ(DL, TLI, DT, AC),
         HasFP32DenormalFlush(SIModeRegisterDefaults(F, ST).FP32Denormals ==
                              DenormalMode::getPreserveSign()) {}
 
@@ -259,11 +264,15 @@ class AMDGPUCodeGenPrepareImpl
   bool visitAddrSpaceCastInst(AddrSpaceCastInst &I);
 
   bool visitIntrinsicInst(IntrinsicInst &I);
+  bool visitWaterfallLastUseIntrinsic(IntrinsicInst &I);
   bool visitFMinLike(IntrinsicInst &I);
   bool visitSqrt(IntrinsicInst &I);
   bool visitLog(FPMathOperator &Log, Intrinsic::ID IID);
   bool visitMbcntLo(IntrinsicInst &I) const;
   bool visitMbcntHi(IntrinsicInst &I) const;
+
+  bool visitGetElementPtrInst(GetElementPtrInst &I);
+
   bool run();
 };
 
@@ -273,6 +282,7 @@ class AMDGPUCodeGenPrepare : public FunctionPass {
   AMDGPUCodeGenPrepare() : FunctionPass(ID) {}
   void getAnalysisUsage(AnalysisUsage &AU) const override {
     AU.addRequired<AssumptionCacheTracker>();
+    AU.addRequired<LoopInfoWrapperPass>();
     AU.addRequired<UniformityInfoWrapperPass>();
     AU.addRequired<TargetLibraryInfoWrapperPass>();
 
@@ -2066,11 +2076,60 @@ bool AMDGPUCodeGenPrepareImpl::visitIntrinsicInst(IntrinsicInst &I) {
     return visitMbcntLo(I);
   case Intrinsic::amdgcn_mbcnt_hi:
     return visitMbcntHi(I);
+  case Intrinsic::amdgcn_waterfall_begin:
+    CurrentWaterfall = nullptr;
+    return false;
+  case Intrinsic::amdgcn_waterfall_end:
+    CurrentWaterfall = &I;
+    return false;
+  case Intrinsic::amdgcn_waterfall_last_use:
+    return visitWaterfallLastUseIntrinsic(I);
   default:
     return false;
   }
 }
 
+bool AMDGPUCodeGenPrepareImpl::visitGetElementPtrInst(GetElementPtrInst &I) {
+  if (!CurrentWaterfall || UA.isUniform(&I))
+    return false;
+  if (I.getParent() != CurrentWaterfall->getParent())
+    return false;
+
+  // Divergent GEP within a waterfall region will introduce a nested waterfall.
+  // This will lead to bad or broken code gen.
+  // Pointer likely became non-uniform due to sinking into a divergent loop.
+  // Make sure LICM can run on the relevant loops and hope it can
+  // hoist/canonicalize the pointer.
+  LLVM_DEBUG(dbgs() << "Divergent GEP found in waterfall = " << I << "\n");
+  Value *PtrVal = I.getPointerOperand()->stripPointerCasts();
+  Instruction *PtrInst = dyn_cast<Instruction>(PtrVal);
+  if (!PtrInst)
+    return false;
+
+  Loop *CurrentLoop = LI.getLoopFor(PtrInst->getParent());
+  while (CurrentLoop) {
+    MDNode *LoopID = CurrentLoop->getLoopID();
+    MDNode *NewLoopID = makePostTransformationMetadata(
+        I.getContext(), LoopID, {"llvm.licm.disable"}, {});
+    CurrentLoop->setLoopID(NewLoopID);
+    CurrentLoop = CurrentLoop->getParentLoop();
+  }
+
+  return true;
+}
+
+bool AMDGPUCodeGenPrepareImpl::visitWaterfallLastUseIntrinsic(
+    IntrinsicInst &I) {
+  auto *Token = I.getOperand(0);
+  for (auto *U : I.users()) {
+    auto *UI = cast<Instruction>(U);
+    BasicBlock::iterator InsertPt = std::next(UI->getIterator());
+    IRBuilder<> Builder(I.getParent(), InsertPt);
+    Builder.CreateIntrinsic(Intrinsic::amdgcn_waterfall_loop_end, {}, {Token});
+  }
+  return true;
+}
+
 /// Match the core sequence in the fract pattern (x - floor(x), which doesn't
 /// need to consider edge case handling.
 Value *AMDGPUCodeGenPrepareImpl::matchFractPatImpl(Value &FractSrc,
@@ -2274,7 +2333,8 @@ bool AMDGPUCodeGenPrepare::runOnFunction(Function &F) {
   const DominatorTree *DT = DTWP ? &DTWP->getDomTree() : nullptr;
   const UniformityInfo &UA =
       getAnalysis<UniformityInfoWrapperPass>().getUniformityInfo();
-  return AMDGPUCodeGenPrepareImpl(F, TM, TLI, AC, DT, UA).run();
+  const LoopInfo &LI = getAnalysis<LoopInfoWrapperPass>().getLoopInfo();
+  return AMDGPUCodeGenPrepareImpl(F, TM, TLI, AC, DT, UA, LI).run();
 }
 
 PreservedAnalyses AMDGPUCodeGenPreparePass::run(Function &F,
@@ -2284,7 +2344,8 @@ PreservedAnalyses AMDGPUCodeGenPreparePass::run(Function &F,
   AssumptionCache *AC = &FAM.getResult<AssumptionAnalysis>(F);
   const DominatorTree *DT = FAM.getCachedResult<DominatorTreeAnalysis>(F);
   const UniformityInfo &UA = FAM.getResult<UniformityInfoAnalysis>(F);
-  AMDGPUCodeGenPrepareImpl Impl(F, ATM, TLI, AC, DT, UA);
+  const LoopInfo &LI = FAM.getResult<LoopAnalysis>(F);
+  AMDGPUCodeGenPrepareImpl Impl(F, ATM, TLI, AC, DT, UA, LI);
   if (!Impl.run())
     return PreservedAnalyses::all();
   PreservedAnalyses PA = PreservedAnalyses::none();
@@ -2296,6 +2357,7 @@ PreservedAnalyses AMDGPUCodeGenPreparePass::run(Function &F,
 INITIALIZE_PASS_BEGIN(AMDGPUCodeGenPrepare, DEBUG_TYPE,
                       "AMDGPU IR optimizations", false, false)
 INITIALIZE_PASS_DEPENDENCY(AssumptionCacheTracker)
+INITIALIZE_PASS_DEPENDENCY(LoopInfoWrapperPass)
 INITIALIZE_PASS_DEPENDENCY(TargetLibraryInfoWrapperPass)
 INITIALIZE_PASS_DEPENDENCY(UniformityInfoWrapperPass)
 INITIALIZE_PASS_END(AMDGPUCodeGenPrepare, DEBUG_TYPE, "AMDGPU IR optimizations",
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUInstCombineIntrinsic.cpp b/llvm/lib/Target/AMDGPU/AMDGPUInstCombineIntrinsic.cpp
index 463d63a88f690..f9af55a1eaf94 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUInstCombineIntrinsic.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUInstCombineIntrinsic.cpp
@@ -136,6 +136,9 @@ static std::optional<Instruction *> modifyIntrinsicCall(
   NewCall->copyMetadata(OldIntr);
   if (isa<FPMathOperator>(NewCall))
     NewCall->copyFastMathFlags(&OldIntr);
+  // Copy attributes
+  AttributeList OldAttrList = OldIntr.getAttributes();
+  NewCall->setAttributes(OldAttrList);
 
   // Erase and replace uses
   if (!InstToReplace.getType()->isVoidTy())
@@ -1747,6 +1750,22 @@ GCNTTIImpl::instCombineIntrinsic(InstCombiner &IC, IntrinsicInst &II) const {
       return IC.replaceInstUsesWith(II, ConstantInt::getFalse(II.getType()));
     break;
   }
+  case Intrinsic::amdgcn_waterfall_begin: {
+    Value *Index = II.getArgOperand(1);
+    // If there is a previous waterfall begin with the same index, we can remove
+    // this one.
+    IntrinsicInst *PrevII;
+    for (Value *Token = II.getArgOperand(0);
+         (PrevII = dyn_cast<IntrinsicInst>(Token)) &&
+         PrevII->getIntrinsicID() == Intrinsic::amdgcn_waterfall_begin;
+         Token = PrevII->getArgOperand(0)) {
+      if (Index == PrevII->getArgOperand(1)) {
+        IC.replaceInstUsesWith(II, II.getArgOperand(0));
+        return IC.eraseInstFromFunction(II);
+      }
+    }
+    break;
+  }
   case Intrinsic::amdgcn_make_buffer_rsrc: {
     Value *Src = II.getArgOperand(0);
     if (isa<PoisonValue>(Src))
@@ -2050,6 +2069,8 @@ static Value *simplifyAMDGCNMemoryIntrinsicDemanded(InstCombiner &IC,
       IC.Builder.CreateIntrinsic(II.getIntrinsicID(), OverloadTys, Args);
   NewCall->takeName(&II);
   NewCall->copyMetadata(II);
+  AttributeList OldAttrList = II.getAttributes();
+  NewCall->setAttributes(OldAttrList);
 
   if (IsLoad) {
     if (NewNumElts == 1) {
@@ -2163,6 +2184,11 @@ std::optional<Value *> GCNTTIImpl::simplifyDemandedVectorEltsIntrinsic(
   case Intrinsic::amdgcn_struct_tbuffer_load:
   case Intrinsic::amdgcn_struct_ptr_tbuffer_load:
     return simplifyAMDGCNMemoryIntrinsicDemanded(IC, II, DemandedElts);
+  case Intrinsic::amdgcn_waterfall_end:
+    // Propagate demanded elements through to the value being returned by the
+    // waterfall loop.
+    SimplifyAndSetOp(&II, 1, DemandedElts, UndefElts);
+    break;
   default: {
     if (getAMDGPUImageDMaskIntrinsic(II.getIntrinsicID())) {
       return simplifyAMDGCNMemoryIntrinsicDemanded(IC, II, DemandedElts, 0);
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUPassRegistry.def b/llvm/lib/Target/AMDGPU/AMDGPUPassRegistry.def
index 8a046e83548cc..6f95ef64e1f22 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUPassRegistry.def
+++ b/llvm/lib/Target/AMDGPU/AMDGPUPassRegistry.def
@@ -135,6 +135,7 @@ MACHINE_FUNCTION_PASS("si-form-memory-clauses", SIFormMemoryClausesPass())
 MACHINE_FUNCTION_PASS("si-i1-copies", SILowerI1CopiesPass())
 MACHINE_FUNCTION_PASS("si-insert-hard-clauses", SIInsertHardClausesPass())
 MACHINE_FUNCTION_PASS("si-insert-waitcnts", SIInsertWaitcntsPass())
+MACHINE_FUNCTION_PASS("si-insert-waterfall", SIInsertWaterfallPass())
 MACHINE_FUNCTION_PASS("si-late-branch-lowering", SILateBranchLoweringPass())
 MACHINE_FUNCTION_PASS("si-load-store-opt", SILoadStoreOptimizerPass())
 MACHINE_FUNCTION_PASS("si-lower-control-flow", SILowerControlFlowPass())
diff --git a/llvm/lib/Target/AMDGPU/AMDGPURegBankCombiner.cpp b/llvm/lib/Target/AMDGPU/AMDGPURegBankCombiner.cpp
index 990d879c2bf09..e2df70b3964ca 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPURegBankCombiner.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPURegBankCombiner.cpp
@@ -20,11 +20,14 @@
 #include "llvm/CodeGen/GlobalISel/Combiner.h"
 #include "llvm/CodeGen/GlobalISel/CombinerHelper.h"
 #include "llvm/CodeGen/GlobalISel/CombinerInfo.h"
+#include "llvm/CodeGen/GlobalISel/GIMatchTableExecutor.h"
 #include "llvm/CodeGen/GlobalISel/GIMatchTableExecutorImpl.h"
 #include "llvm/CodeGen/GlobalISel/GISelValueTracking.h"
+#include "llvm/CodeGen/GlobalISel/GenericMachineInstrs.h"
 #include "llvm/CodeGen/GlobalISel/MIPatternMatch.h"
 #include "llvm/CodeGen/MachineDominators.h"
 #include "llvm/CodeGen/TargetPassConfig.h"
+#include "llvm/IR/IntrinsicsAMDGPU.h"
 #include "llvm/Target/TargetMachine.h"
 
 #define GET_GICOMBINER_DEPS
@@ -89,6 +92,19 @@ class AMDGPURegBankCombinerImpl : public Combiner {
 
   void applyCanonicalizeZextShiftAmt(MachineInstr &MI, MachineInstr &Ext) const;
 
+  struct RmUniformWFMatchInfo {
+    Register Dst, WFReplaceReg;
+  };
+
+  // Combine support to remove amdgcn_waterfall_readfirstlane or
+  // amdgcn_waterfall_begin intrinsics if the index is determined to be
+  // uniform. The SIInsertWaterfall pass can handle their removal and in some
+  // cases remove the waterfall altogether
+  bool matchRmUniformWF(MachineInstr &MI,
+                        RmUniformWFMatchInfo &MatchInfo) const;
+  void applyRmUniformWF(MachineInstr &MI,
+                        RmUniformWFMatchInfo &MatchInfo) const;
+
   bool combineD16Load(MachineInstr &MI) const;
   bool applyD16Load(unsigned D16Opc, MachineInstr &DstMI,
                     MachineInstr *SmallLoad, Register ToOverwriteD16) const;
@@ -396,6 +412,44 @@ void AMDGPURegBankCombinerImpl::applyCanonicalizeZextShiftAmt(
   MI.eraseFromParent();
 }
 
+bool AMDGPURegBankCombinerImpl::matchRmUniformWF(
+    MachineInstr &MI, RmUniformWFMatchInfo &MatchInfo) const {
+  auto IntrID = cast<GIntrinsic>(MI).getIntrinsicID();
+  Register Dst, WFReplaceReg;
+
+  switch (IntrID) {
+  case Intrinsic::amdgcn_waterfall_readfirstlane: {
+    Register IdxReg = MI.getOperand(3).getReg();
+    Register IdxSrcReg = getSrcRegIgnoringCopies(IdxReg, MRI);
+    if (!isVgprRegBank(IdxSrcReg)) {
+      WFReplaceReg = IdxSrcReg;
+      break;
+    }
+    return false;
+  }
+  case Intrinsic::amdgcn_waterfall_begin: {
+    Register IdxReg = MI.getOperand(3).getReg();
+    if (!isVgprRegBank(getSrcRegIgnoringCopies(IdxReg, MRI))) {
+      WFReplaceReg = MI.getOperand(2).getReg();
+      break;
+    }
+    return false;
+  }
+  default:
+    return false;
+  }
+
+  Dst = MI.getOperand(0).getReg();
+  MatchInfo = {Dst, WFReplaceReg};
+  return true;
+}
+
+void AMDGPURegBankCombinerImpl::applyRmUniformWF(
+    MachineInstr &MI, RmUniformWFMatchInfo &MatchInfo) const {
+  MI.eraseFromParent();
+  Helper.replaceRegWith(MRI, MatchInfo.Dst, MatchInfo.WFReplaceReg);
+}
+
 bool AMDGPURegBankCombinerImpl::combineD16Load(MachineInstr &MI) const {
   Register Dst;
   MachineInstr *Load, *SextLoad;
diff --git a/llvm/lib/Target/AMDGPU/AMDGPURegisterBankInfo.cpp b/llvm/lib/Target/AMDGPU/AMDGPURegisterBankInfo.cpp
index ebdd709c34f08..3b2afb3e9da8c 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPURegisterBankInfo.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPURegisterBankInfo.cpp
@@ -5520,6 +5520,56 @@ AMDGPURegisterBankInfo::getInstrMapping(const MachineInstr &MI) const {
       OpdsMapping[1] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size);
       break;
     }
+    case Intrinsic::amdgcn_waterfall_begin: {
+      unsigned SizeDst = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
+      unsigned SizeSrc1 = getSizeInBits(MI.getOperand(2).getReg(), MRI, *TRI);
+      unsigned SizeSrc2 = getSizeInBits(MI.getOperand(3).getReg(), MRI, *TRI);
+      OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, SizeDst);
+      OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, SizeSrc1);
+      OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, SizeSrc2);
+      break;
+    }
+    case Intrinsic::amdgcn_waterfall_readfirstlane: {
+      unsigned SizeDst = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
+      unsigned SizeSrc1 = getSizeInBits(MI.getOperand(2).getReg(), MRI, *TRI);
+      unsigned SizeSrc2 = getSizeInBits(MI.getOperand(3).getReg(), MRI, *TRI);
+      OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, SizeDst);
+      OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, SizeSrc1);
+      OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, SizeSrc2);
+      break;
+    }
+    case Intrinsic::amdgcn_waterfall_end: {
+      unsigned SizeDst = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
+      unsigned SizeSrc1 = getSizeInBits(MI.getOperand(2).getReg(), MRI, *TRI);
+      unsigned SizeSrc2 = getSizeInBits(MI.getOperand(3).getReg(), MRI, *TRI);
+      OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, SizeDst);
+      OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, SizeSrc1);
+      OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, SizeSrc2);
+      break;
+    }
+    case Intrinsic::amdgcn_waterfall_last_use: {
+      unsigned SizeDst = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
+      unsigned SizeSrc1 = getSizeInBits(MI.getOperand(2).getReg(), MRI, *TRI);
+      unsigned SizeSrc2 = getSizeInBits(MI.getOperand(3).getReg(), MRI, *TRI);
+      OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, SizeDst);
+      OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, SizeSrc1);
+      OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, SizeSrc2);
+      break;
+    }
+    case Intrinsic::amdgcn_waterfall_last_use_vgpr: {
+      unsigned SizeDst = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
+      unsigned SizeSrc1 = getSizeInBits(MI.getOperand(2).getReg(), MRI, *TRI);
+      unsigned SizeSrc2 = getSizeInBits(MI.getOperand(3).getReg(), MRI, *TRI);
+      OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, SizeDst);
+      OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, SizeSrc1);
+      OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, SizeSrc2);
+      break;
+    }
+    case Intrinsic::amdgcn_waterfall_loop_end: {
+      unsigned SizeSrc1 = getSizeInBits(MI.getOperand(1).getReg(), MRI, *TRI);
+      OpdsMapping[1] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, SizeSrc1);
+      break;
+    }
     case Intrinsic::amdgcn_ds_gws_init:
     case Intrinsic::amdgcn_ds_gws_barrier:
     case Intrinsic::amdgcn_ds_gws_sema_br: {
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUSearchableTables.td b/llvm/lib/Target/AMDGPU/AMDGPUSearchableTables.td
index 4cb200cd51e51..96aa8b661aab7 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUSearchableTables.td
+++ b/llvm/lib/Target/AMDGPU/AMDGPUSearchableTables.td
@@ -336,6 +336,7 @@ def : SourceOfDivergence<int_amdgcn_permlane16_swap>;
 def : SourceOfDivergence<int_amdgcn_permlane32_swap>;
 def : SourceOfDivergence<int_amdgcn_set_inactive>;
 def : SourceOfDivergence<int_amdgcn_set_inactive_chain_arg>;
+def : SourceOfDivergence<int_amdgcn_waterfall_end>;
 
 foreach intr = AMDGPUMFMAIntrinsics908 in
 def : SourceOfDivergence<intr>;
@@ -401,6 +402,9 @@ def : AlwaysUniform<int_amdgcn_icmp>;
 def : AlwaysUniform<int_amdgcn_fcmp>;
 def : AlwaysUniform<int_amdgcn_ballot>;
 def : AlwaysUniform<int_amdgcn_if_break>;
+def : AlwaysUniform<int_amdgcn_waterfall_readfirstlane>;
+def : AlwaysUniform<int_amdgcn_waterfall_begin>;
+def : AlwaysUniform<int_amdgcn_waterfall_last_use>;
 def : AlwaysUniform<int_amdgcn_cluster_workgroup_id_x>;
 def : AlwaysUniform<int_amdgcn_cluster_workgroup_id_y>;
 def : AlwaysUniform<int_amdgcn_cluster_workgroup_id_z>;
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUTargetMachine.cpp b/llvm/lib/Target/AMDGPU/AMDGPUTargetMachine.cpp
index e08345f5acbc5..73201faaf31b9 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUTargetMachine.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUTargetMachine.cpp
@@ -52,6 +52,7 @@
 #include "SIFixVGPRCopies.h"
 #include "SIFoldOperands.h"
 #include "SIFormMemoryClauses.h"
+#include "SIInsertWaterfall.h"
 #include "SILoadStoreOptimizer.h"
 #include "SILowerControlFlow.h"
 #include "SILowerSGPRSpills.h"
@@ -673,6 +674,7 @@ extern "C" LLVM_ABI LLVM_EXTERNAL_VISIBILITY void LLVMInitializeAMDGPUTarget() {
   initializeSIFixSGPRCopiesLegacyPass(*PR);
   initializeSIFixVGPRCopiesLegacyPass(*PR);
   initializeSIFoldOperandsLegacyPass(*PR);
+  initializeSIInsertWaterfallPass(*PR);
   initializeSIPeepholeSDWALegacyPass(*PR);
   initializeSIShrinkInstructionsLegacyPass(*PR);
   initializeSIOptimizeExecMaskingPreRALegacyPass(*PR);
@@ -1709,6 +1711,7 @@ void GCNPassConfig::addFastRegAlloc() {
 }
 
 void GCNPassConfig::addPreRegAlloc() {
+  addPass(createSIInsertWaterfallPass());
   if (getOptLevel() != CodeGenOptLevel::None)
     addPass(&AMDGPUPrepareAGPRAllocLegacyID);
 }
@@ -2523,6 +2526,7 @@ Error AMDGPUCodeGenPassBuilder::addOptimizedRegAlloc(
 }
 
 void AMDGPUCodeGenPassBuilder::addPreRegAlloc(PassManagerWrapper &PMW) const {
+  addMachineFunctionPass(SIInsertWaterfallPass(), PMW);
   if (getOptLevel() != CodeGenOptLevel::None)
     addMachineFunctionPass(AMDGPUPrepareAGPRAllocPass(), PMW);
 }
diff --git a/llvm/lib/Target/AMDGPU/CMakeLists.txt b/llvm/lib/Target/AMDGPU/CMakeLists.txt
index 2f4883bfc8245..28f85cee5a67f 100644
--- a/llvm/lib/Target/AMDGPU/CMakeLists.txt
+++ b/llvm/lib/Target/AMDGPU/CMakeLists.txt
@@ -164,6 +164,7 @@ add_llvm_target(AMDGPUCodeGen
   SIFrameLowering.cpp
   SIInsertHardClauses.cpp
   SIInsertWaitcnts.cpp
+  SIInsertWaterfall.cpp
   SIInstrInfo.cpp
   SIISelLowering.cpp
   SILateBranchLowering.cpp
diff --git a/llvm/lib/Target/AMDGPU/SIISelLowering.cpp b/llvm/lib/Target/AMDGPU/SIISelLowering.cpp
index 752ea02119e03..9ad308bc70b14 100644
--- a/llvm/lib/Target/AMDGPU/SIISelLowering.cpp
+++ b/llvm/lib/Target/AMDGPU/SIISelLowering.cpp
@@ -11627,6 +11627,27 @@ SDValue SITargetLowering::LowerINTRINSIC_W_CHAIN(SDValue Op,
     DAG.setNodeMemRefs(NewNode, {MemRef});
     return SDValue(NewNode, 0);
   }
+  case Intrinsic::amdgcn_waterfall_readfirstlane: {
+    if (!Op->getOperand(3)->isDivergent()) {
+      // If waterfall_readfirstlane is uniform, it can be removed
+      SDValue InChain = Op.getOperand(0);
+      SDValue OutChain = Op.getValue(1);
+      DAG.ReplaceAllUsesOfValueWith(OutChain, InChain);
+      DAG.ReplaceAllUsesOfValueWith(Op.getValue(0), Op.getOperand(3));
+      return SDValue();
+    }
+    return Op;
+  }
+  case Intrinsic::amdgcn_waterfall_begin: {
+    // If the index in a waterfall.begin is uniform, it can be removed
+    if (!Op->getOperand(3)->isDivergent()) {
+      SDValue InChain = Op.getOperand(0);
+      SDValue OutChain = Op.getValue(1);
+      DAG.ReplaceAllUsesOfValueWith(OutChain, InChain);
+      DAG.ReplaceAllUsesOfValueWith(Op.getValue(0), Op.getOperand(2));
+    }
+    return Op;
+  }
   case Intrinsic::amdgcn_global_atomic_fmin_num:
   case Intrinsic::amdgcn_global_atomic_fmax_num:
   case Intrinsic::amdgcn_flat_atomic_fmin_num:
diff --git a/llvm/lib/Target/AMDGPU/SIInsertWaterfall.cpp b/llvm/lib/Target/AMDGPU/SIInsertWaterfall.cpp
new file mode 100644
index 0000000000000..f0f0662c1b28b
--- /dev/null
+++ b/llvm/lib/Target/AMDGPU/SIInsertWaterfall.cpp
@@ -0,0 +1,896 @@
+//===- SIInsertWaterfall.cpp - insert waterall loops at intrinsic markers -===//
+//
+//                     The LLVM Compiler Infrastructure
+//
+// This file is distributed under the University of Illinois Open Source
+// License. See LICENSE.TXT for details.
+//
+//===----------------------------------------------------------------------===//
+//
+/// \file
+/// Replace 3 intrinsics used to mark waterfall regions with actual waterfall
+/// loops. This is done at MachineIR level rather than LLVM-IR due to the use of
+/// exec mask in this operation.
+///
+//===----------------------------------------------------------------------===//
+
+#include "SIInsertWaterfall.h"
+#include "AMDGPU.h"
+#include "GCNSubtarget.h"
+#include "MCTargetDesc/AMDGPUMCTargetDesc.h"
+#include "SIInstrInfo.h"
+#include "llvm/CodeGen/MachineFunctionPass.h"
+
+using namespace llvm;
+
+#define DEBUG_TYPE "si-insert-waterfall"
+
+namespace {
+
+static unsigned getWFBeginSize(const unsigned Opcode) {
+  switch (Opcode) {
+  case AMDGPU::SI_WATERFALL_BEGIN_V1:
+    return 1;
+  case AMDGPU::SI_WATERFALL_BEGIN_V2:
+    return 2;
+  case AMDGPU::SI_WATERFALL_BEGIN_V4:
+    return 4;
+  case AMDGPU::SI_WATERFALL_BEGIN_V8:
+    return 8;
+  default:
+    break;
+  }
+
+  return 0; // Not SI_WATERFALL_BEGIN_*
+}
+
+static unsigned getWFRFLSize(const unsigned Opcode) {
+  switch (Opcode) {
+  case AMDGPU::SI_WATERFALL_READFIRSTLANE_V1:
+    return 1;
+  case AMDGPU::SI_WATERFALL_READFIRSTLANE_V2:
+    return 2;
+  case AMDGPU::SI_WATERFALL_READFIRSTLANE_V4:
+    return 4;
+  case AMDGPU::SI_WATERFALL_READFIRSTLANE_V8:
+    return 8;
+  default:
+    break;
+  }
+
+  return 0; // Not SI_WATERFALL_READFIRSTLANE_*
+}
+
+static unsigned getWFEndSize(const unsigned Opcode) {
+  switch (Opcode) {
+  case AMDGPU::SI_WATERFALL_END_V1:
+    return 1;
+  case AMDGPU::SI_WATERFALL_END_V2:
+    return 2;
+  case AMDGPU::SI_WATERFALL_END_V4:
+    return 4;
+  case AMDGPU::SI_WATERFALL_END_V8:
+    return 8;
+  default:
+    break;
+  }
+
+  return 0; // Not SI_WATERFALL_END_*
+}
+
+static unsigned getWFLastUseSize(const unsigned Opcode) {
+  switch (Opcode) {
+  case AMDGPU::SI_WATERFALL_LAST_USE_V1:
+  case AMDGPU::SI_WATERFALL_LAST_USE_V1_V:
+    return 1;
+  case AMDGPU::SI_WATERFALL_LAST_USE_V2:
+  case AMDGPU::SI_WATERFALL_LAST_USE_V2_V:
+    return 2;
+  case AMDGPU::SI_WATERFALL_LAST_USE_V4:
+  case AMDGPU::SI_WATERFALL_LAST_USE_V4_V:
+    return 4;
+  case AMDGPU::SI_WATERFALL_LAST_USE_V8:
+  case AMDGPU::SI_WATERFALL_LAST_USE_V8_V:
+    return 8;
+  default:
+    break;
+  }
+
+  return 0; // Not SI_WATERFALL_LAST_USE_*
+}
+
+static bool isWFLastUseVGPR(const unsigned Opcode) {
+  switch (Opcode) {
+  case AMDGPU::SI_WATERFALL_LAST_USE_V1_V:
+  case AMDGPU::SI_WATERFALL_LAST_USE_V2_V:
+  case AMDGPU::SI_WATERFALL_LAST_USE_V4_V:
+  case AMDGPU::SI_WATERFALL_LAST_USE_V8_V:
+    return true;
+  default:
+    break;
+  }
+
+  return false;
+}
+
+static bool isWFLoopEnd(const unsigned Opcode) {
+  return Opcode == AMDGPU::SI_WATERFALL_LOOP_END;
+}
+
+static void readFirstLaneReg(MachineBasicBlock &MBB, MachineRegisterInfo *MRI,
+                             const SIRegisterInfo *RI, const SIInstrInfo *TII,
+                             MachineBasicBlock::iterator &I, const DebugLoc &DL,
+                             Register RFLReg, Register RFLSrcReg,
+                             const MachineOperand &RFLSrcOp) {
+  auto RFLRegRC = MRI->getRegClass(RFLReg);
+  uint32_t RegSize = RI->getRegSizeInBits(*RFLRegRC) / 32;
+  assert(RI->hasVGPRs(MRI->getRegClass(RFLSrcReg)) &&
+         "unexpected uniform operand for readfirstlane");
+
+  if (RegSize == 1) {
+    MRI->constrainRegClass(RFLReg, &AMDGPU::SReg_32_XM0RegClass);
+    BuildMI(MBB, I, DL, TII->get(AMDGPU::V_READFIRSTLANE_B32), RFLReg)
+        .addReg(RFLSrcReg, getUndefRegState(RFLSrcOp.isUndef()),
+                RFLSrcOp.getSubReg());
+  } else {
+    SmallVector<Register, 8> TRegs;
+    for (unsigned i = 0; i < RegSize; ++i) {
+      Register TReg = MRI->createVirtualRegister(&AMDGPU::SReg_32_XM0RegClass);
+      BuildMI(MBB, I, DL, TII->get(AMDGPU::V_READFIRSTLANE_B32), TReg)
+          .addReg(RFLSrcReg, {}, RI->getSubRegFromChannel(i));
+      TRegs.push_back(TReg);
+    }
+    MachineInstrBuilder MIB =
+        BuildMI(MBB, I, DL, TII->get(AMDGPU::REG_SEQUENCE), RFLReg);
+    for (unsigned i = 0; i < RegSize; ++i) {
+      MIB.addReg(TRegs[i]);
+      MIB.addImm(RI->getSubRegFromChannel(i));
+    }
+  }
+}
+
+// Check if operand is uniform by checking:
+// 1. Trivially detectable as operand in SGPR
+// 2. Direct def is from an SGPR->VGPR copy (which may happen if assumed
+// non-uniform value
+//    turns out to be uniform)
+static Register getUniformOperandReplacementReg(MachineRegisterInfo *MRI,
+                                                const SIRegisterInfo *RI,
+                                                Register Reg) {
+  auto RegRC = MRI->getRegClass(Reg);
+  if (!RI->hasVGPRs(RegRC)) {
+    return Reg;
+  }
+
+  // Check for operand def being a copy from SGPR
+  MachineInstr *DefMI = MRI->getVRegDef(Reg);
+  if (DefMI->isFullCopy()) {
+    auto const &DefSrcOp = DefMI->getOperand(1);
+    if (DefSrcOp.isReg() && DefSrcOp.getReg().isVirtual()) {
+      Register ReplaceReg = DefSrcOp.getReg();
+      if (!RI->hasVGPRs(MRI->getRegClass(ReplaceReg)))
+        return ReplaceReg;
+    }
+  }
+  return AMDGPU::NoRegister;
+}
+
+static void compareIdxUsingCmpx(
+    MachineBasicBlock &MBB, MachineRegisterInfo *MRI, const SIRegisterInfo *RI,
+    const SIInstrInfo *TII, MachineBasicBlock::iterator &I, const DebugLoc &DL,
+    Register CurrentIdxReg, const MachineOperand &IndexOp, bool IsWave32) {
+
+  unsigned CmpxEqU32Opc = (IsWave32 ? AMDGPU::V_CMPX_EQ_U32_nosdst_e32
+                                    : AMDGPU::V_CMPX_EQ_U32_nosdst_e64);
+
+  Register IndexReg = IndexOp.getReg();
+  const TargetRegisterClass *IndexRC =
+      RI->getRegClassForOperandReg(*MRI, IndexOp);
+
+  // Iterate over the index in dword chunks
+  uint32_t RegSize = RI->getRegSizeInBits(*IndexRC) / 32;
+  if (RegSize == 1) {
+    BuildMI(MBB, I, DL, TII->get(CmpxEqU32Opc))
+        .addReg(CurrentIdxReg)
+        .addReg(IndexReg, {}, IndexOp.getSubReg());
+
+  } else {
+    for (unsigned Idx = 0; Idx < RegSize; ++Idx) {
+      BuildMI(MBB, I, DL, TII->get(CmpxEqU32Opc))
+          .addReg(CurrentIdxReg, {}, RI->getSubRegFromChannel(Idx))
+          .addReg(IndexReg, {}, RI->getSubRegFromChannel(Idx));
+    }
+  }
+}
+
+static Register compareIdx(MachineBasicBlock &MBB, MachineRegisterInfo *MRI,
+                           const SIRegisterInfo *RI, const SIInstrInfo *TII,
+                           MachineBasicBlock::iterator &I, const DebugLoc &DL,
+                           Register CurrentIdxReg,
+                           const MachineOperand &IndexOp, Register CondReg,
+                           bool IsWave32) {
+  // Iterate over the index in dword chunks and'ing the result with the
+  // CondReg
+  // Optionally CondReg is passed in from a previous compareIdx call
+  Register IndexReg = IndexOp.getReg();
+  auto IndexRC = RI->getRegClassForOperandReg(*MRI, IndexOp);
+  unsigned AndOpc = IsWave32 ? AMDGPU::S_AND_B32 : AMDGPU::S_AND_B64;
+  const auto *BoolXExecRC = TII->getRegisterInfo().getWaveMaskRegClass();
+
+  uint32_t RegSize = RI->getRegSizeInBits(*IndexRC) / 32;
+
+  if (RegSize == 1) {
+    Register TReg = MRI->createVirtualRegister(BoolXExecRC);
+    BuildMI(MBB, I, DL, TII->get(AMDGPU::V_CMP_EQ_U32_e64), TReg)
+        .addReg(CurrentIdxReg)
+        .addReg(IndexReg, {}, IndexOp.getSubReg());
+
+    if (CondReg != AMDGPU::NoRegister) {
+      Register TReg2 = MRI->createVirtualRegister(BoolXExecRC);
+      BuildMI(MBB, I, DL, TII->get(AndOpc), TReg2).addReg(CondReg).addReg(TReg);
+      CondReg = TReg2;
+    } else {
+      CondReg = TReg;
+    }
+  } else {
+    unsigned StartCount;
+    Register TReg;
+    if (CondReg != AMDGPU::NoRegister) {
+      TReg = CondReg;
+      StartCount = 0;
+    } else {
+      TReg = MRI->createVirtualRegister(BoolXExecRC);
+      BuildMI(MBB, I, DL, TII->get(AMDGPU::V_CMP_EQ_U32_e64), TReg)
+          .addReg(CurrentIdxReg, {}, AMDGPU::sub0)
+          .addReg(IndexReg, {}, AMDGPU::sub0);
+      StartCount = 1;
+    }
+
+    for (unsigned i = StartCount; i < RegSize; ++i) {
+      Register TReg2 = MRI->createVirtualRegister(BoolXExecRC);
+      BuildMI(MBB, I, DL, TII->get(AMDGPU::V_CMP_EQ_U32_e64), TReg2)
+          .addReg(CurrentIdxReg, {}, RI->getSubRegFromChannel(i))
+          .addReg(IndexReg, {}, RI->getSubRegFromChannel(i));
+      Register TReg3 = MRI->createVirtualRegister(BoolXExecRC);
+      BuildMI(MBB, I, DL, TII->get(AndOpc), TReg3).addReg(TReg).addReg(TReg2);
+      TReg = TReg3;
+    }
+    CondReg = TReg;
+  }
+  return CondReg;
+}
+
+// Replace all registers From with To.
+// Also handles From and To MachineOperands having sub registers.
+// Note: MRI->replaceRegWith doesn't handle sub registers since it is
+// register based and subreg is carried on the operand.
+static void replaceRegIncSubReg(const MachineRegisterInfo *MRI,
+                                const TargetRegisterInfo *TRI,
+                                const MachineOperand *From,
+                                const MachineOperand *To) {
+  for (auto &O : make_early_inc_range(MRI->reg_operands(From->getReg())))
+    O.substVirtReg(To->getReg(), To->getSubReg(), *TRI);
+}
+
+class SIInsertWaterfall : public MachineFunctionPass {
+private:
+  struct WaterfallWorkitem {
+    const SIInstrInfo *TII;
+    const MachineRegisterInfo *MRI;
+    Register TokReg; // This is always the token from the last begin intrinsic
+    MachineInstr *Final;
+    bool hasVGPRLastUse;
+
+    std::vector<MachineInstr *> BeginList;
+    std::vector<MachineInstr *> RFLList;
+    std::vector<MachineInstr *> EndList;
+    std::vector<MachineInstr *> LastUseList;
+    std::vector<MachineInstr *> LoopEndList;
+
+    // List of corresponding init, newdst and phi registers used in loop for
+    // end pseudos
+    std::vector<std::pair<MachineOperand *, MachineOperand *>> EndRegs;
+    std::vector<Register> RFLRegs;
+
+    WaterfallWorkitem() = default;
+    WaterfallWorkitem(MachineInstr *_Begin, const SIInstrInfo *_TII,
+                      MachineRegisterInfo *_MRI)
+        : TII(_TII), MRI(_MRI), Final(nullptr), hasVGPRLastUse(false) {
+
+      auto TokMO = TII->getNamedOperand(*_Begin, AMDGPU::OpName::tok_ret);
+
+      assert(tokIsStart(TII->getNamedOperand(*_Begin, AMDGPU::OpName::tok)) &&
+             "first begin does not have an undefined input token as expected");
+      assert(TokMO &&
+             "Unable to extract tok operand from SI_WATERFALL_BEGIN pseudo op");
+
+      BeginList.push_back(_Begin);
+      TokReg = TokMO->getReg();
+    }
+
+    WaterfallWorkitem(const SIInstrInfo *_TII, MachineRegisterInfo *_MRI)
+        : TII(_TII), MRI(_MRI), TokReg(AMDGPU::NoRegister), Final(nullptr) {}
+
+    void processCandidate(MachineInstr *Cand) {
+      unsigned Opcode = Cand->getOpcode();
+      // Trivially end any waterfall intrinsic instructions
+      if (getWFBeginSize(Opcode) || getWFRFLSize(Opcode) ||
+          getWFEndSize(Opcode) || getWFLastUseSize(Opcode) ||
+          isWFLoopEnd(Opcode)) {
+        // TODO: A new waterfall clause shouldn't overlap with any uses
+        // tagged by a last_use intrinsic
+        return;
+      }
+
+      // Iterate over the LastUseList to determine if this instruction has a
+      // later use of a tagged last_use
+      for (auto Use : LastUseList) {
+        auto UseMO = TII->getNamedOperand(*Use, AMDGPU::OpName::dst);
+        Register UseReg = UseMO->getReg();
+
+        if (Cand->findRegisterUseOperand(UseReg, /*TRI=*/nullptr))
+          Final = Cand;
+      }
+    }
+
+    MachineInstr *getDefInstr(const MachineOperand *MO) const {
+      if (MO->isReg() && MRI->hasOneDef(MO->getReg())) {
+        return (*MRI->def_begin(MO->getReg())).getParent();
+      }
+      return nullptr;
+    }
+
+    bool tokIsStart(const MachineOperand *MO) const {
+      MachineInstr *defInstr = getDefInstr(MO);
+      if (defInstr && defInstr->getOpcode() == AMDGPU::S_MOV_B32) {
+        auto CopySrcOp = TII->getNamedOperand(*defInstr, AMDGPU::OpName::src0);
+        if (CopySrcOp && CopySrcOp->isImm()) {
+          if (CopySrcOp->getImm() == 0)
+            return true;
+        }
+      }
+      return false;
+    }
+
+    bool addCandidate(MachineInstr *Cand) {
+      unsigned Opcode = Cand->getOpcode();
+
+      assert((getWFBeginSize(Opcode) || getWFRFLSize(Opcode) ||
+              getWFEndSize(Opcode) || getWFLastUseSize(Opcode) ||
+              isWFLoopEnd(Opcode)) &&
+             "expected a waterfall instruction in addCandidate");
+
+      auto CandTokMO = TII->getNamedOperand(*Cand, AMDGPU::OpName::tok);
+      // There are a couple of scenarios at this point:
+      // 1. Standard - there's already been a begin that's been processed and
+      //               set up the WaterfallWorkItem. In which case the token is
+      //               valid and needs to be checked to ensure well- formed
+      //               waterfall groups.
+      // 2. Begins removed - begins were uniform - and all of them have been
+      //               removed. Need to process the rest of the instructions
+      //               in the group, and verify that they have a undefined
+      //               token
+      if (TokReg == AMDGPU::NoRegister) {
+        // All begins have been removed - continue to process the rest of the
+        // grouping ready for them to be removed in the next stage
+        assert(!getWFBeginSize(Opcode) &&
+               "unexpected begin instruction for addCandidate");
+        assert(tokIsStart(CandTokMO) &&
+               "waterfall group with no begin doesn't have undef tok input");
+
+        TokReg = CandTokMO->getReg();
+      }
+      if (CandTokMO->getReg() == TokReg) {
+        if (getWFBeginSize(Opcode)) {
+          auto TokRetMO = TII->getNamedOperand(*Cand, AMDGPU::OpName::tok_ret);
+          assert(TokRetMO && "Unable to extract tok_ret operand from "
+                             "SI_WATERFALL_BEGIN pseudo op");
+          BeginList.push_back(Cand);
+          TokReg = TokRetMO->getReg();
+          return true;
+        } else if (getWFRFLSize(Opcode)) {
+          RFLList.push_back(Cand);
+          return true;
+        } else if (getWFEndSize(Opcode)) {
+          EndList.push_back(Cand);
+          Final = Cand;
+          return true;
+        } else if (getWFLastUseSize(Opcode)) {
+          LastUseList.push_back(Cand);
+          if (isWFLastUseVGPR(Opcode))
+            hasVGPRLastUse = true;
+          return true;
+        } else if (isWFLoopEnd(Opcode)) {
+          LoopEndList.push_back(Cand);
+          return true;
+        } else {
+          report_fatal_error("Unknown opcode, expected waterfall intrinsic");
+        }
+      }
+      LLVM_DEBUG(dbgs() << "malformed waterfall instruction group");
+      return false;
+    }
+
+    void eraseFromParent() {
+      for (auto BeginMI : BeginList)
+        BeginMI->eraseFromParent();
+      for (auto RFLMI : RFLList)
+        RFLMI->eraseFromParent();
+      for (auto EndMI : EndList)
+        EndMI->eraseFromParent();
+      for (auto LUMI : LastUseList)
+        LUMI->eraseFromParent();
+      for (auto LEMI : LoopEndList)
+        LEMI->eraseFromParent();
+    }
+  };
+
+  std::vector<WaterfallWorkitem> Worklist;
+
+  const GCNSubtarget *ST;
+  const SIInstrInfo *TII;
+  MachineRegisterInfo *MRI;
+  const SIRegisterInfo *RI;
+
+public:
+  static char ID;
+
+  SIInsertWaterfall() : MachineFunctionPass(ID) {
+    initializeSIInsertWaterfallPass(*PassRegistry::getPassRegistry());
+  }
+
+  void getAnalysisUsage(AnalysisUsage &AU) const override {
+    MachineFunctionPass::getAnalysisUsage(AU);
+  }
+
+  bool removeRedundantWaterfall(WaterfallWorkitem &Item);
+  bool processWaterfall(MachineBasicBlock &MBB);
+
+  Register getToken(MachineInstr *MI);
+
+  bool runOnMachineFunction(MachineFunction &MF) override;
+};
+
+} // End anonymous namespace.
+
+INITIALIZE_PASS(SIInsertWaterfall, DEBUG_TYPE, "SI Insert waterfalls", false,
+                false)
+
+char SIInsertWaterfall::ID = 0;
+
+char &llvm::SIInsertWaterfallID = SIInsertWaterfall::ID;
+
+FunctionPass *llvm::createSIInsertWaterfallPass() {
+  return new SIInsertWaterfall;
+}
+
+PreservedAnalyses
+SIInsertWaterfallPass::run(MachineFunction &MF,
+                           MachineFunctionAnalysisManager &MFAM) {
+  SIInsertWaterfall Impl;
+  if (!Impl.runOnMachineFunction(MF))
+    return PreservedAnalyses::all();
+  return PreservedAnalyses::none();
+}
+
+bool SIInsertWaterfall::removeRedundantWaterfall(WaterfallWorkitem &Item) {
+  // In some cases, the waterfall is actually redundant
+  // If all the readfirstlane intrinsics are actually for uniform values and
+  // the token used in the begin/end isn't used in anything else the waterfall
+  // can be removed.
+  // Alternatively, prior passes may have removed the readfirstlane intrinsics
+  // altogether, in this case the begin/end intrinsics are now redundant and can
+  // also be removed.
+  // The readfirstlane intrinsics are replaced with the uniform source value,
+  // the loop is removed and the defs in the end intrinsics are just replaced
+  // with the input operands
+  // We can also have cases where the begins are all removed (all the indices
+  // were actually uniform).
+
+  // First step is to identify any readfirstlane intrinsics that are actually
+  // uniform - unless there are no begin instructions, in which case we always
+  // remove
+  bool LoopRemoved = false;
+  unsigned Removed = 0;
+  std::vector<MachineInstr *> NewRFLList;
+  std::vector<MachineInstr *> ToRemoveRFLList;
+
+  for (auto RFLMI : Item.RFLList) {
+    auto RFLSrcOp = TII->getNamedOperand(*RFLMI, AMDGPU::OpName::src);
+    auto RFLDstOp = TII->getNamedOperand(*RFLMI, AMDGPU::OpName::dst);
+    Register RFLSrcReg = RFLSrcOp->getReg();
+    Register RFLDstReg = RFLDstOp->getReg();
+
+    Register ReplaceReg = getUniformOperandReplacementReg(MRI, RI, RFLSrcReg);
+    if (ReplaceReg != AMDGPU::NoRegister) {
+      MRI->replaceRegWith(RFLDstReg, ReplaceReg);
+      Removed++;
+      ToRemoveRFLList.push_back(RFLMI);
+    } else if (RFLDstOp->isDead()) {
+      Removed++;
+      ToRemoveRFLList.push_back(RFLMI);
+    } else {
+      NewRFLList.push_back(RFLMI);
+    }
+  }
+
+  // Note: this test also returns true when there are NO RFL intrinsics, the
+  // case where a prior pass has removed all of them and the loop is now
+  // redundant
+  // Also check for the special case where there are no RFL intrinsics, but
+  // there are some last.use with VGPR uses.
+  if (!Item.BeginList.size() ||
+      (Removed == Item.RFLList.size() && !Item.hasVGPRLastUse)) {
+    // Removed all of the RFLs
+    // We can remove the waterfall loop entirely
+
+    // Protocol is to replace all dst operands for the waterfall_end intrinsics
+    // with their src operands. Replace all last_use dst operands with their src
+    // operands We don't need to check that the loop index isn't used anywhere
+    // as the protocol for waterfall intrinsics is to only use the begin index
+    // via a readfirstlane intrinsic anyway (which should also be removed) Any
+    // problems due to errors in this pass around loop removal will be picked up
+    // later by e.g. use before def errors
+    LLVM_DEBUG(
+        dbgs()
+        << "detected case for waterfall loop removal - already all uniform\n");
+    for (auto EndMI : Item.EndList) {
+      auto EndDstOp = TII->getNamedOperand(*EndMI, AMDGPU::OpName::dst);
+      auto EndSrcOp = TII->getNamedOperand(*EndMI, AMDGPU::OpName::src);
+      replaceRegIncSubReg(MRI, RI, EndDstOp, EndSrcOp);
+    }
+    for (auto LUMI : Item.LastUseList) {
+      auto LUDstOp = TII->getNamedOperand(*LUMI, AMDGPU::OpName::dst);
+      auto LUSrcOp = TII->getNamedOperand(*LUMI, AMDGPU::OpName::src);
+      replaceRegIncSubReg(MRI, RI, LUDstOp, LUSrcOp);
+    }
+    // If all the begins were removed, we have to replace the RFL with actual
+    // RFL, these will show up in the NewRLFList
+    for (auto RFLMI : NewRFLList) {
+      auto DstReg = TII->getNamedOperand(*RFLMI, AMDGPU::OpName::dst)->getReg();
+      auto SrcOp = TII->getNamedOperand(*RFLMI, AMDGPU::OpName::src);
+      Register SrcReg = SrcOp->getReg();
+
+      MachineBasicBlock::iterator RFLInsert(RFLMI);
+      readFirstLaneReg(*RFLMI->getParent(), MRI, RI, TII, RFLInsert,
+                       RFLMI->getDebugLoc(), DstReg, SrcReg, *SrcOp);
+    }
+
+    Item.eraseFromParent();
+
+    LoopRemoved = true;
+
+  } else if (Removed) {
+    LLVM_DEBUG(dbgs() << "Removed " << Removed
+                      << " waterfall rfl intrinsics due to being uniform - "
+                         "updating remaining rfl list\n");
+    // TODO: there's an opportunity to pull the DAG involving the (removed) rfls
+    // out of the waterfall loop
+    Item.RFLList.clear();
+    std::copy(NewRFLList.begin(), NewRFLList.end(),
+              std::back_inserter(Item.RFLList));
+    for (auto RFLMI : ToRemoveRFLList)
+      RFLMI->eraseFromParent();
+  }
+
+  return LoopRemoved;
+}
+
+bool SIInsertWaterfall::processWaterfall(MachineBasicBlock &MBB) {
+  bool Changed = false;
+  MachineFunction &MF = *MBB.getParent();
+  MachineBasicBlock *CurrMBB = &MBB;
+
+  // Firstly we check that there are at least 3 related waterfall instructions
+  // for this begin
+  // SI_WATERFALL_BEGIN [ SI_WATERFALL_BEGIN ]*
+  // [ SI_WATERFALL_READFIRSTLANE ]+ [ SI_WATERFALL_END ]+ If there are multiple
+  // waterfall loops they must also be disjoint
+
+  for (WaterfallWorkitem &Item : Worklist) {
+    LLVM_DEBUG(
+        if (Item.BeginList.size()) {
+          dbgs() << "Processing " << *Item.BeginList[0] << "\n";
+
+          for (auto RUse = MRI->use_begin(Item.TokReg), RSE = MRI->use_end();
+               RUse != RSE; ++RUse) {
+            MachineInstr *RUseMI = RUse->getParent();
+            assert((CurrMBB->getNumber() == RUseMI->getParent()->getNumber()) &&
+                   "Linked WATERFALL pseudo ops found in different BBs");
+          }
+        } else { dbgs() << "Processing redundant waterfall\n"; });
+
+    if (removeRedundantWaterfall(Item)) {
+      Changed = true;
+      continue;
+    }
+
+    assert((Item.RFLList.size() || Item.hasVGPRLastUse) &&
+           (Item.EndList.size() || Item.LastUseList.size()) &&
+           "SI_WATERFALL* pseudo instruction group must have at least 1 of "
+           "each type");
+
+    // Insert the waterfall loop code around the identified region of
+    // instructions
+    // Loop starts at the last SI_WATERFALL_BEGIN
+    // SI_WATERFALL_READFIRSTLANE is replaced with appropriate readfirstlane
+    // instructions OR is removed
+    // if the readfirstlane is using the same index as the SI_WATERFALL_BEGIN
+    // Loop is ended after the last SI_WATERFALL_END and these instructions are
+    // removed with the src replacing all dst uses
+    typedef struct {
+      const MachineOperand *Index;
+      const TargetRegisterClass *IndexRC;
+      const TargetRegisterClass *IndexSRC;
+      Register CurrentIdxReg;
+    } IdxInfo;
+
+    std::vector<IdxInfo> IndexList;
+#ifndef NDEBUG
+    bool IsUniform = true;
+#endif
+    for (auto BeginMI : Item.BeginList) {
+      IdxInfo CurrIdx;
+      CurrIdx.Index = TII->getNamedOperand(*(BeginMI), AMDGPU::OpName::idx);
+      CurrIdx.IndexRC = RI->getRegClassForOperandReg(*MRI, *CurrIdx.Index);
+      CurrIdx.IndexSRC = RI->getEquivalentSGPRClass(CurrIdx.IndexRC);
+      if (CurrIdx.IndexSRC == &AMDGPU::SGPR_32RegClass)
+        CurrIdx.IndexSRC = &AMDGPU::SReg_32_XM0RegClass;
+
+      IndexList.push_back(CurrIdx);
+
+      LLVM_DEBUG(if (RI->hasVGPRs(CurrIdx.IndexRC)) IsUniform = false;);
+    }
+
+    LLVM_DEBUG(if (IsUniform) {
+      // Waterfall loop index is uniform! Loop can be removed
+      // TODO:: Implement loop removal
+      dbgs() << "Uniform loop detected - waterfall loop is redundant\n";
+    });
+
+    MachineBasicBlock::iterator I(Item.BeginList.back());
+    const DebugLoc &DL = Item.BeginList[0]->getDebugLoc();
+
+    // Initialize the register we accumulate the result into, which is the
+    // target of any SI_WATERFALL_END instruction
+    for (auto EndMI : Item.EndList)
+      Item.EndRegs.emplace_back(
+          TII->getNamedOperand(*EndMI, AMDGPU::OpName::dst),
+          TII->getNamedOperand(*EndMI, AMDGPU::OpName::src));
+    for (auto LUMI : Item.LastUseList) {
+      auto LUSrc = TII->getNamedOperand(*LUMI, AMDGPU::OpName::src);
+      auto LUDst = TII->getNamedOperand(*LUMI, AMDGPU::OpName::dst);
+      replaceRegIncSubReg(MRI, RI, LUDst, LUSrc);
+    }
+
+    // EXEC mask handling
+    Register Exec = ST->isWave32() ? AMDGPU::EXEC_LO : AMDGPU::EXEC;
+    unsigned SaveExecOpc = ST->isWave32() ? AMDGPU::S_AND_SAVEEXEC_B32
+                                          : AMDGPU::S_AND_SAVEEXEC_B64;
+    unsigned XorTermOpc =
+        ST->isWave32() ? AMDGPU::S_XOR_B32_term : AMDGPU::S_XOR_B64_term;
+    unsigned MovOpc = ST->isWave32() ? AMDGPU::S_MOV_B32 : AMDGPU::S_MOV_B64;
+    unsigned AndNotWRExecOpc = ST->isWave32() ? AMDGPU::S_ANDN2_WREXEC_B32
+                                              : AMDGPU::S_ANDN2_WREXEC_B64;
+
+    // Emit [v_cmpx_eq] and [s_andn2_wwrexec] when these instructions are
+    // available.
+    // TODO: Accurately detect the availability of [s_andn2_wrexec] instruction
+    // in the target. For now, use the same condition as for the detection
+    // [v_cmpx_eq].
+    auto UseNewExecInstructions = ST->hasNoSdstCMPX();
+
+    const auto *BoolXExecRC = TII->getRegisterInfo().getWaveMaskRegClass();
+
+    MachineBasicBlock &LoopHeaderBB = *MF.CreateMachineBasicBlock();
+    MachineBasicBlock &LoopBB = *MF.CreateMachineBasicBlock();
+    MachineBasicBlock &RemainderBB = *MF.CreateMachineBasicBlock();
+    MachineFunction::iterator MBBI(*CurrMBB);
+    ++MBBI;
+
+    MF.insert(MBBI, &LoopHeaderBB);
+    MF.insert(MBBI, &LoopBB);
+    MF.insert(MBBI, &RemainderBB);
+
+    LoopHeaderBB.addSuccessor(&LoopBB);
+    LoopBB.addSuccessor(&LoopBB);
+    LoopBB.addSuccessor(&RemainderBB);
+
+    Register SaveExec = MRI->createVirtualRegister(BoolXExecRC);
+    Register TmpExec = MRI->createVirtualRegister(BoolXExecRC);
+
+    // Put TmpExec and SaveExec in the loop header.
+    MachineBasicBlock::iterator LH = LoopHeaderBB.begin();
+
+    if (UseNewExecInstructions) {
+      // Initialize TmpExec with the current EXEC mask.
+      // Represents remaining threads to process.
+      BuildMI(LoopHeaderBB, LH, DL, TII->get(MovOpc), TmpExec).addReg(Exec);
+    } else {
+      BuildMI(LoopHeaderBB, LH, DL, TII->get(TargetOpcode::IMPLICIT_DEF),
+              TmpExec);
+    }
+
+    // Save the EXEC mask
+    BuildMI(LoopHeaderBB, LH, DL, TII->get(MovOpc), SaveExec).addReg(Exec);
+
+    // Move all instructions from the SI_WATERFALL_BEGIN to the last
+    // SI_WATERFALL_END or last use tagged from SI_WATERFALL_LAST_USE
+    // into the new LoopBB
+    MachineBasicBlock::iterator SpliceE(Item.Final);
+    ++SpliceE;
+    LoopBB.splice(LoopBB.begin(), CurrMBB, I, SpliceE);
+
+    // Iterate over the instructions inserted into the loop
+    // Need to unset any kill flag on any uses as now this is a loop that is no
+    // longer valid
+    for (MachineInstr &MI : LoopBB)
+      MI.clearKillInfo();
+
+    RemainderBB.transferSuccessorsAndUpdatePHIs(CurrMBB);
+    RemainderBB.splice(RemainderBB.begin(), CurrMBB, SpliceE, CurrMBB->end());
+    MachineBasicBlock::iterator E(Item.Final);
+    ++E;
+
+    CurrMBB->addSuccessor(&LoopHeaderBB);
+
+    MachineBasicBlock::iterator J = LoopBB.begin();
+
+    Register PhiExec = MRI->createVirtualRegister(BoolXExecRC);
+    Register NewExec = MRI->createVirtualRegister(BoolXExecRC);
+
+    for (auto &CurrIdx : IndexList)
+      CurrIdx.CurrentIdxReg = MRI->createVirtualRegister(CurrIdx.IndexSRC);
+
+    BuildMI(LoopBB, J, DL, TII->get(TargetOpcode::PHI), PhiExec)
+        .addReg(TmpExec)
+        .addMBB(&LoopHeaderBB)
+        .addReg(NewExec)
+        .addMBB(&LoopBB);
+
+    // Get the next index to use from the first enabled lane
+    for (auto &CurrIdx : IndexList)
+      readFirstLaneReg(LoopBB, MRI, RI, TII, J, DL, CurrIdx.CurrentIdxReg,
+                       CurrIdx.Index->getReg(), *CurrIdx.Index);
+
+    // Also process the readlane pseudo ops - if readfirstlane is using the
+    // index then just replace with the CurrentIdxReg instead
+    for (auto RFLMI : Item.RFLList) {
+      auto RFLSrcOp = TII->getNamedOperand(*RFLMI, AMDGPU::OpName::src);
+      auto RFLDstOp = TII->getNamedOperand(*RFLMI, AMDGPU::OpName::dst);
+      Register RFLSrcReg = RFLSrcOp->getReg();
+      Register RFLDstReg = RFLDstOp->getReg();
+
+      bool MatchedIdx = false;
+      for (auto &CurrIdx : IndexList) {
+        if (RFLSrcReg == CurrIdx.Index->getReg()) {
+          // Use the CurrentIdxReg for this
+          Item.RFLRegs.push_back(CurrIdx.CurrentIdxReg);
+          MRI->replaceRegWith(RFLDstReg, CurrIdx.CurrentIdxReg);
+          MatchedIdx = true;
+          break;
+        }
+      }
+      if (!MatchedIdx) {
+        Item.RFLRegs.push_back(RFLDstReg);
+        // Insert function to expand to required size here
+        MachineBasicBlock::iterator RFLInsert(RFLMI);
+        readFirstLaneReg(LoopBB, MRI, RI, TII, RFLInsert, DL, RFLDstReg,
+                         RFLSrcReg, *RFLSrcOp);
+      }
+    }
+
+    // Compare the just read idx value to all possible idx values, and update
+    // EXEC
+    if (UseNewExecInstructions) {
+      for (auto &CurrIdx : IndexList)
+        compareIdxUsingCmpx(LoopBB, MRI, RI, TII, J, DL, CurrIdx.CurrentIdxReg,
+                            *CurrIdx.Index, ST->isWave32());
+      MRI->setSimpleHint(NewExec, PhiExec);
+    } else {
+      Register CondReg = AMDGPU::NoRegister;
+      for (auto &CurrIdx : IndexList)
+        CondReg = compareIdx(LoopBB, MRI, RI, TII, J, DL, CurrIdx.CurrentIdxReg,
+                             *CurrIdx.Index, CondReg, ST->isWave32());
+
+      // Update EXEC, save the original EXEC value to VCC
+      BuildMI(LoopBB, J, DL, TII->get(SaveExecOpc), NewExec)
+          .addReg(CondReg, RegState::Kill);
+
+      MRI->setSimpleHint(NewExec, CondReg);
+    }
+    // TODO: Conditional branch here to loop header as potential optimization?
+
+    // Copy the just read value into the destination
+    // Handle cases where sub registers are involved
+    for (auto EndReg : Item.EndRegs) {
+      MachineBasicBlock::iterator EndInsert(Item.Final);
+      BuildMI(LoopBB, EndInsert, DL, TII->get(AMDGPU::COPY))
+          .addReg(EndReg.first->getReg(), RegState::Define,
+                  EndReg.first->getSubReg())
+          .addReg(EndReg.second->getReg(), {}, EndReg.second->getSubReg());
+    }
+
+    // Update EXEC, switch all done bits to 0 and all todo bits to 1.
+    if (UseNewExecInstructions) {
+      BuildMI(LoopBB, E, DL, TII->get(AndNotWRExecOpc), NewExec)
+          .addReg(PhiExec);
+    } else {
+      BuildMI(LoopBB, E, DL, TII->get(XorTermOpc), Exec)
+          .addReg(Exec)
+          .addReg(NewExec);
+    }
+
+    // Loop back if there are still variants to cover
+    BuildMI(LoopBB, E, DL, TII->get(AMDGPU::SI_WATERFALL_LOOP)).addMBB(&LoopBB);
+
+    MachineBasicBlock::iterator First = RemainderBB.begin();
+    BuildMI(RemainderBB, First, DL, TII->get(MovOpc), Exec).addReg(SaveExec);
+
+    Item.eraseFromParent();
+
+    // To process subsequent waterfall groups, update CurrMBB to the RemainderBB
+    CurrMBB = &RemainderBB;
+
+    Changed = true;
+  }
+  return Changed;
+}
+
+Register SIInsertWaterfall::getToken(MachineInstr *MI) {
+  auto CandTokMO = TII->getNamedOperand(*MI, AMDGPU::OpName::tok);
+  return CandTokMO->isReg() ? CandTokMO->getReg() : AMDGPU::NoRegister;
+}
+
+bool SIInsertWaterfall::runOnMachineFunction(MachineFunction &MF) {
+  bool Changed = false;
+
+  ST = &MF.getSubtarget<GCNSubtarget>();
+  TII = ST->getInstrInfo();
+  MRI = &MF.getRegInfo();
+  RI = ST->getRegisterInfo();
+
+  for (MachineBasicBlock &MBB : MF) {
+    Worklist.clear();
+    bool StartNew = true;
+
+    for (MachineInstr &MI : MBB) {
+      unsigned Opcode = MI.getOpcode();
+
+      if (getWFBeginSize(Opcode)) {
+        if (StartNew) {
+          Worklist.push_back(WaterfallWorkitem(&MI, TII, MRI));
+          StartNew = false;
+        } else {
+          if (!Worklist.back().addCandidate(&MI)) {
+            llvm_unreachable("Incorrect SI_WATERFALL_* groups");
+          }
+        }
+      } else if (getWFRFLSize(Opcode) || getWFEndSize(Opcode) ||
+                 getWFLastUseSize(Opcode) || isWFLoopEnd(Opcode)) {
+        // On to the body of the group intrinsics,
+
+        // Tag StartNew as true if we encounter another begin
+        StartNew = true;
+
+        if (!Worklist.size() || getToken(&MI) != Worklist.back().TokReg) {
+          // There's no associated begin for these body intrinsics
+          // That means it's either an error - or all the begin intrinsics
+          // were removed due to being uniform
+          // Set up a WorkItem so we can process this correctly
+          Worklist.push_back(WaterfallWorkitem(TII, MRI));
+        }
+
+        if (!Worklist.back().addCandidate(&MI)) {
+          llvm_unreachable("Overlapping SI_WATERFALL_* groups");
+        }
+      } else {
+        if (Worklist.size())
+          Worklist.back().processCandidate(&MI);
+      }
+    }
+    Changed |= processWaterfall(MBB);
+  }
+
+  return Changed;
+}
diff --git a/llvm/lib/Target/AMDGPU/SIInsertWaterfall.h b/llvm/lib/Target/AMDGPU/SIInsertWaterfall.h
new file mode 100644
index 0000000000000..406df5a8b3b67
--- /dev/null
+++ b/llvm/lib/Target/AMDGPU/SIInsertWaterfall.h
@@ -0,0 +1,22 @@
+//===- SIInsertWaterfall.h --------------------------------------*- C++ -*-===//
+//
+// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
+// See https://llvm.org/LICENSE.txt for license information.
+// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
+//
+//===----------------------------------------------------------------------===//
+
+#ifndef LLVM_LIB_TARGET_AMDGPU_SIINSERTWATERFALL_H
+#define LLVM_LIB_TARGET_AMDGPU_SIINSERTWATERFALL_H
+
+#include "llvm/CodeGen/MachinePassManager.h"
+
+namespace llvm {
+class SIInsertWaterfallPass : public PassInfoMixin<SIInsertWaterfallPass> {
+public:
+  PreservedAnalyses run(MachineFunction &MF,
+                        MachineFunctionAnalysisManager &MFAM);
+};
+} // namespace llvm
+
+#endif // LLVM_LIB_TARGET_AMDGPU_SIINSERTWATERFALL_H
diff --git a/llvm/lib/Target/AMDGPU/SIInstructions.td b/llvm/lib/Target/AMDGPU/SIInstructions.td
index 008fdbaddfd22..c927eb85300eb 100644
--- a/llvm/lib/Target/AMDGPU/SIInstructions.td
+++ b/llvm/lib/Target/AMDGPU/SIInstructions.td
@@ -1377,6 +1377,60 @@ def : GCNPat <
   (COPY $src) // Return the SGPRs representing i1 src
 >;
 
+// Waterfall pseudo ops - these are effectively specialized control flow pseudo
+// instructions
+// They are tagged as defining SCC since the loop that gets inserted will usually clobber it
+let UseNamedOperandTable = 1, Defs = [SCC] in {
+
+class SI_WATERFALL_BEGIN<RegisterClass idx_rc> : SPseudoInstSI <
+  (outs SReg_32:$tok_ret),
+  (ins SReg_32:$tok, idx_rc:$idx)>;
+
+class SI_WATERFALL_READFIRSTLANE<RegisterClass dst_rc, RegisterClass src_rc> : SPseudoInstSI <
+  (outs dst_rc:$dst),
+  (ins SReg_32:$tok, src_rc:$src)>;
+
+class SI_WATERFALL_END<RegisterClass rc> : SPseudoInstSI <
+  (outs rc:$dst),
+  (ins SReg_32:$tok, rc:$src)>;
+
+class SI_WATERFALL_LAST_USE<RegisterClass rc> : SPseudoInstSI <
+  (outs rc:$dst),
+  (ins SReg_32:$tok, rc:$src)>;
+
+} // End UseNamedOperandTable = 1
+
+def SI_WATERFALL_BEGIN_V1 : SI_WATERFALL_BEGIN<VGPR_32>;
+def SI_WATERFALL_BEGIN_V2 : SI_WATERFALL_BEGIN<VReg_64>;
+def SI_WATERFALL_BEGIN_V4 : SI_WATERFALL_BEGIN<VReg_128>;
+def SI_WATERFALL_BEGIN_V8 : SI_WATERFALL_BEGIN<VReg_256>;
+
+def SI_WATERFALL_READFIRSTLANE_V1 : SI_WATERFALL_READFIRSTLANE<SReg_32_XM0, VGPR_32>;
+def SI_WATERFALL_READFIRSTLANE_V2 : SI_WATERFALL_READFIRSTLANE<SReg_64, VReg_64>;
+def SI_WATERFALL_READFIRSTLANE_V4 : SI_WATERFALL_READFIRSTLANE<SGPR_128, VReg_128>;
+def SI_WATERFALL_READFIRSTLANE_V8 : SI_WATERFALL_READFIRSTLANE<SReg_256, VReg_256>;
+
+def SI_WATERFALL_END_V1 : SI_WATERFALL_END<VGPR_32>;
+def SI_WATERFALL_END_V2 : SI_WATERFALL_END<VReg_64>;
+def SI_WATERFALL_END_V4 : SI_WATERFALL_END<VReg_128>;
+def SI_WATERFALL_END_V8 : SI_WATERFALL_END<VReg_256>;
+
+def SI_WATERFALL_LAST_USE_V1   : SI_WATERFALL_LAST_USE<SGPR_32>;
+def SI_WATERFALL_LAST_USE_V1_V : SI_WATERFALL_LAST_USE<VGPR_32>;
+def SI_WATERFALL_LAST_USE_V2   : SI_WATERFALL_LAST_USE<SReg_64>;
+def SI_WATERFALL_LAST_USE_V2_V : SI_WATERFALL_LAST_USE<VReg_64>;
+def SI_WATERFALL_LAST_USE_V4   : SI_WATERFALL_LAST_USE<SGPR_128>;
+def SI_WATERFALL_LAST_USE_V4_V : SI_WATERFALL_LAST_USE<SReg_128>;
+def SI_WATERFALL_LAST_USE_V8   : SI_WATERFALL_LAST_USE<SReg_256>;
+def SI_WATERFALL_LAST_USE_V8_V : SI_WATERFALL_LAST_USE<VReg_256>;
+
+let Defs = [SCC],
+    UseNamedOperandTable = 1 in
+def SI_WATERFALL_LOOP_END : SPseudoInstSI <
+  (outs),
+  (ins SReg_32:$tok)
+>;
+
 //===----------------------------------------------------------------------===//
 // VOP1 Patterns
 //===----------------------------------------------------------------------===//
@@ -2674,6 +2728,59 @@ def : GCNPat <
                    (V_RCP_IFLAG_F32_e32 (V_CVT_F32_U32_e32 $src0))))
 >;
 
+multiclass SI_WATERFALL_Pattern<ValueType dvt, ValueType svt, string VecSize> {
+  def : GCNPat<(dvt(int_amdgcn_waterfall_readfirstlane i32:$tok, svt:$src)),
+               (!cast<Instruction>("SI_WATERFALL_READFIRSTLANE_"#VecSize)
+                       i32:$tok,
+                   svt:$src)>;
+}
+
+multiclass SI_WATERFALL_S_Pattern<list<ValueType> vts, string VecSize> {
+  foreach vt = vts in {
+    def : GCNPat<(i32(int_amdgcn_waterfall_begin i32:$tok, vt:$idx)),
+                 (!cast<Instruction>("SI_WATERFALL_BEGIN_"#VecSize) i32:$tok,
+                     vt:$idx)>;
+
+    def : GCNPat<(vt(int_amdgcn_waterfall_readfirstlane i32:$tok, vt:$src)),
+                 (!cast<Instruction>("SI_WATERFALL_READFIRSTLANE_"#VecSize)
+                         i32:$tok,
+                     vt:$src)>;
+
+    def : GCNPat<(vt(int_amdgcn_waterfall_end i32:$tok, vt:$src)),
+                 (!cast<Instruction>("SI_WATERFALL_END_"#VecSize) i32:$tok,
+                     vt:$src)>;
+
+    def : GCNPat<(vt(int_amdgcn_waterfall_last_use i32:$tok, vt:$src)),
+                 (!cast<Instruction>("SI_WATERFALL_LAST_USE_"#VecSize) i32:$tok,
+                     vt:$src)>;
+
+    def : GCNPat<(vt(int_amdgcn_waterfall_last_use_vgpr i32:$tok, vt:$src)),
+                 (!cast<Instruction>("SI_WATERFALL_LAST_USE_"#VecSize#"_V")
+                         i32:$tok,
+                     vt:$src)>;
+  }
+}
+
+defm : SI_WATERFALL_S_Pattern<Reg16Types.types, "V1">;
+defm : SI_WATERFALL_S_Pattern<Reg32Types.types, "V1">;
+defm : SI_WATERFALL_S_Pattern<Reg64Types.types, "V2">;
+defm : SI_WATERFALL_S_Pattern<Reg128Types.types, "V4">;
+defm : SI_WATERFALL_S_Pattern<Reg256Types.types, "V8">;
+
+defm : SI_WATERFALL_Pattern <i16, f16, "V1">;
+defm : SI_WATERFALL_Pattern <v2i16, v2f16, "V1">;
+defm : SI_WATERFALL_Pattern <v4i16, v4f16, "V2">;
+
+defm : SI_WATERFALL_Pattern <i32, f32, "V1">;
+defm : SI_WATERFALL_Pattern <v2i32, v2f32, "V2">;
+defm : SI_WATERFALL_Pattern <v4i32, v4f32, "V4">;
+defm : SI_WATERFALL_Pattern <v8i32, v8f32, "V8">;
+
+def : GCNPat<
+    (int_amdgcn_waterfall_loop_end i32:$tok),
+    (SI_WATERFALL_LOOP_END i32:$tok)
+>;
+
 //===----------------------------------------------------------------------===//
 // VOP3 Patterns
 //===----------------------------------------------------------------------===//
diff --git a/llvm/test/Analysis/UniformityAnalysis/AMDGPU/MIR/llvm.amdgcn.waterfall.mir b/llvm/test/Analysis/UniformityAnalysis/AMDGPU/MIR/llvm.amdgcn.waterfall.mir
new file mode 100644
index 0000000000000..930d926b76b79
--- /dev/null
+++ b/llvm/test/Analysis/UniformityAnalysis/AMDGPU/MIR/llvm.amdgcn.waterfall.mir
@@ -0,0 +1,51 @@
+# RUN: llc -mtriple=amdgcn-- -mcpu=gfx1010 -run-pass=print-machine-uniformity -o - %s 2>&1 | FileCheck %s
+
+# CHECK-LABEL: MachineUniformityInfo for function: @test_waterfall_readlane
+# CHECK-NOT: DIVERGENT: %0: %0:_(s32) = COPY $sgpr0
+# CHECK-NOT: DIVERGENT: %1: %1:_(s32) = COPY $sgpr1
+# CHECK-NOT: DIVERGENT: %2: %2:_(p1) = G_MERGE_VALUES %0:_(s32), %1:_(s32)
+# CHECK-NOT: DIVERGENT: %3: %3:_(s32) = COPY $sgpr2
+# CHECK-NOT: DIVERGENT: %4: %4:_(s32) = COPY $sgpr3
+# CHECK-NOT: DIVERGENT: %5: %5:_(p1) = G_MERGE_VALUES %3:_(s32), %4:_(s32)
+# CHECK:     DIVERGENT: %6: %6:_(s32) = COPY $vgpr0
+# CHECK:     DIVERGENT: %7: %7:_(s64) = G_SEXT %6:_(s32)
+# CHECK-NOT: DIVERGENT: %8: %8:_(s32) = G_CONSTANT i32 3
+# CHECK:     DIVERGENT: %9: %9:_(s64) = G_SHL %7:_, %8:_(s32)
+# CHECK:     DIVERGENT: %10: %10:_(p1) = G_PTR_ADD %5:_, %9:_(s64)
+# CHECK:     DIVERGENT: %11: %11:_(<2 x s32>) = G_LOAD %10:_(p1) :: (load (<2 x s32>), addrspace 1)
+# CHECK-NOT: DIVERGENT: %12: %12:_(s32) = G_CONSTANT i32 0
+# CHECK:     DIVERGENT: %13: %13:_(s32), %14:_(s32) = G_UNMERGE_VALUES %11:_(<2 x s32>)
+# CHECK:     DIVERGENT: %14: %13:_(s32), %14:_(s32) = G_UNMERGE_VALUES %11:_(<2 x s32>)
+# CHECK-NOT: DIVERGENT:  %15: %15:_(s32) = G_INTRINSIC_CONVERGENT_W_SIDE_EFFECTS intrinsic(@llvm.amdgcn.waterfall.begin), %12:_(s32), %14:_(s32)
+# CHECK-NOT: DIVERGENT:  %16: %16:_(s32) = G_INTRINSIC_CONVERGENT_W_SIDE_EFFECTS intrinsic(@llvm.amdgcn.waterfall.readfirstlane), %15:_(s32), %14:_(s32)
+# CHECK-NOT: DIVERGENT:  %17: %17:_(s32) = G_INTRINSIC_CONVERGENT intrinsic(@llvm.amdgcn.readlane), %13:_(s32), %16:_(s32)
+# CHECK:     DIVERGENT:  %18: %18:_(s32) = G_INTRINSIC_CONVERGENT_W_SIDE_EFFECTS intrinsic(@llvm.amdgcn.waterfall.end), %15:_(s32), %17:_(s32)
+
+---
+name:            test_waterfall_readlane
+tracksRegLiveness: true
+body:             |
+  bb.0:
+    liveins: $sgpr0, $sgpr1, $sgpr2, $sgpr3, $vgpr0
+
+    %3:_(s32) = COPY $sgpr0
+    %4:_(s32) = COPY $sgpr1
+    %0:_(p1) = G_MERGE_VALUES %3(s32), %4(s32)
+    %5:_(s32) = COPY $sgpr2
+    %6:_(s32) = COPY $sgpr3
+    %1:_(p1) = G_MERGE_VALUES %5(s32), %6(s32)
+    %2:_(s32) = COPY $vgpr0
+    %8:_(s64) = G_SEXT %2(s32)
+    %25:_(s32) = G_CONSTANT i32 3
+    %10:_(s64) = G_SHL %8, %25(s32)
+    %11:_(p1) = G_PTR_ADD %1, %10(s64)
+    %13:_(<2 x s32>) = G_LOAD %11(p1) :: (load (<2 x s32>), addrspace 1)
+    %15:_(s32) = G_CONSTANT i32 0
+    %23:_(s32), %24:_(s32) = G_UNMERGE_VALUES %13(<2 x s32>)
+    %18:_(s32) = G_INTRINSIC_CONVERGENT_W_SIDE_EFFECTS intrinsic(@llvm.amdgcn.waterfall.begin), %15(s32), %24(s32)
+    %19:_(s32) = G_INTRINSIC_CONVERGENT_W_SIDE_EFFECTS intrinsic(@llvm.amdgcn.waterfall.readfirstlane), %18(s32), %24(s32)
+    %20:_(s32) = G_INTRINSIC_CONVERGENT intrinsic(@llvm.amdgcn.readlane), %23(s32), %19(s32)
+    %21:_(s32) = G_INTRINSIC_CONVERGENT_W_SIDE_EFFECTS intrinsic(@llvm.amdgcn.waterfall.end), %18(s32), %20(s32)
+    G_STORE %21(s32), %0(p1) :: (store (s32), addrspace 1)
+    S_ENDPGM 0
+...
diff --git a/llvm/test/Analysis/UniformityAnalysis/AMDGPU/llvm.amdgcn.waterfall.ll b/llvm/test/Analysis/UniformityAnalysis/AMDGPU/llvm.amdgcn.waterfall.ll
new file mode 100644
index 0000000000000..7b934e85b68ab
--- /dev/null
+++ b/llvm/test/Analysis/UniformityAnalysis/AMDGPU/llvm.amdgcn.waterfall.ll
@@ -0,0 +1,29 @@
+; RUN: opt -mtriple amdgcn-- -passes='print<uniformity>' -disable-output %s 2>&1 | FileCheck %s
+
+; CHECK-LABEL: UniformityInfo for function 'test_waterfall_readlane':
+; CHECK:     DIVERGENT:   %gep.in = getelementptr <2 x i32>, ptr addrspace(1) %in, i32 %tid
+; CHECK:     DIVERGENT:   %args = load <2 x i32>, ptr addrspace(1) %gep.in, align 8
+; CHECK:     DIVERGENT:   %value = extractelement <2 x i32> %args, i32 0
+; CHECK:     DIVERGENT:   %lane = extractelement <2 x i32> %args, i32 1
+; CHECK-NOT: DIVERGENT:   %wf_token = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %lane)
+; CHECK-NOT: DIVERGENT:   %readlane = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %wf_token, i32 %lane)
+; CHECK-NOT: DIVERGENT:   %readlane1 = call i32 @llvm.amdgcn.readlane(i32 %value, i32 %readlane)
+; CHECK:     DIVERGENT:   %readlane2 = call i32 @llvm.amdgcn.waterfall.end.i32(i32 %wf_token, i32 %readlane1)
+; CHECK:     DIVERGENT:   store i32 %readlane2, ptr addrspace(1) %out, align 4
+define amdgpu_ps void @test_waterfall_readlane(i32 addrspace(1)* inreg %out, <2 x i32> addrspace(1)* inreg %in, i32 %tid) #1 {
+  %gep.in = getelementptr <2 x i32>, <2 x i32> addrspace(1)* %in, i32 %tid
+  %args = load <2 x i32>, <2 x i32> addrspace(1)* %gep.in
+  %value = extractelement <2 x i32> %args, i32 0
+  %lane = extractelement <2 x i32> %args, i32 1
+  %wf_token = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %lane)
+  %readlane = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %wf_token, i32 %lane)
+  %readlane1 = call i32 @llvm.amdgcn.readlane(i32 %value, i32 %readlane)
+  %readlane2 = call i32 @llvm.amdgcn.waterfall.end.i32(i32 %wf_token, i32 %readlane1)
+  store i32 %readlane2, i32 addrspace(1)* %out, align 4
+  ret void
+}
+
+declare i32 @llvm.amdgcn.waterfall.begin.i32(i32, i32)
+declare i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32, i32)
+declare i32 @llvm.amdgcn.readlane(i32, i32)
+declare i32 @llvm.amdgcn.waterfall.end.i32(i32, i32)
diff --git a/llvm/test/CodeGen/AMDGPU/amdgcn.waterfall.atomic.opt.ll b/llvm/test/CodeGen/AMDGPU/amdgcn.waterfall.atomic.opt.ll
new file mode 100644
index 0000000000000..2f63293aaa09c
--- /dev/null
+++ b/llvm/test/CodeGen/AMDGPU/amdgcn.waterfall.atomic.opt.ll
@@ -0,0 +1,406 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py
+; RUN: llc -march=amdgcn -mcpu=gfx1010 -mattr=-wavefrontsize32,+wavefrontsize64 -verify-machineinstrs < %s | FileCheck -check-prefixes=GFX10 %s
+; RUN: llc -march=amdgcn -mcpu=gfx1100 -mattr=-wavefrontsize32,+wavefrontsize64 -verify-machineinstrs < %s | FileCheck -check-prefixes=GFX11 %s
+
+define dllexport amdgpu_cs void @atomic_add_in_wf(ptr addrspace(1) %arg, i32 inreg %arg1, ptr addrspace(4) inreg noundef %arg2) #0 {
+; GFX10-LABEL: atomic_add_in_wf:
+; GFX10:       ; %bb.0: ; %bb
+; GFX10-NEXT:    s_ashr_i32 s3, s0, 31
+; GFX10-NEXT:    v_add_co_u32 v0, vcc, v0, s0
+; GFX10-NEXT:    v_add_co_ci_u32_e32 v1, vcc, s3, v1, vcc
+; GFX10-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-NEXT:    global_load_dword v0, v[0:1], off offset:4
+; GFX10-NEXT:  .LBB0_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-NEXT:    v_readfirstlane_b32 s0, v0
+; GFX10-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-NEXT:    v_cmpx_eq_u32_e64 s0, v0
+; GFX10-NEXT:    s_ashr_i32 s3, s0, 31
+; GFX10-NEXT:    s_add_u32 s6, s1, s0
+; GFX10-NEXT:    s_addc_u32 s7, s2, s3
+; GFX10-NEXT:    v_mov_b32_e32 v0, 1
+; GFX10-NEXT:    s_load_dwordx4 s[8:11], s[6:7], 0x0
+; GFX10-NEXT:    v_mov_b32_e32 v1, 0
+; GFX10-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-NEXT:    buffer_atomic_add v0, v1, s[8:11], 0 idxen glc
+; GFX10-NEXT:    s_andn2_wrexec_b64 s[4:5], s[4:5]
+; GFX10-NEXT:    ; implicit-def: $vgpr0
+; GFX10-NEXT:    s_cbranch_execnz .LBB0_1
+; GFX10-NEXT:  ; %bb.2:
+; GFX10-NEXT:    s_endpgm
+;
+; GFX11-LABEL: atomic_add_in_wf:
+; GFX11:       ; %bb.0: ; %bb
+; GFX11-NEXT:    v_add_co_u32 v0, vcc, v0, s0
+; GFX11-NEXT:    s_ashr_i32 s3, s0, 31
+; GFX11-NEXT:    s_mov_b64 s[4:5], exec
+; GFX11-NEXT:    s_mov_b64 s[6:7], exec
+; GFX11-NEXT:    v_add_co_ci_u32_e64 v1, null, s3, v1, vcc
+; GFX11-NEXT:    global_load_b32 v0, v[0:1], off offset:4
+; GFX11-NEXT:  .LBB0_1: ; =>This Inner Loop Header: Depth=1
+; GFX11-NEXT:    s_waitcnt vmcnt(0)
+; GFX11-NEXT:    v_readfirstlane_b32 s0, v0
+; GFX11-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX11-NEXT:    v_cmpx_eq_u32_e64 s0, v0
+; GFX11-NEXT:    s_ashr_i32 s3, s0, 31
+; GFX11-NEXT:    s_add_u32 s6, s1, s0
+; GFX11-NEXT:    s_addc_u32 s7, s2, s3
+; GFX11-NEXT:    v_mov_b32_e32 v0, 1
+; GFX11-NEXT:    s_load_b128 s[8:11], s[6:7], 0x0
+; GFX11-NEXT:    v_mov_b32_e32 v1, 0
+; GFX11-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX11-NEXT:    buffer_atomic_add_u32 v0, v1, s[8:11], 0 idxen glc
+; GFX11-NEXT:    s_and_not1_wrexec_b64 s[4:5], s[4:5]
+; GFX11-NEXT:    ; implicit-def: $vgpr0
+; GFX11-NEXT:    s_cbranch_execnz .LBB0_1
+; GFX11-NEXT:  ; %bb.2:
+; GFX11-NEXT:    s_endpgm
+bb:
+  %getelementptr = getelementptr i8, ptr addrspace(1) %arg, i32 %arg1
+  %load = load <2 x i32>, ptr addrspace(1) %getelementptr, align 8
+  %extractelement = extractelement <2 x i32> %load, i32 0
+  %extractelement3 = extractelement <2 x i32> %load, i32 1
+  %call = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %extractelement3)
+  %call4 = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %call, i32 %extractelement3)
+  %sext = sext i32 %call4 to i64
+  %getelementptr5 = getelementptr i8, ptr addrspace(4) %arg2, i64 %sext
+  %load6 = load <4 x i32>, ptr addrspace(4) %getelementptr5, align 4
+  %call7 = call i32 @llvm.amdgcn.struct.buffer.atomic.add.i32(i32 1, <4 x i32> %load6, i32 0, i32 0, i32 0, i32 0)
+  %call8 = call i32 @llvm.amdgcn.waterfall.end.i32(i32 %call, i32 %call7)
+  ret void
+}
+
+define dllexport amdgpu_cs void @atomic_add_before(ptr addrspace(1) %arg, i32 inreg %arg1, ptr addrspace(4) inreg noundef %arg2, i32 inreg %arg3) #0 {
+; GFX10-LABEL: atomic_add_before:
+; GFX10:       ; %bb.0: ; %bb
+; GFX10-NEXT:    s_ashr_i32 s6, s0, 31
+; GFX10-NEXT:    v_add_co_u32 v0, vcc, v0, s0
+; GFX10-NEXT:    v_add_co_ci_u32_e32 v1, vcc, s6, v1, vcc
+; GFX10-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-NEXT:    global_load_dwordx2 v[0:1], v[0:1], off
+; GFX10-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-NEXT:    v_mbcnt_lo_u32_b32 v0, s4, 0
+; GFX10-NEXT:    v_mbcnt_hi_u32_b32 v0, s5, v0
+; GFX10-NEXT:    v_cmp_eq_u32_e32 vcc, 0, v0
+; GFX10-NEXT:    s_and_saveexec_b64 s[6:7], vcc
+; GFX10-NEXT:    s_cbranch_execz .LBB1_2
+; GFX10-NEXT:  ; %bb.1:
+; GFX10-NEXT:    s_ashr_i32 s0, s3, 31
+; GFX10-NEXT:    s_add_u32 s12, s1, s3
+; GFX10-NEXT:    s_addc_u32 s13, s2, s0
+; GFX10-NEXT:    s_bcnt1_i32_b64 s0, s[4:5]
+; GFX10-NEXT:    s_load_dwordx4 s[8:11], s[12:13], 0x0
+; GFX10-NEXT:    v_mov_b32_e32 v0, s0
+; GFX10-NEXT:    v_mov_b32_e32 v2, 0
+; GFX10-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-NEXT:    buffer_atomic_add v0, v2, s[8:11], 0 idxen
+; GFX10-NEXT:  .LBB1_2:
+; GFX10-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-NEXT:    s_or_b64 exec, exec, s[6:7]
+; GFX10-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-NEXT:  .LBB1_3: ; =>This Inner Loop Header: Depth=1
+; GFX10-NEXT:    v_readfirstlane_b32 s0, v1
+; GFX10-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-NEXT:    v_cmpx_eq_u32_e64 s0, v1
+; GFX10-NEXT:    s_ashr_i32 s3, s0, 31
+; GFX10-NEXT:    s_add_u32 s6, s1, s0
+; GFX10-NEXT:    s_addc_u32 s7, s2, s3
+; GFX10-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-NEXT:    v_mov_b32_e32 v0, 1
+; GFX10-NEXT:    s_load_dwordx4 s[8:11], s[6:7], 0x0
+; GFX10-NEXT:    v_mov_b32_e32 v1, 0
+; GFX10-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-NEXT:    buffer_atomic_add v0, v1, s[8:11], 0 idxen glc
+; GFX10-NEXT:    s_andn2_wrexec_b64 s[4:5], s[4:5]
+; GFX10-NEXT:    ; implicit-def: $vgpr0_vgpr1
+; GFX10-NEXT:    s_cbranch_execnz .LBB1_3
+; GFX10-NEXT:  ; %bb.4:
+; GFX10-NEXT:    s_endpgm
+;
+; GFX11-LABEL: atomic_add_before:
+; GFX11:       ; %bb.0: ; %bb
+; GFX11-NEXT:    v_add_co_u32 v0, vcc, v0, s0
+; GFX11-NEXT:    s_ashr_i32 s6, s0, 31
+; GFX11-NEXT:    s_mov_b64 s[4:5], exec
+; GFX11-NEXT:    v_add_co_ci_u32_e64 v1, null, s6, v1, vcc
+; GFX11-NEXT:    s_mov_b64 s[6:7], exec
+; GFX11-NEXT:    global_load_b64 v[0:1], v[0:1], off
+; GFX11-NEXT:    s_waitcnt vmcnt(0)
+; GFX11-NEXT:    v_mbcnt_lo_u32_b32 v0, s4, 0
+; GFX11-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX11-NEXT:    v_mbcnt_hi_u32_b32 v0, s5, v0
+; GFX11-NEXT:    v_cmpx_eq_u32_e32 0, v0
+; GFX11-NEXT:    s_cbranch_execz .LBB1_2
+; GFX11-NEXT:  ; %bb.1:
+; GFX11-NEXT:    s_ashr_i32 s0, s3, 31
+; GFX11-NEXT:    s_add_u32 s8, s1, s3
+; GFX11-NEXT:    s_addc_u32 s9, s2, s0
+; GFX11-NEXT:    s_bcnt1_i32_b64 s0, s[4:5]
+; GFX11-NEXT:    s_load_b128 s[8:11], s[8:9], 0x0
+; GFX11-NEXT:    v_mov_b32_e32 v0, s0
+; GFX11-NEXT:    v_mov_b32_e32 v2, 0
+; GFX11-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX11-NEXT:    buffer_atomic_add_u32 v0, v2, s[8:11], 0 idxen
+; GFX11-NEXT:  .LBB1_2:
+; GFX11-NEXT:    s_or_b64 exec, exec, s[6:7]
+; GFX11-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX11-NEXT:    s_mov_b64 s[4:5], exec
+; GFX11-NEXT:    s_mov_b64 s[6:7], exec
+; GFX11-NEXT:  .LBB1_3: ; =>This Inner Loop Header: Depth=1
+; GFX11-NEXT:    v_readfirstlane_b32 s0, v1
+; GFX11-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX11-NEXT:    v_cmpx_eq_u32_e64 s0, v1
+; GFX11-NEXT:    s_ashr_i32 s3, s0, 31
+; GFX11-NEXT:    s_add_u32 s6, s1, s0
+; GFX11-NEXT:    s_addc_u32 s7, s2, s3
+; GFX11-NEXT:    s_waitcnt vmcnt(0)
+; GFX11-NEXT:    v_mov_b32_e32 v0, 1
+; GFX11-NEXT:    s_load_b128 s[8:11], s[6:7], 0x0
+; GFX11-NEXT:    v_mov_b32_e32 v1, 0
+; GFX11-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX11-NEXT:    buffer_atomic_add_u32 v0, v1, s[8:11], 0 idxen glc
+; GFX11-NEXT:    s_and_not1_wrexec_b64 s[4:5], s[4:5]
+; GFX11-NEXT:    ; implicit-def: $vgpr0_vgpr1
+; GFX11-NEXT:    s_cbranch_execnz .LBB1_3
+; GFX11-NEXT:  ; %bb.4:
+; GFX11-NEXT:    s_endpgm
+bb:
+  %getelementptr = getelementptr i8, ptr addrspace(1) %arg, i32 %arg1
+  %load = load <2 x i32>, ptr addrspace(1) %getelementptr, align 8
+  %extractelement = extractelement <2 x i32> %load, i32 0
+  %extractelement4 = extractelement <2 x i32> %load, i32 1
+  %sext = sext i32 %arg3 to i64
+  %getelementptr5 = getelementptr i8, ptr addrspace(4) %arg2, i64 %sext
+  %load6 = load <4 x i32>, ptr addrspace(4) %getelementptr5, align 4
+  %call = call i32 @llvm.amdgcn.struct.buffer.atomic.add.i32(i32 1, <4 x i32> %load6, i32 0, i32 0, i32 0, i32 0)
+  %call7 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %extractelement4)
+  %call8 = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %call7, i32 %extractelement4)
+  %sext9 = sext i32 %call8 to i64
+  %getelementptr10 = getelementptr i8, ptr addrspace(4) %arg2, i64 %sext9
+  %load11 = load <4 x i32>, ptr addrspace(4) %getelementptr10, align 4
+  %call12 = call i32 @llvm.amdgcn.struct.buffer.atomic.add.i32(i32 1, <4 x i32> %load11, i32 0, i32 0, i32 0, i32 0)
+  %call13 = call i32 @llvm.amdgcn.waterfall.end.i32(i32 %call7, i32 %call12)
+  ret void
+}
+
+define dllexport amdgpu_cs void @atomic_add_after(ptr addrspace(1) %arg, i32 inreg %arg1, ptr addrspace(4) inreg noundef %arg2, i32 inreg %arg3) #0 {
+; GFX10-LABEL: atomic_add_after:
+; GFX10:       ; %bb.0: ; %bb
+; GFX10-NEXT:    s_ashr_i32 s4, s0, 31
+; GFX10-NEXT:    v_add_co_u32 v0, vcc, v0, s0
+; GFX10-NEXT:    v_add_co_ci_u32_e32 v1, vcc, s4, v1, vcc
+; GFX10-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-NEXT:    global_load_dword v1, v[0:1], off offset:4
+; GFX10-NEXT:  .LBB2_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-NEXT:    v_readfirstlane_b32 s0, v1
+; GFX10-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-NEXT:    v_cmpx_eq_u32_e64 s0, v1
+; GFX10-NEXT:    s_ashr_i32 s8, s0, 31
+; GFX10-NEXT:    s_add_u32 s12, s1, s0
+; GFX10-NEXT:    s_addc_u32 s13, s2, s8
+; GFX10-NEXT:    v_mov_b32_e32 v1, 1
+; GFX10-NEXT:    s_load_dwordx4 s[8:11], s[12:13], 0x0
+; GFX10-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-NEXT:    buffer_atomic_add v1, v0, s[8:11], 0 idxen glc
+; GFX10-NEXT:    s_andn2_wrexec_b64 s[6:7], s[6:7]
+; GFX10-NEXT:    ; implicit-def: $vgpr1
+; GFX10-NEXT:    s_cbranch_execnz .LBB2_1
+; GFX10-NEXT:  ; %bb.2:
+; GFX10-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-NEXT:    s_mov_b64 exec, s[4:5]
+; GFX10-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-NEXT:    v_mbcnt_lo_u32_b32 v1, s4, 0
+; GFX10-NEXT:    v_mbcnt_hi_u32_b32 v1, s5, v1
+; GFX10-NEXT:    v_cmp_eq_u32_e32 vcc, 0, v1
+; GFX10-NEXT:    s_and_saveexec_b64 s[6:7], vcc
+; GFX10-NEXT:    s_cbranch_execz .LBB2_4
+; GFX10-NEXT:  ; %bb.3:
+; GFX10-NEXT:    s_ashr_i32 s0, s3, 31
+; GFX10-NEXT:    s_add_u32 s6, s1, s3
+; GFX10-NEXT:    s_addc_u32 s7, s2, s0
+; GFX10-NEXT:    s_bcnt1_i32_b64 s4, s[4:5]
+; GFX10-NEXT:    s_load_dwordx4 s[0:3], s[6:7], 0x0
+; GFX10-NEXT:    v_mov_b32_e32 v1, s4
+; GFX10-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-NEXT:    buffer_atomic_add v1, v0, s[0:3], 0 idxen
+; GFX10-NEXT:  .LBB2_4:
+; GFX10-NEXT:    s_endpgm
+;
+; GFX11-LABEL: atomic_add_after:
+; GFX11:       ; %bb.0: ; %bb
+; GFX11-NEXT:    v_add_co_u32 v0, vcc, v0, s0
+; GFX11-NEXT:    s_ashr_i32 s4, s0, 31
+; GFX11-NEXT:    s_mov_b64 s[6:7], exec
+; GFX11-NEXT:    v_add_co_ci_u32_e64 v1, null, s4, v1, vcc
+; GFX11-NEXT:    s_mov_b64 s[4:5], exec
+; GFX11-NEXT:    global_load_b32 v1, v[0:1], off offset:4
+; GFX11-NEXT:  .LBB2_1: ; =>This Inner Loop Header: Depth=1
+; GFX11-NEXT:    s_waitcnt vmcnt(0)
+; GFX11-NEXT:    v_readfirstlane_b32 s0, v1
+; GFX11-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX11-NEXT:    v_cmpx_eq_u32_e64 s0, v1
+; GFX11-NEXT:    s_ashr_i32 s9, s0, 31
+; GFX11-NEXT:    s_add_u32 s8, s1, s0
+; GFX11-NEXT:    s_addc_u32 s9, s2, s9
+; GFX11-NEXT:    v_mov_b32_e32 v1, 1
+; GFX11-NEXT:    s_load_b128 s[8:11], s[8:9], 0x0
+; GFX11-NEXT:    v_mov_b32_e32 v0, 0
+; GFX11-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX11-NEXT:    buffer_atomic_add_u32 v1, v0, s[8:11], 0 idxen glc
+; GFX11-NEXT:    s_and_not1_wrexec_b64 s[6:7], s[6:7]
+; GFX11-NEXT:    ; implicit-def: $vgpr1
+; GFX11-NEXT:    s_cbranch_execnz .LBB2_1
+; GFX11-NEXT:  ; %bb.2:
+; GFX11-NEXT:    s_mov_b64 exec, s[4:5]
+; GFX11-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(SKIP_3) | instid1(VALU_DEP_1)
+; GFX11-NEXT:    s_mov_b64 s[4:5], exec
+; GFX11-NEXT:    s_mov_b64 s[6:7], exec
+; GFX11-NEXT:    s_waitcnt vmcnt(0)
+; GFX11-NEXT:    v_mbcnt_lo_u32_b32 v1, s4, 0
+; GFX11-NEXT:    v_mbcnt_hi_u32_b32 v1, s5, v1
+; GFX11-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX11-NEXT:    v_cmpx_eq_u32_e32 0, v1
+; GFX11-NEXT:    s_cbranch_execz .LBB2_4
+; GFX11-NEXT:  ; %bb.3:
+; GFX11-NEXT:    s_ashr_i32 s6, s3, 31
+; GFX11-NEXT:    s_add_u32 s0, s1, s3
+; GFX11-NEXT:    s_addc_u32 s1, s2, s6
+; GFX11-NEXT:    s_bcnt1_i32_b64 s4, s[4:5]
+; GFX11-NEXT:    s_load_b128 s[0:3], s[0:1], 0x0
+; GFX11-NEXT:    v_mov_b32_e32 v1, s4
+; GFX11-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX11-NEXT:    buffer_atomic_add_u32 v1, v0, s[0:3], 0 idxen
+; GFX11-NEXT:  .LBB2_4:
+; GFX11-NEXT:    s_endpgm
+bb:
+  %getelementptr = getelementptr i8, ptr addrspace(1) %arg, i32 %arg1
+  %load = load <2 x i32>, ptr addrspace(1) %getelementptr, align 8
+  %extractelement = extractelement <2 x i32> %load, i32 0
+  %extractelement4 = extractelement <2 x i32> %load, i32 1
+  %call7 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %extractelement4)
+  %call8 = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %call7, i32 %extractelement4)
+  %sext9 = sext i32 %call8 to i64
+  %getelementptr10 = getelementptr i8, ptr addrspace(4) %arg2, i64 %sext9
+  %load11 = load <4 x i32>, ptr addrspace(4) %getelementptr10, align 4
+  %call12 = call i32 @llvm.amdgcn.struct.buffer.atomic.add.i32(i32 1, <4 x i32> %load11, i32 0, i32 0, i32 0, i32 0)
+  %call13 = call i32 @llvm.amdgcn.waterfall.end.i32(i32 %call7, i32 %call12)
+  %sext = sext i32 %arg3 to i64
+  %getelementptr5 = getelementptr i8, ptr addrspace(4) %arg2, i64 %sext
+  %load6 = load <4 x i32>, ptr addrspace(4) %getelementptr5, align 4
+  %call = call i32 @llvm.amdgcn.struct.buffer.atomic.add.i32(i32 1, <4 x i32> %load6, i32 0, i32 0, i32 0, i32 0)
+  ret void
+}
+
+define dllexport amdgpu_cs void @atomic_add_in_wf_partial(ptr addrspace(1) %arg, i32 inreg %arg1, ptr addrspace(4) inreg noundef %arg2, i32 inreg %arg3) #0 {
+; GFX10-LABEL: atomic_add_in_wf_partial:
+; GFX10:       ; %bb.0: ; %bb
+; GFX10-NEXT:    s_ashr_i32 s10, s0, 31
+; GFX10-NEXT:    s_ashr_i32 s4, s3, 31
+; GFX10-NEXT:    s_add_u32 s8, s1, s3
+; GFX10-NEXT:    s_addc_u32 s9, s2, s4
+; GFX10-NEXT:    v_add_co_u32 v0, vcc, v0, s0
+; GFX10-NEXT:    s_load_dwordx4 s[4:7], s[8:9], 0x0
+; GFX10-NEXT:    v_add_co_ci_u32_e32 v1, vcc, s10, v1, vcc
+; GFX10-NEXT:    v_mov_b32_e32 v2, 1
+; GFX10-NEXT:    v_mov_b32_e32 v3, 0
+; GFX10-NEXT:    global_load_dwordx2 v[0:1], v[0:1], off
+; GFX10-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-NEXT:    buffer_atomic_add v2, v3, s[4:7], 0 idxen
+; GFX10-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-NEXT:  .LBB3_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-NEXT:    v_readfirstlane_b32 s0, v0
+; GFX10-NEXT:    v_readfirstlane_b32 s3, v1
+; GFX10-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-NEXT:    v_cmpx_eq_u32_e64 s0, v0
+; GFX10-NEXT:    v_cmpx_eq_u32_e64 s3, v1
+; GFX10-NEXT:    s_ashr_i32 s3, s0, 31
+; GFX10-NEXT:    s_add_u32 s6, s1, s0
+; GFX10-NEXT:    s_addc_u32 s7, s2, s3
+; GFX10-NEXT:    s_load_dwordx4 s[8:11], s[6:7], 0x0
+; GFX10-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-NEXT:    buffer_atomic_add v2, v3, s[8:11], 0 idxen glc
+; GFX10-NEXT:    s_andn2_wrexec_b64 s[4:5], s[4:5]
+; GFX10-NEXT:    ; implicit-def: $vgpr0_vgpr1
+; GFX10-NEXT:    ; implicit-def: $vgpr2
+; GFX10-NEXT:    ; implicit-def: $vgpr3
+; GFX10-NEXT:    s_cbranch_execnz .LBB3_1
+; GFX10-NEXT:  ; %bb.2:
+; GFX10-NEXT:    s_endpgm
+;
+; GFX11-LABEL: atomic_add_in_wf_partial:
+; GFX11:       ; %bb.0: ; %bb
+; GFX11-NEXT:    s_ashr_i32 s8, s0, 31
+; GFX11-NEXT:    s_ashr_i32 s5, s3, 31
+; GFX11-NEXT:    s_add_u32 s4, s1, s3
+; GFX11-NEXT:    s_addc_u32 s5, s2, s5
+; GFX11-NEXT:    v_add_co_u32 v0, vcc, v0, s0
+; GFX11-NEXT:    s_load_b128 s[4:7], s[4:5], 0x0
+; GFX11-NEXT:    v_mov_b32_e32 v2, 1
+; GFX11-NEXT:    v_mov_b32_e32 v3, 0
+; GFX11-NEXT:    v_add_co_ci_u32_e64 v1, null, s8, v1, vcc
+; GFX11-NEXT:    global_load_b64 v[0:1], v[0:1], off
+; GFX11-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX11-NEXT:    buffer_atomic_add_u32 v2, v3, s[4:7], 0 idxen
+; GFX11-NEXT:    s_mov_b64 s[4:5], exec
+; GFX11-NEXT:    s_mov_b64 s[6:7], exec
+; GFX11-NEXT:  .LBB3_1: ; =>This Inner Loop Header: Depth=1
+; GFX11-NEXT:    s_waitcnt vmcnt(0)
+; GFX11-NEXT:    v_readfirstlane_b32 s0, v0
+; GFX11-NEXT:    v_readfirstlane_b32 s3, v1
+; GFX11-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX11-NEXT:    v_cmpx_eq_u32_e64 s0, v0
+; GFX11-NEXT:    v_cmpx_eq_u32_e64 s3, v1
+; GFX11-NEXT:    s_ashr_i32 s3, s0, 31
+; GFX11-NEXT:    s_add_u32 s6, s1, s0
+; GFX11-NEXT:    s_addc_u32 s7, s2, s3
+; GFX11-NEXT:    s_load_b128 s[8:11], s[6:7], 0x0
+; GFX11-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX11-NEXT:    buffer_atomic_add_u32 v2, v3, s[8:11], 0 idxen glc
+; GFX11-NEXT:    s_and_not1_wrexec_b64 s[4:5], s[4:5]
+; GFX11-NEXT:    ; implicit-def: $vgpr0_vgpr1
+; GFX11-NEXT:    ; implicit-def: $vgpr2
+; GFX11-NEXT:    ; implicit-def: $vgpr3
+; GFX11-NEXT:    s_cbranch_execnz .LBB3_1
+; GFX11-NEXT:  ; %bb.2:
+; GFX11-NEXT:    s_endpgm
+bb:
+  %getelementptr = getelementptr i8, ptr addrspace(1) %arg, i32 %arg1
+  %load = load <2 x i32>, ptr addrspace(1) %getelementptr, align 8
+  %extractelement = extractelement <2 x i32> %load, i32 0
+  %extractelement4 = extractelement <2 x i32> %load, i32 1
+  %token = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %extractelement)
+  %sext = sext i32 %arg3 to i64
+  %getelementptr5 = getelementptr i8, ptr addrspace(4) %arg2, i64 %sext
+  %load6 = load <4 x i32>, ptr addrspace(4) %getelementptr5, align 4
+  %call = call i32 @llvm.amdgcn.struct.buffer.atomic.add.i32(i32 1, <4 x i32> %load6, i32 0, i32 0, i32 0, i32 0)
+  %token2 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 %token, i32 %extractelement4)
+  %call8 = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %token2, i32 %extractelement4)
+  %sext9 = sext i32 %call8 to i64
+  %getelementptr10 = getelementptr i8, ptr addrspace(4) %arg2, i64 %sext9
+  %load11 = load <4 x i32>, ptr addrspace(4) %getelementptr10, align 4
+  %call12 = call i32 @llvm.amdgcn.struct.buffer.atomic.add.i32(i32 1, <4 x i32> %load11, i32 0, i32 0, i32 0, i32 0)
+  %call13 = call i32 @llvm.amdgcn.waterfall.end.i32(i32 %token2, i32 %call12)
+  ret void
+}
+
+
+
+; Function Attrs: nocallback nofree nounwind willreturn
+declare i32 @llvm.amdgcn.struct.buffer.atomic.add.i32(i32, <4 x i32>, i32, i32, i32, i32 immarg) #3
+; Function Attrs: convergent nounwind
+declare i32 @llvm.amdgcn.waterfall.begin.i32(i32, i32) #4
+; Function Attrs: convergent nounwind
+declare i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32, i32) #4
+; Function Attrs: convergent nounwind
+declare i32 @llvm.amdgcn.waterfall.end.i32(i32, i32) #4
+
+attributes #0 = { nounwind readnone convergent }
diff --git a/llvm/test/CodeGen/AMDGPU/llc-pipeline-npm.ll b/llvm/test/CodeGen/AMDGPU/llc-pipeline-npm.ll
index c49b2b927bd31..0cbd3dae53ba8 100644
--- a/llvm/test/CodeGen/AMDGPU/llc-pipeline-npm.ll
+++ b/llvm/test/CodeGen/AMDGPU/llc-pipeline-npm.ll
@@ -57,6 +57,7 @@
 ; GCN-O0-NEXT: localstackalloc)))
 ; GCN-O0-NEXT: require<reg-usage>
 ; GCN-O0-NEXT: cgscc(function(machine-function(reg-usage-propagation
+; GCN-O0-NEXT: si-insert-waterfall
 ; GCN-O0-NEXT: phi-node-elimination
 ; GCN-O0-NEXT: si-lower-control-flow
 ; GCN-O0-NEXT: two-address-instruction
@@ -186,6 +187,7 @@
 ; GCN-O2-NEXT: si-shrink-instructions)))
 ; GCN-O2-NEXT: require<reg-usage>
 ; GCN-O2-NEXT: cgscc(function(machine-function(reg-usage-propagation
+; GCN-O2-NEXT: si-insert-waterfall
 ; GCN-O2-NEXT: amdgpu-prepare-agpr-alloc
 ; GCN-O2-NEXT: detect-dead-lanes
 ; GCN-O2-NEXT: dead-mi-elimination
@@ -355,6 +357,7 @@
 ; GCN-O3-NEXT: si-shrink-instructions)))
 ; GCN-O3-NEXT: require<reg-usage>
 ; GCN-O3-NEXT: cgscc(function(machine-function(reg-usage-propagation
+; GCN-O3-NEXT: si-insert-waterfall
 ; GCN-O3-NEXT: amdgpu-prepare-agpr-alloc
 ; GCN-O3-NEXT: detect-dead-lanes
 ; GCN-O3-NEXT: dead-mi-elimination
diff --git a/llvm/test/CodeGen/AMDGPU/llc-pipeline.ll b/llvm/test/CodeGen/AMDGPU/llc-pipeline.ll
index 070c873798647..e5ae5fdaff6ab 100644
--- a/llvm/test/CodeGen/AMDGPU/llc-pipeline.ll
+++ b/llvm/test/CodeGen/AMDGPU/llc-pipeline.ll
@@ -109,6 +109,7 @@
 ; GCN-O0-NEXT:        Finalize ISel and expand pseudo-instructions
 ; GCN-O0-NEXT:        Local Stack Slot Allocation
 ; GCN-O0-NEXT:        Register Usage Information Propagation
+; GCN-O0-NEXT:        SI Insert waterfalls
 ; GCN-O0-NEXT:        Eliminate PHI nodes for register allocation
 ; GCN-O0-NEXT:        SI Lower control flow pseudo instructions
 ; GCN-O0-NEXT:        Two-Address instruction pass
@@ -343,6 +344,7 @@
 ; GCN-O1-NEXT:        Remove dead machine instructions
 ; GCN-O1-NEXT:        SI Shrink Instructions
 ; GCN-O1-NEXT:        Register Usage Information Propagation
+; GCN-O1-NEXT:        SI Insert waterfalls
 ; GCN-O1-NEXT:        AMDGPU Prepare AGPR Alloc
 ; GCN-O1-NEXT:        Detect Dead Lanes
 ; GCN-O1-NEXT:        Remove dead machine instructions
@@ -351,6 +353,7 @@
 ; GCN-O1-NEXT:        Remove unreachable machine basic blocks
 ; GCN-O1-NEXT:        Live Variable Analysis
 ; GCN-O1-NEXT:        MachineDominator Tree Construction
+; GCN-O1-NEXT:        Machine Natural Loop Construction
 ; GCN-O1-NEXT:        SI Optimize VGPR LiveRange
 ; GCN-O1-NEXT:        Eliminate PHI nodes for register allocation
 ; GCN-O1-NEXT:        SI Lower control flow pseudo instructions
@@ -361,6 +364,7 @@
 ; GCN-O1-NEXT:        Register Coalescer
 ; GCN-O1-NEXT:        Rename Disconnected Subregister Components
 ; GCN-O1-NEXT:        Rewrite Partial Register Uses
+; GCN-O1-NEXT:        Machine Block Frequency Analysis
 ; GCN-O1-NEXT:        Machine Instruction Scheduler
 ; GCN-O1-NEXT:        SI Whole Quad Mode
 ; GCN-O1-NEXT:        SI optimize exec mask operations pre-RA
@@ -661,6 +665,7 @@
 ; GCN-O1-OPTS-NEXT:        Remove dead machine instructions
 ; GCN-O1-OPTS-NEXT:        SI Shrink Instructions
 ; GCN-O1-OPTS-NEXT:        Register Usage Information Propagation
+; GCN-O1-OPTS-NEXT:        SI Insert waterfalls
 ; GCN-O1-OPTS-NEXT:        AMDGPU Prepare AGPR Alloc
 ; GCN-O1-OPTS-NEXT:        Detect Dead Lanes
 ; GCN-O1-OPTS-NEXT:        Remove dead machine instructions
@@ -668,6 +673,8 @@
 ; GCN-O1-OPTS-NEXT:        Process Implicit Definitions
 ; GCN-O1-OPTS-NEXT:        Remove unreachable machine basic blocks
 ; GCN-O1-OPTS-NEXT:        Live Variable Analysis
+; GCN-O1-OPTS-NEXT:        MachineDominator Tree Construction
+; GCN-O1-OPTS-NEXT:        Machine Natural Loop Construction
 ; GCN-O1-OPTS-NEXT:        SI Optimize VGPR LiveRange
 ; GCN-O1-OPTS-NEXT:        Eliminate PHI nodes for register allocation
 ; GCN-O1-OPTS-NEXT:        SI Lower control flow pseudo instructions
@@ -678,6 +685,7 @@
 ; GCN-O1-OPTS-NEXT:        Register Coalescer
 ; GCN-O1-OPTS-NEXT:        Rename Disconnected Subregister Components
 ; GCN-O1-OPTS-NEXT:        Rewrite Partial Register Uses
+; GCN-O1-OPTS-NEXT:        Machine Block Frequency Analysis
 ; GCN-O1-OPTS-NEXT:        Machine Instruction Scheduler
 ; GCN-O1-OPTS-NEXT:        AMDGPU Pre-RA optimizations
 ; GCN-O1-OPTS-NEXT:        SI Whole Quad Mode
@@ -983,6 +991,7 @@
 ; GCN-O2-NEXT:        Remove dead machine instructions
 ; GCN-O2-NEXT:        SI Shrink Instructions
 ; GCN-O2-NEXT:        Register Usage Information Propagation
+; GCN-O2-NEXT:        SI Insert waterfalls
 ; GCN-O2-NEXT:        AMDGPU Prepare AGPR Alloc
 ; GCN-O2-NEXT:        Detect Dead Lanes
 ; GCN-O2-NEXT:        Remove dead machine instructions
@@ -990,6 +999,8 @@
 ; GCN-O2-NEXT:        Process Implicit Definitions
 ; GCN-O2-NEXT:        Remove unreachable machine basic blocks
 ; GCN-O2-NEXT:        Live Variable Analysis
+; GCN-O2-NEXT:        MachineDominator Tree Construction
+; GCN-O2-NEXT:        Machine Natural Loop Construction
 ; GCN-O2-NEXT:        SI Optimize VGPR LiveRange
 ; GCN-O2-NEXT:        Eliminate PHI nodes for register allocation
 ; GCN-O2-NEXT:        SI Lower control flow pseudo instructions
@@ -1000,6 +1011,7 @@
 ; GCN-O2-NEXT:        Register Coalescer
 ; GCN-O2-NEXT:        Rename Disconnected Subregister Components
 ; GCN-O2-NEXT:        Rewrite Partial Register Uses
+; GCN-O2-NEXT:        Machine Block Frequency Analysis
 ; GCN-O2-NEXT:        Machine Instruction Scheduler
 ; GCN-O2-NEXT:        AMDGPU Pre-RA optimizations
 ; GCN-O2-NEXT:        SI Whole Quad Mode
@@ -1319,6 +1331,7 @@
 ; GCN-O3-NEXT:        Remove dead machine instructions
 ; GCN-O3-NEXT:        SI Shrink Instructions
 ; GCN-O3-NEXT:        Register Usage Information Propagation
+; GCN-O3-NEXT:        SI Insert waterfalls
 ; GCN-O3-NEXT:        AMDGPU Prepare AGPR Alloc
 ; GCN-O3-NEXT:        Detect Dead Lanes
 ; GCN-O3-NEXT:        Remove dead machine instructions
@@ -1326,6 +1339,8 @@
 ; GCN-O3-NEXT:        Process Implicit Definitions
 ; GCN-O3-NEXT:        Remove unreachable machine basic blocks
 ; GCN-O3-NEXT:        Live Variable Analysis
+; GCN-O3-NEXT:        MachineDominator Tree Construction
+; GCN-O3-NEXT:        Machine Natural Loop Construction
 ; GCN-O3-NEXT:        SI Optimize VGPR LiveRange
 ; GCN-O3-NEXT:        Eliminate PHI nodes for register allocation
 ; GCN-O3-NEXT:        SI Lower control flow pseudo instructions
@@ -1336,6 +1351,7 @@
 ; GCN-O3-NEXT:        Register Coalescer
 ; GCN-O3-NEXT:        Rename Disconnected Subregister Components
 ; GCN-O3-NEXT:        Rewrite Partial Register Uses
+; GCN-O3-NEXT:        Machine Block Frequency Analysis
 ; GCN-O3-NEXT:        Machine Instruction Scheduler
 ; GCN-O3-NEXT:        AMDGPU Pre-RA optimizations
 ; GCN-O3-NEXT:        SI Whole Quad Mode
diff --git a/llvm/test/CodeGen/AMDGPU/llvm.amdgcn.waterfall.ll b/llvm/test/CodeGen/AMDGPU/llvm.amdgcn.waterfall.ll
new file mode 100644
index 0000000000000..4ceb4245c77e9
--- /dev/null
+++ b/llvm/test/CodeGen/AMDGPU/llvm.amdgcn.waterfall.ll
@@ -0,0 +1,12656 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py
+; RUN: llc -global-isel=0 -march=amdgcn -mcpu=fiji -verify-machineinstrs < %s | FileCheck -check-prefixes=PRE-GFX10,VI,VI-SDAG %s
+; RUN: llc -global-isel=1 -march=amdgcn -mcpu=fiji -verify-machineinstrs < %s | FileCheck -check-prefixes=PRE-GFX10,VI,VI-GISEL %s
+; RUN: llc -global-isel=0 -march=amdgcn -mcpu=gfx900 -verify-machineinstrs < %s | FileCheck -check-prefixes=PRE-GFX10,GFX9,GFX9-SDAG %s
+; RUN: llc -global-isel=1 -march=amdgcn -mcpu=gfx900 -verify-machineinstrs < %s | FileCheck -check-prefixes=PRE-GFX10,GFX9,GFX9-GISEL %s
+; RUN: llc -global-isel=0 -march=amdgcn -mcpu=gfx1010 -mattr=+wavefrontsize32,-wavefrontsize64 -verify-machineinstrs < %s | FileCheck -check-prefixes=GFX10GFX11-SDAG,GFX10,GFX10-32,GFX10-32-SDAG %s
+; RUN: llc -global-isel=1 -march=amdgcn -mcpu=gfx1010 -mattr=+wavefrontsize32,-wavefrontsize64 -verify-machineinstrs < %s | FileCheck -check-prefixes=GFX10,GFX10-32,GFX10-32-GISEL %s
+; RUN: llc -global-isel=0 -march=amdgcn -mcpu=gfx1010 -mattr=-wavefrontsize32,+wavefrontsize64 -verify-machineinstrs < %s | FileCheck -check-prefixes=GFX10GFX11-SDAG,GFX10,GFX10-64,GFX10-64-SDAG %s
+; RUN: llc -global-isel=1 -march=amdgcn -mcpu=gfx1010 -mattr=-wavefrontsize32,+wavefrontsize64 -verify-machineinstrs < %s | FileCheck -check-prefixes=GFX10,GFX10-64,GFX10-64-GISEL %s
+; RUN: llc -global-isel=0 -march=amdgcn -mcpu=gfx1150 -mattr=-wavefrontsize32,+wavefrontsize64 -verify-machineinstrs < %s | FileCheck -check-prefixes=GFX10GFX11-SDAG,GFX1150,GFX1150-SDAG %s
+; RUN: llc -global-isel=1 -march=amdgcn -mcpu=gfx1150 -mattr=-wavefrontsize32,+wavefrontsize64 -verify-machineinstrs < %s | FileCheck -check-prefixes=GFX1150,GFX1150-GISEL %s
+; RUN: llc -global-isel=0 -march=amdgcn -mcpu=gfx1200 -mattr=-wavefrontsize32,+wavefrontsize64 -verify-machineinstrs < %s | FileCheck -check-prefixes=GFX12,GFX12-SDAG %s
+; RUN: llc -global-isel=1 -march=amdgcn -mcpu=gfx1200 -mattr=-wavefrontsize32,+wavefrontsize64 -verify-machineinstrs < %s | FileCheck -check-prefixes=GFX12,GFX12-GISEL %s
+
+ at Lds = addrspace(3) global [16384 x i32] undef
+
+define amdgpu_ps void @test_waterfall_readlane(i32 addrspace(1)* inreg %out, <2 x i32> addrspace(1)* inreg %in, i32 %tid) #1 {
+; VI-SDAG-LABEL: test_waterfall_readlane:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; VI-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 3, v[0:1]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v2, s3
+; VI-SDAG-NEXT:    v_add_u32_e32 v0, vcc, s2, v0
+; VI-SDAG-NEXT:    v_addc_u32_e32 v1, vcc, v2, v1, vcc
+; VI-SDAG-NEXT:    flat_load_dwordx2 v[2:3], v[0:1]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v0, s0
+; VI-SDAG-NEXT:    v_mov_b32_e32 v1, s1
+; VI-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; VI-SDAG-NEXT:  .LBB0_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s4, v3
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[2:3], s4, v3
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; VI-SDAG-NEXT:    s_nop 1
+; VI-SDAG-NEXT:    v_readlane_b32 s4, v2, s4
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, s4
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr2_vgpr3
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB0_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; VI-SDAG-NEXT:    flat_store_dword v[0:1], v4
+; VI-SDAG-NEXT:    s_endpgm
+;
+; VI-GISEL-LABEL: test_waterfall_readlane:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; VI-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 3, v[0:1]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v2, s2
+; VI-GISEL-NEXT:    v_mov_b32_e32 v3, s3
+; VI-GISEL-NEXT:    v_add_u32_e32 v0, vcc, v2, v0
+; VI-GISEL-NEXT:    v_addc_u32_e32 v1, vcc, v3, v1, vcc
+; VI-GISEL-NEXT:    flat_load_dwordx2 v[0:1], v[0:1]
+; VI-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; VI-GISEL-NEXT:  .LBB0_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s6, v1
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[4:5], s6, v1
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[4:5], s[4:5]
+; VI-GISEL-NEXT:    s_nop 1
+; VI-GISEL-NEXT:    v_readlane_b32 s6, v0, s6
+; VI-GISEL-NEXT:    v_mov_b32_e32 v2, s6
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr0_vgpr1
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[4:5]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB0_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v0, s0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v1, s1
+; VI-GISEL-NEXT:    flat_store_dword v[0:1], v2
+; VI-GISEL-NEXT:    s_endpgm
+;
+; GFX9-SDAG-LABEL: test_waterfall_readlane:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX9-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 3, v[0:1]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v2, s3
+; GFX9-SDAG-NEXT:    v_add_co_u32_e32 v0, vcc, s2, v0
+; GFX9-SDAG-NEXT:    v_addc_co_u32_e32 v1, vcc, v2, v1, vcc
+; GFX9-SDAG-NEXT:    global_load_dwordx2 v[0:1], v[0:1], off
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v2, 0
+; GFX9-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX9-SDAG-NEXT:  .LBB0_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s6, v1
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[4:5], s6, v1
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[4:5], s[4:5]
+; GFX9-SDAG-NEXT:    s_nop 1
+; GFX9-SDAG-NEXT:    v_readlane_b32 s6, v0, s6
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v3, s6
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr0_vgpr1
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[4:5]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB0_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX9-SDAG-NEXT:    global_store_dword v2, v3, s[0:1]
+; GFX9-SDAG-NEXT:    s_endpgm
+;
+; GFX9-GISEL-LABEL: test_waterfall_readlane:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX9-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 3, v[0:1]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v2, s2
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v3, s3
+; GFX9-GISEL-NEXT:    v_add_co_u32_e32 v0, vcc, v2, v0
+; GFX9-GISEL-NEXT:    v_addc_co_u32_e32 v1, vcc, v3, v1, vcc
+; GFX9-GISEL-NEXT:    global_load_dwordx2 v[0:1], v[0:1], off
+; GFX9-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX9-GISEL-NEXT:  .LBB0_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s6, v1
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[4:5], s6, v1
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[4:5], s[4:5]
+; GFX9-GISEL-NEXT:    s_nop 1
+; GFX9-GISEL-NEXT:    v_readlane_b32 s6, v0, s6
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v2, s6
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr0_vgpr1
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[4:5]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB0_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX9-GISEL-NEXT:    global_store_dword v0, v2, s[0:1]
+; GFX9-GISEL-NEXT:    s_endpgm
+;
+; GFX10-32-SDAG-LABEL: test_waterfall_readlane:
+; GFX10-32-SDAG:       ; %bb.0:
+; GFX10-32-SDAG-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v2, 0
+; GFX10-32-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 3, v[0:1]
+; GFX10-32-SDAG-NEXT:    v_add_co_u32 v0, vcc_lo, s2, v0
+; GFX10-32-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, s3, v1, vcc_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s3, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s2, exec_lo
+; GFX10-32-SDAG-NEXT:    global_load_dwordx2 v[0:1], v[0:1], off
+; GFX10-32-SDAG-NEXT:  .LBB0_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s4, v1
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s4, v1
+; GFX10-32-SDAG-NEXT:    v_readlane_b32 s4, v0, s4
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v3, s4
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s3, s3
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr0_vgpr1
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB0_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s2
+; GFX10-32-SDAG-NEXT:    global_store_dword v2, v3, s[0:1]
+; GFX10-32-SDAG-NEXT:    s_endpgm
+;
+; GFX10-32-GISEL-LABEL: test_waterfall_readlane:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v2, s2
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v3, s3
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s3, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s2, exec_lo
+; GFX10-32-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 3, v[0:1]
+; GFX10-32-GISEL-NEXT:    v_add_co_u32 v0, vcc_lo, v2, v0
+; GFX10-32-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, v3, v1, vcc_lo
+; GFX10-32-GISEL-NEXT:    global_load_dwordx2 v[0:1], v[0:1], off
+; GFX10-32-GISEL-NEXT:  .LBB0_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s4, v1
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s4, v1
+; GFX10-32-GISEL-NEXT:    v_readlane_b32 s4, v0, s4
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v2, s4
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s3, s3
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr0_vgpr1
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB0_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s2
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-GISEL-NEXT:    global_store_dword v0, v2, s[0:1]
+; GFX10-32-GISEL-NEXT:    s_endpgm
+;
+; GFX10-64-SDAG-LABEL: test_waterfall_readlane:
+; GFX10-64-SDAG:       ; %bb.0:
+; GFX10-64-SDAG-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v2, 0
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 3, v[0:1]
+; GFX10-64-SDAG-NEXT:    v_add_co_u32 v0, vcc, s2, v0
+; GFX10-64-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc, s3, v1, vcc
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-SDAG-NEXT:    global_load_dwordx2 v[0:1], v[0:1], off
+; GFX10-64-SDAG-NEXT:  .LBB0_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s6, v1
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v1
+; GFX10-64-SDAG-NEXT:    v_readlane_b32 s6, v0, s6
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v3, s6
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[4:5], s[4:5]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr0_vgpr1
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB0_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX10-64-SDAG-NEXT:    global_store_dword v2, v3, s[0:1]
+; GFX10-64-SDAG-NEXT:    s_endpgm
+;
+; GFX10-64-GISEL-LABEL: test_waterfall_readlane:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v2, s2
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v3, s3
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 3, v[0:1]
+; GFX10-64-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX10-64-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc, v3, v1, vcc
+; GFX10-64-GISEL-NEXT:    global_load_dwordx2 v[0:1], v[0:1], off
+; GFX10-64-GISEL-NEXT:  .LBB0_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s6, v1
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s6, v1
+; GFX10-64-GISEL-NEXT:    v_readlane_b32 s6, v0, s6
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v2, s6
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[4:5], s[4:5]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr0_vgpr1
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB0_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-GISEL-NEXT:    global_store_dword v0, v2, s[0:1]
+; GFX10-64-GISEL-NEXT:    s_endpgm
+;
+; GFX1150-SDAG-LABEL: test_waterfall_readlane:
+; GFX1150-SDAG:       ; %bb.0:
+; GFX1150-SDAG-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v2, 0
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 3, v[0:1]
+; GFX1150-SDAG-NEXT:    v_add_co_u32 v0, vcc, s2, v0
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s3, v1, vcc
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-SDAG-NEXT:    global_load_b64 v[0:1], v[0:1], off
+; GFX1150-SDAG-NEXT:  .LBB0_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s6, v1
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_1) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v1
+; GFX1150-SDAG-NEXT:    v_readlane_b32 s6, v0, s6
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v3, s6
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[4:5], s[4:5]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr0_vgpr1
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB0_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX1150-SDAG-NEXT:    global_store_b32 v2, v3, s[0:1]
+; GFX1150-SDAG-NEXT:    s_endpgm
+;
+; GFX1150-GISEL-LABEL: test_waterfall_readlane:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v2, s2
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v3, s3
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 3, v[0:1]
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX1150-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v3, v1, vcc
+; GFX1150-GISEL-NEXT:    global_load_b64 v[0:1], v[0:1], off
+; GFX1150-GISEL-NEXT:  .LBB0_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s6, v1
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_1) | instid1(VALU_DEP_1)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s6, v1
+; GFX1150-GISEL-NEXT:    v_readlane_b32 s6, v0, s6
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v2, s6
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[4:5], s[4:5]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr0_vgpr1
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB0_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-GISEL-NEXT:    global_store_b32 v0, v2, s[0:1]
+; GFX1150-GISEL-NEXT:    s_endpgm
+;
+; GFX12-SDAG-LABEL: test_waterfall_readlane:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v2, 0
+; GFX12-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_lshlrev_b64_e32 v[0:1], 3, v[0:1]
+; GFX12-SDAG-NEXT:    v_add_co_u32 v0, vcc, s2, v0
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s3, v1, vcc
+; GFX12-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-SDAG-NEXT:    global_load_b64 v[0:1], v[0:1], off
+; GFX12-SDAG-NEXT:  .LBB0_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x0
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s6, v1
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_2) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v1
+; GFX12-SDAG-NEXT:    v_readlane_b32 s6, v0, s6
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v3, s6
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[4:5], s[4:5]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr0_vgpr1
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB0_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX12-SDAG-NEXT:    global_store_b32 v2, v3, s[0:1]
+; GFX12-SDAG-NEXT:    s_endpgm
+;
+; GFX12-GISEL-LABEL: test_waterfall_readlane:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v2, s2
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v3, s3
+; GFX12-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-GISEL-NEXT:    v_lshlrev_b64_e32 v[0:1], 3, v[0:1]
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX12-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v3, v1, vcc
+; GFX12-GISEL-NEXT:    global_load_b64 v[0:1], v[0:1], off
+; GFX12-GISEL-NEXT:  .LBB0_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x0
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s6, v1
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_2) | instid1(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s6, v1
+; GFX12-GISEL-NEXT:    v_readlane_b32 s6, v0, s6
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v2, s6
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[4:5], s[4:5]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr0_vgpr1
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB0_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-GISEL-NEXT:    global_store_b32 v0, v2, s[0:1]
+; GFX12-GISEL-NEXT:    s_endpgm
+  %gep.in = getelementptr <2 x i32>, <2 x i32> addrspace(1)* %in, i32 %tid
+  %args = load <2 x i32>, <2 x i32> addrspace(1)* %gep.in
+  %value = extractelement <2 x i32> %args, i32 0
+  %lane = extractelement <2 x i32> %args, i32 1
+  %wf_token = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %lane)
+  %readlane = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %wf_token, i32 %lane)
+  %readlane1 = call i32 @llvm.amdgcn.readlane(i32 %value, i32 %readlane)
+  %readlane2 = call i32 @llvm.amdgcn.waterfall.end.i32(i32 %wf_token, i32 %readlane1)
+  ; This store instruction should be outside the waterfall loop and the value
+  ; being stored generated incrementally in the loop itself
+  store i32 %readlane2, i32 addrspace(1)* %out, align 4
+
+  ret void
+}
+
+define amdgpu_ps <4 x float> @test_waterfall_non_uniform_img(<8 x i32> addrspace(4)* inreg %in, i32 %index, float %s, <4 x i32> inreg %samp) #1 {
+; VI-SDAG-LABEL: test_waterfall_non_uniform_img:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; VI-SDAG-NEXT:    s_wqm_b64 exec, exec
+; VI-SDAG-NEXT:    s_mov_b32 s7, s5
+; VI-SDAG-NEXT:    s_mov_b32 s6, s4
+; VI-SDAG-NEXT:    s_mov_b32 s5, s3
+; VI-SDAG-NEXT:    s_mov_b32 s4, s2
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, v1
+; VI-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; VI-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; VI-SDAG-NEXT:  .LBB1_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s10, v5
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[12:13], s10, v5
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[20:21], s[12:13]
+; VI-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; VI-SDAG-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; VI-SDAG-NEXT:    s_add_u32 s10, s0, s10
+; VI-SDAG-NEXT:    s_addc_u32 s11, s1, s11
+; VI-SDAG-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr5
+; VI-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; VI-SDAG-NEXT:    image_sample v[0:3], v4, s[12:19], s[4:7] dmask:0xf
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr4
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[20:21]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB1_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[2:3]
+; VI-SDAG-NEXT:    s_and_b64 exec, exec, s[8:9]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    ; return to shader part epilog
+;
+; VI-GISEL-LABEL: test_waterfall_non_uniform_img:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; VI-GISEL-NEXT:    s_wqm_b64 exec, exec
+; VI-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, v1
+; VI-GISEL-NEXT:    s_mov_b32 s8, s2
+; VI-GISEL-NEXT:    s_mov_b32 s9, s3
+; VI-GISEL-NEXT:    s_mov_b32 s10, s4
+; VI-GISEL-NEXT:    s_mov_b32 s11, s5
+; VI-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; VI-GISEL-NEXT:  .LBB1_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s4, v5
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[12:13], s4, v5
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[20:21], s[12:13]
+; VI-GISEL-NEXT:    s_ashr_i32 s5, s4, 31
+; VI-GISEL-NEXT:    s_lshl_b64 s[4:5], s[4:5], 5
+; VI-GISEL-NEXT:    s_add_u32 s4, s0, s4
+; VI-GISEL-NEXT:    s_addc_u32 s5, s1, s5
+; VI-GISEL-NEXT:    s_load_dwordx8 s[12:19], s[4:5], 0x0
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr5
+; VI-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; VI-GISEL-NEXT:    image_sample v[0:3], v4, s[12:19], s[8:11] dmask:0xf
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr4
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[20:21]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB1_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; VI-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX9-SDAG-LABEL: test_waterfall_non_uniform_img:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX9-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX9-SDAG-NEXT:    s_mov_b32 s7, s5
+; GFX9-SDAG-NEXT:    s_mov_b32 s6, s4
+; GFX9-SDAG-NEXT:    s_mov_b32 s5, s3
+; GFX9-SDAG-NEXT:    s_mov_b32 s4, s2
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v4, v1
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX9-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX9-SDAG-NEXT:  .LBB1_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s10, v5
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[12:13], s10, v5
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[20:21], s[12:13]
+; GFX9-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX9-SDAG-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX9-SDAG-NEXT:    s_add_u32 s10, s0, s10
+; GFX9-SDAG-NEXT:    s_addc_u32 s11, s1, s11
+; GFX9-SDAG-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX9-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX9-SDAG-NEXT:    image_sample v[0:3], v4, s[12:19], s[4:7] dmask:0xf
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[20:21]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB1_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX9-SDAG-NEXT:    s_and_b64 exec, exec, s[8:9]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX9-GISEL-LABEL: test_waterfall_non_uniform_img:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX9-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v4, v1
+; GFX9-GISEL-NEXT:    s_mov_b32 s8, s2
+; GFX9-GISEL-NEXT:    s_mov_b32 s9, s3
+; GFX9-GISEL-NEXT:    s_mov_b32 s10, s4
+; GFX9-GISEL-NEXT:    s_mov_b32 s11, s5
+; GFX9-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX9-GISEL-NEXT:  .LBB1_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s4, v5
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[12:13], s4, v5
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[20:21], s[12:13]
+; GFX9-GISEL-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX9-GISEL-NEXT:    s_lshl_b64 s[4:5], s[4:5], 5
+; GFX9-GISEL-NEXT:    s_add_u32 s4, s0, s4
+; GFX9-GISEL-NEXT:    s_addc_u32 s5, s1, s5
+; GFX9-GISEL-NEXT:    s_load_dwordx8 s[12:19], s[4:5], 0x0
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX9-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX9-GISEL-NEXT:    image_sample v[0:3], v4, s[12:19], s[8:11] dmask:0xf
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[20:21]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB1_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX9-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-SDAG-LABEL: test_waterfall_non_uniform_img:
+; GFX10-32-SDAG:       ; %bb.0:
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s8, exec_lo
+; GFX10-32-SDAG-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v4, v1
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s7, s5
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s6, s4
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s5, s3
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s4, s2
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s3, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s2, exec_lo
+; GFX10-32-SDAG-NEXT:  .LBB1_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s10, v5
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s10, v5
+; GFX10-32-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX10-32-SDAG-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX10-32-SDAG-NEXT:    s_add_u32 s10, s0, s10
+; GFX10-32-SDAG-NEXT:    s_addc_u32 s11, s1, s11
+; GFX10-32-SDAG-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
+; GFX10-32-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-32-SDAG-NEXT:    image_sample v[0:3], v4, s[12:19], s[4:7] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s3, s3
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB1_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s2
+; GFX10-32-SDAG-NEXT:    s_and_b32 exec_lo, exec_lo, s8
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-GISEL-LABEL: test_waterfall_non_uniform_img:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s6, exec_lo
+; GFX10-32-GISEL-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v4, v1
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s8, s2
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s9, s3
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s10, s4
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s11, s5
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s3, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s2, exec_lo
+; GFX10-32-GISEL-NEXT:  .LBB1_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s4, v5
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s4, v5
+; GFX10-32-GISEL-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX10-32-GISEL-NEXT:    s_lshl_b64 s[4:5], s[4:5], 5
+; GFX10-32-GISEL-NEXT:    s_add_u32 s4, s0, s4
+; GFX10-32-GISEL-NEXT:    s_addc_u32 s5, s1, s5
+; GFX10-32-GISEL-NEXT:    s_load_dwordx8 s[12:19], s[4:5], 0x0
+; GFX10-32-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-32-GISEL-NEXT:    image_sample v[0:3], v4, s[12:19], s[8:11] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s3, s3
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB1_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s2
+; GFX10-32-GISEL-NEXT:    s_and_b32 exec_lo, exec_lo, s6
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-SDAG-LABEL: test_waterfall_non_uniform_img:
+; GFX10-64-SDAG:       ; %bb.0:
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX10-64-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v4, v1
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s7, s5
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s6, s4
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s5, s3
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s4, s2
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-SDAG-NEXT:  .LBB1_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s12, v5
+; GFX10-64-SDAG-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX10-64-SDAG-NEXT:    s_lshl_b64 s[12:13], s[12:13], 5
+; GFX10-64-SDAG-NEXT:    s_add_u32 s12, s0, s12
+; GFX10-64-SDAG-NEXT:    s_addc_u32 s13, s1, s13
+; GFX10-64-SDAG-NEXT:    s_load_dwordx8 s[12:19], s[12:13], 0x0
+; GFX10-64-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-64-SDAG-NEXT:    image_sample v[0:3], v4, s[12:19], s[4:7] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[10:11], s[10:11]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB1_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX10-64-SDAG-NEXT:    s_and_b64 exec, exec, s[8:9]
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-GISEL-LABEL: test_waterfall_non_uniform_img:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v4, v1
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s8, s2
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s9, s3
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s10, s4
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s11, s5
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-GISEL-NEXT:  .LBB1_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s12, v5
+; GFX10-64-GISEL-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX10-64-GISEL-NEXT:    s_lshl_b64 s[12:13], s[12:13], 5
+; GFX10-64-GISEL-NEXT:    s_add_u32 s12, s0, s12
+; GFX10-64-GISEL-NEXT:    s_addc_u32 s13, s1, s13
+; GFX10-64-GISEL-NEXT:    s_load_dwordx8 s[12:19], s[12:13], 0x0
+; GFX10-64-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-64-GISEL-NEXT:    image_sample v[0:3], v4, s[12:19], s[8:11] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[4:5], s[4:5]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB1_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX10-64-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX1150-SDAG-LABEL: test_waterfall_non_uniform_img:
+; GFX1150-SDAG:       ; %bb.0:
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX1150-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v4, v1
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX1150-SDAG-NEXT:    s_mov_b32 s7, s5
+; GFX1150-SDAG-NEXT:    s_mov_b32 s6, s4
+; GFX1150-SDAG-NEXT:    s_mov_b32 s5, s3
+; GFX1150-SDAG-NEXT:    s_mov_b32 s4, s2
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-SDAG-NEXT:  .LBB1_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_1) | instid1(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s12, v5
+; GFX1150-SDAG-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX1150-SDAG-NEXT:    s_lshl_b64 s[12:13], s[12:13], 5
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_add_u32 s12, s0, s12
+; GFX1150-SDAG-NEXT:    s_addc_u32 s13, s1, s13
+; GFX1150-SDAG-NEXT:    s_load_b256 s[12:19], s[12:13], 0x0
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0) lgkmcnt(0)
+; GFX1150-SDAG-NEXT:    image_sample v[0:3], v4, s[12:19], s[4:7] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[10:11], s[10:11]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB1_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_and_b64 exec, exec, s[8:9]
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX1150-GISEL-LABEL: test_waterfall_non_uniform_img:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v4, v1
+; GFX1150-GISEL-NEXT:    s_mov_b32 s8, s2
+; GFX1150-GISEL-NEXT:    s_mov_b32 s9, s3
+; GFX1150-GISEL-NEXT:    s_mov_b32 s10, s4
+; GFX1150-GISEL-NEXT:    s_mov_b32 s11, s5
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-GISEL-NEXT:  .LBB1_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_1) | instid1(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s12, v5
+; GFX1150-GISEL-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX1150-GISEL-NEXT:    s_lshl_b64 s[12:13], s[12:13], 5
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_add_u32 s12, s0, s12
+; GFX1150-GISEL-NEXT:    s_addc_u32 s13, s1, s13
+; GFX1150-GISEL-NEXT:    s_load_b256 s[12:19], s[12:13], 0x0
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0) lgkmcnt(0)
+; GFX1150-GISEL-NEXT:    image_sample v[0:3], v4, s[12:19], s[8:11] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[4:5], s[4:5]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB1_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX12-SDAG-LABEL: test_waterfall_non_uniform_img:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX12-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v4, v1
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX12-SDAG-NEXT:    s_mov_b32 s7, s5
+; GFX12-SDAG-NEXT:    s_mov_b32 s6, s4
+; GFX12-SDAG-NEXT:    s_mov_b32 s5, s3
+; GFX12-SDAG-NEXT:    s_mov_b32 s4, s2
+; GFX12-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; GFX12-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-SDAG-NEXT:  .LBB1_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s12, v5
+; GFX12-SDAG-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    s_lshl_b64 s[12:13], s[12:13], 5
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    s_add_nc_u64 s[12:13], s[0:1], s[12:13]
+; GFX12-SDAG-NEXT:    s_load_b256 s[12:19], s[12:13], 0x0
+; GFX12-SDAG-NEXT:    s_wait_kmcnt 0x0
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    image_sample v[0:3], v4, s[12:19], s[4:7] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[10:11], s[10:11]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB1_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-SDAG-NEXT:    s_and_b64 exec, exec, s[8:9]
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX12-GISEL-LABEL: test_waterfall_non_uniform_img:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v4, v1
+; GFX12-GISEL-NEXT:    s_mov_b32 s8, s2
+; GFX12-GISEL-NEXT:    s_mov_b32 s9, s3
+; GFX12-GISEL-NEXT:    s_mov_b32 s10, s4
+; GFX12-GISEL-NEXT:    s_mov_b32 s11, s5
+; GFX12-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-GISEL-NEXT:  .LBB1_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s12, v5
+; GFX12-GISEL-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_lshl_b64 s[12:13], s[12:13], 5
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_add_co_u32 s12, s0, s12
+; GFX12-GISEL-NEXT:    s_add_co_ci_u32 s13, s1, s13
+; GFX12-GISEL-NEXT:    s_load_b256 s[12:19], s[12:13], 0x0
+; GFX12-GISEL-NEXT:    s_wait_kmcnt 0x0
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    image_sample v[0:3], v4, s[12:19], s[8:11] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[4:5], s[4:5]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB1_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    ; return to shader part epilog
+  %wf_token = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %index)
+  %s_idx = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %wf_token, i32 %index)
+  %ptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4) * %ptr, align 32
+  %r = call <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32(i32 15, float %s, <8 x i32> %rsrc, <4 x i32> %samp, i1 0, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %wf_token, <4 x float> %r)
+
+  ret <4 x float> %r1
+}
+
+define amdgpu_ps <4 x float> @test_waterfall_non_uniform_img_single_read(<8 x i32> addrspace(4)* inreg %in, i32 %index, float %s, <4 x i32> inreg %samp) #1 {
+; VI-SDAG-LABEL: test_waterfall_non_uniform_img_single_read:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; VI-SDAG-NEXT:    s_wqm_b64 exec, exec
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; VI-SDAG-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; VI-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; VI-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[4:5]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v2, s1
+; VI-SDAG-NEXT:    v_add_u32_e32 v0, vcc, s0, v0
+; VI-SDAG-NEXT:    v_addc_u32_e32 v1, vcc, v2, v1, vcc
+; VI-SDAG-NEXT:    v_add_u32_e32 v2, vcc, 16, v0
+; VI-SDAG-NEXT:    v_addc_u32_e32 v3, vcc, 0, v1, vcc
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[7:10], v[0:1]
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[11:14], v[2:3]
+; VI-SDAG-NEXT:    s_mov_b32 s7, s5
+; VI-SDAG-NEXT:    s_mov_b32 s6, s4
+; VI-SDAG-NEXT:    s_mov_b32 s5, s3
+; VI-SDAG-NEXT:    s_mov_b32 s4, s2
+; VI-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:  .LBB2_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s2, v4
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v4
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s12, v7
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s13, v8
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s14, v9
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s15, v10
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s16, v11
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s17, v12
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s18, v13
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s19, v14
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr4
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; VI-SDAG-NEXT:    s_nop 4
+; VI-SDAG-NEXT:    image_sample v[0:3], v6, s[12:19], s[4:7] dmask:0xf
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr6
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB2_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; VI-SDAG-NEXT:    s_and_b64 exec, exec, s[8:9]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    ; return to shader part epilog
+;
+; VI-GISEL-LABEL: test_waterfall_non_uniform_img_single_read:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; VI-GISEL-NEXT:    s_wqm_b64 exec, exec
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; VI-GISEL-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; VI-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; VI-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[4:5]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; VI-GISEL-NEXT:    v_add_u32_e32 v0, vcc, v2, v0
+; VI-GISEL-NEXT:    v_addc_u32_e32 v1, vcc, v3, v1, vcc
+; VI-GISEL-NEXT:    v_add_u32_e32 v2, vcc, 16, v0
+; VI-GISEL-NEXT:    v_addc_u32_e32 v3, vcc, 0, v1, vcc
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[7:10], v[0:1]
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[11:14], v[2:3]
+; VI-GISEL-NEXT:    s_mov_b32 s8, s2
+; VI-GISEL-NEXT:    s_mov_b32 s9, s3
+; VI-GISEL-NEXT:    s_mov_b32 s10, s4
+; VI-GISEL-NEXT:    s_mov_b32 s11, s5
+; VI-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:  .LBB2_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s2, v4
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v4
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s12, v7
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s13, v8
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s14, v9
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s15, v10
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s16, v11
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s17, v12
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s18, v13
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s19, v14
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr4
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; VI-GISEL-NEXT:    s_nop 4
+; VI-GISEL-NEXT:    image_sample v[0:3], v6, s[12:19], s[8:11] dmask:0xf
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr6
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB2_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; VI-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX9-SDAG-LABEL: test_waterfall_non_uniform_img_single_read:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX9-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX9-SDAG-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX9-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[4:5]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v2, s1
+; GFX9-SDAG-NEXT:    v_add_co_u32_e32 v0, vcc, s0, v0
+; GFX9-SDAG-NEXT:    v_addc_co_u32_e32 v1, vcc, v2, v1, vcc
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[11:14], v[0:1], off offset:16
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[7:10], v[0:1], off
+; GFX9-SDAG-NEXT:    s_mov_b32 s7, s5
+; GFX9-SDAG-NEXT:    s_mov_b32 s6, s4
+; GFX9-SDAG-NEXT:    s_mov_b32 s5, s3
+; GFX9-SDAG-NEXT:    s_mov_b32 s4, s2
+; GFX9-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:  .LBB2_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s2, v4
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v4
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s12, v7
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s13, v8
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s14, v9
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s15, v10
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s16, v11
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s17, v12
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s18, v13
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s19, v14
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX9-SDAG-NEXT:    s_nop 4
+; GFX9-SDAG-NEXT:    image_sample v[0:3], v6, s[12:19], s[4:7] dmask:0xf
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB2_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX9-SDAG-NEXT:    s_and_b64 exec, exec, s[8:9]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX9-GISEL-LABEL: test_waterfall_non_uniform_img_single_read:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX9-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX9-GISEL-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX9-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[4:5]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX9-GISEL-NEXT:    v_add_co_u32_e32 v0, vcc, v2, v0
+; GFX9-GISEL-NEXT:    v_addc_co_u32_e32 v1, vcc, v3, v1, vcc
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[7:10], v[0:1], off
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[11:14], v[0:1], off offset:16
+; GFX9-GISEL-NEXT:    s_mov_b32 s8, s2
+; GFX9-GISEL-NEXT:    s_mov_b32 s9, s3
+; GFX9-GISEL-NEXT:    s_mov_b32 s10, s4
+; GFX9-GISEL-NEXT:    s_mov_b32 s11, s5
+; GFX9-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:  .LBB2_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s2, v4
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v4
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s12, v7
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s13, v8
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s14, v9
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s15, v10
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s16, v11
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s17, v12
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s18, v13
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s19, v14
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX9-GISEL-NEXT:    s_nop 4
+; GFX9-GISEL-NEXT:    image_sample v[0:3], v6, s[12:19], s[8:11] dmask:0xf
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB2_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX9-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-SDAG-LABEL: test_waterfall_non_uniform_img_single_read:
+; GFX10-32-SDAG:       ; %bb.0:
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s8, exec_lo
+; GFX10-32-SDAG-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s7, s5
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s6, s4
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s5, s3
+; GFX10-32-SDAG-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s4, s2
+; GFX10-32-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[4:5]
+; GFX10-32-SDAG-NEXT:    v_add_co_u32 v0, vcc_lo, s0, v0
+; GFX10-32-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, s1, v1, vcc_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-SDAG-NEXT:    s_clause 0x1
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[11:14], v[0:1], off offset:16
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[7:10], v[0:1], off
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:  .LBB2_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s2, v4
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s2, v4
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s12, v7
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s13, v8
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s14, v9
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s15, v10
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s16, v11
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s17, v12
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s18, v13
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s19, v14
+; GFX10-32-SDAG-NEXT:    image_sample v[0:3], v6, s[12:19], s[4:7] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB2_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-SDAG-NEXT:    s_and_b32 exec_lo, exec_lo, s8
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-GISEL-LABEL: test_waterfall_non_uniform_img_single_read:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s6, exec_lo
+; GFX10-32-GISEL-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s8, s2
+; GFX10-32-GISEL-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s9, s3
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s10, s4
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s11, s5
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[4:5]
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-GISEL-NEXT:    v_add_co_u32 v0, vcc_lo, v2, v0
+; GFX10-32-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, v3, v1, vcc_lo
+; GFX10-32-GISEL-NEXT:    s_clause 0x1
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[7:10], v[0:1], off
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[11:14], v[0:1], off offset:16
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:  .LBB2_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s2, v4
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s2, v4
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s12, v7
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s13, v8
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s14, v9
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s15, v10
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s16, v11
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s17, v12
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s18, v13
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s19, v14
+; GFX10-32-GISEL-NEXT:    image_sample v[0:3], v6, s[12:19], s[8:11] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB2_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-GISEL-NEXT:    s_and_b32 exec_lo, exec_lo, s6
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-SDAG-LABEL: test_waterfall_non_uniform_img_single_read:
+; GFX10-64-SDAG:       ; %bb.0:
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX10-64-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s7, s5
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s6, s4
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s5, s3
+; GFX10-64-SDAG-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s4, s2
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[4:5]
+; GFX10-64-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX10-64-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc, s1, v1, vcc
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-SDAG-NEXT:    s_clause 0x1
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[11:14], v[0:1], off offset:16
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[7:10], v[0:1], off
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:  .LBB2_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s10, v4
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s10, v4
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s12, v7
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s13, v8
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s14, v9
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s15, v10
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s16, v11
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s17, v12
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s18, v13
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s19, v14
+; GFX10-64-SDAG-NEXT:    image_sample v[0:3], v6, s[12:19], s[4:7] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB2_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-SDAG-NEXT:    s_and_b64 exec, exec, s[8:9]
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-GISEL-LABEL: test_waterfall_non_uniform_img_single_read:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s8, s2
+; GFX10-64-GISEL-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s9, s3
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s10, s4
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s11, s5
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[4:5]
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX10-64-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc, v3, v1, vcc
+; GFX10-64-GISEL-NEXT:    s_clause 0x1
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[7:10], v[0:1], off
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[11:14], v[0:1], off offset:16
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:  .LBB2_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s4, v4
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v4
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s12, v7
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s13, v8
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s14, v9
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s15, v10
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s16, v11
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s17, v12
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s18, v13
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s19, v14
+; GFX10-64-GISEL-NEXT:    image_sample v[0:3], v6, s[12:19], s[8:11] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB2_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX1150-SDAG-LABEL: test_waterfall_non_uniform_img_single_read:
+; GFX1150-SDAG:       ; %bb.0:
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX1150-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX1150-SDAG-NEXT:    s_mov_b32 s7, s5
+; GFX1150-SDAG-NEXT:    s_mov_b32 s6, s4
+; GFX1150-SDAG-NEXT:    s_mov_b32 s5, s3
+; GFX1150-SDAG-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; GFX1150-SDAG-NEXT:    s_mov_b32 s4, s2
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[4:5]
+; GFX1150-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s1, v1, vcc
+; GFX1150-SDAG-NEXT:    s_clause 0x1
+; GFX1150-SDAG-NEXT:    global_load_b128 v[11:14], v[0:1], off offset:16
+; GFX1150-SDAG-NEXT:    global_load_b128 v[7:10], v[0:1], off
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:  .LBB2_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s10, v4
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s10, v4
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s12, v7
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s13, v8
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s14, v9
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s15, v10
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s16, v11
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s17, v12
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s18, v13
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s19, v14
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    image_sample v[0:3], v6, s[12:19], s[4:7] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB2_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_and_b64 exec, exec, s[8:9]
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX1150-GISEL-LABEL: test_waterfall_non_uniform_img_single_read:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX1150-GISEL-NEXT:    s_mov_b32 s8, s2
+; GFX1150-GISEL-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; GFX1150-GISEL-NEXT:    s_mov_b32 s9, s3
+; GFX1150-GISEL-NEXT:    s_mov_b32 s10, s4
+; GFX1150-GISEL-NEXT:    s_mov_b32 s11, s5
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[4:5]
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX1150-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v3, v1, vcc
+; GFX1150-GISEL-NEXT:    s_clause 0x1
+; GFX1150-GISEL-NEXT:    global_load_b128 v[7:10], v[0:1], off
+; GFX1150-GISEL-NEXT:    global_load_b128 v[11:14], v[0:1], off offset:16
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:  .LBB2_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s4, v4
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v4
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s12, v7
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s13, v8
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s14, v9
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s15, v10
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s16, v11
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s17, v12
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s18, v13
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s19, v14
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    image_sample v[0:3], v6, s[12:19], s[8:11] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB2_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX12-SDAG-LABEL: test_waterfall_non_uniform_img_single_read:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX12-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX12-SDAG-NEXT:    s_mov_b32 s7, s5
+; GFX12-SDAG-NEXT:    s_mov_b32 s6, s4
+; GFX12-SDAG-NEXT:    s_mov_b32 s5, s3
+; GFX12-SDAG-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; GFX12-SDAG-NEXT:    s_mov_b32 s4, s2
+; GFX12-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_lshlrev_b64_e32 v[0:1], 5, v[4:5]
+; GFX12-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s1, v1, vcc
+; GFX12-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-SDAG-NEXT:    s_clause 0x1
+; GFX12-SDAG-NEXT:    global_load_b128 v[11:14], v[0:1], off offset:16
+; GFX12-SDAG-NEXT:    global_load_b128 v[7:10], v[0:1], off
+; GFX12-SDAG-NEXT:  .LBB2_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s10, v4
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s10, v4
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x0
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s12, v7
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s13, v8
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s14, v9
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s15, v10
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s16, v11
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s17, v12
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s18, v13
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s19, v14
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    image_sample v[0:3], v6, s[12:19], s[4:7] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB2_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-SDAG-NEXT:    s_and_b64 exec, exec, s[8:9]
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX12-GISEL-LABEL: test_waterfall_non_uniform_img_single_read:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX12-GISEL-NEXT:    s_mov_b32 s8, s2
+; GFX12-GISEL-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; GFX12-GISEL-NEXT:    s_mov_b32 s9, s3
+; GFX12-GISEL-NEXT:    s_mov_b32 s10, s4
+; GFX12-GISEL-NEXT:    s_mov_b32 s11, s5
+; GFX12-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-GISEL-NEXT:    v_lshlrev_b64_e32 v[0:1], 5, v[4:5]
+; GFX12-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX12-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v3, v1, vcc
+; GFX12-GISEL-NEXT:    s_clause 0x1
+; GFX12-GISEL-NEXT:    global_load_b128 v[7:10], v[0:1], off
+; GFX12-GISEL-NEXT:    global_load_b128 v[11:14], v[0:1], off offset:16
+; GFX12-GISEL-NEXT:  .LBB2_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s4, v4
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v4
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s12, v7
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s13, v8
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s14, v9
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s15, v10
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x0
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s16, v11
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s17, v12
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s18, v13
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s19, v14
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    image_sample v[0:3], v6, s[12:19], s[8:11] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB2_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    ; return to shader part epilog
+  %ptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %index
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4) * %ptr, align 32
+  %wf_token = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %index)
+  %s_rsrc = call <8 x i32> @llvm.amdgcn.waterfall.readfirstlane.v8i32.v8i32(i32 %wf_token, <8 x i32> %rsrc)
+  %r = call <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32(i32 15, float %s, <8 x i32> %s_rsrc, <4 x i32> %samp, i1 0, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %wf_token, <4 x float> %r)
+
+  ret <4 x float> %r1
+}
+
+define amdgpu_ps void @test_multiple_groups(i32 addrspace(1)* inreg %out1, i32 addrspace(1)* inreg %out2, i32 %idx1, i32 %idx2, i32 %val) #1 {
+; VI-SDAG-LABEL: test_multiple_groups:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    v_mov_b32_e32 v3, s2
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, s3
+; VI-SDAG-NEXT:    v_mov_b32_e32 v5, s0
+; VI-SDAG-NEXT:    v_mov_b32_e32 v6, s1
+; VI-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; VI-SDAG-NEXT:  .LBB3_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s4, v0
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[2:3], s4, v0
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; VI-SDAG-NEXT:    s_nop 1
+; VI-SDAG-NEXT:    v_readlane_b32 s4, v2, s4
+; VI-SDAG-NEXT:    v_mov_b32_e32 v7, s4
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr0
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB3_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; VI-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; VI-SDAG-NEXT:    flat_store_dword v[5:6], v7
+; VI-SDAG-NEXT:  .LBB3_3: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s4, v1
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[2:3], s4, v1
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; VI-SDAG-NEXT:    s_nop 1
+; VI-SDAG-NEXT:    v_readlane_b32 s4, v2, s4
+; VI-SDAG-NEXT:    v_mov_b32_e32 v0, s4
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr1
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr2
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB3_3
+; VI-SDAG-NEXT:  ; %bb.4:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; VI-SDAG-NEXT:    flat_store_dword v[3:4], v0
+; VI-SDAG-NEXT:    s_endpgm
+;
+; VI-GISEL-LABEL: test_multiple_groups:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; VI-GISEL-NEXT:  .LBB3_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s8, v0
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[6:7], s8, v0
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[6:7], s[6:7]
+; VI-GISEL-NEXT:    s_nop 1
+; VI-GISEL-NEXT:    v_readlane_b32 s8, v2, s8
+; VI-GISEL-NEXT:    v_mov_b32_e32 v3, s8
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr0
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[6:7]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB3_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[4:5]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v5, s1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, s0
+; VI-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; VI-GISEL-NEXT:    flat_store_dword v[4:5], v3
+; VI-GISEL-NEXT:  .LBB3_3: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s6, v1
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[4:5], s6, v1
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[4:5], s[4:5]
+; VI-GISEL-NEXT:    s_nop 1
+; VI-GISEL-NEXT:    v_readlane_b32 s6, v2, s6
+; VI-GISEL-NEXT:    v_mov_b32_e32 v0, s6
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr1
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr2
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[4:5]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB3_3
+; VI-GISEL-NEXT:  ; %bb.4:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v1, s2
+; VI-GISEL-NEXT:    v_mov_b32_e32 v2, s3
+; VI-GISEL-NEXT:    flat_store_dword v[1:2], v0
+; VI-GISEL-NEXT:    s_endpgm
+;
+; GFX9-SDAG-LABEL: test_multiple_groups:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v3, 0
+; GFX9-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX9-SDAG-NEXT:  .LBB3_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s8, v0
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[6:7], s8, v0
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[6:7], s[6:7]
+; GFX9-SDAG-NEXT:    s_nop 1
+; GFX9-SDAG-NEXT:    v_readlane_b32 s8, v2, s8
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v4, s8
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr0
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[6:7]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB3_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[4:5]
+; GFX9-SDAG-NEXT:    global_store_dword v3, v4, s[0:1]
+; GFX9-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX9-SDAG-NEXT:  .LBB3_3: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s6, v1
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[4:5], s6, v1
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[4:5], s[4:5]
+; GFX9-SDAG-NEXT:    s_nop 1
+; GFX9-SDAG-NEXT:    v_readlane_b32 s6, v2, s6
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v0, s6
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr1
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr2
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[4:5]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB3_3
+; GFX9-SDAG-NEXT:  ; %bb.4:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX9-SDAG-NEXT:    global_store_dword v3, v0, s[2:3]
+; GFX9-SDAG-NEXT:    s_endpgm
+;
+; GFX9-GISEL-LABEL: test_multiple_groups:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX9-GISEL-NEXT:  .LBB3_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s8, v0
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[6:7], s8, v0
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[6:7], s[6:7]
+; GFX9-GISEL-NEXT:    s_nop 1
+; GFX9-GISEL-NEXT:    v_readlane_b32 s8, v2, s8
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v3, s8
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr0
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[6:7]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB3_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[4:5]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX9-GISEL-NEXT:    global_store_dword v0, v3, s[0:1]
+; GFX9-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX9-GISEL-NEXT:  .LBB3_3: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s6, v1
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[4:5], s6, v1
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[4:5], s[4:5]
+; GFX9-GISEL-NEXT:    s_nop 1
+; GFX9-GISEL-NEXT:    v_readlane_b32 s6, v2, s6
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v3, s6
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr1
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr2
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[4:5]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB3_3
+; GFX9-GISEL-NEXT:  ; %bb.4:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX9-GISEL-NEXT:    global_store_dword v0, v3, s[2:3]
+; GFX9-GISEL-NEXT:    s_endpgm
+;
+; GFX10-32-SDAG-LABEL: test_multiple_groups:
+; GFX10-32-SDAG:       ; %bb.0:
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v3, 0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s5, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-SDAG-NEXT:  .LBB3_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s6, v0
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s6, v0
+; GFX10-32-SDAG-NEXT:    v_readlane_b32 s6, v2, s6
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v4, s6
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s5, s5
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr0
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB3_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s4
+; GFX10-32-SDAG-NEXT:    global_store_dword v3, v4, s[0:1]
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-SDAG-NEXT:  .LBB3_3: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s4, v1
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s4, v1
+; GFX10-32-SDAG-NEXT:    v_readlane_b32 s4, v2, s4
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v0, s4
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr1
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr2
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB3_3
+; GFX10-32-SDAG-NEXT:  ; %bb.4:
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-SDAG-NEXT:    global_store_dword v3, v0, s[2:3]
+; GFX10-32-SDAG-NEXT:    s_endpgm
+;
+; GFX10-32-GISEL-LABEL: test_multiple_groups:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s5, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-GISEL-NEXT:  .LBB3_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s6, v0
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s6, v0
+; GFX10-32-GISEL-NEXT:    v_readlane_b32 s6, v2, s6
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v3, s6
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s5, s5
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr0
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB3_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s4
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-GISEL-NEXT:    global_store_dword v0, v3, s[0:1]
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-GISEL-NEXT:  .LBB3_3: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s4, v1
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s4, v1
+; GFX10-32-GISEL-NEXT:    v_readlane_b32 s4, v2, s4
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v3, s4
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr1
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr2
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB3_3
+; GFX10-32-GISEL-NEXT:  ; %bb.4:
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-GISEL-NEXT:    global_store_dword v0, v3, s[2:3]
+; GFX10-32-GISEL-NEXT:    s_endpgm
+;
+; GFX10-64-SDAG-LABEL: test_multiple_groups:
+; GFX10-64-SDAG:       ; %bb.0:
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v3, 0
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-SDAG-NEXT:  .LBB3_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s8, v0
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s8, v0
+; GFX10-64-SDAG-NEXT:    v_readlane_b32 s8, v2, s8
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v4, s8
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[6:7], s[6:7]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr0
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB3_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[4:5]
+; GFX10-64-SDAG-NEXT:    global_store_dword v3, v4, s[0:1]
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-SDAG-NEXT:  .LBB3_3: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s6, v1
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v1
+; GFX10-64-SDAG-NEXT:    v_readlane_b32 s6, v2, s6
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v0, s6
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[4:5], s[4:5]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr1
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr2
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB3_3
+; GFX10-64-SDAG-NEXT:  ; %bb.4:
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-SDAG-NEXT:    global_store_dword v3, v0, s[2:3]
+; GFX10-64-SDAG-NEXT:    s_endpgm
+;
+; GFX10-64-GISEL-LABEL: test_multiple_groups:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-GISEL-NEXT:  .LBB3_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s8, v0
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s8, v0
+; GFX10-64-GISEL-NEXT:    v_readlane_b32 s8, v2, s8
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v3, s8
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[6:7], s[6:7]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr0
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB3_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[4:5]
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-GISEL-NEXT:    global_store_dword v0, v3, s[0:1]
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-GISEL-NEXT:  .LBB3_3: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s6, v1
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s6, v1
+; GFX10-64-GISEL-NEXT:    v_readlane_b32 s6, v2, s6
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v3, s6
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[4:5], s[4:5]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr1
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr2
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB3_3
+; GFX10-64-GISEL-NEXT:  ; %bb.4:
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-GISEL-NEXT:    global_store_dword v0, v3, s[2:3]
+; GFX10-64-GISEL-NEXT:    s_endpgm
+;
+; GFX1150-SDAG-LABEL: test_multiple_groups:
+; GFX1150-SDAG:       ; %bb.0:
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v3, 0
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-SDAG-NEXT:  .LBB3_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s8, v0
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_1) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s8, v0
+; GFX1150-SDAG-NEXT:    v_readlane_b32 s8, v2, s8
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v4, s8
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[6:7], s[6:7]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr0
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB3_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[4:5]
+; GFX1150-SDAG-NEXT:    global_store_b32 v3, v4, s[0:1]
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-SDAG-NEXT:  .LBB3_3: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s6, v1
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_1) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v1
+; GFX1150-SDAG-NEXT:    v_readlane_b32 s6, v2, s6
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v0, s6
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[4:5], s[4:5]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr1
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr2
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB3_3
+; GFX1150-SDAG-NEXT:  ; %bb.4:
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-SDAG-NEXT:    global_store_b32 v3, v0, s[2:3]
+; GFX1150-SDAG-NEXT:    s_endpgm
+;
+; GFX1150-GISEL-LABEL: test_multiple_groups:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-GISEL-NEXT:  .LBB3_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s8, v0
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_1) | instid1(VALU_DEP_1)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s8, v0
+; GFX1150-GISEL-NEXT:    v_readlane_b32 s8, v2, s8
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v3, s8
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[6:7], s[6:7]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr0
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB3_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[4:5]
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-GISEL-NEXT:    global_store_b32 v0, v3, s[0:1]
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-GISEL-NEXT:  .LBB3_3: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s6, v1
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_1) | instid1(VALU_DEP_1)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s6, v1
+; GFX1150-GISEL-NEXT:    v_readlane_b32 s6, v2, s6
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v3, s6
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[4:5], s[4:5]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr1
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr2
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB3_3
+; GFX1150-GISEL-NEXT:  ; %bb.4:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-GISEL-NEXT:    global_store_b32 v0, v3, s[2:3]
+; GFX1150-GISEL-NEXT:    s_endpgm
+;
+; GFX12-SDAG-LABEL: test_multiple_groups:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v3, 0
+; GFX12-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-SDAG-NEXT:  .LBB3_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s8, v0
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_2) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s8, v0
+; GFX12-SDAG-NEXT:    v_readlane_b32 s8, v2, s8
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v4, s8
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[6:7], s[6:7]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr0
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB3_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[4:5]
+; GFX12-SDAG-NEXT:    global_store_b32 v3, v4, s[0:1]
+; GFX12-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-SDAG-NEXT:  .LBB3_3: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s6, v1
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_2) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v1
+; GFX12-SDAG-NEXT:    v_readlane_b32 s6, v2, s6
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v0, s6
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[4:5], s[4:5]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr1
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr2
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB3_3
+; GFX12-SDAG-NEXT:  ; %bb.4:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-SDAG-NEXT:    global_store_b32 v3, v0, s[2:3]
+; GFX12-SDAG-NEXT:    s_endpgm
+;
+; GFX12-GISEL-LABEL: test_multiple_groups:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-GISEL-NEXT:  .LBB3_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s8, v0
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_2) | instid1(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s8, v0
+; GFX12-GISEL-NEXT:    v_readlane_b32 s8, v2, s8
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v3, s8
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[6:7], s[6:7]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr0
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB3_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[4:5]
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-GISEL-NEXT:    global_store_b32 v0, v3, s[0:1]
+; GFX12-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-GISEL-NEXT:  .LBB3_3: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s6, v1
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_2) | instid1(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s6, v1
+; GFX12-GISEL-NEXT:    v_readlane_b32 s6, v2, s6
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v3, s6
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[4:5], s[4:5]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr1
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr2
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB3_3
+; GFX12-GISEL-NEXT:  ; %bb.4:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-GISEL-NEXT:    global_store_b32 v0, v3, s[2:3]
+; GFX12-GISEL-NEXT:    s_endpgm
+  %wf_token = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %idx1)
+  %readlane1 = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %wf_token, i32 %idx1)
+  %readlane1.1 = call i32 @llvm.amdgcn.readlane(i32 %val, i32 %readlane1)
+  %readlane1.2 = call i32 @llvm.amdgcn.waterfall.end.i32(i32 %wf_token, i32 %readlane1.1)
+  ; This store instruction should be outside the waterfall loop and the value
+  ; being stored generated incrementally in the loop itself
+  store i32 %readlane1.2, i32 addrspace(1)* %out1, align 4
+
+  %wf_token2 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %idx2)
+  %readlane2 = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %wf_token2, i32 %idx2)
+  %readlane2.1 = call i32 @llvm.amdgcn.readlane(i32 %val, i32 %readlane2)
+  %readlane2.2 = call i32 @llvm.amdgcn.waterfall.end.i32(i32 %wf_token2, i32 %readlane2.1)
+  store i32 %readlane2.2, i32 addrspace(1)* %out2, align 4
+
+  ret void
+}
+
+
+define amdgpu_ps <4 x float> @test_waterfall_non_uniform_img_multi_rl(<8 x i32> addrspace(4)* inreg %in, <4 x i32> addrspace(4)* inreg %samp_in, i32 %index, float %s, i32 %val) #1 {
+; VI-SDAG-LABEL: test_waterfall_non_uniform_img_multi_rl:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; VI-SDAG-NEXT:    s_wqm_b64 exec, exec
+; VI-SDAG-NEXT:    v_mov_b32_e32 v5, v2
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, v1
+; VI-SDAG-NEXT:    v_mov_b32_e32 v6, v0
+; VI-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; VI-SDAG-NEXT:  .LBB4_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s8, v6
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[10:11], s8, v6
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[20:21], s[10:11]
+; VI-SDAG-NEXT:    s_ashr_i32 s9, s8, 31
+; VI-SDAG-NEXT:    s_lshl_b64 s[8:9], s[8:9], 5
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s10, v5
+; VI-SDAG-NEXT:    s_add_u32 s8, s0, s8
+; VI-SDAG-NEXT:    s_addc_u32 s9, s1, s9
+; VI-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; VI-SDAG-NEXT:    s_lshl_b64 s[10:11], s[10:11], 4
+; VI-SDAG-NEXT:    s_add_u32 s16, s2, s10
+; VI-SDAG-NEXT:    s_addc_u32 s17, s3, s11
+; VI-SDAG-NEXT:    s_load_dwordx8 s[8:15], s[8:9], 0x0
+; VI-SDAG-NEXT:    s_load_dwordx4 s[16:19], s[16:17], 0x0
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr6
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr5
+; VI-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; VI-SDAG-NEXT:    image_sample v[0:3], v4, s[8:15], s[16:19] dmask:0xf
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr4
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[20:21]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB4_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[6:7]
+; VI-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    ; return to shader part epilog
+;
+; VI-GISEL-LABEL: test_waterfall_non_uniform_img_multi_rl:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; VI-GISEL-NEXT:    s_wqm_b64 exec, exec
+; VI-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, v1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v6, v2
+; VI-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; VI-GISEL-NEXT:  .LBB4_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s8, v5
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[10:11], s8, v5
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[20:21], s[10:11]
+; VI-GISEL-NEXT:    s_ashr_i32 s9, s8, 31
+; VI-GISEL-NEXT:    s_lshl_b64 s[8:9], s[8:9], 5
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s10, v6
+; VI-GISEL-NEXT:    s_add_u32 s8, s0, s8
+; VI-GISEL-NEXT:    s_addc_u32 s9, s1, s9
+; VI-GISEL-NEXT:    s_ashr_i32 s11, s10, 31
+; VI-GISEL-NEXT:    s_lshl_b64 s[10:11], s[10:11], 4
+; VI-GISEL-NEXT:    s_add_u32 s16, s2, s10
+; VI-GISEL-NEXT:    s_addc_u32 s17, s3, s11
+; VI-GISEL-NEXT:    s_load_dwordx8 s[8:15], s[8:9], 0x0
+; VI-GISEL-NEXT:    s_load_dwordx4 s[16:19], s[16:17], 0x0
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr5
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr6
+; VI-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; VI-GISEL-NEXT:    image_sample v[0:3], v4, s[8:15], s[16:19] dmask:0xf
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr4
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[20:21]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB4_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[6:7]
+; VI-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX9-SDAG-LABEL: test_waterfall_non_uniform_img_multi_rl:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX9-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v5, v2
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v4, v1
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v6, v0
+; GFX9-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX9-SDAG-NEXT:  .LBB4_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s8, v6
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[10:11], s8, v6
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[20:21], s[10:11]
+; GFX9-SDAG-NEXT:    s_ashr_i32 s9, s8, 31
+; GFX9-SDAG-NEXT:    s_lshl_b64 s[8:9], s[8:9], 5
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s10, v5
+; GFX9-SDAG-NEXT:    s_add_u32 s22, s0, s8
+; GFX9-SDAG-NEXT:    s_addc_u32 s23, s1, s9
+; GFX9-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX9-SDAG-NEXT:    s_lshl_b64 s[8:9], s[10:11], 4
+; GFX9-SDAG-NEXT:    s_add_u32 s24, s2, s8
+; GFX9-SDAG-NEXT:    s_addc_u32 s25, s3, s9
+; GFX9-SDAG-NEXT:    s_load_dwordx8 s[8:15], s[22:23], 0x0
+; GFX9-SDAG-NEXT:    s_load_dwordx4 s[16:19], s[24:25], 0x0
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX9-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX9-SDAG-NEXT:    image_sample v[0:3], v4, s[8:15], s[16:19] dmask:0xf
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[20:21]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB4_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX9-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX9-GISEL-LABEL: test_waterfall_non_uniform_img_multi_rl:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX9-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v4, v1
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v6, v2
+; GFX9-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX9-GISEL-NEXT:  .LBB4_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[10:11], s8, v5
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[20:21], s[10:11]
+; GFX9-GISEL-NEXT:    s_ashr_i32 s9, s8, 31
+; GFX9-GISEL-NEXT:    s_lshl_b64 s[8:9], s[8:9], 5
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s10, v6
+; GFX9-GISEL-NEXT:    s_add_u32 s22, s0, s8
+; GFX9-GISEL-NEXT:    s_addc_u32 s23, s1, s9
+; GFX9-GISEL-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX9-GISEL-NEXT:    s_lshl_b64 s[8:9], s[10:11], 4
+; GFX9-GISEL-NEXT:    s_add_u32 s24, s2, s8
+; GFX9-GISEL-NEXT:    s_addc_u32 s25, s3, s9
+; GFX9-GISEL-NEXT:    s_load_dwordx8 s[8:15], s[22:23], 0x0
+; GFX9-GISEL-NEXT:    s_load_dwordx4 s[16:19], s[24:25], 0x0
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX9-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX9-GISEL-NEXT:    image_sample v[0:3], v4, s[8:15], s[16:19] dmask:0xf
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[20:21]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB4_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX9-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-SDAG-LABEL: test_waterfall_non_uniform_img_multi_rl:
+; GFX10-32-SDAG:       ; %bb.0:
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-SDAG-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v5, v2
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v4, v1
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v6, v0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s6, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s5, exec_lo
+; GFX10-32-SDAG-NEXT:  .LBB4_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s8, v6
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s8, v6
+; GFX10-32-SDAG-NEXT:    s_ashr_i32 s9, s8, 31
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s10, v5
+; GFX10-32-SDAG-NEXT:    s_lshl_b64 s[8:9], s[8:9], 5
+; GFX10-32-SDAG-NEXT:    s_add_u32 s20, s0, s8
+; GFX10-32-SDAG-NEXT:    s_addc_u32 s21, s1, s9
+; GFX10-32-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX10-32-SDAG-NEXT:    s_lshl_b64 s[8:9], s[10:11], 4
+; GFX10-32-SDAG-NEXT:    s_add_u32 s22, s2, s8
+; GFX10-32-SDAG-NEXT:    s_addc_u32 s23, s3, s9
+; GFX10-32-SDAG-NEXT:    s_load_dwordx8 s[8:15], s[20:21], 0x0
+; GFX10-32-SDAG-NEXT:    s_load_dwordx4 s[16:19], s[22:23], 0x0
+; GFX10-32-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-32-SDAG-NEXT:    image_sample v[0:3], v4, s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s6, s6
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB4_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s5
+; GFX10-32-SDAG-NEXT:    s_and_b32 exec_lo, exec_lo, s4
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-GISEL-LABEL: test_waterfall_non_uniform_img_multi_rl:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-GISEL-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v4, v1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v6, v2
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s6, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s5, exec_lo
+; GFX10-32-GISEL-NEXT:  .LBB4_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s8, v5
+; GFX10-32-GISEL-NEXT:    s_ashr_i32 s9, s8, 31
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s10, v6
+; GFX10-32-GISEL-NEXT:    s_lshl_b64 s[8:9], s[8:9], 5
+; GFX10-32-GISEL-NEXT:    s_add_u32 s20, s0, s8
+; GFX10-32-GISEL-NEXT:    s_addc_u32 s21, s1, s9
+; GFX10-32-GISEL-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX10-32-GISEL-NEXT:    s_lshl_b64 s[8:9], s[10:11], 4
+; GFX10-32-GISEL-NEXT:    s_add_u32 s22, s2, s8
+; GFX10-32-GISEL-NEXT:    s_addc_u32 s23, s3, s9
+; GFX10-32-GISEL-NEXT:    s_load_dwordx8 s[8:15], s[20:21], 0x0
+; GFX10-32-GISEL-NEXT:    s_load_dwordx4 s[16:19], s[22:23], 0x0
+; GFX10-32-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-32-GISEL-NEXT:    image_sample v[0:3], v4, s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s6, s6
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB4_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s5
+; GFX10-32-GISEL-NEXT:    s_and_b32 exec_lo, exec_lo, s4
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-SDAG-LABEL: test_waterfall_non_uniform_img_multi_rl:
+; GFX10-64-SDAG:       ; %bb.0:
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v5, v2
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v4, v1
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v6, v0
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-SDAG-NEXT:  .LBB4_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s10, v6
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s10, v6
+; GFX10-64-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX10-64-SDAG-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX10-64-SDAG-NEXT:    s_add_u32 s10, s0, s10
+; GFX10-64-SDAG-NEXT:    s_addc_u32 s11, s1, s11
+; GFX10-64-SDAG-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX10-64-SDAG-NEXT:    s_lshl_b64 s[12:13], s[12:13], 4
+; GFX10-64-SDAG-NEXT:    s_add_u32 s24, s2, s12
+; GFX10-64-SDAG-NEXT:    s_addc_u32 s25, s3, s13
+; GFX10-64-SDAG-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
+; GFX10-64-SDAG-NEXT:    s_load_dwordx4 s[20:23], s[24:25], 0x0
+; GFX10-64-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-64-SDAG-NEXT:    image_sample v[0:3], v4, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[8:9], s[8:9]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB4_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX10-64-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-GISEL-LABEL: test_waterfall_non_uniform_img_multi_rl:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v4, v1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v6, v2
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-GISEL-NEXT:  .LBB4_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s10, v5
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s10, v5
+; GFX10-64-GISEL-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s12, v6
+; GFX10-64-GISEL-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX10-64-GISEL-NEXT:    s_add_u32 s10, s0, s10
+; GFX10-64-GISEL-NEXT:    s_addc_u32 s11, s1, s11
+; GFX10-64-GISEL-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX10-64-GISEL-NEXT:    s_lshl_b64 s[12:13], s[12:13], 4
+; GFX10-64-GISEL-NEXT:    s_add_u32 s24, s2, s12
+; GFX10-64-GISEL-NEXT:    s_addc_u32 s25, s3, s13
+; GFX10-64-GISEL-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
+; GFX10-64-GISEL-NEXT:    s_load_dwordx4 s[20:23], s[24:25], 0x0
+; GFX10-64-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-64-GISEL-NEXT:    image_sample v[0:3], v4, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[8:9], s[8:9]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB4_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX10-64-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX1150-SDAG-LABEL: test_waterfall_non_uniform_img_multi_rl:
+; GFX1150-SDAG:       ; %bb.0:
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v5, v2
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v4, v1
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v6, v0
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-SDAG-NEXT:  .LBB4_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s10, v6
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s10, v6
+; GFX1150-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX1150-SDAG-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(SKIP_2) | instid1(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_add_u32 s10, s0, s10
+; GFX1150-SDAG-NEXT:    s_addc_u32 s11, s1, s11
+; GFX1150-SDAG-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX1150-SDAG-NEXT:    s_lshl_b64 s[12:13], s[12:13], 4
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_add_u32 s20, s2, s12
+; GFX1150-SDAG-NEXT:    s_addc_u32 s21, s3, s13
+; GFX1150-SDAG-NEXT:    s_load_b256 s[12:19], s[10:11], 0x0
+; GFX1150-SDAG-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0) lgkmcnt(0)
+; GFX1150-SDAG-NEXT:    image_sample v[0:3], v4, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[8:9], s[8:9]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB4_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX1150-GISEL-LABEL: test_waterfall_non_uniform_img_multi_rl:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v4, v1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v6, v2
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-GISEL-NEXT:  .LBB4_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s10, v5
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_3) | instid1(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s10, v5
+; GFX1150-GISEL-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s12, v6
+; GFX1150-GISEL-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX1150-GISEL-NEXT:    s_add_u32 s10, s0, s10
+; GFX1150-GISEL-NEXT:    s_addc_u32 s11, s1, s11
+; GFX1150-GISEL-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_lshl_b64 s[12:13], s[12:13], 4
+; GFX1150-GISEL-NEXT:    s_add_u32 s20, s2, s12
+; GFX1150-GISEL-NEXT:    s_addc_u32 s21, s3, s13
+; GFX1150-GISEL-NEXT:    s_load_b256 s[12:19], s[10:11], 0x0
+; GFX1150-GISEL-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0) lgkmcnt(0)
+; GFX1150-GISEL-NEXT:    image_sample v[0:3], v4, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[8:9], s[8:9]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB4_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX12-SDAG-LABEL: test_waterfall_non_uniform_img_multi_rl:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v5, v2
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v4, v1
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v6, v0
+; GFX12-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX12-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-SDAG-NEXT:  .LBB4_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_1) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s10, v6
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s10, v6
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX12-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    s_add_nc_u64 s[10:11], s[0:1], s[10:11]
+; GFX12-SDAG-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(SALU_CYCLE_1)
+; GFX12-SDAG-NEXT:    s_lshl_b64 s[12:13], s[12:13], 4
+; GFX12-SDAG-NEXT:    s_add_nc_u64 s[20:21], s[2:3], s[12:13]
+; GFX12-SDAG-NEXT:    s_load_b256 s[12:19], s[10:11], 0x0
+; GFX12-SDAG-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX12-SDAG-NEXT:    s_wait_kmcnt 0x0
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    image_sample v[0:3], v4, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[8:9], s[8:9]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB4_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX12-GISEL-LABEL: test_waterfall_non_uniform_img_multi_rl:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v4, v1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v6, v2
+; GFX12-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-GISEL-NEXT:  .LBB4_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s10, v5
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s10, v5
+; GFX12-GISEL-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s12, v6
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_add_co_u32 s10, s0, s10
+; GFX12-GISEL-NEXT:    s_add_co_ci_u32 s11, s1, s11
+; GFX12-GISEL-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_lshl_b64 s[12:13], s[12:13], 4
+; GFX12-GISEL-NEXT:    s_add_co_u32 s20, s2, s12
+; GFX12-GISEL-NEXT:    s_add_co_ci_u32 s21, s3, s13
+; GFX12-GISEL-NEXT:    s_load_b256 s[12:19], s[10:11], 0x0
+; GFX12-GISEL-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX12-GISEL-NEXT:    s_wait_kmcnt 0x0
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    image_sample v[0:3], v4, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[8:9], s[8:9]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB4_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    ; return to shader part epilog
+  %wf_token = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %index)
+  %s_idx = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %wf_token, i32 %index)
+  %s_idx2 = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %wf_token, i32 %val)
+  %ptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx
+  %ptr2 = getelementptr <4 x i32>, <4 x i32> addrspace(4)* %samp_in, i32 %s_idx2
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4) * %ptr, align 32
+  %samp = load <4 x i32>, <4 x i32> addrspace(4) * %ptr2, align 32
+  %r = call <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32(i32 15, float %s, <8 x i32> %rsrc, <4 x i32> %samp, i1 0, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %wf_token, <4 x float> %r)
+
+  ret <4 x float> %r1
+}
+
+define amdgpu_ps <4 x float> @test_waterfall_non_uni_img_2_idx(<8 x i32> addrspace(4)* inreg %in, <4 x i32> addrspace(4)* inreg %samp_in, i32 %index1, i32 %index2, float %s) #1 {
+; VI-SDAG-LABEL: test_waterfall_non_uni_img_2_idx:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; VI-SDAG-NEXT:    s_wqm_b64 exec, exec
+; VI-SDAG-NEXT:    v_mov_b32_e32 v6, v2
+; VI-SDAG-NEXT:    v_mov_b32_e32 v5, v1
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; VI-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; VI-SDAG-NEXT:  .LBB5_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s12, v4
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s13, v5
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[8:9], s12, v4
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[10:11], s13, v5
+; VI-SDAG-NEXT:    s_and_b64 s[8:9], s[8:9], s[10:11]
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[20:21], s[8:9]
+; VI-SDAG-NEXT:    s_ashr_i32 s9, s12, 31
+; VI-SDAG-NEXT:    s_mov_b32 s8, s12
+; VI-SDAG-NEXT:    s_lshl_b64 s[8:9], s[8:9], 5
+; VI-SDAG-NEXT:    s_add_u32 s8, s0, s8
+; VI-SDAG-NEXT:    s_addc_u32 s9, s1, s9
+; VI-SDAG-NEXT:    s_ashr_i32 s11, s13, 31
+; VI-SDAG-NEXT:    s_mov_b32 s10, s13
+; VI-SDAG-NEXT:    s_lshl_b64 s[10:11], s[10:11], 4
+; VI-SDAG-NEXT:    s_add_u32 s16, s2, s10
+; VI-SDAG-NEXT:    s_addc_u32 s17, s3, s11
+; VI-SDAG-NEXT:    s_load_dwordx8 s[8:15], s[8:9], 0x0
+; VI-SDAG-NEXT:    s_load_dwordx4 s[16:19], s[16:17], 0x0
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr4_vgpr5
+; VI-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; VI-SDAG-NEXT:    image_sample v[0:3], v6, s[8:15], s[16:19] dmask:0xf
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr6
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[20:21]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB5_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[6:7]
+; VI-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    ; return to shader part epilog
+;
+; VI-GISEL-LABEL: test_waterfall_non_uni_img_2_idx:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; VI-GISEL-NEXT:    s_wqm_b64 exec, exec
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v5, v1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v6, v2
+; VI-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; VI-GISEL-NEXT:  .LBB5_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s12, v4
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s13, v5
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[8:9], s12, v4
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[10:11], s13, v5
+; VI-GISEL-NEXT:    s_and_b64 s[8:9], s[8:9], s[10:11]
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[20:21], s[8:9]
+; VI-GISEL-NEXT:    s_ashr_i32 s9, s12, 31
+; VI-GISEL-NEXT:    s_mov_b32 s8, s12
+; VI-GISEL-NEXT:    s_lshl_b64 s[8:9], s[8:9], 5
+; VI-GISEL-NEXT:    s_add_u32 s8, s0, s8
+; VI-GISEL-NEXT:    s_addc_u32 s9, s1, s9
+; VI-GISEL-NEXT:    s_ashr_i32 s11, s13, 31
+; VI-GISEL-NEXT:    s_mov_b32 s10, s13
+; VI-GISEL-NEXT:    s_lshl_b64 s[10:11], s[10:11], 4
+; VI-GISEL-NEXT:    s_add_u32 s16, s2, s10
+; VI-GISEL-NEXT:    s_addc_u32 s17, s3, s11
+; VI-GISEL-NEXT:    s_load_dwordx8 s[8:15], s[8:9], 0x0
+; VI-GISEL-NEXT:    s_load_dwordx4 s[16:19], s[16:17], 0x0
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr4_vgpr5
+; VI-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; VI-GISEL-NEXT:    image_sample v[0:3], v6, s[8:15], s[16:19] dmask:0xf
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr6
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[20:21]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB5_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[6:7]
+; VI-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX9-SDAG-LABEL: test_waterfall_non_uni_img_2_idx:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX9-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v6, v2
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v5, v1
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX9-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX9-SDAG-NEXT:  .LBB5_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s12, v4
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s13, v5
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[8:9], s12, v4
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[10:11], s13, v5
+; GFX9-SDAG-NEXT:    s_and_b64 s[8:9], s[8:9], s[10:11]
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[20:21], s[8:9]
+; GFX9-SDAG-NEXT:    s_ashr_i32 s9, s12, 31
+; GFX9-SDAG-NEXT:    s_mov_b32 s8, s12
+; GFX9-SDAG-NEXT:    s_lshl_b64 s[8:9], s[8:9], 5
+; GFX9-SDAG-NEXT:    s_add_u32 s22, s0, s8
+; GFX9-SDAG-NEXT:    s_addc_u32 s23, s1, s9
+; GFX9-SDAG-NEXT:    s_ashr_i32 s9, s13, 31
+; GFX9-SDAG-NEXT:    s_mov_b32 s8, s13
+; GFX9-SDAG-NEXT:    s_lshl_b64 s[8:9], s[8:9], 4
+; GFX9-SDAG-NEXT:    s_add_u32 s24, s2, s8
+; GFX9-SDAG-NEXT:    s_addc_u32 s25, s3, s9
+; GFX9-SDAG-NEXT:    s_load_dwordx8 s[8:15], s[22:23], 0x0
+; GFX9-SDAG-NEXT:    s_load_dwordx4 s[16:19], s[24:25], 0x0
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr4_vgpr5
+; GFX9-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX9-SDAG-NEXT:    image_sample v[0:3], v6, s[8:15], s[16:19] dmask:0xf
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[20:21]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB5_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX9-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX9-GISEL-LABEL: test_waterfall_non_uni_img_2_idx:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX9-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v5, v1
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v6, v2
+; GFX9-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX9-GISEL-NEXT:  .LBB5_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s12, v4
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s13, v5
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[8:9], s12, v4
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[10:11], s13, v5
+; GFX9-GISEL-NEXT:    s_and_b64 s[8:9], s[8:9], s[10:11]
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[20:21], s[8:9]
+; GFX9-GISEL-NEXT:    s_ashr_i32 s9, s12, 31
+; GFX9-GISEL-NEXT:    s_mov_b32 s8, s12
+; GFX9-GISEL-NEXT:    s_lshl_b64 s[8:9], s[8:9], 5
+; GFX9-GISEL-NEXT:    s_add_u32 s22, s0, s8
+; GFX9-GISEL-NEXT:    s_addc_u32 s23, s1, s9
+; GFX9-GISEL-NEXT:    s_ashr_i32 s9, s13, 31
+; GFX9-GISEL-NEXT:    s_mov_b32 s8, s13
+; GFX9-GISEL-NEXT:    s_lshl_b64 s[8:9], s[8:9], 4
+; GFX9-GISEL-NEXT:    s_add_u32 s24, s2, s8
+; GFX9-GISEL-NEXT:    s_addc_u32 s25, s3, s9
+; GFX9-GISEL-NEXT:    s_load_dwordx8 s[8:15], s[22:23], 0x0
+; GFX9-GISEL-NEXT:    s_load_dwordx4 s[16:19], s[24:25], 0x0
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr4_vgpr5
+; GFX9-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX9-GISEL-NEXT:    image_sample v[0:3], v6, s[8:15], s[16:19] dmask:0xf
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[20:21]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB5_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX9-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-SDAG-LABEL: test_waterfall_non_uni_img_2_idx:
+; GFX10-32-SDAG:       ; %bb.0:
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-SDAG-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v6, v2
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v5, v1
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s6, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s5, exec_lo
+; GFX10-32-SDAG-NEXT:  .LBB5_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s7, v4
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s10, v5
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s7, v4
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s10, v5
+; GFX10-32-SDAG-NEXT:    s_ashr_i32 s9, s7, 31
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s8, s7
+; GFX10-32-SDAG-NEXT:    s_lshl_b64 s[8:9], s[8:9], 5
+; GFX10-32-SDAG-NEXT:    s_add_u32 s20, s0, s8
+; GFX10-32-SDAG-NEXT:    s_addc_u32 s21, s1, s9
+; GFX10-32-SDAG-NEXT:    s_ashr_i32 s9, s10, 31
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s8, s10
+; GFX10-32-SDAG-NEXT:    s_lshl_b64 s[8:9], s[8:9], 4
+; GFX10-32-SDAG-NEXT:    s_add_u32 s22, s2, s8
+; GFX10-32-SDAG-NEXT:    s_addc_u32 s23, s3, s9
+; GFX10-32-SDAG-NEXT:    s_load_dwordx8 s[8:15], s[20:21], 0x0
+; GFX10-32-SDAG-NEXT:    s_load_dwordx4 s[16:19], s[22:23], 0x0
+; GFX10-32-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-32-SDAG-NEXT:    image_sample v[0:3], v6, s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s6, s6
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr4_vgpr5
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB5_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s5
+; GFX10-32-SDAG-NEXT:    s_and_b32 exec_lo, exec_lo, s4
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-GISEL-LABEL: test_waterfall_non_uni_img_2_idx:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-GISEL-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v5, v1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v6, v2
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s6, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s5, exec_lo
+; GFX10-32-GISEL-NEXT:  .LBB5_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s7, v4
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s10, v5
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s7, v4
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s10, v5
+; GFX10-32-GISEL-NEXT:    s_ashr_i32 s9, s7, 31
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s8, s7
+; GFX10-32-GISEL-NEXT:    s_lshl_b64 s[8:9], s[8:9], 5
+; GFX10-32-GISEL-NEXT:    s_add_u32 s20, s0, s8
+; GFX10-32-GISEL-NEXT:    s_addc_u32 s21, s1, s9
+; GFX10-32-GISEL-NEXT:    s_ashr_i32 s9, s10, 31
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s8, s10
+; GFX10-32-GISEL-NEXT:    s_lshl_b64 s[8:9], s[8:9], 4
+; GFX10-32-GISEL-NEXT:    s_add_u32 s22, s2, s8
+; GFX10-32-GISEL-NEXT:    s_addc_u32 s23, s3, s9
+; GFX10-32-GISEL-NEXT:    s_load_dwordx8 s[8:15], s[20:21], 0x0
+; GFX10-32-GISEL-NEXT:    s_load_dwordx4 s[16:19], s[22:23], 0x0
+; GFX10-32-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-32-GISEL-NEXT:    image_sample v[0:3], v6, s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s6, s6
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr4_vgpr5
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB5_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s5
+; GFX10-32-GISEL-NEXT:    s_and_b32 exec_lo, exec_lo, s4
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-SDAG-LABEL: test_waterfall_non_uni_img_2_idx:
+; GFX10-64-SDAG:       ; %bb.0:
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v6, v2
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v5, v1
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-SDAG-NEXT:  .LBB5_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s10, v4
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s10, v4
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s12, v5
+; GFX10-64-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX10-64-SDAG-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX10-64-SDAG-NEXT:    s_add_u32 s10, s0, s10
+; GFX10-64-SDAG-NEXT:    s_addc_u32 s11, s1, s11
+; GFX10-64-SDAG-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX10-64-SDAG-NEXT:    s_lshl_b64 s[12:13], s[12:13], 4
+; GFX10-64-SDAG-NEXT:    s_add_u32 s24, s2, s12
+; GFX10-64-SDAG-NEXT:    s_addc_u32 s25, s3, s13
+; GFX10-64-SDAG-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
+; GFX10-64-SDAG-NEXT:    s_load_dwordx4 s[20:23], s[24:25], 0x0
+; GFX10-64-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-64-SDAG-NEXT:    image_sample v[0:3], v6, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[8:9], s[8:9]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr4_vgpr5
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB5_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX10-64-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-GISEL-LABEL: test_waterfall_non_uni_img_2_idx:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v5, v1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v6, v2
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-GISEL-NEXT:  .LBB5_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s10, v4
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s10, v4
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s12, v5
+; GFX10-64-GISEL-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX10-64-GISEL-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX10-64-GISEL-NEXT:    s_add_u32 s10, s0, s10
+; GFX10-64-GISEL-NEXT:    s_addc_u32 s11, s1, s11
+; GFX10-64-GISEL-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX10-64-GISEL-NEXT:    s_lshl_b64 s[12:13], s[12:13], 4
+; GFX10-64-GISEL-NEXT:    s_add_u32 s24, s2, s12
+; GFX10-64-GISEL-NEXT:    s_addc_u32 s25, s3, s13
+; GFX10-64-GISEL-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
+; GFX10-64-GISEL-NEXT:    s_load_dwordx4 s[20:23], s[24:25], 0x0
+; GFX10-64-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-64-GISEL-NEXT:    image_sample v[0:3], v6, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[8:9], s[8:9]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr4_vgpr5
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB5_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX10-64-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX1150-SDAG-LABEL: test_waterfall_non_uni_img_2_idx:
+; GFX1150-SDAG:       ; %bb.0:
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v6, v2
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v5, v1
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-SDAG-NEXT:  .LBB5_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_1) | instid1(VALU_DEP_2)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s10, v4
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s10, v4
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(SKIP_1) | instid1(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s12, v5
+; GFX1150-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX1150-SDAG-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(SKIP_2) | instid1(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_add_u32 s10, s0, s10
+; GFX1150-SDAG-NEXT:    s_addc_u32 s11, s1, s11
+; GFX1150-SDAG-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX1150-SDAG-NEXT:    s_lshl_b64 s[12:13], s[12:13], 4
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_add_u32 s20, s2, s12
+; GFX1150-SDAG-NEXT:    s_addc_u32 s21, s3, s13
+; GFX1150-SDAG-NEXT:    s_load_b256 s[12:19], s[10:11], 0x0
+; GFX1150-SDAG-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0) lgkmcnt(0)
+; GFX1150-SDAG-NEXT:    image_sample v[0:3], v6, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[8:9], s[8:9]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr4_vgpr5
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB5_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX1150-GISEL-LABEL: test_waterfall_non_uni_img_2_idx:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v5, v1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v6, v2
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-GISEL-NEXT:  .LBB5_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s10, v4
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s10, v4
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s12, v5
+; GFX1150-GISEL-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX1150-GISEL-NEXT:    s_add_u32 s10, s0, s10
+; GFX1150-GISEL-NEXT:    s_addc_u32 s11, s1, s11
+; GFX1150-GISEL-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_lshl_b64 s[12:13], s[12:13], 4
+; GFX1150-GISEL-NEXT:    s_add_u32 s20, s2, s12
+; GFX1150-GISEL-NEXT:    s_addc_u32 s21, s3, s13
+; GFX1150-GISEL-NEXT:    s_load_b256 s[12:19], s[10:11], 0x0
+; GFX1150-GISEL-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0) lgkmcnt(0)
+; GFX1150-GISEL-NEXT:    image_sample v[0:3], v6, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[8:9], s[8:9]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr4_vgpr5
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB5_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX12-SDAG-LABEL: test_waterfall_non_uni_img_2_idx:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v6, v2
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v5, v1
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX12-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-SDAG-NEXT:  .LBB5_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_2) | instid1(VALU_DEP_2)
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s10, v4
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s10, v4
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s12, v5
+; GFX12-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX12-SDAG-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX12-SDAG-NEXT:    s_lshl_b64 s[12:13], s[12:13], 4
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    s_add_nc_u64 s[10:11], s[0:1], s[10:11]
+; GFX12-SDAG-NEXT:    s_add_nc_u64 s[20:21], s[2:3], s[12:13]
+; GFX12-SDAG-NEXT:    s_load_b256 s[12:19], s[10:11], 0x0
+; GFX12-SDAG-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX12-SDAG-NEXT:    s_wait_kmcnt 0x0
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    image_sample v[0:3], v6, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[8:9], s[8:9]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr4_vgpr5
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB5_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX12-GISEL-LABEL: test_waterfall_non_uni_img_2_idx:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v5, v1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v6, v2
+; GFX12-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-GISEL-NEXT:  .LBB5_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s10, v4
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s10, v4
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s12, v5
+; GFX12-GISEL-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_add_co_u32 s10, s0, s10
+; GFX12-GISEL-NEXT:    s_add_co_ci_u32 s11, s1, s11
+; GFX12-GISEL-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_lshl_b64 s[12:13], s[12:13], 4
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_add_co_u32 s20, s2, s12
+; GFX12-GISEL-NEXT:    s_add_co_ci_u32 s21, s3, s13
+; GFX12-GISEL-NEXT:    s_load_b256 s[12:19], s[10:11], 0x0
+; GFX12-GISEL-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX12-GISEL-NEXT:    s_wait_kmcnt 0x0
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    image_sample v[0:3], v6, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[8:9], s[8:9]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr4_vgpr5
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB5_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    ; return to shader part epilog
+  %t_idx = insertelement <2 x i32> undef, i32 %index1, i32 0
+  %combined_idx = insertelement <2 x i32> %t_idx, i32 %index2, i32 1
+  %wf_token = call i32 @llvm.amdgcn.waterfall.begin.v2i32(i32 0, <2 x i32> %combined_idx)
+  %s_c_idx = call <2 x i32> @llvm.amdgcn.waterfall.readfirstlane.v2i32.v2i32(i32 %wf_token, <2 x i32> %combined_idx)
+  %s_idx1 = extractelement <2 x i32> %s_c_idx, i32 0
+  %s_idx2 = extractelement <2 x i32> %s_c_idx, i32 1
+  %ptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx1
+  %ptr2 = getelementptr <4 x i32>, <4 x i32> addrspace(4)* %samp_in, i32 %s_idx2
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4) * %ptr, align 32
+  %samp = load <4 x i32>, <4 x i32> addrspace(4) * %ptr2, align 32
+  %r = call <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32(i32 15, float %s, <8 x i32> %rsrc, <4 x i32> %samp, i1 0, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %wf_token, <4 x float> %r)
+
+  ret <4 x float> %r1
+}
+
+define amdgpu_ps void @test_waterfall_non_uniform_img_single_store(<8 x i32> addrspace(4)* inreg %in, i32 %index, i32 %s, <4 x float> %data) #1 {
+; VI-SDAG-LABEL: test_waterfall_non_uniform_img_single_store:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; VI-SDAG-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; VI-SDAG-NEXT:    v_lshlrev_b64 v[7:8], 5, v[0:1]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v1, s1
+; VI-SDAG-NEXT:    v_add_u32_e32 v7, vcc, s0, v7
+; VI-SDAG-NEXT:    v_addc_u32_e32 v8, vcc, v1, v8, vcc
+; VI-SDAG-NEXT:    v_add_u32_e32 v11, vcc, 16, v7
+; VI-SDAG-NEXT:    v_addc_u32_e32 v12, vcc, 0, v8, vcc
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[7:10], v[7:8]
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[11:14], v[11:12]
+; VI-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:  .LBB6_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s0, v0
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[0:1], s0, v0
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[8:9], s[0:1]
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s0, v7
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s1, v8
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s2, v9
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s3, v10
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s4, v11
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s5, v12
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s6, v13
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s7, v14
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr0
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; VI-SDAG-NEXT:    s_nop 4
+; VI-SDAG-NEXT:    image_store v[2:5], v6, s[0:7] dmask:0xf unorm
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr2_vgpr3_vgpr4_vgpr5
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr6
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[8:9]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB6_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_endpgm
+;
+; VI-GISEL-LABEL: test_waterfall_non_uniform_img_single_store:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; VI-GISEL-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; VI-GISEL-NEXT:    v_lshlrev_b64 v[7:8], 5, v[0:1]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v10, s1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v9, s0
+; VI-GISEL-NEXT:    v_add_u32_e32 v7, vcc, v9, v7
+; VI-GISEL-NEXT:    v_addc_u32_e32 v8, vcc, v10, v8, vcc
+; VI-GISEL-NEXT:    v_add_u32_e32 v11, vcc, 16, v7
+; VI-GISEL-NEXT:    v_addc_u32_e32 v12, vcc, 0, v8, vcc
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[7:10], v[7:8]
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[11:14], v[11:12]
+; VI-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:  .LBB6_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s0, v0
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[0:1], s0, v0
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[8:9], s[0:1]
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s0, v7
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s1, v8
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s2, v9
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s3, v10
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s4, v11
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s5, v12
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s6, v13
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s7, v14
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr0
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; VI-GISEL-NEXT:    s_nop 4
+; VI-GISEL-NEXT:    image_store v[2:5], v6, s[0:7] dmask:0xf unorm
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr2_vgpr3_vgpr4_vgpr5
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr6
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[8:9]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB6_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_endpgm
+;
+; GFX9-SDAG-LABEL: test_waterfall_non_uniform_img_single_store:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX9-SDAG-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX9-SDAG-NEXT:    v_lshlrev_b64 v[7:8], 5, v[0:1]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v1, s1
+; GFX9-SDAG-NEXT:    v_add_co_u32_e32 v15, vcc, s0, v7
+; GFX9-SDAG-NEXT:    v_addc_co_u32_e32 v16, vcc, v1, v8, vcc
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[11:14], v[15:16], off offset:16
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[7:10], v[15:16], off
+; GFX9-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:  .LBB6_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s0, v0
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[0:1], s0, v0
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[8:9], s[0:1]
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s0, v7
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s1, v8
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s2, v9
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s3, v10
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s4, v11
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s5, v12
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s6, v13
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s7, v14
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr0
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX9-SDAG-NEXT:    s_nop 4
+; GFX9-SDAG-NEXT:    image_store v[2:5], v6, s[0:7] dmask:0xf unorm
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr2_vgpr3_vgpr4_vgpr5
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[8:9]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB6_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_endpgm
+;
+; GFX9-GISEL-LABEL: test_waterfall_non_uniform_img_single_store:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX9-GISEL-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX9-GISEL-NEXT:    v_lshlrev_b64 v[7:8], 5, v[0:1]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v10, s1
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v9, s0
+; GFX9-GISEL-NEXT:    v_add_co_u32_e32 v15, vcc, v9, v7
+; GFX9-GISEL-NEXT:    v_addc_co_u32_e32 v16, vcc, v10, v8, vcc
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[7:10], v[15:16], off
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[11:14], v[15:16], off offset:16
+; GFX9-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:  .LBB6_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s0, v0
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[0:1], s0, v0
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[8:9], s[0:1]
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s0, v7
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s1, v8
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s2, v9
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s3, v10
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s4, v11
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s5, v12
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s6, v13
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s7, v14
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr0
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX9-GISEL-NEXT:    s_nop 4
+; GFX9-GISEL-NEXT:    image_store v[2:5], v6, s[0:7] dmask:0xf unorm
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr2_vgpr3_vgpr4_vgpr5
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[8:9]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB6_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_endpgm
+;
+; GFX10-32-SDAG-LABEL: test_waterfall_non_uniform_img_single_store:
+; GFX10-32-SDAG:       ; %bb.0:
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX10-32-SDAG-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX10-32-SDAG-NEXT:    v_lshlrev_b64 v[7:8], 5, v[0:1]
+; GFX10-32-SDAG-NEXT:    v_add_co_u32 v15, vcc_lo, s0, v7
+; GFX10-32-SDAG-NEXT:    v_add_co_ci_u32_e32 v16, vcc_lo, s1, v8, vcc_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-SDAG-NEXT:    s_clause 0x1
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[11:14], v[15:16], off offset:16
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[7:10], v[15:16], off
+; GFX10-32-SDAG-NEXT:  .LBB6_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s1, v0
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s1, v0
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s4, v7
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s5, v8
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s6, v9
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s7, v10
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s8, v11
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s9, v12
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s10, v13
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s11, v14
+; GFX10-32-SDAG-NEXT:    image_store v[2:5], v6, s[4:11] dmask:0xf dim:SQ_RSRC_IMG_1D unorm
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s0, s0
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr0
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr2_vgpr3_vgpr4_vgpr5
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB6_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_endpgm
+;
+; GFX10-32-GISEL-LABEL: test_waterfall_non_uniform_img_single_store:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX10-32-GISEL-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v10, s1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v9, s0
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-GISEL-NEXT:    v_lshlrev_b64 v[7:8], 5, v[0:1]
+; GFX10-32-GISEL-NEXT:    v_add_co_u32 v15, vcc_lo, v9, v7
+; GFX10-32-GISEL-NEXT:    v_add_co_ci_u32_e32 v16, vcc_lo, v10, v8, vcc_lo
+; GFX10-32-GISEL-NEXT:    s_clause 0x1
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[7:10], v[15:16], off
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[11:14], v[15:16], off offset:16
+; GFX10-32-GISEL-NEXT:  .LBB6_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s1, v0
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s1, v0
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s4, v7
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s5, v8
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s6, v9
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s7, v10
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s8, v11
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s9, v12
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s10, v13
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s11, v14
+; GFX10-32-GISEL-NEXT:    image_store v[2:5], v6, s[4:11] dmask:0xf dim:SQ_RSRC_IMG_1D unorm
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s0, s0
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr0
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr2_vgpr3_vgpr4_vgpr5
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB6_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_endpgm
+;
+; GFX10-64-SDAG-LABEL: test_waterfall_non_uniform_img_single_store:
+; GFX10-64-SDAG:       ; %bb.0:
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX10-64-SDAG-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-SDAG-NEXT:    v_lshlrev_b64 v[7:8], 5, v[0:1]
+; GFX10-64-SDAG-NEXT:    v_add_co_u32 v15, vcc, s0, v7
+; GFX10-64-SDAG-NEXT:    v_add_co_ci_u32_e32 v16, vcc, s1, v8, vcc
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-SDAG-NEXT:    s_clause 0x1
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[11:14], v[15:16], off offset:16
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[7:10], v[15:16], off
+; GFX10-64-SDAG-NEXT:  .LBB6_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s2, v0
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s2, v0
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s4, v7
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s5, v8
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s6, v9
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s7, v10
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s8, v11
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s9, v12
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s10, v13
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s11, v14
+; GFX10-64-SDAG-NEXT:    image_store v[2:5], v6, s[4:11] dmask:0xf dim:SQ_RSRC_IMG_1D unorm
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[0:1], s[0:1]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr0
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr2_vgpr3_vgpr4_vgpr5
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB6_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_endpgm
+;
+; GFX10-64-GISEL-LABEL: test_waterfall_non_uniform_img_single_store:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX10-64-GISEL-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v10, s1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v9, s0
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-GISEL-NEXT:    v_lshlrev_b64 v[7:8], 5, v[0:1]
+; GFX10-64-GISEL-NEXT:    v_add_co_u32 v15, vcc, v9, v7
+; GFX10-64-GISEL-NEXT:    v_add_co_ci_u32_e32 v16, vcc, v10, v8, vcc
+; GFX10-64-GISEL-NEXT:    s_clause 0x1
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[7:10], v[15:16], off
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[11:14], v[15:16], off offset:16
+; GFX10-64-GISEL-NEXT:  .LBB6_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s2, v0
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s2, v0
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s4, v7
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s5, v8
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s6, v9
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s7, v10
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s8, v11
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s9, v12
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s10, v13
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s11, v14
+; GFX10-64-GISEL-NEXT:    image_store v[2:5], v6, s[4:11] dmask:0xf dim:SQ_RSRC_IMG_1D unorm
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[0:1], s[0:1]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr0
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr2_vgpr3_vgpr4_vgpr5
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB6_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_endpgm
+;
+; GFX1150-SDAG-LABEL: test_waterfall_non_uniform_img_single_store:
+; GFX1150-SDAG:       ; %bb.0:
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX1150-SDAG-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_lshlrev_b64 v[7:8], 5, v[0:1]
+; GFX1150-SDAG-NEXT:    v_add_co_u32 v7, vcc, s0, v7
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_add_co_ci_u32_e64 v8, null, s1, v8, vcc
+; GFX1150-SDAG-NEXT:    s_clause 0x1
+; GFX1150-SDAG-NEXT:    global_load_b128 v[11:14], v[7:8], off offset:16
+; GFX1150-SDAG-NEXT:    global_load_b128 v[7:10], v[7:8], off
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-SDAG-NEXT:  .LBB6_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s2, v0
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s2, v0
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s4, v7
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s5, v8
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s6, v9
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s7, v10
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s8, v11
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s9, v12
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s10, v13
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s11, v14
+; GFX1150-SDAG-NEXT:    image_store v[2:5], v6, s[4:11] dmask:0xf dim:SQ_RSRC_IMG_1D unorm
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[0:1], s[0:1]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr0
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr2_vgpr3_vgpr4_vgpr5
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB6_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_endpgm
+;
+; GFX1150-GISEL-LABEL: test_waterfall_non_uniform_img_single_store:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX1150-GISEL-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v10, s1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v9, s0
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-GISEL-NEXT:    v_lshlrev_b64 v[7:8], 5, v[0:1]
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-GISEL-NEXT:    v_add_co_u32 v11, vcc, v9, v7
+; GFX1150-GISEL-NEXT:    v_add_co_ci_u32_e64 v12, null, v10, v8, vcc
+; GFX1150-GISEL-NEXT:    s_clause 0x1
+; GFX1150-GISEL-NEXT:    global_load_b128 v[7:10], v[11:12], off
+; GFX1150-GISEL-NEXT:    global_load_b128 v[11:14], v[11:12], off offset:16
+; GFX1150-GISEL-NEXT:  .LBB6_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s2, v0
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s2, v0
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s4, v7
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s5, v8
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s6, v9
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s7, v10
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s8, v11
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s9, v12
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s10, v13
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s11, v14
+; GFX1150-GISEL-NEXT:    image_store v[2:5], v6, s[4:11] dmask:0xf dim:SQ_RSRC_IMG_1D unorm
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[0:1], s[0:1]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr0
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr2_vgpr3_vgpr4_vgpr5
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB6_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_endpgm
+;
+; GFX12-SDAG-LABEL: test_waterfall_non_uniform_img_single_store:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX12-SDAG-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX12-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_lshlrev_b64_e32 v[7:8], 5, v[0:1]
+; GFX12-SDAG-NEXT:    v_add_co_u32 v7, vcc, s0, v7
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_add_co_ci_u32_e64 v8, null, s1, v8, vcc
+; GFX12-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-SDAG-NEXT:    s_clause 0x1
+; GFX12-SDAG-NEXT:    global_load_b128 v[11:14], v[7:8], off offset:16
+; GFX12-SDAG-NEXT:    global_load_b128 v[7:10], v[7:8], off
+; GFX12-SDAG-NEXT:  .LBB6_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s2, v0
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s2, v0
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x0
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s4, v7
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s5, v8
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s6, v9
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s7, v10
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s8, v11
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s9, v12
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s10, v13
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s11, v14
+; GFX12-SDAG-NEXT:    image_store v[2:5], v6, s[4:11] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[0:1], s[0:1]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr0
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr2_vgpr3_vgpr4_vgpr5
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB6_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_endpgm
+;
+; GFX12-GISEL-LABEL: test_waterfall_non_uniform_img_single_store:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX12-GISEL-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v10, s1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v9, s0
+; GFX12-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-GISEL-NEXT:    v_lshlrev_b64_e32 v[7:8], 5, v[0:1]
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_add_co_u32 v11, vcc, v9, v7
+; GFX12-GISEL-NEXT:    v_add_co_ci_u32_e64 v12, null, v10, v8, vcc
+; GFX12-GISEL-NEXT:    s_clause 0x1
+; GFX12-GISEL-NEXT:    global_load_b128 v[7:10], v[11:12], off
+; GFX12-GISEL-NEXT:    global_load_b128 v[11:14], v[11:12], off offset:16
+; GFX12-GISEL-NEXT:  .LBB6_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s2, v0
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s2, v0
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s4, v7
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s5, v8
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s6, v9
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s7, v10
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x0
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s8, v11
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s9, v12
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s10, v13
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s11, v14
+; GFX12-GISEL-NEXT:    image_store v[2:5], v6, s[4:11] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[0:1], s[0:1]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr0
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr2_vgpr3_vgpr4_vgpr5
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr6
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB6_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_endpgm
+  %ptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %index
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4) * %ptr, align 32
+  %wf_token = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %index)
+  %s_rsrc = call <8 x i32> @llvm.amdgcn.waterfall.readfirstlane.v8i32.v8i32(i32 %wf_token, <8 x i32> %rsrc)
+  %s_rsrc_use = call <8 x i32> @llvm.amdgcn.waterfall.last.use.v8i32(i32 %wf_token, <8 x i32> %s_rsrc)
+  call void @llvm.amdgcn.image.store.1d.v4f32.i32(<4 x float> %data, i32 15, i32 %s, <8 x i32> %s_rsrc_use, i32 0, i32 0)
+
+  ret void
+}
+
+define amdgpu_ps void @test_remove_waterfall_last_use(<8 x i32> addrspace(4)* inreg %in, i32 %index, i32 %s, <4 x float> %data) #1 {
+; PRE-GFX10-LABEL: test_remove_waterfall_last_use:
+; PRE-GFX10:       ; %bb.0:
+; PRE-GFX10-NEXT:    s_load_dwordx8 s[0:7], s[0:1], 0x0
+; PRE-GFX10-NEXT:    s_waitcnt lgkmcnt(0)
+; PRE-GFX10-NEXT:    image_store v[2:5], v1, s[0:7] dmask:0xf unorm
+; PRE-GFX10-NEXT:    s_endpgm
+;
+; GFX10-LABEL: test_remove_waterfall_last_use:
+; GFX10:       ; %bb.0:
+; GFX10-NEXT:    s_load_dwordx8 s[0:7], s[0:1], 0x0
+; GFX10-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-NEXT:    image_store v[2:5], v1, s[0:7] dmask:0xf dim:SQ_RSRC_IMG_1D unorm
+; GFX10-NEXT:    s_endpgm
+;
+; GFX1150-LABEL: test_remove_waterfall_last_use:
+; GFX1150:       ; %bb.0:
+; GFX1150-NEXT:    s_load_b256 s[0:7], s[0:1], 0x0
+; GFX1150-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX1150-NEXT:    image_store v[2:5], v1, s[0:7] dmask:0xf dim:SQ_RSRC_IMG_1D unorm
+; GFX1150-NEXT:    s_endpgm
+;
+; GFX12-LABEL: test_remove_waterfall_last_use:
+; GFX12:       ; %bb.0:
+; GFX12-NEXT:    s_load_b256 s[0:7], s[0:1], 0x0
+; GFX12-NEXT:    s_wait_kmcnt 0x0
+; GFX12-NEXT:    image_store v[2:5], v1, s[0:7] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-NEXT:    s_endpgm
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4) * %in, align 32
+  %wf_token = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %index)
+  %s_rsrc = call <8 x i32> @llvm.amdgcn.waterfall.readfirstlane.v8i32.v8i32(i32 %wf_token, <8 x i32> %rsrc)
+  %s_rsrc_use = call <8 x i32> @llvm.amdgcn.waterfall.last.use.v8i32(i32 %wf_token, <8 x i32> %s_rsrc)
+  call void @llvm.amdgcn.image.store.1d.v4f32.i32(<4 x float> %data, i32 15, i32 %s, <8 x i32> %s_rsrc_use, i32 0, i32 0)
+
+  ret void
+}
+
+define amdgpu_ps <4 x float> @test_remove_waterfall_multi_rl(<8 x i32> addrspace(4)* inreg %in, <4 x i32> addrspace(4)* inreg %samp_in, i32 %index, float %s, i32 inreg %val1, i32 inreg %val2) #1 {
+; VI-LABEL: test_remove_waterfall_multi_rl:
+; VI:       ; %bb.0:
+; VI-NEXT:    s_mov_b64 s[12:13], exec
+; VI-NEXT:    s_mov_b32 s6, s5
+; VI-NEXT:    s_wqm_b64 exec, exec
+; VI-NEXT:    s_ashr_i32 s5, s4, 31
+; VI-NEXT:    s_lshl_b64 s[4:5], s[4:5], 5
+; VI-NEXT:    s_add_u32 s0, s0, s4
+; VI-NEXT:    s_addc_u32 s1, s1, s5
+; VI-NEXT:    s_ashr_i32 s7, s6, 31
+; VI-NEXT:    s_lshl_b64 s[4:5], s[6:7], 4
+; VI-NEXT:    s_add_u32 s8, s2, s4
+; VI-NEXT:    s_addc_u32 s9, s3, s5
+; VI-NEXT:    s_load_dwordx8 s[0:7], s[0:1], 0x0
+; VI-NEXT:    s_load_dwordx4 s[8:11], s[8:9], 0x0
+; VI-NEXT:    s_and_b64 exec, exec, s[12:13]
+; VI-NEXT:    s_waitcnt lgkmcnt(0)
+; VI-NEXT:    image_sample v[0:3], v1, s[0:7], s[8:11] dmask:0xf
+; VI-NEXT:    s_waitcnt vmcnt(0)
+; VI-NEXT:    ; return to shader part epilog
+;
+; GFX9-LABEL: test_remove_waterfall_multi_rl:
+; GFX9:       ; %bb.0:
+; GFX9-NEXT:    s_mov_b64 s[12:13], exec
+; GFX9-NEXT:    s_mov_b32 s6, s5
+; GFX9-NEXT:    s_wqm_b64 exec, exec
+; GFX9-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX9-NEXT:    s_lshl_b64 s[4:5], s[4:5], 5
+; GFX9-NEXT:    s_add_u32 s14, s0, s4
+; GFX9-NEXT:    s_addc_u32 s15, s1, s5
+; GFX9-NEXT:    s_ashr_i32 s7, s6, 31
+; GFX9-NEXT:    s_lshl_b64 s[0:1], s[6:7], 4
+; GFX9-NEXT:    s_add_u32 s16, s2, s0
+; GFX9-NEXT:    s_addc_u32 s17, s3, s1
+; GFX9-NEXT:    s_load_dwordx8 s[0:7], s[14:15], 0x0
+; GFX9-NEXT:    s_load_dwordx4 s[8:11], s[16:17], 0x0
+; GFX9-NEXT:    s_and_b64 exec, exec, s[12:13]
+; GFX9-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX9-NEXT:    image_sample v[0:3], v1, s[0:7], s[8:11] dmask:0xf
+; GFX9-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-LABEL: test_remove_waterfall_multi_rl:
+; GFX10-32:       ; %bb.0:
+; GFX10-32-NEXT:    s_mov_b32 s16, exec_lo
+; GFX10-32-NEXT:    s_mov_b32 s6, s5
+; GFX10-32-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX10-32-NEXT:    s_lshl_b64 s[4:5], s[4:5], 5
+; GFX10-32-NEXT:    s_add_u32 s12, s0, s4
+; GFX10-32-NEXT:    s_addc_u32 s13, s1, s5
+; GFX10-32-NEXT:    s_ashr_i32 s7, s6, 31
+; GFX10-32-NEXT:    s_lshl_b64 s[0:1], s[6:7], 4
+; GFX10-32-NEXT:    s_add_u32 s14, s2, s0
+; GFX10-32-NEXT:    s_addc_u32 s15, s3, s1
+; GFX10-32-NEXT:    s_load_dwordx8 s[0:7], s[12:13], 0x0
+; GFX10-32-NEXT:    s_load_dwordx4 s[8:11], s[14:15], 0x0
+; GFX10-32-NEXT:    s_and_b32 exec_lo, exec_lo, s16
+; GFX10-32-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-32-NEXT:    image_sample v[0:3], v1, s[0:7], s[8:11] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-32-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-LABEL: test_remove_waterfall_multi_rl:
+; GFX10-64:       ; %bb.0:
+; GFX10-64-NEXT:    s_mov_b64 s[12:13], exec
+; GFX10-64-NEXT:    s_mov_b32 s6, s5
+; GFX10-64-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX10-64-NEXT:    s_lshl_b64 s[4:5], s[4:5], 5
+; GFX10-64-NEXT:    s_add_u32 s14, s0, s4
+; GFX10-64-NEXT:    s_addc_u32 s15, s1, s5
+; GFX10-64-NEXT:    s_ashr_i32 s7, s6, 31
+; GFX10-64-NEXT:    s_lshl_b64 s[0:1], s[6:7], 4
+; GFX10-64-NEXT:    s_add_u32 s16, s2, s0
+; GFX10-64-NEXT:    s_addc_u32 s17, s3, s1
+; GFX10-64-NEXT:    s_load_dwordx8 s[0:7], s[14:15], 0x0
+; GFX10-64-NEXT:    s_load_dwordx4 s[8:11], s[16:17], 0x0
+; GFX10-64-NEXT:    s_and_b64 exec, exec, s[12:13]
+; GFX10-64-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-64-NEXT:    image_sample v[0:3], v1, s[0:7], s[8:11] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-64-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-NEXT:    ; return to shader part epilog
+;
+; GFX1150-LABEL: test_remove_waterfall_multi_rl:
+; GFX1150:       ; %bb.0:
+; GFX1150-NEXT:    s_mov_b64 s[12:13], exec
+; GFX1150-NEXT:    s_mov_b32 s6, s5
+; GFX1150-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX1150-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(SALU_CYCLE_1)
+; GFX1150-NEXT:    s_lshl_b64 s[4:5], s[4:5], 5
+; GFX1150-NEXT:    s_add_u32 s0, s0, s4
+; GFX1150-NEXT:    s_addc_u32 s1, s1, s5
+; GFX1150-NEXT:    s_ashr_i32 s7, s6, 31
+; GFX1150-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(SALU_CYCLE_1)
+; GFX1150-NEXT:    s_lshl_b64 s[4:5], s[6:7], 4
+; GFX1150-NEXT:    s_add_u32 s8, s2, s4
+; GFX1150-NEXT:    s_addc_u32 s9, s3, s5
+; GFX1150-NEXT:    s_load_b256 s[0:7], s[0:1], 0x0
+; GFX1150-NEXT:    s_load_b128 s[8:11], s[8:9], 0x0
+; GFX1150-NEXT:    s_and_b64 exec, exec, s[12:13]
+; GFX1150-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX1150-NEXT:    image_sample v[0:3], v1, s[0:7], s[8:11] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX1150-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-NEXT:    ; return to shader part epilog
+;
+; GFX12-SDAG-LABEL: test_remove_waterfall_multi_rl:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    s_mov_b64 s[12:13], exec
+; GFX12-SDAG-NEXT:    s_mov_b32 s6, s5
+; GFX12-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX12-SDAG-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX12-SDAG-NEXT:    s_ashr_i32 s7, s6, 31
+; GFX12-SDAG-NEXT:    s_lshl_b64 s[4:5], s[4:5], 5
+; GFX12-SDAG-NEXT:    s_lshl_b64 s[6:7], s[6:7], 4
+; GFX12-SDAG-NEXT:    s_add_nc_u64 s[0:1], s[0:1], s[4:5]
+; GFX12-SDAG-NEXT:    s_add_nc_u64 s[8:9], s[2:3], s[6:7]
+; GFX12-SDAG-NEXT:    s_load_b256 s[0:7], s[0:1], 0x0
+; GFX12-SDAG-NEXT:    s_load_b128 s[8:11], s[8:9], 0x0
+; GFX12-SDAG-NEXT:    s_and_b64 exec, exec, s[12:13]
+; GFX12-SDAG-NEXT:    s_wait_kmcnt 0x0
+; GFX12-SDAG-NEXT:    image_sample v[0:3], v1, s[0:7], s[8:11] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX12-GISEL-LABEL: test_remove_waterfall_multi_rl:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    s_mov_b64 s[12:13], exec
+; GFX12-GISEL-NEXT:    s_mov_b32 s6, s5
+; GFX12-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX12-GISEL-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_lshl_b64 s[4:5], s[4:5], 5
+; GFX12-GISEL-NEXT:    s_add_co_u32 s0, s0, s4
+; GFX12-GISEL-NEXT:    s_add_co_ci_u32 s1, s1, s5
+; GFX12-GISEL-NEXT:    s_ashr_i32 s7, s6, 31
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_lshl_b64 s[4:5], s[6:7], 4
+; GFX12-GISEL-NEXT:    s_add_co_u32 s8, s2, s4
+; GFX12-GISEL-NEXT:    s_add_co_ci_u32 s9, s3, s5
+; GFX12-GISEL-NEXT:    s_load_b256 s[0:7], s[0:1], 0x0
+; GFX12-GISEL-NEXT:    s_load_b128 s[8:11], s[8:9], 0x0
+; GFX12-GISEL-NEXT:    s_and_b64 exec, exec, s[12:13]
+; GFX12-GISEL-NEXT:    s_wait_kmcnt 0x0
+; GFX12-GISEL-NEXT:    image_sample v[0:3], v1, s[0:7], s[8:11] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    ; return to shader part epilog
+  %wf_token = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %index)
+  %s_idx = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %wf_token, i32 %val1)
+  %s_idx2 = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %wf_token, i32 %val2)
+  %ptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx
+  %ptr2 = getelementptr <4 x i32>, <4 x i32> addrspace(4)* %samp_in, i32 %s_idx2
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4) * %ptr, align 32
+  %samp = load <4 x i32>, <4 x i32> addrspace(4) * %ptr2, align 32
+  %r = call <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32(i32 15, float %s, <8 x i32> %rsrc, <4 x i32> %samp, i1 0, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %wf_token, <4 x float> %r)
+
+  ret <4 x float> %r1
+}
+
+define amdgpu_ps <4 x float> @test_keep_waterfall_multi_rl(<8 x i32> addrspace(4)* inreg %in, <4 x i32> addrspace(4)* inreg %samp_in, i32 %index, float %s, i32 inreg %val) #1 {
+; VI-SDAG-LABEL: test_keep_waterfall_multi_rl:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; VI-SDAG-NEXT:    s_wqm_b64 exec, exec
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, v1
+; VI-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; VI-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; VI-SDAG-NEXT:  .LBB9_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s10, v5
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[12:13], s10, v5
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[24:25], s[12:13]
+; VI-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; VI-SDAG-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; VI-SDAG-NEXT:    s_add_u32 s10, s0, s10
+; VI-SDAG-NEXT:    s_addc_u32 s11, s1, s11
+; VI-SDAG-NEXT:    s_ashr_i32 s5, s4, 31
+; VI-SDAG-NEXT:    s_lshl_b64 s[12:13], s[4:5], 4
+; VI-SDAG-NEXT:    s_add_u32 s20, s2, s12
+; VI-SDAG-NEXT:    s_addc_u32 s21, s3, s13
+; VI-SDAG-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
+; VI-SDAG-NEXT:    s_load_dwordx4 s[20:23], s[20:21], 0x0
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr5
+; VI-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; VI-SDAG-NEXT:    image_sample v[0:3], v4, s[12:19], s[20:23] dmask:0xf
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr4
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[24:25]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB9_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[8:9]
+; VI-SDAG-NEXT:    s_and_b64 exec, exec, s[6:7]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    ; return to shader part epilog
+;
+; VI-GISEL-LABEL: test_keep_waterfall_multi_rl:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; VI-GISEL-NEXT:    s_wqm_b64 exec, exec
+; VI-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, v1
+; VI-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; VI-GISEL-NEXT:  .LBB9_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s10, v5
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[12:13], s10, v5
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[24:25], s[12:13]
+; VI-GISEL-NEXT:    s_ashr_i32 s11, s10, 31
+; VI-GISEL-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; VI-GISEL-NEXT:    s_add_u32 s10, s0, s10
+; VI-GISEL-NEXT:    s_addc_u32 s11, s1, s11
+; VI-GISEL-NEXT:    s_ashr_i32 s5, s4, 31
+; VI-GISEL-NEXT:    s_lshl_b64 s[12:13], s[4:5], 4
+; VI-GISEL-NEXT:    s_add_u32 s20, s2, s12
+; VI-GISEL-NEXT:    s_addc_u32 s21, s3, s13
+; VI-GISEL-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
+; VI-GISEL-NEXT:    s_load_dwordx4 s[20:23], s[20:21], 0x0
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr5
+; VI-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; VI-GISEL-NEXT:    image_sample v[0:3], v4, s[12:19], s[20:23] dmask:0xf
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr4
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[24:25]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB9_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[8:9]
+; VI-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX9-SDAG-LABEL: test_keep_waterfall_multi_rl:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX9-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v4, v1
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX9-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX9-SDAG-NEXT:  .LBB9_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s10, v5
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[12:13], s10, v5
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[24:25], s[12:13]
+; GFX9-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX9-SDAG-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX9-SDAG-NEXT:    s_add_u32 s10, s0, s10
+; GFX9-SDAG-NEXT:    s_addc_u32 s11, s1, s11
+; GFX9-SDAG-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX9-SDAG-NEXT:    s_lshl_b64 s[12:13], s[4:5], 4
+; GFX9-SDAG-NEXT:    s_add_u32 s26, s2, s12
+; GFX9-SDAG-NEXT:    s_addc_u32 s27, s3, s13
+; GFX9-SDAG-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
+; GFX9-SDAG-NEXT:    s_load_dwordx4 s[20:23], s[26:27], 0x0
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX9-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX9-SDAG-NEXT:    image_sample v[0:3], v4, s[12:19], s[20:23] dmask:0xf
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[24:25]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB9_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[8:9]
+; GFX9-SDAG-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX9-GISEL-LABEL: test_keep_waterfall_multi_rl:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX9-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v4, v1
+; GFX9-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX9-GISEL-NEXT:  .LBB9_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s10, v5
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[12:13], s10, v5
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[24:25], s[12:13]
+; GFX9-GISEL-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX9-GISEL-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX9-GISEL-NEXT:    s_add_u32 s10, s0, s10
+; GFX9-GISEL-NEXT:    s_addc_u32 s11, s1, s11
+; GFX9-GISEL-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX9-GISEL-NEXT:    s_lshl_b64 s[12:13], s[4:5], 4
+; GFX9-GISEL-NEXT:    s_add_u32 s26, s2, s12
+; GFX9-GISEL-NEXT:    s_addc_u32 s27, s3, s13
+; GFX9-GISEL-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
+; GFX9-GISEL-NEXT:    s_load_dwordx4 s[20:23], s[26:27], 0x0
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX9-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX9-GISEL-NEXT:    image_sample v[0:3], v4, s[12:19], s[20:23] dmask:0xf
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[24:25]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB9_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[8:9]
+; GFX9-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-SDAG-LABEL: test_keep_waterfall_multi_rl:
+; GFX10-32-SDAG:       ; %bb.0:
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s6, exec_lo
+; GFX10-32-SDAG-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v4, v1
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s8, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s7, exec_lo
+; GFX10-32-SDAG-NEXT:  .LBB9_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s10, v5
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s10, v5
+; GFX10-32-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX10-32-SDAG-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX10-32-SDAG-NEXT:    s_add_u32 s10, s0, s10
+; GFX10-32-SDAG-NEXT:    s_addc_u32 s11, s1, s11
+; GFX10-32-SDAG-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX10-32-SDAG-NEXT:    s_lshl_b64 s[12:13], s[4:5], 4
+; GFX10-32-SDAG-NEXT:    s_add_u32 s24, s2, s12
+; GFX10-32-SDAG-NEXT:    s_addc_u32 s25, s3, s13
+; GFX10-32-SDAG-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
+; GFX10-32-SDAG-NEXT:    s_load_dwordx4 s[20:23], s[24:25], 0x0
+; GFX10-32-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-32-SDAG-NEXT:    image_sample v[0:3], v4, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s8, s8
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB9_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s7
+; GFX10-32-SDAG-NEXT:    s_and_b32 exec_lo, exec_lo, s6
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-GISEL-LABEL: test_keep_waterfall_multi_rl:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s6, exec_lo
+; GFX10-32-GISEL-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v4, v1
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s8, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s7, exec_lo
+; GFX10-32-GISEL-NEXT:  .LBB9_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s10, v5
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s10, v5
+; GFX10-32-GISEL-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX10-32-GISEL-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX10-32-GISEL-NEXT:    s_add_u32 s10, s0, s10
+; GFX10-32-GISEL-NEXT:    s_addc_u32 s11, s1, s11
+; GFX10-32-GISEL-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX10-32-GISEL-NEXT:    s_lshl_b64 s[12:13], s[4:5], 4
+; GFX10-32-GISEL-NEXT:    s_add_u32 s24, s2, s12
+; GFX10-32-GISEL-NEXT:    s_addc_u32 s25, s3, s13
+; GFX10-32-GISEL-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
+; GFX10-32-GISEL-NEXT:    s_load_dwordx4 s[20:23], s[24:25], 0x0
+; GFX10-32-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-32-GISEL-NEXT:    image_sample v[0:3], v4, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s8, s8
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB9_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s7
+; GFX10-32-GISEL-NEXT:    s_and_b32 exec_lo, exec_lo, s6
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-SDAG-LABEL: test_keep_waterfall_multi_rl:
+; GFX10-64-SDAG:       ; %bb.0:
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v4, v1
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX10-64-SDAG-NEXT:  .LBB9_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s12, v5
+; GFX10-64-SDAG-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX10-64-SDAG-NEXT:    s_lshl_b64 s[12:13], s[12:13], 5
+; GFX10-64-SDAG-NEXT:    s_add_u32 s24, s0, s12
+; GFX10-64-SDAG-NEXT:    s_addc_u32 s25, s1, s13
+; GFX10-64-SDAG-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX10-64-SDAG-NEXT:    s_lshl_b64 s[12:13], s[4:5], 4
+; GFX10-64-SDAG-NEXT:    s_add_u32 s26, s2, s12
+; GFX10-64-SDAG-NEXT:    s_addc_u32 s27, s3, s13
+; GFX10-64-SDAG-NEXT:    s_load_dwordx8 s[12:19], s[24:25], 0x0
+; GFX10-64-SDAG-NEXT:    s_load_dwordx4 s[20:23], s[26:27], 0x0
+; GFX10-64-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-64-SDAG-NEXT:    image_sample v[0:3], v4, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[10:11], s[10:11]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB9_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[8:9]
+; GFX10-64-SDAG-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-GISEL-LABEL: test_keep_waterfall_multi_rl:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v4, v1
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[10:11], exec
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX10-64-GISEL-NEXT:  .LBB9_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s12, v5
+; GFX10-64-GISEL-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX10-64-GISEL-NEXT:    s_lshl_b64 s[12:13], s[12:13], 5
+; GFX10-64-GISEL-NEXT:    s_add_u32 s24, s0, s12
+; GFX10-64-GISEL-NEXT:    s_addc_u32 s25, s1, s13
+; GFX10-64-GISEL-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX10-64-GISEL-NEXT:    s_lshl_b64 s[12:13], s[4:5], 4
+; GFX10-64-GISEL-NEXT:    s_add_u32 s26, s2, s12
+; GFX10-64-GISEL-NEXT:    s_addc_u32 s27, s3, s13
+; GFX10-64-GISEL-NEXT:    s_load_dwordx8 s[12:19], s[24:25], 0x0
+; GFX10-64-GISEL-NEXT:    s_load_dwordx4 s[20:23], s[26:27], 0x0
+; GFX10-64-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-64-GISEL-NEXT:    image_sample v[0:3], v4, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[10:11], s[10:11]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB9_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[8:9]
+; GFX10-64-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX1150-SDAG-LABEL: test_keep_waterfall_multi_rl:
+; GFX1150-SDAG:       ; %bb.0:
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v4, v1
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX1150-SDAG-NEXT:  .LBB9_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s12, v5
+; GFX1150-SDAG-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_lshl_b64 s[12:13], s[12:13], 5
+; GFX1150-SDAG-NEXT:    s_add_u32 s12, s0, s12
+; GFX1150-SDAG-NEXT:    s_addc_u32 s13, s1, s13
+; GFX1150-SDAG-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_lshl_b64 s[14:15], s[4:5], 4
+; GFX1150-SDAG-NEXT:    s_add_u32 s20, s2, s14
+; GFX1150-SDAG-NEXT:    s_addc_u32 s21, s3, s15
+; GFX1150-SDAG-NEXT:    s_load_b256 s[12:19], s[12:13], 0x0
+; GFX1150-SDAG-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0) lgkmcnt(0)
+; GFX1150-SDAG-NEXT:    image_sample v[0:3], v4, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[10:11], s[10:11]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB9_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[8:9]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX1150-GISEL-LABEL: test_keep_waterfall_multi_rl:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v4, v1
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[10:11], exec
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX1150-GISEL-NEXT:  .LBB9_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s12, v5
+; GFX1150-GISEL-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_lshl_b64 s[12:13], s[12:13], 5
+; GFX1150-GISEL-NEXT:    s_add_u32 s12, s0, s12
+; GFX1150-GISEL-NEXT:    s_addc_u32 s13, s1, s13
+; GFX1150-GISEL-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_lshl_b64 s[14:15], s[4:5], 4
+; GFX1150-GISEL-NEXT:    s_add_u32 s20, s2, s14
+; GFX1150-GISEL-NEXT:    s_addc_u32 s21, s3, s15
+; GFX1150-GISEL-NEXT:    s_load_b256 s[12:19], s[12:13], 0x0
+; GFX1150-GISEL-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0) lgkmcnt(0)
+; GFX1150-GISEL-NEXT:    image_sample v[0:3], v4, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[10:11], s[10:11]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB9_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[8:9]
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX12-SDAG-LABEL: test_keep_waterfall_multi_rl:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v4, v1
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX12-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; GFX12-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX12-SDAG-NEXT:  .LBB9_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_1) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s12, v5
+; GFX12-SDAG-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX12-SDAG-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    s_lshl_b64 s[12:13], s[12:13], 5
+; GFX12-SDAG-NEXT:    s_lshl_b64 s[14:15], s[4:5], 4
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    s_add_nc_u64 s[12:13], s[0:1], s[12:13]
+; GFX12-SDAG-NEXT:    s_add_nc_u64 s[20:21], s[2:3], s[14:15]
+; GFX12-SDAG-NEXT:    s_load_b256 s[12:19], s[12:13], 0x0
+; GFX12-SDAG-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX12-SDAG-NEXT:    s_wait_kmcnt 0x0
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    image_sample v[0:3], v4, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[10:11], s[10:11]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB9_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[8:9]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-SDAG-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX12-GISEL-LABEL: test_keep_waterfall_multi_rl:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v4, v1
+; GFX12-GISEL-NEXT:    s_mov_b64 s[10:11], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX12-GISEL-NEXT:  .LBB9_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(SKIP_1) | instid1(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s12, v5
+; GFX12-GISEL-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_lshl_b64 s[12:13], s[12:13], 5
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_add_co_u32 s12, s0, s12
+; GFX12-GISEL-NEXT:    s_add_co_ci_u32 s13, s1, s13
+; GFX12-GISEL-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_lshl_b64 s[14:15], s[4:5], 4
+; GFX12-GISEL-NEXT:    s_add_co_u32 s20, s2, s14
+; GFX12-GISEL-NEXT:    s_add_co_ci_u32 s21, s3, s15
+; GFX12-GISEL-NEXT:    s_load_b256 s[12:19], s[12:13], 0x0
+; GFX12-GISEL-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX12-GISEL-NEXT:    s_wait_kmcnt 0x0
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    image_sample v[0:3], v4, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[10:11], s[10:11]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB9_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[8:9]
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    ; return to shader part epilog
+  %wf_token = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %index)
+  %s_idx = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %wf_token, i32 %index)
+  %s_idx2 = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %wf_token, i32 %val)
+  %ptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx
+  %ptr2 = getelementptr <4 x i32>, <4 x i32> addrspace(4)* %samp_in, i32 %s_idx2
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4) * %ptr, align 32
+  %samp = load <4 x i32>, <4 x i32> addrspace(4) * %ptr2, align 32
+  %r = call <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32(i32 15, float %s, <8 x i32> %rsrc, <4 x i32> %samp, i1 0, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %wf_token, <4 x float> %r)
+
+  ret <4 x float> %r1
+}
+
+define amdgpu_ps void @test_waterfall_sample_with_kill(<8 x i32> addrspace(4)* inreg %in, <4 x i32> addrspace(4)* inreg %samp_in, i32 %index, float %s, i32 inreg %val) #1 {
+; VI-LABEL: test_waterfall_sample_with_kill:
+; VI:       ; %bb.0:
+; VI-NEXT:    s_mov_b64 s[6:7], exec
+; VI-NEXT:    s_wqm_b64 exec, exec
+; VI-NEXT:    s_mov_b64 s[8:9], exec
+; VI-NEXT:  .LBB10_1: ; =>This Inner Loop Header: Depth=1
+; VI-NEXT:    v_readfirstlane_b32 s10, v0
+; VI-NEXT:    v_cmp_eq_u32_e64 s[12:13], s10, v0
+; VI-NEXT:    s_and_saveexec_b64 s[24:25], s[12:13]
+; VI-NEXT:    s_ashr_i32 s11, s10, 31
+; VI-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; VI-NEXT:    s_add_u32 s10, s0, s10
+; VI-NEXT:    s_addc_u32 s11, s1, s11
+; VI-NEXT:    s_ashr_i32 s5, s4, 31
+; VI-NEXT:    s_lshl_b64 s[12:13], s[4:5], 4
+; VI-NEXT:    s_add_u32 s20, s2, s12
+; VI-NEXT:    s_addc_u32 s21, s3, s13
+; VI-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
+; VI-NEXT:    s_load_dwordx4 s[20:23], s[20:21], 0x0
+; VI-NEXT:    ; implicit-def: $vgpr0
+; VI-NEXT:    s_waitcnt lgkmcnt(0)
+; VI-NEXT:    image_sample v[2:5], v1, s[12:19], s[20:23] dmask:0xf
+; VI-NEXT:    ; implicit-def: $vgpr1
+; VI-NEXT:    s_xor_b64 exec, exec, s[24:25]
+; VI-NEXT:    s_cbranch_execnz .LBB10_1
+; VI-NEXT:  ; %bb.2:
+; VI-NEXT:    s_mov_b64 exec, s[8:9]
+; VI-NEXT:    s_and_b64 exec, exec, s[6:7]
+; VI-NEXT:    s_waitcnt vmcnt(0)
+; VI-NEXT:    v_cmp_gt_f32_e32 vcc, 0, v2
+; VI-NEXT:    s_and_saveexec_b64 s[0:1], vcc
+; VI-NEXT:    s_xor_b64 s[0:1], exec, s[0:1]
+; VI-NEXT:    s_cbranch_execz .LBB10_5
+; VI-NEXT:  ; %bb.3: ; %.kill
+; VI-NEXT:    s_andn2_b64 s[6:7], s[6:7], exec
+; VI-NEXT:    s_cbranch_scc0 .LBB10_6
+; VI-NEXT:  ; %bb.4: ; %.kill
+; VI-NEXT:    s_mov_b64 exec, 0
+; VI-NEXT:  .LBB10_5: ; %.exit
+; VI-NEXT:    s_or_b64 exec, exec, s[0:1]
+; VI-NEXT:    v_mov_b32_e32 v0, 0
+; VI-NEXT:    exp mrt0, v0, off, off, off done vm
+; VI-NEXT:    s_endpgm
+; VI-NEXT:  .LBB10_6:
+; VI-NEXT:    s_mov_b64 exec, 0
+; VI-NEXT:    exp null, off, off, off, off done vm
+; VI-NEXT:    s_endpgm
+;
+; GFX9-LABEL: test_waterfall_sample_with_kill:
+; GFX9:       ; %bb.0:
+; GFX9-NEXT:    s_mov_b64 s[6:7], exec
+; GFX9-NEXT:    s_wqm_b64 exec, exec
+; GFX9-NEXT:    s_mov_b64 s[8:9], exec
+; GFX9-NEXT:  .LBB10_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-NEXT:    v_readfirstlane_b32 s10, v0
+; GFX9-NEXT:    v_cmp_eq_u32_e64 s[12:13], s10, v0
+; GFX9-NEXT:    s_and_saveexec_b64 s[24:25], s[12:13]
+; GFX9-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX9-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX9-NEXT:    s_add_u32 s10, s0, s10
+; GFX9-NEXT:    s_addc_u32 s11, s1, s11
+; GFX9-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX9-NEXT:    s_lshl_b64 s[12:13], s[4:5], 4
+; GFX9-NEXT:    s_add_u32 s26, s2, s12
+; GFX9-NEXT:    s_addc_u32 s27, s3, s13
+; GFX9-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
+; GFX9-NEXT:    s_load_dwordx4 s[20:23], s[26:27], 0x0
+; GFX9-NEXT:    ; implicit-def: $vgpr0
+; GFX9-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX9-NEXT:    image_sample v[2:5], v1, s[12:19], s[20:23] dmask:0xf
+; GFX9-NEXT:    ; implicit-def: $vgpr1
+; GFX9-NEXT:    s_xor_b64 exec, exec, s[24:25]
+; GFX9-NEXT:    s_cbranch_execnz .LBB10_1
+; GFX9-NEXT:  ; %bb.2:
+; GFX9-NEXT:    s_mov_b64 exec, s[8:9]
+; GFX9-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX9-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-NEXT:    v_cmp_gt_f32_e32 vcc, 0, v2
+; GFX9-NEXT:    s_and_saveexec_b64 s[0:1], vcc
+; GFX9-NEXT:    s_xor_b64 s[0:1], exec, s[0:1]
+; GFX9-NEXT:    s_cbranch_execz .LBB10_5
+; GFX9-NEXT:  ; %bb.3: ; %.kill
+; GFX9-NEXT:    s_andn2_b64 s[6:7], s[6:7], exec
+; GFX9-NEXT:    s_cbranch_scc0 .LBB10_6
+; GFX9-NEXT:  ; %bb.4: ; %.kill
+; GFX9-NEXT:    s_mov_b64 exec, 0
+; GFX9-NEXT:  .LBB10_5: ; %.exit
+; GFX9-NEXT:    s_or_b64 exec, exec, s[0:1]
+; GFX9-NEXT:    v_mov_b32_e32 v0, 0
+; GFX9-NEXT:    exp mrt0, v0, off, off, off done vm
+; GFX9-NEXT:    s_endpgm
+; GFX9-NEXT:  .LBB10_6:
+; GFX9-NEXT:    s_mov_b64 exec, 0
+; GFX9-NEXT:    exp null, off, off, off, off done vm
+; GFX9-NEXT:    s_endpgm
+;
+; GFX10-32-LABEL: test_waterfall_sample_with_kill:
+; GFX10-32:       ; %bb.0:
+; GFX10-32-NEXT:    s_mov_b32 s6, exec_lo
+; GFX10-32-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-NEXT:    s_mov_b32 s8, exec_lo
+; GFX10-32-NEXT:    s_mov_b32 s7, exec_lo
+; GFX10-32-NEXT:  .LBB10_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-NEXT:    v_readfirstlane_b32 s10, v0
+; GFX10-32-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-NEXT:    v_cmpx_eq_u32_e32 s10, v0
+; GFX10-32-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX10-32-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX10-32-NEXT:    s_add_u32 s10, s0, s10
+; GFX10-32-NEXT:    s_addc_u32 s11, s1, s11
+; GFX10-32-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX10-32-NEXT:    s_lshl_b64 s[12:13], s[4:5], 4
+; GFX10-32-NEXT:    s_add_u32 s24, s2, s12
+; GFX10-32-NEXT:    s_addc_u32 s25, s3, s13
+; GFX10-32-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
+; GFX10-32-NEXT:    s_load_dwordx4 s[20:23], s[24:25], 0x0
+; GFX10-32-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-32-NEXT:    image_sample v[2:5], v1, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-32-NEXT:    s_andn2_wrexec_b32 s8, s8
+; GFX10-32-NEXT:    ; implicit-def: $vgpr0
+; GFX10-32-NEXT:    ; implicit-def: $vgpr1
+; GFX10-32-NEXT:    s_cbranch_execnz .LBB10_1
+; GFX10-32-NEXT:  ; %bb.2:
+; GFX10-32-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-NEXT:    s_mov_b32 exec_lo, s7
+; GFX10-32-NEXT:    s_and_b32 exec_lo, exec_lo, s6
+; GFX10-32-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-NEXT:    v_cmp_gt_f32_e32 vcc_lo, 0, v2
+; GFX10-32-NEXT:    s_and_saveexec_b32 s0, vcc_lo
+; GFX10-32-NEXT:    s_xor_b32 s0, exec_lo, s0
+; GFX10-32-NEXT:    s_cbranch_execz .LBB10_5
+; GFX10-32-NEXT:  ; %bb.3: ; %.kill
+; GFX10-32-NEXT:    s_andn2_b32 s6, s6, exec_lo
+; GFX10-32-NEXT:    s_cbranch_scc0 .LBB10_6
+; GFX10-32-NEXT:  ; %bb.4: ; %.kill
+; GFX10-32-NEXT:    s_mov_b32 exec_lo, 0
+; GFX10-32-NEXT:  .LBB10_5: ; %.exit
+; GFX10-32-NEXT:    s_or_b32 exec_lo, exec_lo, s0
+; GFX10-32-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-NEXT:    exp mrt0, v0, off, off, off done vm
+; GFX10-32-NEXT:    s_endpgm
+; GFX10-32-NEXT:  .LBB10_6:
+; GFX10-32-NEXT:    s_mov_b32 exec_lo, 0
+; GFX10-32-NEXT:    exp null, off, off, off, off done vm
+; GFX10-32-NEXT:    s_endpgm
+;
+; GFX10-64-LABEL: test_waterfall_sample_with_kill:
+; GFX10-64:       ; %bb.0:
+; GFX10-64-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-NEXT:    s_mov_b64 s[10:11], exec
+; GFX10-64-NEXT:    s_mov_b64 s[8:9], exec
+; GFX10-64-NEXT:  .LBB10_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-NEXT:    v_readfirstlane_b32 s12, v0
+; GFX10-64-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-NEXT:    v_cmpx_eq_u32_e64 s12, v0
+; GFX10-64-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX10-64-NEXT:    s_lshl_b64 s[12:13], s[12:13], 5
+; GFX10-64-NEXT:    s_add_u32 s24, s0, s12
+; GFX10-64-NEXT:    s_addc_u32 s25, s1, s13
+; GFX10-64-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX10-64-NEXT:    s_lshl_b64 s[12:13], s[4:5], 4
+; GFX10-64-NEXT:    s_add_u32 s26, s2, s12
+; GFX10-64-NEXT:    s_addc_u32 s27, s3, s13
+; GFX10-64-NEXT:    s_load_dwordx8 s[12:19], s[24:25], 0x0
+; GFX10-64-NEXT:    s_load_dwordx4 s[20:23], s[26:27], 0x0
+; GFX10-64-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-64-NEXT:    image_sample v[2:5], v1, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-64-NEXT:    s_andn2_wrexec_b64 s[10:11], s[10:11]
+; GFX10-64-NEXT:    ; implicit-def: $vgpr0
+; GFX10-64-NEXT:    ; implicit-def: $vgpr1
+; GFX10-64-NEXT:    s_cbranch_execnz .LBB10_1
+; GFX10-64-NEXT:  ; %bb.2:
+; GFX10-64-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-NEXT:    s_mov_b64 exec, s[8:9]
+; GFX10-64-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX10-64-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-NEXT:    v_cmp_gt_f32_e32 vcc, 0, v2
+; GFX10-64-NEXT:    s_and_saveexec_b64 s[0:1], vcc
+; GFX10-64-NEXT:    s_xor_b64 s[0:1], exec, s[0:1]
+; GFX10-64-NEXT:    s_cbranch_execz .LBB10_5
+; GFX10-64-NEXT:  ; %bb.3: ; %.kill
+; GFX10-64-NEXT:    s_andn2_b64 s[6:7], s[6:7], exec
+; GFX10-64-NEXT:    s_cbranch_scc0 .LBB10_6
+; GFX10-64-NEXT:  ; %bb.4: ; %.kill
+; GFX10-64-NEXT:    s_mov_b64 exec, 0
+; GFX10-64-NEXT:  .LBB10_5: ; %.exit
+; GFX10-64-NEXT:    s_or_b64 exec, exec, s[0:1]
+; GFX10-64-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-NEXT:    exp mrt0, v0, off, off, off done vm
+; GFX10-64-NEXT:    s_endpgm
+; GFX10-64-NEXT:  .LBB10_6:
+; GFX10-64-NEXT:    s_mov_b64 exec, 0
+; GFX10-64-NEXT:    exp null, off, off, off, off done vm
+; GFX10-64-NEXT:    s_endpgm
+;
+; GFX1150-LABEL: test_waterfall_sample_with_kill:
+; GFX1150:       ; %bb.0:
+; GFX1150-NEXT:    s_setprio 2
+; GFX1150-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-NEXT:    s_mov_b64 s[10:11], exec
+; GFX1150-NEXT:    s_mov_b64 s[8:9], exec
+; GFX1150-NEXT:  .LBB10_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-NEXT:    v_readfirstlane_b32 s12, v0
+; GFX1150-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_1) | instid1(SALU_CYCLE_1)
+; GFX1150-NEXT:    v_cmpx_eq_u32_e64 s12, v0
+; GFX1150-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX1150-NEXT:    s_lshl_b64 s[12:13], s[12:13], 5
+; GFX1150-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(SKIP_2) | instid1(SALU_CYCLE_1)
+; GFX1150-NEXT:    s_add_u32 s12, s0, s12
+; GFX1150-NEXT:    s_addc_u32 s13, s1, s13
+; GFX1150-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX1150-NEXT:    s_lshl_b64 s[14:15], s[4:5], 4
+; GFX1150-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-NEXT:    s_add_u32 s20, s2, s14
+; GFX1150-NEXT:    s_addc_u32 s21, s3, s15
+; GFX1150-NEXT:    s_load_b256 s[12:19], s[12:13], 0x0
+; GFX1150-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX1150-NEXT:    s_waitcnt vmcnt(0) lgkmcnt(0)
+; GFX1150-NEXT:    image_sample v[2:5], v1, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX1150-NEXT:    s_and_not1_wrexec_b64 s[10:11], s[10:11]
+; GFX1150-NEXT:    ; implicit-def: $vgpr0
+; GFX1150-NEXT:    ; implicit-def: $vgpr1
+; GFX1150-NEXT:    s_cbranch_execnz .LBB10_1
+; GFX1150-NEXT:  ; %bb.2:
+; GFX1150-NEXT:    s_mov_b64 exec, s[8:9]
+; GFX1150-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(SALU_CYCLE_1)
+; GFX1150-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX1150-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-NEXT:    v_cmpx_gt_f32_e32 0, v2
+; GFX1150-NEXT:    s_xor_b64 s[0:1], exec, s[0:1]
+; GFX1150-NEXT:    s_cbranch_execz .LBB10_5
+; GFX1150-NEXT:  ; %bb.3: ; %.kill
+; GFX1150-NEXT:    s_and_not1_b64 s[6:7], s[6:7], exec
+; GFX1150-NEXT:    s_cbranch_scc0 .LBB10_6
+; GFX1150-NEXT:  ; %bb.4: ; %.kill
+; GFX1150-NEXT:    s_mov_b64 exec, 0
+; GFX1150-NEXT:  .LBB10_5: ; %.exit
+; GFX1150-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-NEXT:    s_or_b64 exec, exec, s[0:1]
+; GFX1150-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-NEXT:    exp mrt0, v0, off, off, off done
+; GFX1150-NEXT:    s_setprio 0
+; GFX1150-NEXT:    s_nop 0
+; GFX1150-NEXT:    s_nop 0
+; GFX1150-NEXT:    s_endpgm
+; GFX1150-NEXT:  .LBB10_6:
+; GFX1150-NEXT:    s_mov_b64 exec, 0
+; GFX1150-NEXT:    exp mrt0, off, off, off, off done
+; GFX1150-NEXT:    s_setprio 0
+; GFX1150-NEXT:    s_nop 0
+; GFX1150-NEXT:    s_nop 0
+; GFX1150-NEXT:    s_endpgm
+;
+; GFX12-SDAG-LABEL: test_waterfall_sample_with_kill:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; GFX12-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX12-SDAG-NEXT:  .LBB10_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s12, v0
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s12, v0
+; GFX12-SDAG-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX12-SDAG-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    s_lshl_b64 s[12:13], s[12:13], 5
+; GFX12-SDAG-NEXT:    s_lshl_b64 s[14:15], s[4:5], 4
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    s_add_nc_u64 s[12:13], s[0:1], s[12:13]
+; GFX12-SDAG-NEXT:    s_add_nc_u64 s[20:21], s[2:3], s[14:15]
+; GFX12-SDAG-NEXT:    s_load_b256 s[12:19], s[12:13], 0x0
+; GFX12-SDAG-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX12-SDAG-NEXT:    s_wait_kmcnt 0x0
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    image_sample v[2:5], v1, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[10:11], s[10:11]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr0
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr1
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB10_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[8:9]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(SALU_CYCLE_1)
+; GFX12-SDAG-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX12-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    v_cmpx_gt_f32_e32 0, v2
+; GFX12-SDAG-NEXT:    s_xor_b64 s[0:1], exec, s[0:1]
+; GFX12-SDAG-NEXT:    s_cbranch_execz .LBB10_5
+; GFX12-SDAG-NEXT:  ; %bb.3: ; %.kill
+; GFX12-SDAG-NEXT:    s_and_not1_b64 s[6:7], s[6:7], exec
+; GFX12-SDAG-NEXT:    s_cbranch_scc0 .LBB10_6
+; GFX12-SDAG-NEXT:  ; %bb.4: ; %.kill
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, 0
+; GFX12-SDAG-NEXT:  .LBB10_5: ; %.exit
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-SDAG-NEXT:    s_or_b64 exec, exec, s[0:1]
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-SDAG-NEXT:    export mrt0, v0, off, off, off done
+; GFX12-SDAG-NEXT:    s_endpgm
+; GFX12-SDAG-NEXT:  .LBB10_6:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, 0
+; GFX12-SDAG-NEXT:    export mrt0, off, off, off, off done
+; GFX12-SDAG-NEXT:    s_endpgm
+;
+; GFX12-GISEL-LABEL: test_waterfall_sample_with_kill:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_mov_b64 s[10:11], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX12-GISEL-NEXT:  .LBB10_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s12, v0
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s12, v0
+; GFX12-GISEL-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_lshl_b64 s[12:13], s[12:13], 5
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_add_co_u32 s12, s0, s12
+; GFX12-GISEL-NEXT:    s_add_co_ci_u32 s13, s1, s13
+; GFX12-GISEL-NEXT:    s_ashr_i32 s5, s4, 31
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_lshl_b64 s[14:15], s[4:5], 4
+; GFX12-GISEL-NEXT:    s_add_co_u32 s20, s2, s14
+; GFX12-GISEL-NEXT:    s_add_co_ci_u32 s21, s3, s15
+; GFX12-GISEL-NEXT:    s_load_b256 s[12:19], s[12:13], 0x0
+; GFX12-GISEL-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX12-GISEL-NEXT:    s_wait_kmcnt 0x0
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    image_sample v[2:5], v1, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[10:11], s[10:11]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr0
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr1
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB10_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[8:9]
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX12-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    v_cmpx_gt_f32_e32 0, v2
+; GFX12-GISEL-NEXT:    s_xor_b64 s[0:1], exec, s[0:1]
+; GFX12-GISEL-NEXT:    s_cbranch_execz .LBB10_5
+; GFX12-GISEL-NEXT:  ; %bb.3: ; %.kill
+; GFX12-GISEL-NEXT:    s_and_not1_b64 s[6:7], s[6:7], exec
+; GFX12-GISEL-NEXT:    s_cbranch_scc0 .LBB10_6
+; GFX12-GISEL-NEXT:  ; %bb.4: ; %.kill
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, 0
+; GFX12-GISEL-NEXT:  .LBB10_5: ; %.exit
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_or_b64 exec, exec, s[0:1]
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-GISEL-NEXT:    export mrt0, v0, off, off, off done
+; GFX12-GISEL-NEXT:    s_endpgm
+; GFX12-GISEL-NEXT:  .LBB10_6:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, 0
+; GFX12-GISEL-NEXT:    export mrt0, off, off, off, off done
+; GFX12-GISEL-NEXT:    s_endpgm
+  %wf_token = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %index)
+  %s_idx = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %wf_token, i32 %index)
+  %s_idx2 = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %wf_token, i32 %val)
+  %ptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx
+  %ptr2 = getelementptr <4 x i32>, <4 x i32> addrspace(4)* %samp_in, i32 %s_idx2
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4) * %ptr, align 32
+  %samp = load <4 x i32>, <4 x i32> addrspace(4) * %ptr2, align 32
+  %r = call <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32(i32 15, float %s, <8 x i32> %rsrc, <4 x i32> %samp, i1 0, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %wf_token, <4 x float> %r)
+  %r2 = extractelement <4 x float> %r1, i32 0
+  %cond = fcmp olt float %r2, 0.000000e+00
+  br i1 %cond, label %.kill, label %.exit
+
+.kill:
+  call void @llvm.amdgcn.kill(i1 false)
+  br label %.exit
+
+.exit:
+  call void @llvm.amdgcn.exp.f32(i32 immarg 0, i32 immarg 1, float 0.000000e+00, float undef, float undef, float undef, i1 immarg true, i1 immarg true)
+  ret void
+}
+
+
+define amdgpu_ps <4 x float> @test_waterfall_multi_begin(<8 x i32> addrspace(4)* inreg %in, <4 x i32> addrspace(4)* inreg %s_in,
+; VI-SDAG-LABEL: test_waterfall_multi_begin:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; VI-SDAG-NEXT:    s_wqm_b64 exec, exec
+; VI-SDAG-NEXT:    v_mov_b32_e32 v7, v2
+; VI-SDAG-NEXT:    v_ashrrev_i32_e32 v8, 31, v7
+; VI-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; VI-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; VI-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[7:8]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v2, s1
+; VI-SDAG-NEXT:    v_add_u32_e32 v0, vcc, s0, v0
+; VI-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; VI-SDAG-NEXT:    v_addc_u32_e32 v1, vcc, v2, v1, vcc
+; VI-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[3:4]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, s3
+; VI-SDAG-NEXT:    v_add_u32_e32 v2, vcc, s2, v2
+; VI-SDAG-NEXT:    v_addc_u32_e32 v3, vcc, v4, v3, vcc
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[7:10], v[0:1]
+; VI-SDAG-NEXT:    v_add_u32_e32 v0, vcc, 16, v0
+; VI-SDAG-NEXT:    v_addc_u32_e32 v1, vcc, 0, v1, vcc
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[11:14], v[0:1]
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[15:18], v[2:3]
+; VI-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; VI-SDAG-NEXT:  .LBB11_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s2, v5
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s6, v6
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v5
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[6:7], s6, v6
+; VI-SDAG-NEXT:    s_and_b64 s[2:3], s[2:3], s[6:7]
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(2)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s8, v7
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s9, v8
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s10, v9
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s11, v10
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s12, v11
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s13, v12
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s14, v13
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s15, v14
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s16, v15
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s17, v16
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s18, v17
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s19, v18
+; VI-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr5
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr6
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr15_vgpr16_vgpr17_vgpr18
+; VI-SDAG-NEXT:    s_nop 3
+; VI-SDAG-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[16:19] dmask:0xf
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB11_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; VI-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    ; return to shader part epilog
+;
+; VI-GISEL-LABEL: test_waterfall_multi_begin:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; VI-GISEL-NEXT:    s_wqm_b64 exec, exec
+; VI-GISEL-NEXT:    v_mov_b32_e32 v6, v3
+; VI-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v5, v1
+; VI-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[2:3]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; VI-GISEL-NEXT:    v_add_u32_e32 v0, vcc, v2, v0
+; VI-GISEL-NEXT:    v_ashrrev_i32_e32 v7, 31, v6
+; VI-GISEL-NEXT:    v_addc_u32_e32 v1, vcc, v3, v1, vcc
+; VI-GISEL-NEXT:    v_lshlrev_b64 v[2:3], 4, v[6:7]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v7, s3
+; VI-GISEL-NEXT:    v_mov_b32_e32 v6, s2
+; VI-GISEL-NEXT:    v_add_u32_e32 v2, vcc, v6, v2
+; VI-GISEL-NEXT:    v_addc_u32_e32 v3, vcc, v7, v3, vcc
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[6:9], v[0:1]
+; VI-GISEL-NEXT:    v_add_u32_e32 v0, vcc, 16, v0
+; VI-GISEL-NEXT:    v_addc_u32_e32 v1, vcc, 0, v1, vcc
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[10:13], v[0:1]
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[14:17], v[2:3]
+; VI-GISEL-NEXT:    s_mov_b32 s0, 0
+; VI-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; VI-GISEL-NEXT:  .LBB11_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s1, v4
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s8, v5
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[6:7], s1, v4
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[8:9], s8, v5
+; VI-GISEL-NEXT:    s_and_b64 s[6:7], s[6:7], s[8:9]
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[6:7], s[6:7]
+; VI-GISEL-NEXT:    s_mov_b32 s1, s0
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    v_mov_b32_e32 v0, s0
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s8, v6
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s9, v7
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s10, v8
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s11, v9
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s12, v10
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s13, v11
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s14, v12
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s15, v13
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s16, v14
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s17, v15
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s18, v16
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s19, v17
+; VI-GISEL-NEXT:    v_mov_b32_e32 v1, s1
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr4
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr5
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr14_vgpr15_vgpr16_vgpr17
+; VI-GISEL-NEXT:    s_nop 3
+; VI-GISEL-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[16:19] dmask:0xf
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[6:7]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB11_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; VI-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX9-SDAG-LABEL: test_waterfall_multi_begin:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX9-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v7, v2
+; GFX9-SDAG-NEXT:    v_ashrrev_i32_e32 v8, 31, v7
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX9-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[7:8]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v2, s1
+; GFX9-SDAG-NEXT:    v_add_co_u32_e32 v0, vcc, s0, v0
+; GFX9-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX9-SDAG-NEXT:    v_addc_co_u32_e32 v1, vcc, v2, v1, vcc
+; GFX9-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[3:4]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v4, s3
+; GFX9-SDAG-NEXT:    v_add_co_u32_e32 v2, vcc, s2, v2
+; GFX9-SDAG-NEXT:    v_addc_co_u32_e32 v3, vcc, v4, v3, vcc
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[11:14], v[0:1], off offset:16
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[7:10], v[0:1], off
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[15:18], v[2:3], off
+; GFX9-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX9-SDAG-NEXT:  .LBB11_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s2, v5
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s6, v6
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v5
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[6:7], s6, v6
+; GFX9-SDAG-NEXT:    s_and_b64 s[2:3], s[2:3], s[6:7]
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s8, v7
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s9, v8
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s10, v9
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s11, v10
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s13, v12
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s14, v13
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s15, v14
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s16, v15
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s17, v16
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s18, v17
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s19, v18
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr15_vgpr16_vgpr17_vgpr18
+; GFX9-SDAG-NEXT:    s_nop 3
+; GFX9-SDAG-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[16:19] dmask:0xf
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB11_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX9-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX9-GISEL-LABEL: test_waterfall_multi_begin:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX9-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v6, v3
+; GFX9-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v5, v1
+; GFX9-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[2:3]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX9-GISEL-NEXT:    v_add_co_u32_e32 v0, vcc, v2, v0
+; GFX9-GISEL-NEXT:    v_ashrrev_i32_e32 v7, 31, v6
+; GFX9-GISEL-NEXT:    v_addc_co_u32_e32 v1, vcc, v3, v1, vcc
+; GFX9-GISEL-NEXT:    v_lshlrev_b64 v[2:3], 4, v[6:7]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v7, s3
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v6, s2
+; GFX9-GISEL-NEXT:    v_add_co_u32_e32 v2, vcc, v6, v2
+; GFX9-GISEL-NEXT:    v_addc_co_u32_e32 v3, vcc, v7, v3, vcc
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[6:9], v[0:1], off
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[10:13], v[0:1], off offset:16
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[14:17], v[2:3], off
+; GFX9-GISEL-NEXT:    s_mov_b32 s0, 0
+; GFX9-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX9-GISEL-NEXT:  .LBB11_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s1, v4
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[6:7], s1, v4
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[8:9], s8, v5
+; GFX9-GISEL-NEXT:    s_and_b64 s[6:7], s[6:7], s[8:9]
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[6:7], s[6:7]
+; GFX9-GISEL-NEXT:    s_mov_b32 s1, s0
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v0, s0
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s8, v6
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s9, v7
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s10, v8
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s11, v9
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s12, v10
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s13, v11
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s14, v12
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s15, v13
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s16, v14
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s17, v15
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s18, v16
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s19, v17
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v1, s1
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr14_vgpr15_vgpr16_vgpr17
+; GFX9-GISEL-NEXT:    s_nop 3
+; GFX9-GISEL-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[16:19] dmask:0xf
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[6:7]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB11_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX9-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-SDAG-LABEL: test_waterfall_multi_begin:
+; GFX10-32-SDAG:       ; %bb.0:
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-SDAG-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v7, v2
+; GFX10-32-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX10-32-SDAG-NEXT:    v_ashrrev_i32_e32 v8, 31, v7
+; GFX10-32-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[3:4]
+; GFX10-32-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[7:8]
+; GFX10-32-SDAG-NEXT:    v_add_co_u32 v0, vcc_lo, s0, v0
+; GFX10-32-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, s1, v1, vcc_lo
+; GFX10-32-SDAG-NEXT:    v_add_co_u32 v2, vcc_lo, s2, v2
+; GFX10-32-SDAG-NEXT:    v_add_co_ci_u32_e32 v3, vcc_lo, s3, v3, vcc_lo
+; GFX10-32-SDAG-NEXT:    s_clause 0x1
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[11:14], v[0:1], off offset:16
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[7:10], v[0:1], off
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[15:18], v[2:3], off
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-SDAG-NEXT:  .LBB11_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s2, v5
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s3, v6
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s2, v5
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s3, v6
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s8, v7
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s9, v8
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s10, v9
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s11, v10
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s13, v12
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s14, v13
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s15, v14
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s16, v15
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s17, v16
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s18, v17
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s19, v18
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr15_vgpr16_vgpr17_vgpr18
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB11_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-SDAG-NEXT:    s_and_b32 exec_lo, exec_lo, s4
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-GISEL-LABEL: test_waterfall_multi_begin:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-GISEL-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v6, v3
+; GFX10-32-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v5, v1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v9, s3
+; GFX10-32-GISEL-NEXT:    v_ashrrev_i32_e32 v7, 31, v6
+; GFX10-32-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[2:3]
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v8, s2
+; GFX10-32-GISEL-NEXT:    v_lshlrev_b64 v[6:7], 4, v[6:7]
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-GISEL-NEXT:    v_add_co_u32 v0, vcc_lo, v2, v0
+; GFX10-32-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, v3, v1, vcc_lo
+; GFX10-32-GISEL-NEXT:    v_add_co_u32 v2, vcc_lo, v8, v6
+; GFX10-32-GISEL-NEXT:    v_add_co_ci_u32_e32 v3, vcc_lo, v9, v7, vcc_lo
+; GFX10-32-GISEL-NEXT:    s_clause 0x1
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[6:9], v[0:1], off
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[10:13], v[0:1], off offset:16
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[14:17], v[2:3], off
+; GFX10-32-GISEL-NEXT:  .LBB11_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s2, v4
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s3, v5
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s2, v4
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s3, v5
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s8, v6
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s9, v7
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s10, v8
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s11, v9
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s12, v10
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s13, v11
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s14, v12
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s15, v13
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s16, v14
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s17, v15
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s18, v16
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s19, v17
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr14_vgpr15_vgpr16_vgpr17
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB11_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-GISEL-NEXT:    s_and_b32 exec_lo, exec_lo, s4
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-SDAG-LABEL: test_waterfall_multi_begin:
+; GFX10-64-SDAG:       ; %bb.0:
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v7, v2
+; GFX10-64-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX10-64-SDAG-NEXT:    v_ashrrev_i32_e32 v8, 31, v7
+; GFX10-64-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[3:4]
+; GFX10-64-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[7:8]
+; GFX10-64-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX10-64-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc, s1, v1, vcc
+; GFX10-64-SDAG-NEXT:    v_add_co_u32 v2, vcc, s2, v2
+; GFX10-64-SDAG-NEXT:    v_add_co_ci_u32_e32 v3, vcc, s3, v3, vcc
+; GFX10-64-SDAG-NEXT:    s_clause 0x1
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[11:14], v[0:1], off offset:16
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[7:10], v[0:1], off
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[15:18], v[2:3], off
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-SDAG-NEXT:  .LBB11_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s6, v5
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s7, v6
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v5
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s7, v6
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s8, v7
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s9, v8
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s10, v9
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s11, v10
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s13, v12
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s14, v13
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s15, v14
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s16, v15
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s17, v16
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s18, v17
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s19, v18
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr15_vgpr16_vgpr17_vgpr18
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB11_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-GISEL-LABEL: test_waterfall_multi_begin:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v6, v3
+; GFX10-64-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v5, v1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v9, s3
+; GFX10-64-GISEL-NEXT:    v_ashrrev_i32_e32 v7, 31, v6
+; GFX10-64-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[2:3]
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v8, s2
+; GFX10-64-GISEL-NEXT:    v_lshlrev_b64 v[6:7], 4, v[6:7]
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX10-64-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc, v3, v1, vcc
+; GFX10-64-GISEL-NEXT:    v_add_co_u32 v2, vcc, v8, v6
+; GFX10-64-GISEL-NEXT:    v_add_co_ci_u32_e32 v3, vcc, v9, v7, vcc
+; GFX10-64-GISEL-NEXT:    s_clause 0x1
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[6:9], v[0:1], off
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[10:13], v[0:1], off offset:16
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[14:17], v[2:3], off
+; GFX10-64-GISEL-NEXT:  .LBB11_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s6, v4
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s7, v5
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s6, v4
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s7, v5
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s8, v6
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s9, v7
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s10, v8
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s11, v9
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s12, v10
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s13, v11
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s14, v12
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s15, v13
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s16, v14
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s17, v15
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s18, v16
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s19, v17
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr14_vgpr15_vgpr16_vgpr17
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB11_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX1150-SDAG-LABEL: test_waterfall_multi_begin:
+; GFX1150-SDAG:       ; %bb.0:
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v7, v2
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX1150-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_4) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-SDAG-NEXT:    v_ashrrev_i32_e32 v8, 31, v7
+; GFX1150-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[3:4]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[7:8]
+; GFX1150-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_4)
+; GFX1150-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s1, v1, vcc
+; GFX1150-SDAG-NEXT:    v_add_co_u32 v2, vcc, s2, v2
+; GFX1150-SDAG-NEXT:    s_waitcnt_depctr depctr_va_vcc(0)
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-SDAG-NEXT:    v_add_co_ci_u32_e64 v3, null, s3, v3, vcc
+; GFX1150-SDAG-NEXT:    s_clause 0x1
+; GFX1150-SDAG-NEXT:    global_load_b128 v[11:14], v[0:1], off offset:16
+; GFX1150-SDAG-NEXT:    global_load_b128 v[7:10], v[0:1], off
+; GFX1150-SDAG-NEXT:    global_load_b128 v[15:18], v[2:3], off
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-SDAG-NEXT:  .LBB11_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s6, v5
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s7, v6
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v5
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s7, v6
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s8, v7
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s9, v8
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s10, v9
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s11, v10
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s13, v12
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s14, v13
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s15, v14
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s16, v15
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s17, v16
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s18, v17
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s19, v18
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr15_vgpr16_vgpr17_vgpr18
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB11_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX1150-GISEL-LABEL: test_waterfall_multi_begin:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v6, v3
+; GFX1150-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v5, v1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v9, s3
+; GFX1150-GISEL-NEXT:    v_ashrrev_i32_e32 v7, 31, v6
+; GFX1150-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[2:3]
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v8, s2
+; GFX1150-GISEL-NEXT:    v_lshlrev_b64 v[6:7], 4, v[6:7]
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_3)
+; GFX1150-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v3, v1, vcc
+; GFX1150-GISEL-NEXT:    v_add_co_u32 v2, vcc, v8, v6
+; GFX1150-GISEL-NEXT:    s_waitcnt_depctr depctr_va_vcc(0)
+; GFX1150-GISEL-NEXT:    v_add_co_ci_u32_e64 v3, null, v9, v7, vcc
+; GFX1150-GISEL-NEXT:    s_clause 0x1
+; GFX1150-GISEL-NEXT:    global_load_b128 v[6:9], v[0:1], off
+; GFX1150-GISEL-NEXT:    global_load_b128 v[10:13], v[0:1], off offset:16
+; GFX1150-GISEL-NEXT:    global_load_b128 v[14:17], v[2:3], off
+; GFX1150-GISEL-NEXT:  .LBB11_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s6, v4
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s7, v5
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s6, v4
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s7, v5
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s8, v6
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s9, v7
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s10, v8
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s11, v9
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s12, v10
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s13, v11
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s14, v12
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s15, v13
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s16, v14
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s17, v15
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s18, v16
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s19, v17
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr14_vgpr15_vgpr16_vgpr17
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB11_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX12-SDAG-LABEL: test_waterfall_multi_begin:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v7, v2
+; GFX12-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_4) | instskip(NEXT) | instid1(VALU_DEP_4)
+; GFX12-SDAG-NEXT:    v_ashrrev_i32_e32 v8, 31, v7
+; GFX12-SDAG-NEXT:    v_lshlrev_b64_e32 v[2:3], 4, v[3:4]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_lshlrev_b64_e32 v[0:1], 5, v[7:8]
+; GFX12-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_4)
+; GFX12-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s1, v1, vcc
+; GFX12-SDAG-NEXT:    v_add_co_u32 v2, vcc, s2, v2
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_vcc(0)
+; GFX12-SDAG-NEXT:    v_add_co_ci_u32_e64 v3, null, s3, v3, vcc
+; GFX12-SDAG-NEXT:    s_clause 0x1
+; GFX12-SDAG-NEXT:    global_load_b128 v[11:14], v[0:1], off offset:16
+; GFX12-SDAG-NEXT:    global_load_b128 v[7:10], v[0:1], off
+; GFX12-SDAG-NEXT:    global_load_b128 v[15:18], v[2:3], off
+; GFX12-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-SDAG-NEXT:  .LBB11_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s6, v5
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s7, v6
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v5
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s7, v6
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s8, v7
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s9, v8
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s10, v9
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s11, v10
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s13, v12
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s14, v13
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s15, v14
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x0
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s16, v15
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s17, v16
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s18, v17
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s19, v18
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr15_vgpr16_vgpr17_vgpr18
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB11_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX12-GISEL-LABEL: test_waterfall_multi_begin:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v6, v3
+; GFX12-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v5, v1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v9, s3
+; GFX12-GISEL-NEXT:    v_ashrrev_i32_e32 v7, 31, v6
+; GFX12-GISEL-NEXT:    v_lshlrev_b64_e32 v[0:1], 5, v[2:3]
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v8, s2
+; GFX12-GISEL-NEXT:    v_lshlrev_b64_e32 v[6:7], 4, v[6:7]
+; GFX12-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_3)
+; GFX12-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v3, v1, vcc
+; GFX12-GISEL-NEXT:    v_add_co_u32 v2, vcc, v8, v6
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_vcc(0)
+; GFX12-GISEL-NEXT:    v_add_co_ci_u32_e64 v3, null, v9, v7, vcc
+; GFX12-GISEL-NEXT:    s_clause 0x1
+; GFX12-GISEL-NEXT:    global_load_b128 v[6:9], v[0:1], off
+; GFX12-GISEL-NEXT:    global_load_b128 v[10:13], v[0:1], off offset:16
+; GFX12-GISEL-NEXT:    global_load_b128 v[14:17], v[2:3], off
+; GFX12-GISEL-NEXT:  .LBB11_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s6, v4
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s7, v5
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s6, v4
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s7, v5
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x2
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s8, v6
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s9, v7
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s10, v8
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s11, v9
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s12, v10
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s13, v11
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s14, v12
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s15, v13
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x0
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s16, v14
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s17, v15
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s18, v16
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s19, v17
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr14_vgpr15_vgpr16_vgpr17
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB11_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    ; return to shader part epilog
+                                                         i32 %idx1, i32 %idx2, i32 %s_idx, i32 %s_idx2) #1 {
+  %rptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx
+  %sptr = getelementptr <4 x i32>, <4 x i32> addrspace(4)* %s_in, i32 %s_idx2
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4)* %rptr, align 16
+  %srsrc = load <4 x i32>, <4 x i32> addrspace(4)* %sptr, align 16
+  %tok = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %idx1)
+  %tok1 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 %tok, i32 %idx2)
+  %s_rsrc = call <8 x i32> @llvm.amdgcn.waterfall.readfirstlane.v8i32.v8i32(i32 %tok1, <8 x i32> %rsrc)
+  %s_srsrc = call <4 x i32> @llvm.amdgcn.waterfall.readfirstlane.v4i32.v4i32(i32 %tok1, <4 x i32> %srsrc)
+  %r = call <4 x float> @llvm.amdgcn.image.sample.2d.v4f32.f32(i32 15, float 0.000000e+00, float 0.000000e+00, <8 x i32> %s_rsrc, <4 x i32> %s_srsrc, i1 false, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok1, <4 x float> %r)
+
+  ret <4 x float> %r1
+}
+
+define amdgpu_ps <4 x float> @test_waterfall_full_idx_multi_begin(<8 x i32> addrspace(4)* inreg %in, <4 x i32> addrspace(4)* inreg %s_in,
+; VI-SDAG-LABEL: test_waterfall_full_idx_multi_begin:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; VI-SDAG-NEXT:    s_wqm_b64 exec, exec
+; VI-SDAG-NEXT:    v_mov_b32_e32 v2, v0
+; VI-SDAG-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; VI-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 5, v[2:3]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v0, s1
+; VI-SDAG-NEXT:    v_add_u32_e32 v8, vcc, s0, v2
+; VI-SDAG-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; VI-SDAG-NEXT:    v_addc_u32_e32 v9, vcc, v0, v3, vcc
+; VI-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 4, v[1:2]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v2, s3
+; VI-SDAG-NEXT:    v_add_u32_e32 v0, vcc, s2, v0
+; VI-SDAG-NEXT:    v_addc_u32_e32 v1, vcc, v2, v1, vcc
+; VI-SDAG-NEXT:    v_add_u32_e32 v2, vcc, 16, v8
+; VI-SDAG-NEXT:    v_addc_u32_e32 v3, vcc, 0, v9, vcc
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[4:7], v[8:9]
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[8:11], v[2:3]
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[12:15], v[0:1]
+; VI-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; VI-SDAG-NEXT:  .LBB12_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(2)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s0, v4
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s1, v5
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s2, v6
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[16:17], s0, v4
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s1, v5
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s3, v7
+; VI-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s2, v6
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s4, v8
+; VI-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s3, v7
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s5, v9
+; VI-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s4, v8
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s6, v10
+; VI-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s5, v9
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s7, v11
+; VI-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s6, v10
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s12, v12
+; VI-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s7, v11
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s13, v13
+; VI-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s12, v12
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s14, v14
+; VI-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s13, v13
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s15, v15
+; VI-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s14, v14
+; VI-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s15, v15
+; VI-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[16:17], s[16:17]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; VI-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; VI-SDAG-NEXT:    image_sample v[0:3], v[0:1], s[0:7], s[12:15] dmask:0xf
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr4_vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr12_vgpr13_vgpr14_vgpr15
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[16:17]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB12_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[10:11]
+; VI-SDAG-NEXT:    s_and_b64 exec, exec, s[8:9]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    ; return to shader part epilog
+;
+; VI-GISEL-LABEL: test_waterfall_full_idx_multi_begin:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; VI-GISEL-NEXT:    s_wqm_b64 exec, exec
+; VI-GISEL-NEXT:    v_mov_b32_e32 v2, v1
+; VI-GISEL-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; VI-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[0:1]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, s1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v3, s0
+; VI-GISEL-NEXT:    v_add_u32_e32 v0, vcc, v3, v0
+; VI-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; VI-GISEL-NEXT:    v_addc_u32_e32 v1, vcc, v4, v1, vcc
+; VI-GISEL-NEXT:    v_lshlrev_b64 v[2:3], 4, v[2:3]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v5, s3
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, s2
+; VI-GISEL-NEXT:    v_add_u32_e32 v2, vcc, v4, v2
+; VI-GISEL-NEXT:    v_addc_u32_e32 v3, vcc, v5, v3, vcc
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[4:7], v[0:1]
+; VI-GISEL-NEXT:    v_add_u32_e32 v0, vcc, 16, v0
+; VI-GISEL-NEXT:    v_addc_u32_e32 v1, vcc, 0, v1, vcc
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[8:11], v[0:1]
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[12:15], v[2:3]
+; VI-GISEL-NEXT:    s_mov_b32 s10, 0
+; VI-GISEL-NEXT:    s_mov_b64 s[12:13], exec
+; VI-GISEL-NEXT:  .LBB12_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s0, v4
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s1, v5
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s2, v6
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[14:15], s0, v4
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s1, v5
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s3, v7
+; VI-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s2, v6
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s4, v8
+; VI-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s3, v7
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s5, v9
+; VI-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s4, v8
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s6, v10
+; VI-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s5, v9
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s7, v11
+; VI-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s6, v10
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s16, v12
+; VI-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s7, v11
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s17, v13
+; VI-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s16, v12
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s18, v14
+; VI-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s17, v13
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s19, v15
+; VI-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s18, v14
+; VI-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s19, v15
+; VI-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[14:15], s[14:15]
+; VI-GISEL-NEXT:    s_mov_b32 s11, s10
+; VI-GISEL-NEXT:    v_mov_b32_e32 v0, s10
+; VI-GISEL-NEXT:    v_mov_b32_e32 v1, s11
+; VI-GISEL-NEXT:    image_sample v[0:3], v[0:1], s[0:7], s[16:19] dmask:0xf
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr4_vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr12_vgpr13_vgpr14_vgpr15
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[14:15]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB12_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[12:13]
+; VI-GISEL-NEXT:    s_and_b64 exec, exec, s[8:9]
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX9-SDAG-LABEL: test_waterfall_full_idx_multi_begin:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX9-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v2, v0
+; GFX9-SDAG-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX9-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 5, v[2:3]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v0, s1
+; GFX9-SDAG-NEXT:    v_add_co_u32_e32 v16, vcc, s0, v2
+; GFX9-SDAG-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX9-SDAG-NEXT:    v_addc_co_u32_e32 v17, vcc, v0, v3, vcc
+; GFX9-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 4, v[1:2]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v2, s3
+; GFX9-SDAG-NEXT:    v_add_co_u32_e32 v0, vcc, s2, v0
+; GFX9-SDAG-NEXT:    v_addc_co_u32_e32 v1, vcc, v2, v1, vcc
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[8:11], v[16:17], off offset:16
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[4:7], v[16:17], off
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[12:15], v[0:1], off
+; GFX9-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; GFX9-SDAG-NEXT:  .LBB12_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s0, v4
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s1, v5
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s2, v6
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[16:17], s0, v4
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s1, v5
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s3, v7
+; GFX9-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s2, v6
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s4, v8
+; GFX9-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s3, v7
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s5, v9
+; GFX9-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s4, v8
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s6, v10
+; GFX9-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s5, v9
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s7, v11
+; GFX9-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s6, v10
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s12, v12
+; GFX9-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s7, v11
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s13, v13
+; GFX9-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s12, v12
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s14, v14
+; GFX9-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s13, v13
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s15, v15
+; GFX9-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s14, v14
+; GFX9-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[18:19], s15, v15
+; GFX9-SDAG-NEXT:    s_and_b64 s[16:17], s[16:17], s[18:19]
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[16:17], s[16:17]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; GFX9-SDAG-NEXT:    image_sample v[0:3], v[0:1], s[0:7], s[12:15] dmask:0xf
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr4_vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr12_vgpr13_vgpr14_vgpr15
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[16:17]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB12_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[10:11]
+; GFX9-SDAG-NEXT:    s_and_b64 exec, exec, s[8:9]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX9-GISEL-LABEL: test_waterfall_full_idx_multi_begin:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX9-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v2, v1
+; GFX9-GISEL-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX9-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[0:1]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v4, s1
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v3, s0
+; GFX9-GISEL-NEXT:    v_add_co_u32_e32 v0, vcc, v3, v0
+; GFX9-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX9-GISEL-NEXT:    v_addc_co_u32_e32 v1, vcc, v4, v1, vcc
+; GFX9-GISEL-NEXT:    v_lshlrev_b64 v[2:3], 4, v[2:3]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v5, s3
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v4, s2
+; GFX9-GISEL-NEXT:    v_add_co_u32_e32 v2, vcc, v4, v2
+; GFX9-GISEL-NEXT:    v_addc_co_u32_e32 v3, vcc, v5, v3, vcc
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[4:7], v[0:1], off
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[8:11], v[0:1], off offset:16
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[12:15], v[2:3], off
+; GFX9-GISEL-NEXT:    s_mov_b32 s10, 0
+; GFX9-GISEL-NEXT:    s_mov_b64 s[12:13], exec
+; GFX9-GISEL-NEXT:  .LBB12_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s0, v4
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s1, v5
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s2, v6
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[14:15], s0, v4
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s1, v5
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s3, v7
+; GFX9-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s2, v6
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s4, v8
+; GFX9-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s3, v7
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s5, v9
+; GFX9-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s4, v8
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s6, v10
+; GFX9-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s5, v9
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s7, v11
+; GFX9-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s6, v10
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s16, v12
+; GFX9-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s7, v11
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s17, v13
+; GFX9-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s16, v12
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s18, v14
+; GFX9-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s17, v13
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s19, v15
+; GFX9-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s18, v14
+; GFX9-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[20:21], s19, v15
+; GFX9-GISEL-NEXT:    s_and_b64 s[14:15], s[14:15], s[20:21]
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[14:15], s[14:15]
+; GFX9-GISEL-NEXT:    s_mov_b32 s11, s10
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v0, s10
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v1, s11
+; GFX9-GISEL-NEXT:    image_sample v[0:3], v[0:1], s[0:7], s[16:19] dmask:0xf
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr4_vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr12_vgpr13_vgpr14_vgpr15
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[14:15]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB12_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[12:13]
+; GFX9-GISEL-NEXT:    s_and_b64 exec, exec, s[8:9]
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-SDAG-LABEL: test_waterfall_full_idx_multi_begin:
+; GFX10-32-SDAG:       ; %bb.0:
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-SDAG-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v3, v0
+; GFX10-32-SDAG-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX10-32-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX10-32-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 4, v[1:2]
+; GFX10-32-SDAG-NEXT:    v_lshlrev_b64 v[3:4], 5, v[3:4]
+; GFX10-32-SDAG-NEXT:    v_add_co_u32 v2, vcc_lo, s0, v3
+; GFX10-32-SDAG-NEXT:    v_add_co_ci_u32_e32 v3, vcc_lo, s1, v4, vcc_lo
+; GFX10-32-SDAG-NEXT:    v_add_co_u32 v0, vcc_lo, s2, v0
+; GFX10-32-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, s3, v1, vcc_lo
+; GFX10-32-SDAG-NEXT:    s_clause 0x1
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[8:11], v[2:3], off offset:16
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[4:7], v[2:3], off
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[12:15], v[0:1], off
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-SDAG-NEXT:  .LBB12_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s8, v4
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s9, v5
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s10, v6
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s11, v7
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s12, v8
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s13, v9
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s14, v10
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s15, v11
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s16, v12
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s17, v13
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s18, v14
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s19, v15
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s8, v4
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s9, v5
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s10, v6
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s11, v7
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s12, v8
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s13, v9
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s14, v10
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s15, v11
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s16, v12
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s17, v13
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s18, v14
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s19, v15
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr4_vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr12_vgpr13_vgpr14_vgpr15
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB12_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-SDAG-NEXT:    s_and_b32 exec_lo, exec_lo, s4
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-GISEL-LABEL: test_waterfall_full_idx_multi_begin:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-GISEL-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v2, v1
+; GFX10-32-GISEL-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v5, s1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v4, s0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v7, s3
+; GFX10-32-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX10-32-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[0:1]
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v6, s2
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-GISEL-NEXT:    v_lshlrev_b64 v[2:3], 4, v[2:3]
+; GFX10-32-GISEL-NEXT:    v_add_co_u32 v0, vcc_lo, v4, v0
+; GFX10-32-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, v5, v1, vcc_lo
+; GFX10-32-GISEL-NEXT:    v_add_co_u32 v2, vcc_lo, v6, v2
+; GFX10-32-GISEL-NEXT:    v_add_co_ci_u32_e32 v3, vcc_lo, v7, v3, vcc_lo
+; GFX10-32-GISEL-NEXT:    s_clause 0x1
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[4:7], v[0:1], off
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[8:11], v[0:1], off offset:16
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[12:15], v[2:3], off
+; GFX10-32-GISEL-NEXT:  .LBB12_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s8, v4
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s9, v5
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s10, v6
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s11, v7
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s12, v8
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s13, v9
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s14, v10
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s15, v11
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s16, v12
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s17, v13
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s18, v14
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s19, v15
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s8, v4
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s9, v5
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s10, v6
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s11, v7
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s12, v8
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s13, v9
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s14, v10
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s15, v11
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s16, v12
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s17, v13
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s18, v14
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s19, v15
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr4_vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr12_vgpr13_vgpr14_vgpr15
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB12_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-GISEL-NEXT:    s_and_b32 exec_lo, exec_lo, s4
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-SDAG-LABEL: test_waterfall_full_idx_multi_begin:
+; GFX10-64-SDAG:       ; %bb.0:
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v3, v0
+; GFX10-64-SDAG-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX10-64-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX10-64-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 4, v[1:2]
+; GFX10-64-SDAG-NEXT:    v_lshlrev_b64 v[3:4], 5, v[3:4]
+; GFX10-64-SDAG-NEXT:    v_add_co_u32 v2, vcc, s0, v3
+; GFX10-64-SDAG-NEXT:    v_add_co_ci_u32_e32 v3, vcc, s1, v4, vcc
+; GFX10-64-SDAG-NEXT:    v_add_co_u32 v0, vcc, s2, v0
+; GFX10-64-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc, s3, v1, vcc
+; GFX10-64-SDAG-NEXT:    s_clause 0x1
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[8:11], v[2:3], off offset:16
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[4:7], v[2:3], off
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[12:15], v[0:1], off
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-SDAG-NEXT:  .LBB12_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s8, v4
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s9, v5
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s10, v6
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s11, v7
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s12, v8
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s13, v9
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s14, v10
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s15, v11
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s16, v12
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s17, v13
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s18, v14
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s19, v15
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s8, v4
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s9, v5
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s10, v6
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s11, v7
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s12, v8
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s13, v9
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s14, v10
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s15, v11
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s16, v12
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s17, v13
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s18, v14
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s19, v15
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr4_vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr12_vgpr13_vgpr14_vgpr15
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB12_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-GISEL-LABEL: test_waterfall_full_idx_multi_begin:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v2, v1
+; GFX10-64-GISEL-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v5, s1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v4, s0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v7, s3
+; GFX10-64-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX10-64-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[0:1]
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v6, s2
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-GISEL-NEXT:    v_lshlrev_b64 v[2:3], 4, v[2:3]
+; GFX10-64-GISEL-NEXT:    v_add_co_u32 v0, vcc, v4, v0
+; GFX10-64-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc, v5, v1, vcc
+; GFX10-64-GISEL-NEXT:    v_add_co_u32 v2, vcc, v6, v2
+; GFX10-64-GISEL-NEXT:    v_add_co_ci_u32_e32 v3, vcc, v7, v3, vcc
+; GFX10-64-GISEL-NEXT:    s_clause 0x1
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[4:7], v[0:1], off
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[8:11], v[0:1], off offset:16
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[12:15], v[2:3], off
+; GFX10-64-GISEL-NEXT:  .LBB12_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s8, v4
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s9, v5
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s10, v6
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s11, v7
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s12, v8
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s13, v9
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s14, v10
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s15, v11
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s16, v12
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s17, v13
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s18, v14
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s19, v15
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s8, v4
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s9, v5
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s10, v6
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s11, v7
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s12, v8
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s13, v9
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s14, v10
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s15, v11
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s16, v12
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s17, v13
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s18, v14
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s19, v15
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr4_vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr12_vgpr13_vgpr14_vgpr15
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB12_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX1150-SDAG-LABEL: test_waterfall_full_idx_multi_begin:
+; GFX1150-SDAG:       ; %bb.0:
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v3, v0
+; GFX1150-SDAG-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX1150-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 4, v[1:2]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_lshlrev_b64 v[3:4], 5, v[3:4]
+; GFX1150-SDAG-NEXT:    v_add_co_u32 v2, vcc, s0, v3
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_4)
+; GFX1150-SDAG-NEXT:    v_add_co_ci_u32_e64 v3, null, s1, v4, vcc
+; GFX1150-SDAG-NEXT:    v_add_co_u32 v0, vcc, s2, v0
+; GFX1150-SDAG-NEXT:    s_waitcnt_depctr depctr_va_vcc(0)
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s3, v1, vcc
+; GFX1150-SDAG-NEXT:    s_clause 0x1
+; GFX1150-SDAG-NEXT:    global_load_b128 v[8:11], v[2:3], off offset:16
+; GFX1150-SDAG-NEXT:    global_load_b128 v[4:7], v[2:3], off
+; GFX1150-SDAG-NEXT:    global_load_b128 v[12:15], v[0:1], off
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-SDAG-NEXT:    s_set_inst_prefetch_distance 0x1
+; GFX1150-SDAG-NEXT:  .LBB12_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s8, v4
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s9, v5
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s10, v6
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s11, v7
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s12, v8
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s13, v9
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s14, v10
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s15, v11
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s16, v12
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s17, v13
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s18, v14
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s19, v15
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s8, v4
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s9, v5
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s10, v6
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s11, v7
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s12, v8
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s13, v9
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s14, v10
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s15, v11
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s16, v12
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s17, v13
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s18, v14
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s19, v15
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr4_vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr12_vgpr13_vgpr14_vgpr15
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB12_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_set_inst_prefetch_distance 0x2
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX1150-GISEL-LABEL: test_waterfall_full_idx_multi_begin:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v2, v1
+; GFX1150-GISEL-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v5, s1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v4, s0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v7, s3
+; GFX1150-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX1150-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[0:1]
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v6, s2
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-GISEL-NEXT:    v_lshlrev_b64 v[2:3], 4, v[2:3]
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_3) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-GISEL-NEXT:    v_add_co_u32 v0, vcc, v4, v0
+; GFX1150-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v5, v1, vcc
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_3) | instskip(SKIP_1) | instid1(VALU_DEP_4)
+; GFX1150-GISEL-NEXT:    v_add_co_u32 v2, vcc, v6, v2
+; GFX1150-GISEL-NEXT:    s_waitcnt_depctr depctr_va_vcc(0)
+; GFX1150-GISEL-NEXT:    v_add_co_ci_u32_e64 v3, null, v7, v3, vcc
+; GFX1150-GISEL-NEXT:    s_clause 0x1
+; GFX1150-GISEL-NEXT:    global_load_b128 v[4:7], v[0:1], off
+; GFX1150-GISEL-NEXT:    global_load_b128 v[8:11], v[0:1], off offset:16
+; GFX1150-GISEL-NEXT:    global_load_b128 v[12:15], v[2:3], off
+; GFX1150-GISEL-NEXT:    s_set_inst_prefetch_distance 0x1
+; GFX1150-GISEL-NEXT:  .LBB12_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s8, v4
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s9, v5
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s10, v6
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s11, v7
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s12, v8
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s13, v9
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s14, v10
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s15, v11
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s16, v12
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s17, v13
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s18, v14
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s19, v15
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s8, v4
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s9, v5
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s10, v6
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s11, v7
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s12, v8
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s13, v9
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s14, v10
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s15, v11
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s16, v12
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s17, v13
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s18, v14
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s19, v15
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr4_vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr12_vgpr13_vgpr14_vgpr15
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB12_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_set_inst_prefetch_distance 0x2
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX12-SDAG-LABEL: test_waterfall_full_idx_multi_begin:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v3, v0
+; GFX12-SDAG-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX12-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX12-SDAG-NEXT:    v_lshlrev_b64_e32 v[0:1], 4, v[1:2]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_lshlrev_b64_e32 v[3:4], 5, v[3:4]
+; GFX12-SDAG-NEXT:    v_add_co_u32 v2, vcc, s0, v3
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_4)
+; GFX12-SDAG-NEXT:    v_add_co_ci_u32_e64 v3, null, s1, v4, vcc
+; GFX12-SDAG-NEXT:    v_add_co_u32 v0, vcc, s2, v0
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_vcc(0)
+; GFX12-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s3, v1, vcc
+; GFX12-SDAG-NEXT:    s_clause 0x1
+; GFX12-SDAG-NEXT:    global_load_b128 v[8:11], v[2:3], off offset:16
+; GFX12-SDAG-NEXT:    global_load_b128 v[4:7], v[2:3], off
+; GFX12-SDAG-NEXT:    global_load_b128 v[12:15], v[0:1], off
+; GFX12-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-SDAG-NEXT:  .LBB12_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s8, v4
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s9, v5
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s10, v6
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s11, v7
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s12, v8
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s13, v9
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s14, v10
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s15, v11
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x0
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s16, v12
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s17, v13
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s18, v14
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s19, v15
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s8, v4
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s9, v5
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s10, v6
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s11, v7
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s12, v8
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s13, v9
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s14, v10
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s15, v11
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s16, v12
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s17, v13
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s18, v14
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s19, v15
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr4_vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr12_vgpr13_vgpr14_vgpr15
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB12_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX12-GISEL-LABEL: test_waterfall_full_idx_multi_begin:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v2, v1
+; GFX12-GISEL-NEXT:    v_ashrrev_i32_e32 v1, 31, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v5, s1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v4, s0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v7, s3
+; GFX12-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX12-GISEL-NEXT:    v_lshlrev_b64_e32 v[0:1], 5, v[0:1]
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v6, s2
+; GFX12-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-GISEL-NEXT:    v_lshlrev_b64_e32 v[2:3], 4, v[2:3]
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_3) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_add_co_u32 v0, vcc, v4, v0
+; GFX12-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v5, v1, vcc
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_3) | instskip(SKIP_1) | instid1(VALU_DEP_4)
+; GFX12-GISEL-NEXT:    v_add_co_u32 v2, vcc, v6, v2
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_vcc(0)
+; GFX12-GISEL-NEXT:    v_add_co_ci_u32_e64 v3, null, v7, v3, vcc
+; GFX12-GISEL-NEXT:    s_clause 0x1
+; GFX12-GISEL-NEXT:    global_load_b128 v[4:7], v[0:1], off
+; GFX12-GISEL-NEXT:    global_load_b128 v[8:11], v[0:1], off offset:16
+; GFX12-GISEL-NEXT:    global_load_b128 v[12:15], v[2:3], off
+; GFX12-GISEL-NEXT:  .LBB12_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x2
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s8, v4
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s9, v5
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s10, v6
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s11, v7
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s12, v8
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s13, v9
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s14, v10
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s15, v11
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x0
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s16, v12
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s17, v13
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s18, v14
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s19, v15
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s8, v4
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s9, v5
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s10, v6
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s11, v7
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s12, v8
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s13, v9
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s14, v10
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s15, v11
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s16, v12
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s17, v13
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s18, v14
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s19, v15
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr4_vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr12_vgpr13_vgpr14_vgpr15
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB12_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    ; return to shader part epilog
+                                                                  i32 %s_idx, i32 %s_idx2) #1 {
+  %rptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx
+  %sptr = getelementptr <4 x i32>, <4 x i32> addrspace(4)* %s_in, i32 %s_idx2
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4)* %rptr, align 16
+  %srsrc = load <4 x i32>, <4 x i32> addrspace(4)* %sptr, align 16
+  %tok = call i32 @llvm.amdgcn.waterfall.begin.v8i32(i32 0, <8 x i32> %rsrc)
+  %tok1 = call i32 @llvm.amdgcn.waterfall.begin.v4i32(i32 %tok, <4 x i32> %srsrc)
+  %s_rsrc = call <8 x i32> @llvm.amdgcn.waterfall.readfirstlane.v8i32.v8i32(i32 %tok1, <8 x i32> %rsrc)
+  %s_srsrc = call <4 x i32> @llvm.amdgcn.waterfall.readfirstlane.v4i32.v4i32(i32 %tok1, <4 x i32> %srsrc)
+  %r = call <4 x float> @llvm.amdgcn.image.sample.2d.v4f32.f32(i32 15, float 0.000000e+00, float 0.000000e+00, <8 x i32> %s_rsrc, <4 x i32> %s_srsrc, i1 false, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok1, <4 x float> %r)
+
+  ret <4 x float> %r1
+}
+
+define amdgpu_ps <4 x float> @test_waterfall_multi_begin_uniform_idx_1(<8 x i32> addrspace(4)* inreg %in,
+; VI-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_1:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; VI-SDAG-NEXT:    s_wqm_b64 exec, exec
+; VI-SDAG-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; VI-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 4, v[1:2]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v2, s3
+; VI-SDAG-NEXT:    v_add_u32_e32 v0, vcc, s2, v0
+; VI-SDAG-NEXT:    v_addc_u32_e32 v1, vcc, v2, v1, vcc
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[5:8], v[0:1]
+; VI-SDAG-NEXT:    s_mov_b32 s4, s5
+; VI-SDAG-NEXT:    s_ashr_i32 s5, s5, 31
+; VI-SDAG-NEXT:    s_lshl_b64 s[2:3], s[4:5], 5
+; VI-SDAG-NEXT:    s_add_u32 s0, s0, s2
+; VI-SDAG-NEXT:    s_addc_u32 s1, s1, s3
+; VI-SDAG-NEXT:    s_load_dwordx8 s[0:7], s[0:1], 0x0
+; VI-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; VI-SDAG-NEXT:  .LBB13_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s12, v4
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[12:13], s12, v4
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[16:17], s[12:13]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s12, v5
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s13, v6
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s14, v7
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s15, v8
+; VI-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr4
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8
+; VI-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; VI-SDAG-NEXT:    s_nop 2
+; VI-SDAG-NEXT:    image_sample v[0:3], v[0:1], s[0:7], s[12:15] dmask:0xf
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[16:17]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB13_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[10:11]
+; VI-SDAG-NEXT:    s_and_b64 exec, exec, s[8:9]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    ; return to shader part epilog
+;
+; VI-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_1:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; VI-GISEL-NEXT:    s_wqm_b64 exec, exec
+; VI-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; VI-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 4, v[1:2]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v2, s2
+; VI-GISEL-NEXT:    v_mov_b32_e32 v3, s3
+; VI-GISEL-NEXT:    v_add_u32_e32 v0, vcc, v2, v0
+; VI-GISEL-NEXT:    v_addc_u32_e32 v1, vcc, v3, v1, vcc
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[5:8], v[0:1]
+; VI-GISEL-NEXT:    s_mov_b32 s8, s5
+; VI-GISEL-NEXT:    s_ashr_i32 s9, s5, 31
+; VI-GISEL-NEXT:    s_lshl_b64 s[2:3], s[8:9], 5
+; VI-GISEL-NEXT:    s_add_u32 s0, s0, s2
+; VI-GISEL-NEXT:    s_addc_u32 s1, s1, s3
+; VI-GISEL-NEXT:    s_load_dwordx8 s[8:15], s[0:1], 0x0
+; VI-GISEL-NEXT:    s_mov_b32 s0, 0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v9, s4
+; VI-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; VI-GISEL-NEXT:  .LBB13_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s1, v9
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s16, v4
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[4:5], s1, v9
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[16:17], s16, v4
+; VI-GISEL-NEXT:    s_and_b64 s[4:5], s[4:5], s[16:17]
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[4:5], s[4:5]
+; VI-GISEL-NEXT:    s_mov_b32 s1, s0
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    v_mov_b32_e32 v0, s0
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s16, v5
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s17, v6
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s18, v7
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s19, v8
+; VI-GISEL-NEXT:    v_mov_b32_e32 v1, s1
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr9
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr4
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8
+; VI-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; VI-GISEL-NEXT:    s_nop 2
+; VI-GISEL-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[16:19] dmask:0xf
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[4:5]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB13_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; VI-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX9-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_1:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX9-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX9-SDAG-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX9-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 4, v[1:2]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v2, s3
+; GFX9-SDAG-NEXT:    v_add_co_u32_e32 v0, vcc, s2, v0
+; GFX9-SDAG-NEXT:    v_addc_co_u32_e32 v1, vcc, v2, v1, vcc
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX9-SDAG-NEXT:    s_mov_b32 s4, s5
+; GFX9-SDAG-NEXT:    s_ashr_i32 s5, s5, 31
+; GFX9-SDAG-NEXT:    s_lshl_b64 s[2:3], s[4:5], 5
+; GFX9-SDAG-NEXT:    s_add_u32 s10, s0, s2
+; GFX9-SDAG-NEXT:    s_addc_u32 s11, s1, s3
+; GFX9-SDAG-NEXT:    s_load_dwordx8 s[0:7], s[10:11], 0x0
+; GFX9-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; GFX9-SDAG-NEXT:  .LBB13_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s12, v4
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[12:13], s12, v4
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[16:17], s[12:13]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s13, v6
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s14, v7
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s15, v8
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8
+; GFX9-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX9-SDAG-NEXT:    s_nop 2
+; GFX9-SDAG-NEXT:    image_sample v[0:3], v[0:1], s[0:7], s[12:15] dmask:0xf
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[16:17]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB13_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[10:11]
+; GFX9-SDAG-NEXT:    s_and_b64 exec, exec, s[8:9]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX9-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_1:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX9-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX9-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX9-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 4, v[1:2]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v2, s2
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v3, s3
+; GFX9-GISEL-NEXT:    v_add_co_u32_e32 v0, vcc, v2, v0
+; GFX9-GISEL-NEXT:    v_addc_co_u32_e32 v1, vcc, v3, v1, vcc
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX9-GISEL-NEXT:    s_mov_b32 s8, s5
+; GFX9-GISEL-NEXT:    s_ashr_i32 s9, s5, 31
+; GFX9-GISEL-NEXT:    s_lshl_b64 s[2:3], s[8:9], 5
+; GFX9-GISEL-NEXT:    s_add_u32 s0, s0, s2
+; GFX9-GISEL-NEXT:    s_addc_u32 s1, s1, s3
+; GFX9-GISEL-NEXT:    s_load_dwordx8 s[8:15], s[0:1], 0x0
+; GFX9-GISEL-NEXT:    s_mov_b32 s0, 0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v9, s4
+; GFX9-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX9-GISEL-NEXT:  .LBB13_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s1, v9
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s16, v4
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[4:5], s1, v9
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[16:17], s16, v4
+; GFX9-GISEL-NEXT:    s_and_b64 s[4:5], s[4:5], s[16:17]
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[4:5], s[4:5]
+; GFX9-GISEL-NEXT:    s_mov_b32 s1, s0
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v0, s0
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s16, v5
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s17, v6
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s18, v7
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s19, v8
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v1, s1
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8
+; GFX9-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX9-GISEL-NEXT:    s_nop 2
+; GFX9-GISEL-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[16:19] dmask:0xf
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[4:5]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB13_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX9-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_1:
+; GFX10-32-SDAG:       ; %bb.0:
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s8, exec_lo
+; GFX10-32-SDAG-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-SDAG-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s4, s5
+; GFX10-32-SDAG-NEXT:    s_ashr_i32 s5, s5, 31
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s9, exec_lo
+; GFX10-32-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 4, v[1:2]
+; GFX10-32-SDAG-NEXT:    v_add_co_u32 v0, vcc_lo, s2, v0
+; GFX10-32-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, s3, v1, vcc_lo
+; GFX10-32-SDAG-NEXT:    s_lshl_b64 s[2:3], s[4:5], 5
+; GFX10-32-SDAG-NEXT:    s_add_u32 s10, s0, s2
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX10-32-SDAG-NEXT:    s_addc_u32 s11, s1, s3
+; GFX10-32-SDAG-NEXT:    s_load_dwordx8 s[0:7], s[10:11], 0x0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s10, exec_lo
+; GFX10-32-SDAG-NEXT:  .LBB13_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s11, v4
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s11, v4
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s13, v6
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s14, v7
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s15, v8
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-32-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[0:7], s[12:15] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s10, s10
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB13_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s9
+; GFX10-32-SDAG-NEXT:    s_and_b32 exec_lo, exec_lo, s8
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_1:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s6, exec_lo
+; GFX10-32-GISEL-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s8, s5
+; GFX10-32-GISEL-NEXT:    s_ashr_i32 s9, s5, 31
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v9, s4
+; GFX10-32-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 4, v[1:2]
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v2, s2
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v3, s3
+; GFX10-32-GISEL-NEXT:    s_lshl_b64 s[2:3], s[8:9], 5
+; GFX10-32-GISEL-NEXT:    s_add_u32 s0, s0, s2
+; GFX10-32-GISEL-NEXT:    v_add_co_u32 v0, vcc_lo, v2, v0
+; GFX10-32-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, v3, v1, vcc_lo
+; GFX10-32-GISEL-NEXT:    s_addc_u32 s1, s1, s3
+; GFX10-32-GISEL-NEXT:    s_load_dwordx8 s[8:15], s[0:1], 0x0
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-GISEL-NEXT:  .LBB13_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s2, v9
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s3, v4
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s2, v9
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s3, v4
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s16, v5
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s17, v6
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s18, v7
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s19, v8
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-32-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB13_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-GISEL-NEXT:    s_and_b32 exec_lo, exec_lo, s6
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_1:
+; GFX10-64-SDAG:       ; %bb.0:
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX10-64-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-SDAG-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s4, s5
+; GFX10-64-SDAG-NEXT:    s_ashr_i32 s5, s5, 31
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[12:13], exec
+; GFX10-64-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 4, v[1:2]
+; GFX10-64-SDAG-NEXT:    v_add_co_u32 v0, vcc, s2, v0
+; GFX10-64-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc, s3, v1, vcc
+; GFX10-64-SDAG-NEXT:    s_lshl_b64 s[2:3], s[4:5], 5
+; GFX10-64-SDAG-NEXT:    s_add_u32 s10, s0, s2
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX10-64-SDAG-NEXT:    s_addc_u32 s11, s1, s3
+; GFX10-64-SDAG-NEXT:    s_load_dwordx8 s[0:7], s[10:11], 0x0
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; GFX10-64-SDAG-NEXT:  .LBB13_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s14, v4
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s14, v4
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s16, v5
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s17, v6
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s18, v7
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s19, v8
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-64-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[0:7], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[12:13], s[12:13]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB13_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[10:11]
+; GFX10-64-SDAG-NEXT:    s_and_b64 exec, exec, s[8:9]
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_1:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s8, s5
+; GFX10-64-GISEL-NEXT:    s_ashr_i32 s9, s5, 31
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v9, s4
+; GFX10-64-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 4, v[1:2]
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v2, s2
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v3, s3
+; GFX10-64-GISEL-NEXT:    s_lshl_b64 s[2:3], s[8:9], 5
+; GFX10-64-GISEL-NEXT:    s_add_u32 s0, s0, s2
+; GFX10-64-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX10-64-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc, v3, v1, vcc
+; GFX10-64-GISEL-NEXT:    s_addc_u32 s1, s1, s3
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-GISEL-NEXT:    s_load_dwordx8 s[8:15], s[0:1], 0x0
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-GISEL-NEXT:  .LBB13_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s4, v9
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s5, v4
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v9
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s5, v4
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s16, v5
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s17, v6
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s18, v7
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s19, v8
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-64-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB13_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX1150-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_1:
+; GFX1150-SDAG:       ; %bb.0:
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX1150-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-SDAG-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX1150-SDAG-NEXT:    s_mov_b32 s4, s5
+; GFX1150-SDAG-NEXT:    s_ashr_i32 s5, s5, 31
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[12:13], exec
+; GFX1150-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 4, v[1:2]
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_add_co_u32 v0, vcc, s2, v0
+; GFX1150-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s3, v1, vcc
+; GFX1150-SDAG-NEXT:    s_lshl_b64 s[2:3], s[4:5], 5
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_add_u32 s0, s0, s2
+; GFX1150-SDAG-NEXT:    global_load_b128 v[5:8], v[0:1], off
+; GFX1150-SDAG-NEXT:    s_addc_u32 s1, s1, s3
+; GFX1150-SDAG-NEXT:    s_load_b256 s[0:7], s[0:1], 0x0
+; GFX1150-SDAG-NEXT:  .LBB13_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s14, v4
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s14, v4
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s16, v5
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s17, v6
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s18, v7
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s19, v8
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX1150-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[0:7], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[12:13], s[12:13]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB13_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[10:11]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_and_b64 exec, exec, s[8:9]
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX1150-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_1:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX1150-GISEL-NEXT:    s_mov_b32 s8, s5
+; GFX1150-GISEL-NEXT:    s_ashr_i32 s9, s5, 31
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v9, s4
+; GFX1150-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 4, v[1:2]
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v2, s2
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v3, s3
+; GFX1150-GISEL-NEXT:    s_lshl_b64 s[2:3], s[8:9], 5
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-GISEL-NEXT:    s_add_u32 s0, s0, s2
+; GFX1150-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX1150-GISEL-NEXT:    s_addc_u32 s1, s1, s3
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-GISEL-NEXT:    s_load_b256 s[8:15], s[0:1], 0x0
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v3, v1, vcc
+; GFX1150-GISEL-NEXT:    global_load_b128 v[5:8], v[0:1], off
+; GFX1150-GISEL-NEXT:  .LBB13_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s4, v9
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s5, v4
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v9
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s5, v4
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s16, v5
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s17, v6
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s18, v7
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s19, v8
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX1150-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB13_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX12-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_1:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX12-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX12-SDAG-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-SDAG-NEXT:    s_mov_b32 s4, s5
+; GFX12-SDAG-NEXT:    s_ashr_i32 s5, s5, 31
+; GFX12-SDAG-NEXT:    s_mov_b64 s[12:13], exec
+; GFX12-SDAG-NEXT:    v_lshlrev_b64_e32 v[0:1], 4, v[1:2]
+; GFX12-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_add_co_u32 v0, vcc, s2, v0
+; GFX12-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s3, v1, vcc
+; GFX12-SDAG-NEXT:    s_lshl_b64 s[2:3], s[4:5], 5
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    s_add_nc_u64 s[0:1], s[0:1], s[2:3]
+; GFX12-SDAG-NEXT:    global_load_b128 v[5:8], v[0:1], off
+; GFX12-SDAG-NEXT:    s_load_b256 s[0:7], s[0:1], 0x0
+; GFX12-SDAG-NEXT:  .LBB13_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s14, v4
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s14, v4
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x0
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s16, v5
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s17, v6
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s18, v7
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s19, v8
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-SDAG-NEXT:    s_wait_kmcnt 0x0
+; GFX12-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[0:7], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[12:13], s[12:13]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB13_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[10:11]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-SDAG-NEXT:    s_and_b64 exec, exec, s[8:9]
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX12-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_1:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX12-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-GISEL-NEXT:    s_mov_b32 s8, s5
+; GFX12-GISEL-NEXT:    s_ashr_i32 s9, s5, 31
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v9, s4
+; GFX12-GISEL-NEXT:    v_lshlrev_b64_e32 v[0:1], 4, v[1:2]
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v2, s2
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v3, s3
+; GFX12-GISEL-NEXT:    s_lshl_b64 s[2:3], s[8:9], 5
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_add_co_u32 s0, s0, s2
+; GFX12-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v3, v1, vcc
+; GFX12-GISEL-NEXT:    s_add_co_ci_u32 s1, s1, s3
+; GFX12-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-GISEL-NEXT:    s_load_b256 s[8:15], s[0:1], 0x0
+; GFX12-GISEL-NEXT:    global_load_b128 v[5:8], v[0:1], off
+; GFX12-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-GISEL-NEXT:  .LBB13_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s4, v9
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s5, v4
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v9
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s5, v4
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x0
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s16, v5
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s17, v6
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s18, v7
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s19, v8
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-GISEL-NEXT:    s_wait_kmcnt 0x0
+; GFX12-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB13_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    ; return to shader part epilog
+       <4 x i32> addrspace(4)* inreg %s_in, i32 inreg %idx1, i32 %idx2, i32 inreg %s_idx, i32 %s_idx2) #1 {
+  %rptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx
+  %sptr = getelementptr <4 x i32>, <4 x i32> addrspace(4)* %s_in, i32 %s_idx2
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4)* %rptr, align 16
+  %srsrc = load <4 x i32>, <4 x i32> addrspace(4)* %sptr, align 16
+  %tok = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %idx1)
+  %tok1 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 %tok, i32 %idx2)
+  %s_rsrc = call <8 x i32> @llvm.amdgcn.waterfall.readfirstlane.v8i32.v8i32(i32 %tok1, <8 x i32> %rsrc)
+  %s_srsrc = call <4 x i32> @llvm.amdgcn.waterfall.readfirstlane.v4i32.v4i32(i32 %tok1, <4 x i32> %srsrc)
+  %r = call <4 x float> @llvm.amdgcn.image.sample.2d.v4f32.f32(i32 15, float 0.000000e+00, float 0.000000e+00, <8 x i32> %s_rsrc, <4 x i32> %s_srsrc, i1 false, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok1, <4 x float> %r)
+
+  ret <4 x float> %r1
+}
+
+define amdgpu_ps <4 x float> @test_waterfall_multi_begin_uniform_idx_2(<8 x i32> addrspace(4)* inreg %in,
+; VI-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_2:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; VI-SDAG-NEXT:    s_mov_b32 s4, s5
+; VI-SDAG-NEXT:    s_wqm_b64 exec, exec
+; VI-SDAG-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; VI-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v2, s1
+; VI-SDAG-NEXT:    v_add_u32_e32 v0, vcc, s0, v0
+; VI-SDAG-NEXT:    v_addc_u32_e32 v1, vcc, v2, v1, vcc
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[5:8], v[0:1]
+; VI-SDAG-NEXT:    v_add_u32_e32 v0, vcc, 16, v0
+; VI-SDAG-NEXT:    v_addc_u32_e32 v1, vcc, 0, v1, vcc
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[9:12], v[0:1]
+; VI-SDAG-NEXT:    s_ashr_i32 s5, s5, 31
+; VI-SDAG-NEXT:    s_lshl_b64 s[0:1], s[4:5], 4
+; VI-SDAG-NEXT:    s_add_u32 s0, s2, s0
+; VI-SDAG-NEXT:    s_addc_u32 s1, s3, s1
+; VI-SDAG-NEXT:    s_load_dwordx4 s[0:3], s[0:1], 0x0
+; VI-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; VI-SDAG-NEXT:  .LBB14_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s8, v4
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[8:9], s8, v4
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[16:17], s[8:9]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s8, v5
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s9, v6
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s10, v7
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s11, v8
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s12, v9
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s13, v10
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s14, v11
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s15, v12
+; VI-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr4
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; VI-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; VI-SDAG-NEXT:    s_nop 2
+; VI-SDAG-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[0:3] dmask:0xf
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[16:17]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB14_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[4:5]
+; VI-SDAG-NEXT:    s_and_b64 exec, exec, s[6:7]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    ; return to shader part epilog
+;
+; VI-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_2:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; VI-GISEL-NEXT:    s_wqm_b64 exec, exec
+; VI-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; VI-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; VI-GISEL-NEXT:    v_add_u32_e32 v0, vcc, v2, v0
+; VI-GISEL-NEXT:    v_addc_u32_e32 v1, vcc, v3, v1, vcc
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[5:8], v[0:1]
+; VI-GISEL-NEXT:    v_add_u32_e32 v0, vcc, 16, v0
+; VI-GISEL-NEXT:    v_addc_u32_e32 v1, vcc, 0, v1, vcc
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[9:12], v[0:1]
+; VI-GISEL-NEXT:    s_mov_b32 s8, s5
+; VI-GISEL-NEXT:    s_ashr_i32 s9, s5, 31
+; VI-GISEL-NEXT:    s_lshl_b64 s[0:1], s[8:9], 4
+; VI-GISEL-NEXT:    s_add_u32 s0, s2, s0
+; VI-GISEL-NEXT:    s_addc_u32 s1, s3, s1
+; VI-GISEL-NEXT:    s_load_dwordx4 s[0:3], s[0:1], 0x0
+; VI-GISEL-NEXT:    s_mov_b32 s8, 0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v13, s4
+; VI-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; VI-GISEL-NEXT:  .LBB14_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s9, v4
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s12, v13
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[10:11], s9, v4
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[12:13], s12, v13
+; VI-GISEL-NEXT:    s_and_b64 s[10:11], s[10:11], s[12:13]
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[10:11], s[10:11]
+; VI-GISEL-NEXT:    s_mov_b32 s9, s8
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    v_mov_b32_e32 v0, s8
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s12, v5
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s13, v6
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s14, v7
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s15, v8
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s16, v9
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s17, v10
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s18, v11
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s19, v12
+; VI-GISEL-NEXT:    v_mov_b32_e32 v1, s9
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr4
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr13
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; VI-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; VI-GISEL-NEXT:    s_nop 2
+; VI-GISEL-NEXT:    image_sample v[0:3], v[0:1], s[12:19], s[0:3] dmask:0xf
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[10:11]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB14_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[4:5]
+; VI-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX9-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_2:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX9-SDAG-NEXT:    s_mov_b32 s4, s5
+; GFX9-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX9-SDAG-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX9-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v2, s1
+; GFX9-SDAG-NEXT:    v_add_co_u32_e32 v0, vcc, s0, v0
+; GFX9-SDAG-NEXT:    v_addc_co_u32_e32 v1, vcc, v2, v1, vcc
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[9:12], v[0:1], off offset:16
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX9-SDAG-NEXT:    s_ashr_i32 s5, s5, 31
+; GFX9-SDAG-NEXT:    s_lshl_b64 s[0:1], s[4:5], 4
+; GFX9-SDAG-NEXT:    s_add_u32 s0, s2, s0
+; GFX9-SDAG-NEXT:    s_addc_u32 s1, s3, s1
+; GFX9-SDAG-NEXT:    s_load_dwordx4 s[0:3], s[0:1], 0x0
+; GFX9-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX9-SDAG-NEXT:  .LBB14_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s8, v4
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[8:9], s8, v4
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[16:17], s[8:9]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX9-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX9-SDAG-NEXT:    s_nop 2
+; GFX9-SDAG-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[0:3] dmask:0xf
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[16:17]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB14_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[4:5]
+; GFX9-SDAG-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX9-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_2:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX9-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX9-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX9-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX9-GISEL-NEXT:    v_add_co_u32_e32 v0, vcc, v2, v0
+; GFX9-GISEL-NEXT:    v_addc_co_u32_e32 v1, vcc, v3, v1, vcc
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[9:12], v[0:1], off offset:16
+; GFX9-GISEL-NEXT:    s_mov_b32 s8, s5
+; GFX9-GISEL-NEXT:    s_ashr_i32 s9, s5, 31
+; GFX9-GISEL-NEXT:    s_lshl_b64 s[0:1], s[8:9], 4
+; GFX9-GISEL-NEXT:    s_add_u32 s0, s2, s0
+; GFX9-GISEL-NEXT:    s_addc_u32 s1, s3, s1
+; GFX9-GISEL-NEXT:    s_load_dwordx4 s[0:3], s[0:1], 0x0
+; GFX9-GISEL-NEXT:    s_mov_b32 s8, 0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v13, s4
+; GFX9-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX9-GISEL-NEXT:  .LBB14_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s9, v4
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s12, v13
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[10:11], s9, v4
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[12:13], s12, v13
+; GFX9-GISEL-NEXT:    s_and_b64 s[10:11], s[10:11], s[12:13]
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[10:11], s[10:11]
+; GFX9-GISEL-NEXT:    s_mov_b32 s9, s8
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v0, s8
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s13, v6
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s14, v7
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s15, v8
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s16, v9
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s17, v10
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s18, v11
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s19, v12
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v1, s9
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr13
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX9-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX9-GISEL-NEXT:    s_nop 2
+; GFX9-GISEL-NEXT:    image_sample v[0:3], v[0:1], s[12:19], s[0:3] dmask:0xf
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[10:11]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB14_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[4:5]
+; GFX9-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_2:
+; GFX10-32-SDAG:       ; %bb.0:
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s6, s5
+; GFX10-32-SDAG-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-SDAG-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-32-SDAG-NEXT:    s_ashr_i32 s7, s5, 31
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s5, exec_lo
+; GFX10-32-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; GFX10-32-SDAG-NEXT:    v_add_co_u32 v0, vcc_lo, s0, v0
+; GFX10-32-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, s1, v1, vcc_lo
+; GFX10-32-SDAG-NEXT:    s_lshl_b64 s[0:1], s[6:7], 4
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s6, exec_lo
+; GFX10-32-SDAG-NEXT:    s_add_u32 s0, s2, s0
+; GFX10-32-SDAG-NEXT:    s_clause 0x1
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[9:12], v[0:1], off offset:16
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX10-32-SDAG-NEXT:    s_addc_u32 s1, s3, s1
+; GFX10-32-SDAG-NEXT:    s_load_dwordx4 s[0:3], s[0:1], 0x0
+; GFX10-32-SDAG-NEXT:  .LBB14_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s7, v4
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s7, v4
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-32-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[0:3] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s6, s6
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB14_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s5
+; GFX10-32-SDAG-NEXT:    s_and_b32 exec_lo, exec_lo, s4
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_2:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s6, exec_lo
+; GFX10-32-GISEL-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s8, s5
+; GFX10-32-GISEL-NEXT:    s_ashr_i32 s9, s5, 31
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v13, s4
+; GFX10-32-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX10-32-GISEL-NEXT:    s_lshl_b64 s[0:1], s[8:9], 4
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s5, exec_lo
+; GFX10-32-GISEL-NEXT:    s_add_u32 s0, s2, s0
+; GFX10-32-GISEL-NEXT:    s_addc_u32 s1, s3, s1
+; GFX10-32-GISEL-NEXT:    v_add_co_u32 v0, vcc_lo, v2, v0
+; GFX10-32-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, v3, v1, vcc_lo
+; GFX10-32-GISEL-NEXT:    s_load_dwordx4 s[0:3], s[0:1], 0x0
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-GISEL-NEXT:    s_clause 0x1
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[9:12], v[0:1], off offset:16
+; GFX10-32-GISEL-NEXT:  .LBB14_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s7, v4
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s8, v13
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s7, v4
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s8, v13
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-32-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[0:3] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s5, s5
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr13
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB14_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s4
+; GFX10-32-GISEL-NEXT:    s_and_b32 exec_lo, exec_lo, s6
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_2:
+; GFX10-64-SDAG:       ; %bb.0:
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s4, s5
+; GFX10-64-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-SDAG-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-64-SDAG-NEXT:    s_ashr_i32 s5, s5, 31
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX10-64-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; GFX10-64-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX10-64-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc, s1, v1, vcc
+; GFX10-64-SDAG-NEXT:    s_lshl_b64 s[0:1], s[4:5], 4
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-SDAG-NEXT:    s_add_u32 s0, s2, s0
+; GFX10-64-SDAG-NEXT:    s_clause 0x1
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[9:12], v[0:1], off offset:16
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX10-64-SDAG-NEXT:    s_addc_u32 s1, s3, s1
+; GFX10-64-SDAG-NEXT:    s_load_dwordx4 s[0:3], s[0:1], 0x0
+; GFX10-64-SDAG-NEXT:  .LBB14_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s10, v4
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s10, v4
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s13, v6
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s14, v7
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s15, v8
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s16, v9
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s17, v10
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s18, v11
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s19, v12
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-64-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[12:19], s[0:3] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[8:9], s[8:9]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB14_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[4:5]
+; GFX10-64-SDAG-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_2:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s8, s5
+; GFX10-64-GISEL-NEXT:    s_ashr_i32 s9, s5, 31
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v13, s4
+; GFX10-64-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX10-64-GISEL-NEXT:    s_lshl_b64 s[0:1], s[8:9], 4
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX10-64-GISEL-NEXT:    s_add_u32 s0, s2, s0
+; GFX10-64-GISEL-NEXT:    s_addc_u32 s1, s3, s1
+; GFX10-64-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX10-64-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc, v3, v1, vcc
+; GFX10-64-GISEL-NEXT:    s_load_dwordx4 s[0:3], s[0:1], 0x0
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-GISEL-NEXT:    s_clause 0x1
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[9:12], v[0:1], off offset:16
+; GFX10-64-GISEL-NEXT:  .LBB14_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s10, v4
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s11, v13
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s10, v4
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s11, v13
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s13, v6
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s14, v7
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s15, v8
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s16, v9
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s17, v10
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s18, v11
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s19, v12
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-64-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[12:19], s[0:3] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[8:9], s[8:9]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr13
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB14_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[4:5]
+; GFX10-64-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX1150-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_2:
+; GFX1150-SDAG:       ; %bb.0:
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-SDAG-NEXT:    s_mov_b32 s4, s5
+; GFX1150-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-SDAG-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX1150-SDAG-NEXT:    s_ashr_i32 s5, s5, 31
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; GFX1150-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s1, v1, vcc
+; GFX1150-SDAG-NEXT:    s_clause 0x1
+; GFX1150-SDAG-NEXT:    global_load_b128 v[9:12], v[0:1], off offset:16
+; GFX1150-SDAG-NEXT:    global_load_b128 v[5:8], v[0:1], off
+; GFX1150-SDAG-NEXT:    s_lshl_b64 s[0:1], s[4:5], 4
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-SDAG-NEXT:    s_add_u32 s0, s2, s0
+; GFX1150-SDAG-NEXT:    s_addc_u32 s1, s3, s1
+; GFX1150-SDAG-NEXT:    s_load_b128 s[0:3], s[0:1], 0x0
+; GFX1150-SDAG-NEXT:  .LBB14_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s10, v4
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s10, v4
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s13, v6
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s14, v7
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s15, v8
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s16, v9
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s17, v10
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s18, v11
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s19, v12
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX1150-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[12:19], s[0:3] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[8:9], s[8:9]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB14_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[4:5]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX1150-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_2:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX1150-GISEL-NEXT:    s_mov_b32 s8, s5
+; GFX1150-GISEL-NEXT:    s_ashr_i32 s9, s5, 31
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v13, s4
+; GFX1150-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX1150-GISEL-NEXT:    s_lshl_b64 s[0:1], s[8:9], 4
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX1150-GISEL-NEXT:    s_add_u32 s0, s2, s0
+; GFX1150-GISEL-NEXT:    s_addc_u32 s1, s3, s1
+; GFX1150-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX1150-GISEL-NEXT:    s_load_b128 s[0:3], s[0:1], 0x0
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v3, v1, vcc
+; GFX1150-GISEL-NEXT:    s_clause 0x1
+; GFX1150-GISEL-NEXT:    global_load_b128 v[5:8], v[0:1], off
+; GFX1150-GISEL-NEXT:    global_load_b128 v[9:12], v[0:1], off offset:16
+; GFX1150-GISEL-NEXT:  .LBB14_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s10, v4
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s11, v13
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s10, v4
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s11, v13
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s13, v6
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s14, v7
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s15, v8
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s16, v9
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s17, v10
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s18, v11
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s19, v12
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX1150-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[12:19], s[0:3] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[8:9], s[8:9]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr13
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB14_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[4:5]
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX12-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_2:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-SDAG-NEXT:    s_mov_b32 s4, s5
+; GFX12-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX12-SDAG-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-SDAG-NEXT:    s_ashr_i32 s5, s5, 31
+; GFX12-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_lshlrev_b64_e32 v[0:1], 5, v[1:2]
+; GFX12-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s1, v1, vcc
+; GFX12-SDAG-NEXT:    s_lshl_b64 s[0:1], s[4:5], 4
+; GFX12-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    s_add_nc_u64 s[0:1], s[2:3], s[0:1]
+; GFX12-SDAG-NEXT:    s_clause 0x1
+; GFX12-SDAG-NEXT:    global_load_b128 v[9:12], v[0:1], off offset:16
+; GFX12-SDAG-NEXT:    global_load_b128 v[5:8], v[0:1], off
+; GFX12-SDAG-NEXT:    s_load_b128 s[0:3], s[0:1], 0x0
+; GFX12-SDAG-NEXT:  .LBB14_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s10, v4
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s10, v4
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x0
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s13, v6
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s14, v7
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s15, v8
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s16, v9
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s17, v10
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s18, v11
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s19, v12
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-SDAG-NEXT:    s_wait_kmcnt 0x0
+; GFX12-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[12:19], s[0:3] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[8:9], s[8:9]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB14_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[4:5]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-SDAG-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX12-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_2:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX12-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-GISEL-NEXT:    s_mov_b32 s8, s5
+; GFX12-GISEL-NEXT:    s_ashr_i32 s9, s5, 31
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v13, s4
+; GFX12-GISEL-NEXT:    v_lshlrev_b64_e32 v[0:1], 5, v[1:2]
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX12-GISEL-NEXT:    s_lshl_b64 s[0:1], s[8:9], 4
+; GFX12-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_add_co_u32 s0, s2, s0
+; GFX12-GISEL-NEXT:    s_add_co_ci_u32 s1, s3, s1
+; GFX12-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v3, v1, vcc
+; GFX12-GISEL-NEXT:    s_load_b128 s[0:3], s[0:1], 0x0
+; GFX12-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-GISEL-NEXT:    s_clause 0x1
+; GFX12-GISEL-NEXT:    global_load_b128 v[5:8], v[0:1], off
+; GFX12-GISEL-NEXT:    global_load_b128 v[9:12], v[0:1], off offset:16
+; GFX12-GISEL-NEXT:  .LBB14_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s10, v4
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s11, v13
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s10, v4
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s11, v13
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s12, v5
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s13, v6
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s14, v7
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s15, v8
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x0
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s16, v9
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s17, v10
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s18, v11
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s19, v12
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-GISEL-NEXT:    s_wait_kmcnt 0x0
+; GFX12-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[12:19], s[0:3] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[8:9], s[8:9]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr13
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB14_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[4:5]
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    ; return to shader part epilog
+       <4 x i32> addrspace(4)* inreg %s_in, i32 %idx1, i32 inreg %idx2, i32 %s_idx, i32 inreg %s_idx2) #1 {
+  %rptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx
+  %sptr = getelementptr <4 x i32>, <4 x i32> addrspace(4)* %s_in, i32 %s_idx2
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4)* %rptr, align 16
+  %srsrc = load <4 x i32>, <4 x i32> addrspace(4)* %sptr, align 16
+  %tok = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %idx1)
+  %tok1 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 %tok, i32 %idx2)
+  %s_rsrc = call <8 x i32> @llvm.amdgcn.waterfall.readfirstlane.v8i32.v8i32(i32 %tok1, <8 x i32> %rsrc)
+  %s_srsrc = call <4 x i32> @llvm.amdgcn.waterfall.readfirstlane.v4i32.v4i32(i32 %tok1, <4 x i32> %srsrc)
+  %r = call <4 x float> @llvm.amdgcn.image.sample.2d.v4f32.f32(i32 15, float 0.000000e+00, float 0.000000e+00, <8 x i32> %s_rsrc, <4 x i32> %s_srsrc, i1 false, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok1, <4 x float> %r)
+
+  ret <4 x float> %r1
+}
+
+define amdgpu_ps <4 x float> @test_waterfall_multi_begin_uniform_idx_3(<8 x i32> addrspace(4)* inreg %in,
+; VI-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_3:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; VI-SDAG-NEXT:    s_wqm_b64 exec, exec
+; VI-SDAG-NEXT:    v_mov_b32_e32 v5, v1
+; VI-SDAG-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; VI-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[5:6]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v3, s1
+; VI-SDAG-NEXT:    v_add_u32_e32 v0, vcc, s0, v0
+; VI-SDAG-NEXT:    v_addc_u32_e32 v1, vcc, v3, v1, vcc
+; VI-SDAG-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; VI-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[2:3]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v5, s3
+; VI-SDAG-NEXT:    v_add_u32_e32 v2, vcc, s2, v2
+; VI-SDAG-NEXT:    v_addc_u32_e32 v3, vcc, v5, v3, vcc
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[5:8], v[0:1]
+; VI-SDAG-NEXT:    v_add_u32_e32 v0, vcc, 16, v0
+; VI-SDAG-NEXT:    v_addc_u32_e32 v1, vcc, 0, v1, vcc
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[9:12], v[0:1]
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[13:16], v[2:3]
+; VI-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; VI-SDAG-NEXT:  .LBB15_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s2, v4
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v4
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(2)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s8, v5
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s9, v6
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s10, v7
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s11, v8
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s12, v9
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s13, v10
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s14, v11
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s15, v12
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s16, v13
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s17, v14
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s18, v15
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s19, v16
+; VI-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr4
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; VI-SDAG-NEXT:    s_nop 3
+; VI-SDAG-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[16:19] dmask:0xf
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB15_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; VI-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    ; return to shader part epilog
+;
+; VI-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_3:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; VI-GISEL-NEXT:    s_wqm_b64 exec, exec
+; VI-GISEL-NEXT:    v_mov_b32_e32 v5, v2
+; VI-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; VI-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; VI-GISEL-NEXT:    v_add_u32_e32 v0, vcc, v2, v0
+; VI-GISEL-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; VI-GISEL-NEXT:    v_addc_u32_e32 v1, vcc, v3, v1, vcc
+; VI-GISEL-NEXT:    v_lshlrev_b64 v[2:3], 4, v[5:6]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v6, s3
+; VI-GISEL-NEXT:    v_mov_b32_e32 v5, s2
+; VI-GISEL-NEXT:    v_add_u32_e32 v2, vcc, v5, v2
+; VI-GISEL-NEXT:    v_addc_u32_e32 v3, vcc, v6, v3, vcc
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[5:8], v[0:1]
+; VI-GISEL-NEXT:    v_add_u32_e32 v0, vcc, 16, v0
+; VI-GISEL-NEXT:    v_addc_u32_e32 v1, vcc, 0, v1, vcc
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[9:12], v[0:1]
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[13:16], v[2:3]
+; VI-GISEL-NEXT:    s_mov_b32 s0, 0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v17, s4
+; VI-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; VI-GISEL-NEXT:  .LBB15_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s1, v17
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s8, v4
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[4:5], s1, v17
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[8:9], s8, v4
+; VI-GISEL-NEXT:    s_and_b64 s[4:5], s[4:5], s[8:9]
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[4:5], s[4:5]
+; VI-GISEL-NEXT:    s_mov_b32 s1, s0
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    v_mov_b32_e32 v0, s0
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s8, v5
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s9, v6
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s10, v7
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s11, v8
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s12, v9
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s13, v10
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s14, v11
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s15, v12
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s16, v13
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s17, v14
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s18, v15
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s19, v16
+; VI-GISEL-NEXT:    v_mov_b32_e32 v1, s1
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr17
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr4
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; VI-GISEL-NEXT:    s_nop 3
+; VI-GISEL-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[16:19] dmask:0xf
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[4:5]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB15_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; VI-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX9-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_3:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX9-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v5, v1
+; GFX9-SDAG-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX9-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[5:6]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v3, s1
+; GFX9-SDAG-NEXT:    v_add_co_u32_e32 v0, vcc, s0, v0
+; GFX9-SDAG-NEXT:    v_addc_co_u32_e32 v1, vcc, v3, v1, vcc
+; GFX9-SDAG-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX9-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[2:3]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v5, s3
+; GFX9-SDAG-NEXT:    v_add_co_u32_e32 v2, vcc, s2, v2
+; GFX9-SDAG-NEXT:    v_addc_co_u32_e32 v3, vcc, v5, v3, vcc
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[9:12], v[0:1], off offset:16
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[13:16], v[2:3], off
+; GFX9-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX9-SDAG-NEXT:  .LBB15_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s2, v4
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v4
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX9-SDAG-NEXT:    s_nop 3
+; GFX9-SDAG-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[16:19] dmask:0xf
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB15_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX9-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX9-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_3:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX9-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v5, v2
+; GFX9-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX9-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX9-GISEL-NEXT:    v_add_co_u32_e32 v0, vcc, v2, v0
+; GFX9-GISEL-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX9-GISEL-NEXT:    v_addc_co_u32_e32 v1, vcc, v3, v1, vcc
+; GFX9-GISEL-NEXT:    v_lshlrev_b64 v[2:3], 4, v[5:6]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v6, s3
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v5, s2
+; GFX9-GISEL-NEXT:    v_add_co_u32_e32 v2, vcc, v5, v2
+; GFX9-GISEL-NEXT:    v_addc_co_u32_e32 v3, vcc, v6, v3, vcc
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[9:12], v[0:1], off offset:16
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[13:16], v[2:3], off
+; GFX9-GISEL-NEXT:    s_mov_b32 s0, 0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v17, s4
+; GFX9-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX9-GISEL-NEXT:  .LBB15_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s1, v17
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s8, v4
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[4:5], s1, v17
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[8:9], s8, v4
+; GFX9-GISEL-NEXT:    s_and_b64 s[4:5], s[4:5], s[8:9]
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[4:5], s[4:5]
+; GFX9-GISEL-NEXT:    s_mov_b32 s1, s0
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v0, s0
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v1, s1
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr17
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX9-GISEL-NEXT:    s_nop 3
+; GFX9-GISEL-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[16:19] dmask:0xf
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[4:5]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB15_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX9-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_3:
+; GFX10-32-SDAG:       ; %bb.0:
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-SDAG-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v5, v1
+; GFX10-32-SDAG-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-32-SDAG-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX10-32-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[2:3]
+; GFX10-32-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[5:6]
+; GFX10-32-SDAG-NEXT:    v_add_co_u32 v0, vcc_lo, s0, v0
+; GFX10-32-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, s1, v1, vcc_lo
+; GFX10-32-SDAG-NEXT:    v_add_co_u32 v2, vcc_lo, s2, v2
+; GFX10-32-SDAG-NEXT:    v_add_co_ci_u32_e32 v3, vcc_lo, s3, v3, vcc_lo
+; GFX10-32-SDAG-NEXT:    s_clause 0x1
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[9:12], v[0:1], off offset:16
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[13:16], v[2:3], off
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-SDAG-NEXT:  .LBB15_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s2, v4
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s2, v4
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB15_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-SDAG-NEXT:    s_and_b32 exec_lo, exec_lo, s4
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_3:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s5, exec_lo
+; GFX10-32-GISEL-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v5, v2
+; GFX10-32-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v8, s3
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v7, s2
+; GFX10-32-GISEL-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX10-32-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v17, s4
+; GFX10-32-GISEL-NEXT:    v_lshlrev_b64 v[5:6], 4, v[5:6]
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-GISEL-NEXT:    v_add_co_u32 v0, vcc_lo, v2, v0
+; GFX10-32-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, v3, v1, vcc_lo
+; GFX10-32-GISEL-NEXT:    v_add_co_u32 v2, vcc_lo, v7, v5
+; GFX10-32-GISEL-NEXT:    v_add_co_ci_u32_e32 v3, vcc_lo, v8, v6, vcc_lo
+; GFX10-32-GISEL-NEXT:    s_clause 0x1
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[9:12], v[0:1], off offset:16
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[13:16], v[2:3], off
+; GFX10-32-GISEL-NEXT:  .LBB15_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s2, v17
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s3, v4
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s2, v17
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s3, v4
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr17
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB15_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-GISEL-NEXT:    s_and_b32 exec_lo, exec_lo, s5
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_3:
+; GFX10-64-SDAG:       ; %bb.0:
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v5, v1
+; GFX10-64-SDAG-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-64-SDAG-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX10-64-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[2:3]
+; GFX10-64-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[5:6]
+; GFX10-64-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX10-64-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc, s1, v1, vcc
+; GFX10-64-SDAG-NEXT:    v_add_co_u32 v2, vcc, s2, v2
+; GFX10-64-SDAG-NEXT:    v_add_co_ci_u32_e32 v3, vcc, s3, v3, vcc
+; GFX10-64-SDAG-NEXT:    s_clause 0x1
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[9:12], v[0:1], off offset:16
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[13:16], v[2:3], off
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-SDAG-NEXT:  .LBB15_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s6, v4
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v4
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB15_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_3:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v5, v2
+; GFX10-64-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v8, s3
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v7, s2
+; GFX10-64-GISEL-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX10-64-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v17, s4
+; GFX10-64-GISEL-NEXT:    v_lshlrev_b64 v[5:6], 4, v[5:6]
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX10-64-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc, v3, v1, vcc
+; GFX10-64-GISEL-NEXT:    v_add_co_u32 v2, vcc, v7, v5
+; GFX10-64-GISEL-NEXT:    v_add_co_ci_u32_e32 v3, vcc, v8, v6, vcc
+; GFX10-64-GISEL-NEXT:    s_clause 0x1
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[9:12], v[0:1], off offset:16
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[13:16], v[2:3], off
+; GFX10-64-GISEL-NEXT:  .LBB15_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s4, v17
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s5, v4
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v17
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s5, v4
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr17
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB15_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX1150-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_3:
+; GFX1150-SDAG:       ; %bb.0:
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v5, v1
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX1150-SDAG-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_3) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-SDAG-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX1150-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[2:3]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[5:6]
+; GFX1150-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_4)
+; GFX1150-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s1, v1, vcc
+; GFX1150-SDAG-NEXT:    v_add_co_u32 v2, vcc, s2, v2
+; GFX1150-SDAG-NEXT:    s_waitcnt_depctr depctr_va_vcc(0)
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-SDAG-NEXT:    v_add_co_ci_u32_e64 v3, null, s3, v3, vcc
+; GFX1150-SDAG-NEXT:    s_clause 0x1
+; GFX1150-SDAG-NEXT:    global_load_b128 v[9:12], v[0:1], off offset:16
+; GFX1150-SDAG-NEXT:    global_load_b128 v[5:8], v[0:1], off
+; GFX1150-SDAG-NEXT:    global_load_b128 v[13:16], v[2:3], off
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-SDAG-NEXT:  .LBB15_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s6, v4
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v4
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB15_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX1150-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_3:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v5, v2
+; GFX1150-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v8, s3
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v7, s2
+; GFX1150-GISEL-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX1150-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v17, s4
+; GFX1150-GISEL-NEXT:    v_lshlrev_b64 v[5:6], 4, v[5:6]
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_3)
+; GFX1150-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v3, v1, vcc
+; GFX1150-GISEL-NEXT:    v_add_co_u32 v2, vcc, v7, v5
+; GFX1150-GISEL-NEXT:    s_waitcnt_depctr depctr_va_vcc(0)
+; GFX1150-GISEL-NEXT:    v_add_co_ci_u32_e64 v3, null, v8, v6, vcc
+; GFX1150-GISEL-NEXT:    s_clause 0x1
+; GFX1150-GISEL-NEXT:    global_load_b128 v[5:8], v[0:1], off
+; GFX1150-GISEL-NEXT:    global_load_b128 v[9:12], v[0:1], off offset:16
+; GFX1150-GISEL-NEXT:    global_load_b128 v[13:16], v[2:3], off
+; GFX1150-GISEL-NEXT:  .LBB15_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s4, v17
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s5, v4
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v17
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s5, v4
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr17
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB15_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX12-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_3:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v5, v1
+; GFX12-SDAG-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_3) | instskip(NEXT) | instid1(VALU_DEP_3)
+; GFX12-SDAG-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX12-SDAG-NEXT:    v_lshlrev_b64_e32 v[2:3], 4, v[2:3]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_lshlrev_b64_e32 v[0:1], 5, v[5:6]
+; GFX12-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_4)
+; GFX12-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s1, v1, vcc
+; GFX12-SDAG-NEXT:    v_add_co_u32 v2, vcc, s2, v2
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_vcc(0)
+; GFX12-SDAG-NEXT:    v_add_co_ci_u32_e64 v3, null, s3, v3, vcc
+; GFX12-SDAG-NEXT:    s_clause 0x1
+; GFX12-SDAG-NEXT:    global_load_b128 v[9:12], v[0:1], off offset:16
+; GFX12-SDAG-NEXT:    global_load_b128 v[5:8], v[0:1], off
+; GFX12-SDAG-NEXT:    global_load_b128 v[13:16], v[2:3], off
+; GFX12-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-SDAG-NEXT:  .LBB15_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s6, v4
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v4
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x0
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB15_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX12-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_3:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v5, v2
+; GFX12-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v8, s3
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v7, s2
+; GFX12-GISEL-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX12-GISEL-NEXT:    v_lshlrev_b64_e32 v[0:1], 5, v[1:2]
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v17, s4
+; GFX12-GISEL-NEXT:    v_lshlrev_b64_e32 v[5:6], 4, v[5:6]
+; GFX12-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_3)
+; GFX12-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v3, v1, vcc
+; GFX12-GISEL-NEXT:    v_add_co_u32 v2, vcc, v7, v5
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_vcc(0)
+; GFX12-GISEL-NEXT:    v_add_co_ci_u32_e64 v3, null, v8, v6, vcc
+; GFX12-GISEL-NEXT:    s_clause 0x1
+; GFX12-GISEL-NEXT:    global_load_b128 v[5:8], v[0:1], off
+; GFX12-GISEL-NEXT:    global_load_b128 v[9:12], v[0:1], off offset:16
+; GFX12-GISEL-NEXT:    global_load_b128 v[13:16], v[2:3], off
+; GFX12-GISEL-NEXT:  .LBB15_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s4, v17
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s5, v4
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v17
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s5, v4
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x2
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x0
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr17
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB15_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    ; return to shader part epilog
+       <4 x i32> addrspace(4)* inreg %s_in, i32 inreg %idx1, i32 %idx2, i32 %s_idx, i32 %s_idx2) #1 {
+  %rptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx
+  %sptr = getelementptr <4 x i32>, <4 x i32> addrspace(4)* %s_in, i32 %s_idx2
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4)* %rptr, align 16
+  %srsrc = load <4 x i32>, <4 x i32> addrspace(4)* %sptr, align 16
+  %tok = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %idx1)
+  %tok1 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 %tok, i32 %idx2)
+  %s_rsrc = call <8 x i32> @llvm.amdgcn.waterfall.readfirstlane.v8i32.v8i32(i32 %tok1, <8 x i32> %rsrc)
+  %s_srsrc = call <4 x i32> @llvm.amdgcn.waterfall.readfirstlane.v4i32.v4i32(i32 %tok1, <4 x i32> %srsrc)
+  %r = call <4 x float> @llvm.amdgcn.image.sample.2d.v4f32.f32(i32 15, float 0.000000e+00, float 0.000000e+00, <8 x i32> %s_rsrc, <4 x i32> %s_srsrc, i1 false, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok1, <4 x float> %r)
+
+  ret <4 x float> %r1
+}
+
+define amdgpu_ps <4 x float> @test_waterfall_multi_begin_uniform_idx_4(<8 x i32> addrspace(4)* inreg %in,
+; VI-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_4:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; VI-SDAG-NEXT:    s_wqm_b64 exec, exec
+; VI-SDAG-NEXT:    v_mov_b32_e32 v5, v1
+; VI-SDAG-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; VI-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[5:6]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v3, s1
+; VI-SDAG-NEXT:    v_add_u32_e32 v0, vcc, s0, v0
+; VI-SDAG-NEXT:    v_addc_u32_e32 v1, vcc, v3, v1, vcc
+; VI-SDAG-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; VI-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[2:3]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v5, s3
+; VI-SDAG-NEXT:    v_add_u32_e32 v2, vcc, s2, v2
+; VI-SDAG-NEXT:    v_addc_u32_e32 v3, vcc, v5, v3, vcc
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[5:8], v[0:1]
+; VI-SDAG-NEXT:    v_add_u32_e32 v0, vcc, 16, v0
+; VI-SDAG-NEXT:    v_addc_u32_e32 v1, vcc, 0, v1, vcc
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[9:12], v[0:1]
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[13:16], v[2:3]
+; VI-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; VI-SDAG-NEXT:  .LBB16_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s2, v4
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v4
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(2)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s8, v5
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s9, v6
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s10, v7
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s11, v8
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s12, v9
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s13, v10
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s14, v11
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s15, v12
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s16, v13
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s17, v14
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s18, v15
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s19, v16
+; VI-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr4
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; VI-SDAG-NEXT:    s_nop 3
+; VI-SDAG-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[16:19] dmask:0xf
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB16_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; VI-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    ; return to shader part epilog
+;
+; VI-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_4:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; VI-GISEL-NEXT:    s_wqm_b64 exec, exec
+; VI-GISEL-NEXT:    v_mov_b32_e32 v5, v2
+; VI-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; VI-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; VI-GISEL-NEXT:    v_add_u32_e32 v0, vcc, v2, v0
+; VI-GISEL-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; VI-GISEL-NEXT:    v_addc_u32_e32 v1, vcc, v3, v1, vcc
+; VI-GISEL-NEXT:    v_lshlrev_b64 v[2:3], 4, v[5:6]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v6, s3
+; VI-GISEL-NEXT:    v_mov_b32_e32 v5, s2
+; VI-GISEL-NEXT:    v_add_u32_e32 v2, vcc, v5, v2
+; VI-GISEL-NEXT:    v_addc_u32_e32 v3, vcc, v6, v3, vcc
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[5:8], v[0:1]
+; VI-GISEL-NEXT:    v_add_u32_e32 v0, vcc, 16, v0
+; VI-GISEL-NEXT:    v_addc_u32_e32 v1, vcc, 0, v1, vcc
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[9:12], v[0:1]
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[13:16], v[2:3]
+; VI-GISEL-NEXT:    s_mov_b32 s0, 0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v17, s4
+; VI-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; VI-GISEL-NEXT:  .LBB16_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s1, v4
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s8, v17
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[4:5], s1, v4
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[8:9], s8, v17
+; VI-GISEL-NEXT:    s_and_b64 s[4:5], s[4:5], s[8:9]
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[4:5], s[4:5]
+; VI-GISEL-NEXT:    s_mov_b32 s1, s0
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    v_mov_b32_e32 v0, s0
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s8, v5
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s9, v6
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s10, v7
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s11, v8
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s12, v9
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s13, v10
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s14, v11
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s15, v12
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s16, v13
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s17, v14
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s18, v15
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s19, v16
+; VI-GISEL-NEXT:    v_mov_b32_e32 v1, s1
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr4
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr17
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; VI-GISEL-NEXT:    s_nop 3
+; VI-GISEL-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[16:19] dmask:0xf
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[4:5]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB16_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; VI-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX9-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_4:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX9-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v5, v1
+; GFX9-SDAG-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX9-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[5:6]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v3, s1
+; GFX9-SDAG-NEXT:    v_add_co_u32_e32 v0, vcc, s0, v0
+; GFX9-SDAG-NEXT:    v_addc_co_u32_e32 v1, vcc, v3, v1, vcc
+; GFX9-SDAG-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX9-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[2:3]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v5, s3
+; GFX9-SDAG-NEXT:    v_add_co_u32_e32 v2, vcc, s2, v2
+; GFX9-SDAG-NEXT:    v_addc_co_u32_e32 v3, vcc, v5, v3, vcc
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[9:12], v[0:1], off offset:16
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[13:16], v[2:3], off
+; GFX9-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX9-SDAG-NEXT:  .LBB16_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s2, v4
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v4
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX9-SDAG-NEXT:    s_nop 3
+; GFX9-SDAG-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[16:19] dmask:0xf
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB16_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX9-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX9-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_4:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX9-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v5, v2
+; GFX9-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX9-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX9-GISEL-NEXT:    v_add_co_u32_e32 v0, vcc, v2, v0
+; GFX9-GISEL-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX9-GISEL-NEXT:    v_addc_co_u32_e32 v1, vcc, v3, v1, vcc
+; GFX9-GISEL-NEXT:    v_lshlrev_b64 v[2:3], 4, v[5:6]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v6, s3
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v5, s2
+; GFX9-GISEL-NEXT:    v_add_co_u32_e32 v2, vcc, v5, v2
+; GFX9-GISEL-NEXT:    v_addc_co_u32_e32 v3, vcc, v6, v3, vcc
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[9:12], v[0:1], off offset:16
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[13:16], v[2:3], off
+; GFX9-GISEL-NEXT:    s_mov_b32 s0, 0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v17, s4
+; GFX9-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX9-GISEL-NEXT:  .LBB16_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s1, v4
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s8, v17
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[4:5], s1, v4
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[8:9], s8, v17
+; GFX9-GISEL-NEXT:    s_and_b64 s[4:5], s[4:5], s[8:9]
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[4:5], s[4:5]
+; GFX9-GISEL-NEXT:    s_mov_b32 s1, s0
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v0, s0
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v1, s1
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr17
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX9-GISEL-NEXT:    s_nop 3
+; GFX9-GISEL-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[16:19] dmask:0xf
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[4:5]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB16_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX9-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_4:
+; GFX10-32-SDAG:       ; %bb.0:
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-SDAG-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v5, v1
+; GFX10-32-SDAG-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-32-SDAG-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX10-32-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[2:3]
+; GFX10-32-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[5:6]
+; GFX10-32-SDAG-NEXT:    v_add_co_u32 v0, vcc_lo, s0, v0
+; GFX10-32-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, s1, v1, vcc_lo
+; GFX10-32-SDAG-NEXT:    v_add_co_u32 v2, vcc_lo, s2, v2
+; GFX10-32-SDAG-NEXT:    v_add_co_ci_u32_e32 v3, vcc_lo, s3, v3, vcc_lo
+; GFX10-32-SDAG-NEXT:    s_clause 0x1
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[9:12], v[0:1], off offset:16
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[13:16], v[2:3], off
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-SDAG-NEXT:  .LBB16_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s2, v4
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s2, v4
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB16_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-SDAG-NEXT:    s_and_b32 exec_lo, exec_lo, s4
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_4:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s5, exec_lo
+; GFX10-32-GISEL-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v5, v2
+; GFX10-32-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v8, s3
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v7, s2
+; GFX10-32-GISEL-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX10-32-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v17, s4
+; GFX10-32-GISEL-NEXT:    v_lshlrev_b64 v[5:6], 4, v[5:6]
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-GISEL-NEXT:    v_add_co_u32 v0, vcc_lo, v2, v0
+; GFX10-32-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, v3, v1, vcc_lo
+; GFX10-32-GISEL-NEXT:    v_add_co_u32 v2, vcc_lo, v7, v5
+; GFX10-32-GISEL-NEXT:    v_add_co_ci_u32_e32 v3, vcc_lo, v8, v6, vcc_lo
+; GFX10-32-GISEL-NEXT:    s_clause 0x1
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[9:12], v[0:1], off offset:16
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[13:16], v[2:3], off
+; GFX10-32-GISEL-NEXT:  .LBB16_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s2, v4
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s3, v17
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s2, v4
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s3, v17
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr17
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB16_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-GISEL-NEXT:    s_and_b32 exec_lo, exec_lo, s5
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_4:
+; GFX10-64-SDAG:       ; %bb.0:
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v5, v1
+; GFX10-64-SDAG-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-64-SDAG-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX10-64-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[2:3]
+; GFX10-64-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[5:6]
+; GFX10-64-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX10-64-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc, s1, v1, vcc
+; GFX10-64-SDAG-NEXT:    v_add_co_u32 v2, vcc, s2, v2
+; GFX10-64-SDAG-NEXT:    v_add_co_ci_u32_e32 v3, vcc, s3, v3, vcc
+; GFX10-64-SDAG-NEXT:    s_clause 0x1
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[9:12], v[0:1], off offset:16
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[13:16], v[2:3], off
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-SDAG-NEXT:  .LBB16_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s6, v4
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v4
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB16_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_4:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v5, v2
+; GFX10-64-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v8, s3
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v7, s2
+; GFX10-64-GISEL-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX10-64-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v17, s4
+; GFX10-64-GISEL-NEXT:    v_lshlrev_b64 v[5:6], 4, v[5:6]
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX10-64-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc, v3, v1, vcc
+; GFX10-64-GISEL-NEXT:    v_add_co_u32 v2, vcc, v7, v5
+; GFX10-64-GISEL-NEXT:    v_add_co_ci_u32_e32 v3, vcc, v8, v6, vcc
+; GFX10-64-GISEL-NEXT:    s_clause 0x1
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[5:8], v[0:1], off
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[9:12], v[0:1], off offset:16
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[13:16], v[2:3], off
+; GFX10-64-GISEL-NEXT:  .LBB16_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s4, v4
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s5, v17
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v4
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s5, v17
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr17
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB16_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX1150-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_4:
+; GFX1150-SDAG:       ; %bb.0:
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v5, v1
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX1150-SDAG-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_3) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-SDAG-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX1150-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[2:3]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[5:6]
+; GFX1150-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_4)
+; GFX1150-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s1, v1, vcc
+; GFX1150-SDAG-NEXT:    v_add_co_u32 v2, vcc, s2, v2
+; GFX1150-SDAG-NEXT:    s_waitcnt_depctr depctr_va_vcc(0)
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-SDAG-NEXT:    v_add_co_ci_u32_e64 v3, null, s3, v3, vcc
+; GFX1150-SDAG-NEXT:    s_clause 0x1
+; GFX1150-SDAG-NEXT:    global_load_b128 v[9:12], v[0:1], off offset:16
+; GFX1150-SDAG-NEXT:    global_load_b128 v[5:8], v[0:1], off
+; GFX1150-SDAG-NEXT:    global_load_b128 v[13:16], v[2:3], off
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-SDAG-NEXT:  .LBB16_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s6, v4
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v4
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB16_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX1150-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_4:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v5, v2
+; GFX1150-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v8, s3
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v7, s2
+; GFX1150-GISEL-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX1150-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[1:2]
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v17, s4
+; GFX1150-GISEL-NEXT:    v_lshlrev_b64 v[5:6], 4, v[5:6]
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_3)
+; GFX1150-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v3, v1, vcc
+; GFX1150-GISEL-NEXT:    v_add_co_u32 v2, vcc, v7, v5
+; GFX1150-GISEL-NEXT:    s_waitcnt_depctr depctr_va_vcc(0)
+; GFX1150-GISEL-NEXT:    v_add_co_ci_u32_e64 v3, null, v8, v6, vcc
+; GFX1150-GISEL-NEXT:    s_clause 0x1
+; GFX1150-GISEL-NEXT:    global_load_b128 v[5:8], v[0:1], off
+; GFX1150-GISEL-NEXT:    global_load_b128 v[9:12], v[0:1], off offset:16
+; GFX1150-GISEL-NEXT:    global_load_b128 v[13:16], v[2:3], off
+; GFX1150-GISEL-NEXT:  .LBB16_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s4, v4
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s5, v17
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v4
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s5, v17
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr17
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB16_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX12-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_4:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v5, v1
+; GFX12-SDAG-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_3) | instskip(NEXT) | instid1(VALU_DEP_3)
+; GFX12-SDAG-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX12-SDAG-NEXT:    v_lshlrev_b64_e32 v[2:3], 4, v[2:3]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_lshlrev_b64_e32 v[0:1], 5, v[5:6]
+; GFX12-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_4)
+; GFX12-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s1, v1, vcc
+; GFX12-SDAG-NEXT:    v_add_co_u32 v2, vcc, s2, v2
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_vcc(0)
+; GFX12-SDAG-NEXT:    v_add_co_ci_u32_e64 v3, null, s3, v3, vcc
+; GFX12-SDAG-NEXT:    s_clause 0x1
+; GFX12-SDAG-NEXT:    global_load_b128 v[9:12], v[0:1], off offset:16
+; GFX12-SDAG-NEXT:    global_load_b128 v[5:8], v[0:1], off
+; GFX12-SDAG-NEXT:    global_load_b128 v[13:16], v[2:3], off
+; GFX12-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-SDAG-NEXT:  .LBB16_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s6, v4
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v4
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x0
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr4
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB16_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX12-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_4:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v5, v2
+; GFX12-GISEL-NEXT:    v_ashrrev_i32_e32 v2, 31, v1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v8, s3
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v7, s2
+; GFX12-GISEL-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX12-GISEL-NEXT:    v_lshlrev_b64_e32 v[0:1], 5, v[1:2]
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v17, s4
+; GFX12-GISEL-NEXT:    v_lshlrev_b64_e32 v[5:6], 4, v[5:6]
+; GFX12-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_3)
+; GFX12-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v3, v1, vcc
+; GFX12-GISEL-NEXT:    v_add_co_u32 v2, vcc, v7, v5
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_vcc(0)
+; GFX12-GISEL-NEXT:    v_add_co_ci_u32_e64 v3, null, v8, v6, vcc
+; GFX12-GISEL-NEXT:    s_clause 0x1
+; GFX12-GISEL-NEXT:    global_load_b128 v[5:8], v[0:1], off
+; GFX12-GISEL-NEXT:    global_load_b128 v[9:12], v[0:1], off offset:16
+; GFX12-GISEL-NEXT:    global_load_b128 v[13:16], v[2:3], off
+; GFX12-GISEL-NEXT:  .LBB16_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s4, v4
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s5, v17
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v4
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s5, v17
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x2
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s8, v5
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s9, v6
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s10, v7
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s13, v10
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s14, v11
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s15, v12
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x0
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s16, v13
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s17, v14
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s18, v15
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s19, v16
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr17
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8_vgpr9_vgpr10_vgpr11_vgpr12
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr13_vgpr14_vgpr15_vgpr16
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB16_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    ; return to shader part epilog
+       <4 x i32> addrspace(4)* inreg %s_in, i32 %idx1, i32 inreg %idx2, i32 %s_idx, i32 %s_idx2) #1 {
+  %rptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx
+  %sptr = getelementptr <4 x i32>, <4 x i32> addrspace(4)* %s_in, i32 %s_idx2
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4)* %rptr, align 16
+  %srsrc = load <4 x i32>, <4 x i32> addrspace(4)* %sptr, align 16
+  %tok = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %idx1)
+  %tok1 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 %tok, i32 %idx2)
+  %s_rsrc = call <8 x i32> @llvm.amdgcn.waterfall.readfirstlane.v8i32.v8i32(i32 %tok1, <8 x i32> %rsrc)
+  %s_srsrc = call <4 x i32> @llvm.amdgcn.waterfall.readfirstlane.v4i32.v4i32(i32 %tok1, <4 x i32> %srsrc)
+  %r = call <4 x float> @llvm.amdgcn.image.sample.2d.v4f32.f32(i32 15, float 0.000000e+00, float 0.000000e+00, <8 x i32> %s_rsrc, <4 x i32> %s_srsrc, i1 false, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok1, <4 x float> %r)
+
+  ret <4 x float> %r1
+}
+
+define amdgpu_ps {<4 x float>,<4 x float>} @test_waterfall_multi_begin_uniform_idx_1_2loops(
+; VI-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_1_2loops:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; VI-SDAG-NEXT:    s_wqm_b64 exec, exec
+; VI-SDAG-NEXT:    v_mov_b32_e32 v7, v2
+; VI-SDAG-NEXT:    v_ashrrev_i32_e32 v8, 31, v7
+; VI-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; VI-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; VI-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[7:8]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v2, s1
+; VI-SDAG-NEXT:    v_add_u32_e32 v0, vcc, s0, v0
+; VI-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; VI-SDAG-NEXT:    v_addc_u32_e32 v1, vcc, v2, v1, vcc
+; VI-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[3:4]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, s3
+; VI-SDAG-NEXT:    v_add_u32_e32 v2, vcc, s2, v2
+; VI-SDAG-NEXT:    v_addc_u32_e32 v3, vcc, v4, v3, vcc
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[11:14], v[0:1]
+; VI-SDAG-NEXT:    v_add_u32_e32 v0, vcc, 16, v0
+; VI-SDAG-NEXT:    v_addc_u32_e32 v1, vcc, 0, v1, vcc
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[15:18], v[0:1]
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[7:10], v[2:3]
+; VI-SDAG-NEXT:    s_mov_b32 s0, 0
+; VI-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:  .LBB17_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s1, v5
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s8, v6
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[6:7], s1, v5
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[8:9], s8, v6
+; VI-SDAG-NEXT:    s_and_b64 s[6:7], s[6:7], s[8:9]
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[6:7], s[6:7]
+; VI-SDAG-NEXT:    s_mov_b32 s1, s0
+; VI-SDAG-NEXT:    v_mov_b32_e32 v20, s1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s8, v11
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s9, v12
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s10, v13
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s11, v14
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s12, v15
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s13, v16
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s14, v17
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s15, v18
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s16, v7
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s17, v8
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s18, v9
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s19, v10
+; VI-SDAG-NEXT:    v_mov_b32_e32 v19, s0
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr5
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr6
+; VI-SDAG-NEXT:    s_nop 3
+; VI-SDAG-NEXT:    image_sample v[0:3], v[19:20], s[8:15], s[16:19] dmask:0xf
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[6:7]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB17_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[2:3]
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s8, v11
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s9, v12
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s10, v13
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s11, v14
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s12, v15
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s13, v16
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s14, v17
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s15, v18
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s0, v7
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s1, v8
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s2, v9
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s3, v10
+; VI-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; VI-SDAG-NEXT:    s_nop 3
+; VI-SDAG-NEXT:    image_sample v[4:7], v[19:20], s[8:15], s[0:3] dmask:0xf
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    ; return to shader part epilog
+;
+; VI-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_1_2loops:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; VI-GISEL-NEXT:    s_wqm_b64 exec, exec
+; VI-GISEL-NEXT:    v_mov_b32_e32 v6, v3
+; VI-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v5, v1
+; VI-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[2:3]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; VI-GISEL-NEXT:    v_add_u32_e32 v0, vcc, v2, v0
+; VI-GISEL-NEXT:    v_ashrrev_i32_e32 v7, 31, v6
+; VI-GISEL-NEXT:    v_addc_u32_e32 v1, vcc, v3, v1, vcc
+; VI-GISEL-NEXT:    v_lshlrev_b64 v[2:3], 4, v[6:7]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v7, s3
+; VI-GISEL-NEXT:    v_mov_b32_e32 v6, s2
+; VI-GISEL-NEXT:    v_add_u32_e32 v2, vcc, v6, v2
+; VI-GISEL-NEXT:    v_addc_u32_e32 v3, vcc, v7, v3, vcc
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[8:11], v[0:1]
+; VI-GISEL-NEXT:    v_add_u32_e32 v0, vcc, 16, v0
+; VI-GISEL-NEXT:    v_addc_u32_e32 v1, vcc, 0, v1, vcc
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[12:15], v[0:1]
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[16:19], v[2:3]
+; VI-GISEL-NEXT:    s_mov_b32 s0, 0
+; VI-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:  .LBB17_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s1, v4
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s10, v5
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[8:9], s1, v4
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[10:11], s10, v5
+; VI-GISEL-NEXT:    s_and_b64 s[8:9], s[8:9], s[10:11]
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[20:21], s[8:9]
+; VI-GISEL-NEXT:    s_mov_b32 s1, s0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v21, s1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s8, v8
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s9, v9
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s10, v10
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s11, v11
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s12, v12
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s13, v13
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s14, v14
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s15, v15
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s16, v16
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s17, v17
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s18, v18
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s19, v19
+; VI-GISEL-NEXT:    v_mov_b32_e32 v20, s0
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr4
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr5
+; VI-GISEL-NEXT:    s_nop 3
+; VI-GISEL-NEXT:    image_sample v[0:3], v[20:21], s[8:15], s[16:19] dmask:0xf
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[20:21]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB17_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v22, s4
+; VI-GISEL-NEXT:    v_mov_b32_e32 v23, s5
+; VI-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; VI-GISEL-NEXT:  .LBB17_3: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s2, v22
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s4, v23
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v22
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[4:5], s4, v23
+; VI-GISEL-NEXT:    s_and_b64 s[2:3], s[2:3], s[4:5]
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s8, v8
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s9, v9
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s10, v10
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s11, v11
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s12, v12
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s13, v13
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s14, v14
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s15, v15
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s16, v16
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s17, v17
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s18, v18
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s19, v19
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr22
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr23
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr16_vgpr17_vgpr18_vgpr19
+; VI-GISEL-NEXT:    s_nop 4
+; VI-GISEL-NEXT:    image_sample v[4:7], v[20:21], s[8:15], s[16:19] dmask:0xf
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr20_vgpr21
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB17_3
+; VI-GISEL-NEXT:  ; %bb.4:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; VI-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX9-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_1_2loops:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX9-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v7, v2
+; GFX9-SDAG-NEXT:    v_ashrrev_i32_e32 v8, 31, v7
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX9-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[7:8]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v2, s1
+; GFX9-SDAG-NEXT:    v_add_co_u32_e32 v0, vcc, s0, v0
+; GFX9-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX9-SDAG-NEXT:    v_addc_co_u32_e32 v1, vcc, v2, v1, vcc
+; GFX9-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[3:4]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v4, s3
+; GFX9-SDAG-NEXT:    v_add_co_u32_e32 v2, vcc, s2, v2
+; GFX9-SDAG-NEXT:    v_addc_co_u32_e32 v3, vcc, v4, v3, vcc
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[11:14], v[0:1], off offset:16
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[15:18], v[0:1], off
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[7:10], v[2:3], off
+; GFX9-SDAG-NEXT:    s_mov_b32 s0, 0
+; GFX9-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:  .LBB17_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s1, v5
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s8, v6
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[6:7], s1, v5
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[8:9], s8, v6
+; GFX9-SDAG-NEXT:    s_and_b64 s[6:7], s[6:7], s[8:9]
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[6:7], s[6:7]
+; GFX9-SDAG-NEXT:    s_mov_b32 s1, s0
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v20, s1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s8, v15
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s9, v16
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s10, v17
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s11, v18
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s13, v12
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s14, v13
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s15, v14
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s16, v7
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s17, v8
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s18, v9
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s19, v10
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v19, s0
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX9-SDAG-NEXT:    s_nop 3
+; GFX9-SDAG-NEXT:    image_sample v[0:3], v[19:20], s[8:15], s[16:19] dmask:0xf
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[6:7]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB17_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s8, v15
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s9, v16
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s10, v17
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s11, v18
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s13, v12
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s14, v13
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s15, v14
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s0, v7
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s1, v8
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s2, v9
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s3, v10
+; GFX9-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX9-SDAG-NEXT:    s_nop 3
+; GFX9-SDAG-NEXT:    image_sample v[4:7], v[19:20], s[8:15], s[0:3] dmask:0xf
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX9-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_1_2loops:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX9-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v6, v3
+; GFX9-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v5, v1
+; GFX9-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[2:3]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX9-GISEL-NEXT:    v_add_co_u32_e32 v0, vcc, v2, v0
+; GFX9-GISEL-NEXT:    v_ashrrev_i32_e32 v7, 31, v6
+; GFX9-GISEL-NEXT:    v_addc_co_u32_e32 v1, vcc, v3, v1, vcc
+; GFX9-GISEL-NEXT:    v_lshlrev_b64 v[2:3], 4, v[6:7]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v7, s3
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v6, s2
+; GFX9-GISEL-NEXT:    v_add_co_u32_e32 v2, vcc, v6, v2
+; GFX9-GISEL-NEXT:    v_addc_co_u32_e32 v3, vcc, v7, v3, vcc
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[8:11], v[0:1], off
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[12:15], v[0:1], off offset:16
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[16:19], v[2:3], off
+; GFX9-GISEL-NEXT:    s_mov_b32 s0, 0
+; GFX9-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:  .LBB17_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s1, v4
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s10, v5
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[8:9], s1, v4
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[10:11], s10, v5
+; GFX9-GISEL-NEXT:    s_and_b64 s[8:9], s[8:9], s[10:11]
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[20:21], s[8:9]
+; GFX9-GISEL-NEXT:    s_mov_b32 s1, s0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v21, s1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s8, v8
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s9, v9
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s10, v10
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s11, v11
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s12, v12
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s13, v13
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s14, v14
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s15, v15
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s16, v16
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s17, v17
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s18, v18
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s19, v19
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v20, s0
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX9-GISEL-NEXT:    s_nop 3
+; GFX9-GISEL-NEXT:    image_sample v[0:3], v[20:21], s[8:15], s[16:19] dmask:0xf
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[20:21]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB17_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v22, s4
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v23, s5
+; GFX9-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX9-GISEL-NEXT:  .LBB17_3: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s2, v22
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s4, v23
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v22
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[4:5], s4, v23
+; GFX9-GISEL-NEXT:    s_and_b64 s[2:3], s[2:3], s[4:5]
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s8, v8
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s9, v9
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s10, v10
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s11, v11
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s12, v12
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s13, v13
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s14, v14
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s15, v15
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s16, v16
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s17, v17
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s18, v18
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s19, v19
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr22
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr23
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr16_vgpr17_vgpr18_vgpr19
+; GFX9-GISEL-NEXT:    s_nop 4
+; GFX9-GISEL-NEXT:    image_sample v[4:7], v[20:21], s[8:15], s[16:19] dmask:0xf
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr20_vgpr21
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB17_3
+; GFX9-GISEL-NEXT:  ; %bb.4:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX9-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_1_2loops:
+; GFX10-32-SDAG:       ; %bb.0:
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-SDAG-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v7, v2
+; GFX10-32-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX10-32-SDAG-NEXT:    v_ashrrev_i32_e32 v8, 31, v7
+; GFX10-32-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[3:4]
+; GFX10-32-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[7:8]
+; GFX10-32-SDAG-NEXT:    v_add_co_u32 v0, vcc_lo, s0, v0
+; GFX10-32-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, s1, v1, vcc_lo
+; GFX10-32-SDAG-NEXT:    v_add_co_u32 v2, vcc_lo, s2, v2
+; GFX10-32-SDAG-NEXT:    v_add_co_ci_u32_e32 v3, vcc_lo, s3, v3, vcc_lo
+; GFX10-32-SDAG-NEXT:    s_clause 0x1
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[11:14], v[0:1], off offset:16
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[15:18], v[0:1], off
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[7:10], v[2:3], off
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:  .LBB17_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s2, v5
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s3, v6
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s2, v5
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s3, v6
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s8, v15
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s9, v16
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s10, v17
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s11, v18
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s13, v12
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s14, v13
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s15, v14
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s16, v7
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s17, v8
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s18, v9
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s19, v10
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v4, 0
+; GFX10-32-SDAG-NEXT:    image_sample v[0:3], [v4, v4], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB17_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s8, v15
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s9, v16
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s10, v17
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s11, v18
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s13, v12
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s14, v13
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s15, v14
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s0, v7
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s1, v8
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s2, v9
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s3, v10
+; GFX10-32-SDAG-NEXT:    s_and_b32 exec_lo, exec_lo, s4
+; GFX10-32-SDAG-NEXT:    image_sample v[4:7], [v4, v4], s[8:15], s[0:3] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_1_2loops:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s6, exec_lo
+; GFX10-32-GISEL-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v6, v3
+; GFX10-32-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v5, v1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v9, s3
+; GFX10-32-GISEL-NEXT:    v_ashrrev_i32_e32 v7, 31, v6
+; GFX10-32-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[2:3]
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v8, s2
+; GFX10-32-GISEL-NEXT:    v_lshlrev_b64 v[6:7], 4, v[6:7]
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-GISEL-NEXT:    v_add_co_u32 v0, vcc_lo, v2, v0
+; GFX10-32-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, v3, v1, vcc_lo
+; GFX10-32-GISEL-NEXT:    v_add_co_u32 v2, vcc_lo, v8, v6
+; GFX10-32-GISEL-NEXT:    v_add_co_ci_u32_e32 v3, vcc_lo, v9, v7, vcc_lo
+; GFX10-32-GISEL-NEXT:    s_clause 0x1
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[8:11], v[0:1], off
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[12:15], v[0:1], off offset:16
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[16:19], v[2:3], off
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:  .LBB17_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s2, v4
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s3, v5
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s2, v4
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s3, v5
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s8, v8
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s9, v9
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s10, v10
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s11, v11
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s12, v12
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s13, v13
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s14, v14
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s15, v15
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s16, v16
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s17, v17
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s18, v18
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s19, v19
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v20, 0
+; GFX10-32-GISEL-NEXT:    image_sample v[0:3], [v20, v20], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB17_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v21, s4
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v22, s5
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-GISEL-NEXT:  .LBB17_3: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s2, v21
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s3, v22
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s2, v21
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s3, v22
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s8, v8
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s9, v9
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s10, v10
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s11, v11
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s12, v12
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s13, v13
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s14, v14
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s15, v15
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s16, v16
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s17, v17
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s18, v18
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s19, v19
+; GFX10-32-GISEL-NEXT:    image_sample v[4:7], [v20, v20], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr21
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr22
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr16_vgpr17_vgpr18_vgpr19
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr20
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB17_3
+; GFX10-32-GISEL-NEXT:  ; %bb.4:
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-GISEL-NEXT:    s_and_b32 exec_lo, exec_lo, s6
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_1_2loops:
+; GFX10-64-SDAG:       ; %bb.0:
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v7, v2
+; GFX10-64-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX10-64-SDAG-NEXT:    v_ashrrev_i32_e32 v8, 31, v7
+; GFX10-64-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[3:4]
+; GFX10-64-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[7:8]
+; GFX10-64-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX10-64-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc, s1, v1, vcc
+; GFX10-64-SDAG-NEXT:    v_add_co_u32 v2, vcc, s2, v2
+; GFX10-64-SDAG-NEXT:    v_add_co_ci_u32_e32 v3, vcc, s3, v3, vcc
+; GFX10-64-SDAG-NEXT:    s_clause 0x1
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[11:14], v[0:1], off offset:16
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[15:18], v[0:1], off
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[7:10], v[2:3], off
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:  .LBB17_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s6, v5
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s7, v6
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v5
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s7, v6
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s8, v15
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s9, v16
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s10, v17
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s11, v18
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s13, v12
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s14, v13
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s15, v14
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s16, v7
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s17, v8
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s18, v9
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s19, v10
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v4, 0
+; GFX10-64-SDAG-NEXT:    image_sample v[0:3], [v4, v4], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB17_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s8, v15
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s9, v16
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s10, v17
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s11, v18
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s13, v12
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s14, v13
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s15, v14
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s0, v7
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s1, v8
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s2, v9
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s3, v10
+; GFX10-64-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX10-64-SDAG-NEXT:    image_sample v[4:7], [v4, v4], s[8:15], s[0:3] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_1_2loops:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v6, v3
+; GFX10-64-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v5, v1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v9, s3
+; GFX10-64-GISEL-NEXT:    v_ashrrev_i32_e32 v7, 31, v6
+; GFX10-64-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[2:3]
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v8, s2
+; GFX10-64-GISEL-NEXT:    v_lshlrev_b64 v[6:7], 4, v[6:7]
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX10-64-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc, v3, v1, vcc
+; GFX10-64-GISEL-NEXT:    v_add_co_u32 v2, vcc, v8, v6
+; GFX10-64-GISEL-NEXT:    v_add_co_ci_u32_e32 v3, vcc, v9, v7, vcc
+; GFX10-64-GISEL-NEXT:    s_clause 0x1
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[8:11], v[0:1], off
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[12:15], v[0:1], off offset:16
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[16:19], v[2:3], off
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:  .LBB17_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s8, v4
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s9, v5
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s8, v4
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s9, v5
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s8, v8
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s9, v9
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s10, v10
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s11, v11
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s12, v12
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s13, v13
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s14, v14
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s15, v15
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s16, v16
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s17, v17
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s18, v18
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s19, v19
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v20, 0
+; GFX10-64-GISEL-NEXT:    image_sample v[0:3], [v20, v20], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB17_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v21, s4
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v22, s5
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-GISEL-NEXT:  .LBB17_3: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s4, v21
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s5, v22
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v21
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s5, v22
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s8, v8
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s9, v9
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s10, v10
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s11, v11
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s12, v12
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s13, v13
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s14, v14
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s15, v15
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s16, v16
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s17, v17
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s18, v18
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s19, v19
+; GFX10-64-GISEL-NEXT:    image_sample v[4:7], [v20, v20], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr21
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr22
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr16_vgpr17_vgpr18_vgpr19
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr20
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB17_3
+; GFX10-64-GISEL-NEXT:  ; %bb.4:
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX1150-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_1_2loops:
+; GFX1150-SDAG:       ; %bb.0:
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v7, v2
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX1150-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_4) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-SDAG-NEXT:    v_ashrrev_i32_e32 v8, 31, v7
+; GFX1150-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[3:4]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[7:8]
+; GFX1150-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_4)
+; GFX1150-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s1, v1, vcc
+; GFX1150-SDAG-NEXT:    v_add_co_u32 v2, vcc, s2, v2
+; GFX1150-SDAG-NEXT:    s_waitcnt_depctr depctr_va_vcc(0)
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-SDAG-NEXT:    v_add_co_ci_u32_e64 v3, null, s3, v3, vcc
+; GFX1150-SDAG-NEXT:    s_clause 0x1
+; GFX1150-SDAG-NEXT:    global_load_b128 v[11:14], v[0:1], off offset:16
+; GFX1150-SDAG-NEXT:    global_load_b128 v[15:18], v[0:1], off
+; GFX1150-SDAG-NEXT:    global_load_b128 v[7:10], v[2:3], off
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:  .LBB17_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s6, v5
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s7, v6
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v5
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s7, v6
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s8, v15
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s9, v16
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s10, v17
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s11, v18
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s13, v12
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s14, v13
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s15, v14
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s16, v7
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s17, v8
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s18, v9
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s19, v10
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v4, 0
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    image_sample v[0:3], [v4, v4], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB17_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s8, v15
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s9, v16
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s10, v17
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s11, v18
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s13, v12
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s14, v13
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s15, v14
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s0, v7
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s1, v8
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s2, v9
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s3, v10
+; GFX1150-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX1150-SDAG-NEXT:    image_sample v[4:7], [v4, v4], s[8:15], s[0:3] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX1150-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_1_2loops:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v6, v3
+; GFX1150-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v5, v1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v9, s3
+; GFX1150-GISEL-NEXT:    v_ashrrev_i32_e32 v7, 31, v6
+; GFX1150-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[2:3]
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v8, s2
+; GFX1150-GISEL-NEXT:    v_lshlrev_b64 v[6:7], 4, v[6:7]
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_3)
+; GFX1150-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v3, v1, vcc
+; GFX1150-GISEL-NEXT:    v_add_co_u32 v2, vcc, v8, v6
+; GFX1150-GISEL-NEXT:    s_waitcnt_depctr depctr_va_vcc(0)
+; GFX1150-GISEL-NEXT:    v_add_co_ci_u32_e64 v3, null, v9, v7, vcc
+; GFX1150-GISEL-NEXT:    s_clause 0x1
+; GFX1150-GISEL-NEXT:    global_load_b128 v[8:11], v[0:1], off
+; GFX1150-GISEL-NEXT:    global_load_b128 v[12:15], v[0:1], off offset:16
+; GFX1150-GISEL-NEXT:    global_load_b128 v[16:19], v[2:3], off
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:  .LBB17_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s8, v4
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s9, v5
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s8, v4
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s9, v5
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s8, v8
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s9, v9
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s10, v10
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s11, v11
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s12, v12
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s13, v13
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s14, v14
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s15, v15
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s16, v16
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s17, v17
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s18, v18
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s19, v19
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v20, 0
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    image_sample v[0:3], [v20, v20], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB17_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v21, s4
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v22, s5
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-GISEL-NEXT:  .LBB17_3: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s4, v21
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s5, v22
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v21
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s5, v22
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s8, v8
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s9, v9
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s10, v10
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s11, v11
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s12, v12
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s13, v13
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s14, v14
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s15, v15
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s16, v16
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s17, v17
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s18, v18
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s19, v19
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    image_sample v[4:7], [v20, v20], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr21
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr22
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr16_vgpr17_vgpr18_vgpr19
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr20
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB17_3
+; GFX1150-GISEL-NEXT:  ; %bb.4:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX12-SDAG-LABEL: test_waterfall_multi_begin_uniform_idx_1_2loops:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v7, v2
+; GFX12-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v5, v0
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_4) | instskip(NEXT) | instid1(VALU_DEP_4)
+; GFX12-SDAG-NEXT:    v_ashrrev_i32_e32 v8, 31, v7
+; GFX12-SDAG-NEXT:    v_lshlrev_b64_e32 v[2:3], 4, v[3:4]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_lshlrev_b64_e32 v[0:1], 5, v[7:8]
+; GFX12-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_4)
+; GFX12-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s1, v1, vcc
+; GFX12-SDAG-NEXT:    v_add_co_u32 v2, vcc, s2, v2
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_vcc(0)
+; GFX12-SDAG-NEXT:    v_add_co_ci_u32_e64 v3, null, s3, v3, vcc
+; GFX12-SDAG-NEXT:    s_clause 0x1
+; GFX12-SDAG-NEXT:    global_load_b128 v[11:14], v[0:1], off offset:16
+; GFX12-SDAG-NEXT:    global_load_b128 v[15:18], v[0:1], off
+; GFX12-SDAG-NEXT:    global_load_b128 v[7:10], v[2:3], off
+; GFX12-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-SDAG-NEXT:  .LBB17_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s6, v5
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s7, v6
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v5
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s7, v6
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s8, v15
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s9, v16
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s10, v17
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s11, v18
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s13, v12
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s14, v13
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s15, v14
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x0
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s16, v7
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s17, v8
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s18, v9
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s19, v10
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v4, 0
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    image_sample v[0:3], [v4, v4], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr6
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB17_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s8, v15
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s9, v16
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s10, v17
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s11, v18
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s13, v12
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s14, v13
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s15, v14
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s0, v7
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s1, v8
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s2, v9
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s3, v10
+; GFX12-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX12-SDAG-NEXT:    image_sample v[4:7], [v4, v4], s[8:15], s[0:3] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX12-GISEL-LABEL: test_waterfall_multi_begin_uniform_idx_1_2loops:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v6, v3
+; GFX12-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v5, v1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v9, s3
+; GFX12-GISEL-NEXT:    v_ashrrev_i32_e32 v7, 31, v6
+; GFX12-GISEL-NEXT:    v_lshlrev_b64_e32 v[0:1], 5, v[2:3]
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v8, s2
+; GFX12-GISEL-NEXT:    v_lshlrev_b64_e32 v[6:7], 4, v[6:7]
+; GFX12-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_3)
+; GFX12-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v3, v1, vcc
+; GFX12-GISEL-NEXT:    v_add_co_u32 v2, vcc, v8, v6
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_vcc(0)
+; GFX12-GISEL-NEXT:    v_add_co_ci_u32_e64 v3, null, v9, v7, vcc
+; GFX12-GISEL-NEXT:    s_clause 0x1
+; GFX12-GISEL-NEXT:    global_load_b128 v[8:11], v[0:1], off
+; GFX12-GISEL-NEXT:    global_load_b128 v[12:15], v[0:1], off offset:16
+; GFX12-GISEL-NEXT:    global_load_b128 v[16:19], v[2:3], off
+; GFX12-GISEL-NEXT:  .LBB17_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s8, v4
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s9, v5
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s8, v4
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s9, v5
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x2
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s8, v8
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s9, v9
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s10, v10
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s11, v11
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s12, v12
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s13, v13
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s14, v14
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s15, v15
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x0
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s16, v16
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s17, v17
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s18, v18
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s19, v19
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v20, 0
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    image_sample v[0:3], [v20, v20], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr4
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB17_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v21, s4
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v22, s5
+; GFX12-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-GISEL-NEXT:  .LBB17_3: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s4, v21
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s5, v22
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v21
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s5, v22
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s8, v8
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s9, v9
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s10, v10
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s11, v11
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s12, v12
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s13, v13
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s14, v14
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s15, v15
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s16, v16
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s17, v17
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s18, v18
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s19, v19
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    image_sample v[4:7], [v20, v20], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr21
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr22
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr8_vgpr9_vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr16_vgpr17_vgpr18_vgpr19
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr20
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB17_3
+; GFX12-GISEL-NEXT:  ; %bb.4:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    ; return to shader part epilog
+       <8 x i32> addrspace(4)* inreg %in, <4 x i32> addrspace(4)* inreg %s_in,
+       i32 %idx1, i32 %idx2, i32 inreg %idx3, i32 inreg%idx4,
+       i32 %s_idx, i32 %s_idx2) #1 {
+  %rptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx
+  %sptr = getelementptr <4 x i32>, <4 x i32> addrspace(4)* %s_in, i32 %s_idx2
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4)* %rptr, align 16
+  %srsrc = load <4 x i32>, <4 x i32> addrspace(4)* %sptr, align 16
+  %tok = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %idx1)
+  %tok1 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 %tok, i32 %idx2)
+  %s_rsrc = call <8 x i32> @llvm.amdgcn.waterfall.readfirstlane.v8i32.v8i32(i32 %tok1, <8 x i32> %rsrc)
+  %s_srsrc = call <4 x i32> @llvm.amdgcn.waterfall.readfirstlane.v4i32.v4i32(i32 %tok1, <4 x i32> %srsrc)
+  %r = call <4 x float> @llvm.amdgcn.image.sample.2d.v4f32.f32(i32 15, float 0.000000e+00, float 0.000000e+00, <8 x i32> %s_rsrc, <4 x i32> %s_srsrc, i1 false, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok1, <4 x float> %r)
+
+  %tok2 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %idx3)
+  %tok3 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 %tok2, i32 %idx4)
+  %s_rsrc1 = call <8 x i32> @llvm.amdgcn.waterfall.readfirstlane.v8i32.v8i32(i32 %tok3, <8 x i32> %rsrc)
+  %s_srsrc1 = call <4 x i32> @llvm.amdgcn.waterfall.readfirstlane.v4i32.v4i32(i32 %tok3, <4 x i32> %srsrc)
+  %r2 = call <4 x float> @llvm.amdgcn.image.sample.2d.v4f32.f32(i32 15, float 0.000000e+00, float 0.000000e+00, <8 x i32> %s_rsrc1, <4 x i32> %s_srsrc1, i1 false, i32 0, i32 0)
+  %r3 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok3, <4 x float> %r2)
+
+  %insert = insertvalue { <4 x float>, <4 x float> } undef, <4 x float> %r1, 0
+  %insert1 = insertvalue { <4 x float>, <4 x float> } %insert, <4 x float> %r3, 1
+  ret {<4 x float> , <4 x float>} %insert1
+}
+
+define amdgpu_gfx i32 @test_indirect_call_vgpr_ptr_arg_and_reuse(i32 %i, i32 %fptr) {
+; PRE-GFX10-LABEL: test_indirect_call_vgpr_ptr_arg_and_reuse:
+; PRE-GFX10:       ; %bb.0:
+; PRE-GFX10-NEXT:    s_waitcnt vmcnt(0) expcnt(0) lgkmcnt(0)
+; PRE-GFX10-NEXT:    s_mov_b32 s34, s33
+; PRE-GFX10-NEXT:    s_mov_b32 s33, s32
+; PRE-GFX10-NEXT:    s_or_saveexec_b64 s[36:37], -1
+; PRE-GFX10-NEXT:    buffer_store_dword v40, off, s[0:3], s33 ; 4-byte Folded Spill
+; PRE-GFX10-NEXT:    s_mov_b64 exec, s[36:37]
+; PRE-GFX10-NEXT:    v_writelane_b32 v40, s34, 8
+; PRE-GFX10-NEXT:    v_writelane_b32 v40, s4, 0
+; PRE-GFX10-NEXT:    v_writelane_b32 v40, s5, 1
+; PRE-GFX10-NEXT:    v_writelane_b32 v40, s6, 2
+; PRE-GFX10-NEXT:    v_writelane_b32 v40, s7, 3
+; PRE-GFX10-NEXT:    v_writelane_b32 v40, s8, 4
+; PRE-GFX10-NEXT:    v_writelane_b32 v40, s9, 5
+; PRE-GFX10-NEXT:    v_writelane_b32 v40, s30, 6
+; PRE-GFX10-NEXT:    s_mov_b32 s5, 0
+; PRE-GFX10-NEXT:    s_mov_b64 s[6:7], exec
+; PRE-GFX10-NEXT:    s_addk_i32 s32, 0x400
+; PRE-GFX10-NEXT:    v_writelane_b32 v40, s31, 7
+; PRE-GFX10-NEXT:  .LBB18_1: ; =>This Inner Loop Header: Depth=1
+; PRE-GFX10-NEXT:    v_readfirstlane_b32 s4, v1
+; PRE-GFX10-NEXT:    v_cmp_eq_u32_e64 s[34:35], s4, v1
+; PRE-GFX10-NEXT:    s_and_saveexec_b64 s[8:9], s[34:35]
+; PRE-GFX10-NEXT:    s_swappc_b64 s[30:31], s[4:5]
+; PRE-GFX10-NEXT:    v_mov_b32_e32 v2, v0
+; PRE-GFX10-NEXT:    ; implicit-def: $vgpr1
+; PRE-GFX10-NEXT:    ; implicit-def: $vgpr0
+; PRE-GFX10-NEXT:    s_xor_b64 exec, exec, s[8:9]
+; PRE-GFX10-NEXT:    s_cbranch_execnz .LBB18_1
+; PRE-GFX10-NEXT:  ; %bb.2:
+; PRE-GFX10-NEXT:    s_mov_b64 exec, s[6:7]
+; PRE-GFX10-NEXT:    v_mov_b32_e32 v0, v2
+; PRE-GFX10-NEXT:    v_readlane_b32 s31, v40, 7
+; PRE-GFX10-NEXT:    v_readlane_b32 s30, v40, 6
+; PRE-GFX10-NEXT:    v_readlane_b32 s9, v40, 5
+; PRE-GFX10-NEXT:    v_readlane_b32 s8, v40, 4
+; PRE-GFX10-NEXT:    v_readlane_b32 s7, v40, 3
+; PRE-GFX10-NEXT:    v_readlane_b32 s6, v40, 2
+; PRE-GFX10-NEXT:    v_readlane_b32 s5, v40, 1
+; PRE-GFX10-NEXT:    v_readlane_b32 s4, v40, 0
+; PRE-GFX10-NEXT:    s_mov_b32 s32, s33
+; PRE-GFX10-NEXT:    v_readlane_b32 s34, v40, 8
+; PRE-GFX10-NEXT:    s_or_saveexec_b64 s[36:37], -1
+; PRE-GFX10-NEXT:    buffer_load_dword v40, off, s[0:3], s33 ; 4-byte Folded Reload
+; PRE-GFX10-NEXT:    s_mov_b64 exec, s[36:37]
+; PRE-GFX10-NEXT:    s_mov_b32 s33, s34
+; PRE-GFX10-NEXT:    s_waitcnt vmcnt(0)
+; PRE-GFX10-NEXT:    s_setpc_b64 s[30:31]
+;
+; GFX10-32-LABEL: test_indirect_call_vgpr_ptr_arg_and_reuse:
+; GFX10-32:       ; %bb.0:
+; GFX10-32-NEXT:    s_waitcnt vmcnt(0) expcnt(0) lgkmcnt(0)
+; GFX10-32-NEXT:    s_mov_b32 s34, s33
+; GFX10-32-NEXT:    s_mov_b32 s33, s32
+; GFX10-32-NEXT:    s_or_saveexec_b32 s35, -1
+; GFX10-32-NEXT:    buffer_store_dword v40, off, s[0:3], s33 ; 4-byte Folded Spill
+; GFX10-32-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-NEXT:    s_mov_b32 exec_lo, s35
+; GFX10-32-NEXT:    v_writelane_b32 v40, s34, 6
+; GFX10-32-NEXT:    s_addk_i32 s32, 0x200
+; GFX10-32-NEXT:    v_writelane_b32 v40, s4, 0
+; GFX10-32-NEXT:    v_writelane_b32 v40, s5, 1
+; GFX10-32-NEXT:    s_mov_b32 s5, 0
+; GFX10-32-NEXT:    v_writelane_b32 v40, s6, 2
+; GFX10-32-NEXT:    s_mov_b32 s6, exec_lo
+; GFX10-32-NEXT:    v_writelane_b32 v40, s7, 3
+; GFX10-32-NEXT:    s_mov_b32 s7, exec_lo
+; GFX10-32-NEXT:    v_writelane_b32 v40, s30, 4
+; GFX10-32-NEXT:    v_writelane_b32 v40, s31, 5
+; GFX10-32-NEXT:  .LBB18_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-NEXT:    v_readfirstlane_b32 s4, v1
+; GFX10-32-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-NEXT:    v_cmpx_eq_u32_e32 s4, v1
+; GFX10-32-NEXT:    s_swappc_b64 s[30:31], s[4:5]
+; GFX10-32-NEXT:    v_mov_b32_e32 v2, v0
+; GFX10-32-NEXT:    s_andn2_wrexec_b32 s7, s7
+; GFX10-32-NEXT:    ; implicit-def: $vgpr1
+; GFX10-32-NEXT:    ; implicit-def: $vgpr0
+; GFX10-32-NEXT:    s_cbranch_execnz .LBB18_1
+; GFX10-32-NEXT:  ; %bb.2:
+; GFX10-32-NEXT:    s_mov_b32 exec_lo, s6
+; GFX10-32-NEXT:    v_mov_b32_e32 v0, v2
+; GFX10-32-NEXT:    v_readlane_b32 s31, v40, 5
+; GFX10-32-NEXT:    v_readlane_b32 s30, v40, 4
+; GFX10-32-NEXT:    v_readlane_b32 s7, v40, 3
+; GFX10-32-NEXT:    v_readlane_b32 s6, v40, 2
+; GFX10-32-NEXT:    v_readlane_b32 s5, v40, 1
+; GFX10-32-NEXT:    v_readlane_b32 s4, v40, 0
+; GFX10-32-NEXT:    s_mov_b32 s32, s33
+; GFX10-32-NEXT:    v_readlane_b32 s34, v40, 6
+; GFX10-32-NEXT:    s_or_saveexec_b32 s35, -1
+; GFX10-32-NEXT:    buffer_load_dword v40, off, s[0:3], s33 ; 4-byte Folded Reload
+; GFX10-32-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-NEXT:    s_mov_b32 exec_lo, s35
+; GFX10-32-NEXT:    s_mov_b32 s33, s34
+; GFX10-32-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-NEXT:    s_setpc_b64 s[30:31]
+;
+; GFX10-64-LABEL: test_indirect_call_vgpr_ptr_arg_and_reuse:
+; GFX10-64:       ; %bb.0:
+; GFX10-64-NEXT:    s_waitcnt vmcnt(0) expcnt(0) lgkmcnt(0)
+; GFX10-64-NEXT:    s_mov_b32 s34, s33
+; GFX10-64-NEXT:    s_mov_b32 s33, s32
+; GFX10-64-NEXT:    s_or_saveexec_b64 s[36:37], -1
+; GFX10-64-NEXT:    buffer_store_dword v40, off, s[0:3], s33 ; 4-byte Folded Spill
+; GFX10-64-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-NEXT:    s_mov_b64 exec, s[36:37]
+; GFX10-64-NEXT:    v_writelane_b32 v40, s34, 8
+; GFX10-64-NEXT:    s_addk_i32 s32, 0x400
+; GFX10-64-NEXT:    v_writelane_b32 v40, s4, 0
+; GFX10-64-NEXT:    v_writelane_b32 v40, s5, 1
+; GFX10-64-NEXT:    s_mov_b32 s5, 0
+; GFX10-64-NEXT:    v_writelane_b32 v40, s6, 2
+; GFX10-64-NEXT:    v_writelane_b32 v40, s7, 3
+; GFX10-64-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-NEXT:    v_writelane_b32 v40, s8, 4
+; GFX10-64-NEXT:    v_writelane_b32 v40, s9, 5
+; GFX10-64-NEXT:    s_mov_b64 s[8:9], exec
+; GFX10-64-NEXT:    v_writelane_b32 v40, s30, 6
+; GFX10-64-NEXT:    v_writelane_b32 v40, s31, 7
+; GFX10-64-NEXT:  .LBB18_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-NEXT:    v_readfirstlane_b32 s4, v1
+; GFX10-64-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-NEXT:    v_cmpx_eq_u32_e64 s4, v1
+; GFX10-64-NEXT:    s_swappc_b64 s[30:31], s[4:5]
+; GFX10-64-NEXT:    v_mov_b32_e32 v2, v0
+; GFX10-64-NEXT:    s_andn2_wrexec_b64 s[8:9], s[8:9]
+; GFX10-64-NEXT:    ; implicit-def: $vgpr1
+; GFX10-64-NEXT:    ; implicit-def: $vgpr0
+; GFX10-64-NEXT:    s_cbranch_execnz .LBB18_1
+; GFX10-64-NEXT:  ; %bb.2:
+; GFX10-64-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX10-64-NEXT:    v_mov_b32_e32 v0, v2
+; GFX10-64-NEXT:    v_readlane_b32 s31, v40, 7
+; GFX10-64-NEXT:    v_readlane_b32 s30, v40, 6
+; GFX10-64-NEXT:    v_readlane_b32 s9, v40, 5
+; GFX10-64-NEXT:    v_readlane_b32 s8, v40, 4
+; GFX10-64-NEXT:    v_readlane_b32 s7, v40, 3
+; GFX10-64-NEXT:    v_readlane_b32 s6, v40, 2
+; GFX10-64-NEXT:    v_readlane_b32 s5, v40, 1
+; GFX10-64-NEXT:    v_readlane_b32 s4, v40, 0
+; GFX10-64-NEXT:    s_mov_b32 s32, s33
+; GFX10-64-NEXT:    v_readlane_b32 s34, v40, 8
+; GFX10-64-NEXT:    s_or_saveexec_b64 s[36:37], -1
+; GFX10-64-NEXT:    buffer_load_dword v40, off, s[0:3], s33 ; 4-byte Folded Reload
+; GFX10-64-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-NEXT:    s_mov_b64 exec, s[36:37]
+; GFX10-64-NEXT:    s_mov_b32 s33, s34
+; GFX10-64-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-NEXT:    s_setpc_b64 s[30:31]
+;
+; GFX1150-LABEL: test_indirect_call_vgpr_ptr_arg_and_reuse:
+; GFX1150:       ; %bb.0:
+; GFX1150-NEXT:    s_waitcnt vmcnt(0) expcnt(0) lgkmcnt(0)
+; GFX1150-NEXT:    s_mov_b32 s0, s33
+; GFX1150-NEXT:    s_mov_b32 s33, s32
+; GFX1150-NEXT:    s_or_saveexec_b64 s[2:3], -1
+; GFX1150-NEXT:    scratch_store_b32 off, v40, s33 ; 4-byte Folded Spill
+; GFX1150-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX1150-NEXT:    v_writelane_b32 v40, s0, 8
+; GFX1150-NEXT:    s_add_i32 s32, s32, 16
+; GFX1150-NEXT:    v_writelane_b32 v40, s4, 0
+; GFX1150-NEXT:    v_writelane_b32 v40, s5, 1
+; GFX1150-NEXT:    s_mov_b32 s5, 0
+; GFX1150-NEXT:    v_writelane_b32 v40, s6, 2
+; GFX1150-NEXT:    v_writelane_b32 v40, s7, 3
+; GFX1150-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-NEXT:    v_writelane_b32 v40, s8, 4
+; GFX1150-NEXT:    v_writelane_b32 v40, s9, 5
+; GFX1150-NEXT:    s_mov_b64 s[8:9], exec
+; GFX1150-NEXT:    v_writelane_b32 v40, s30, 6
+; GFX1150-NEXT:    v_writelane_b32 v40, s31, 7
+; GFX1150-NEXT:  .LBB18_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-NEXT:    v_readfirstlane_b32 s4, v1
+; GFX1150-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-NEXT:    v_cmpx_eq_u32_e64 s4, v1
+; GFX1150-NEXT:    s_swappc_b64 s[30:31], s[4:5]
+; GFX1150-NEXT:    v_mov_b32_e32 v2, v0
+; GFX1150-NEXT:    s_and_not1_wrexec_b64 s[8:9], s[8:9]
+; GFX1150-NEXT:    ; implicit-def: $vgpr1
+; GFX1150-NEXT:    ; implicit-def: $vgpr0
+; GFX1150-NEXT:    s_cbranch_execnz .LBB18_1
+; GFX1150-NEXT:  ; %bb.2:
+; GFX1150-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX1150-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-NEXT:    v_mov_b32_e32 v0, v2
+; GFX1150-NEXT:    v_readlane_b32 s31, v40, 7
+; GFX1150-NEXT:    v_readlane_b32 s30, v40, 6
+; GFX1150-NEXT:    v_readlane_b32 s9, v40, 5
+; GFX1150-NEXT:    v_readlane_b32 s8, v40, 4
+; GFX1150-NEXT:    v_readlane_b32 s7, v40, 3
+; GFX1150-NEXT:    v_readlane_b32 s6, v40, 2
+; GFX1150-NEXT:    v_readlane_b32 s5, v40, 1
+; GFX1150-NEXT:    v_readlane_b32 s4, v40, 0
+; GFX1150-NEXT:    s_mov_b32 s32, s33
+; GFX1150-NEXT:    v_readlane_b32 s0, v40, 8
+; GFX1150-NEXT:    s_or_saveexec_b64 s[2:3], -1
+; GFX1150-NEXT:    scratch_load_b32 v40, off, s33 ; 4-byte Folded Reload
+; GFX1150-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX1150-NEXT:    s_mov_b32 s33, s0
+; GFX1150-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-NEXT:    s_setpc_b64 s[30:31]
+;
+; GFX12-LABEL: test_indirect_call_vgpr_ptr_arg_and_reuse:
+; GFX12:       ; %bb.0:
+; GFX12-NEXT:    s_wait_loadcnt_dscnt 0x0
+; GFX12-NEXT:    s_wait_expcnt 0x0
+; GFX12-NEXT:    s_wait_samplecnt 0x0
+; GFX12-NEXT:    s_wait_bvhcnt 0x0
+; GFX12-NEXT:    s_wait_kmcnt 0x0
+; GFX12-NEXT:    s_mov_b32 s0, s33
+; GFX12-NEXT:    s_mov_b32 s33, s32
+; GFX12-NEXT:    s_or_saveexec_b64 s[2:3], -1
+; GFX12-NEXT:    scratch_store_b32 off, v40, s33 ; 4-byte Folded Spill
+; GFX12-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX12-NEXT:    v_writelane_b32 v40, s0, 8
+; GFX12-NEXT:    s_add_co_i32 s32, s32, 16
+; GFX12-NEXT:    v_writelane_b32 v40, s4, 0
+; GFX12-NEXT:    v_writelane_b32 v40, s5, 1
+; GFX12-NEXT:    s_mov_b32 s5, 0
+; GFX12-NEXT:    v_writelane_b32 v40, s6, 2
+; GFX12-NEXT:    v_writelane_b32 v40, s7, 3
+; GFX12-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-NEXT:    v_writelane_b32 v40, s8, 4
+; GFX12-NEXT:    v_writelane_b32 v40, s9, 5
+; GFX12-NEXT:    s_mov_b64 s[8:9], exec
+; GFX12-NEXT:    v_writelane_b32 v40, s30, 6
+; GFX12-NEXT:    v_writelane_b32 v40, s31, 7
+; GFX12-NEXT:  .LBB18_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-NEXT:    v_readfirstlane_b32 s4, v1
+; GFX12-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-NEXT:    v_cmpx_eq_u32_e64 s4, v1
+; GFX12-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-NEXT:    s_swappc_b64 s[30:31], s[4:5]
+; GFX12-NEXT:    v_mov_b32_e32 v2, v0
+; GFX12-NEXT:    s_and_not1_wrexec_b64 s[8:9], s[8:9]
+; GFX12-NEXT:    ; implicit-def: $vgpr1
+; GFX12-NEXT:    ; implicit-def: $vgpr0
+; GFX12-NEXT:    s_cbranch_execnz .LBB18_1
+; GFX12-NEXT:  ; %bb.2:
+; GFX12-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX12-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-NEXT:    v_mov_b32_e32 v0, v2
+; GFX12-NEXT:    v_readlane_b32 s31, v40, 7
+; GFX12-NEXT:    v_readlane_b32 s30, v40, 6
+; GFX12-NEXT:    v_readlane_b32 s9, v40, 5
+; GFX12-NEXT:    v_readlane_b32 s8, v40, 4
+; GFX12-NEXT:    v_readlane_b32 s7, v40, 3
+; GFX12-NEXT:    v_readlane_b32 s6, v40, 2
+; GFX12-NEXT:    v_readlane_b32 s5, v40, 1
+; GFX12-NEXT:    v_readlane_b32 s4, v40, 0
+; GFX12-NEXT:    s_mov_b32 s32, s33
+; GFX12-NEXT:    v_readlane_b32 s0, v40, 8
+; GFX12-NEXT:    s_or_saveexec_b64 s[2:3], -1
+; GFX12-NEXT:    scratch_load_b32 v40, off, s33 ; 4-byte Folded Reload
+; GFX12-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX12-NEXT:    s_mov_b32 s33, s0
+; GFX12-NEXT:    s_wait_loadcnt 0x0
+; GFX12-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-NEXT:    s_setpc_b64 s[30:31]
+  %tok = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %fptr)
+  %s_fptr = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %tok, i32 %fptr)
+  %ext = zext i32 %s_fptr to i64
+  %f = inttoptr i64 %ext to i32(i32)*
+  %callres = call amdgpu_gfx i32 %f(i32 %i)
+  %r3 = call i32 @llvm.amdgcn.waterfall.end.i32(i32 %tok, i32 %callres)
+  ret i32 %r3
+}
+
+define amdgpu_cs void @ds_write_8(i8 %value, i32 %index) #1 {
+; VI-LABEL: ds_write_8:
+; VI:       ; %bb.0: ; %.entry
+; VI-NEXT:    v_lshlrev_b32_e32 v1, 2, v1
+; VI-NEXT:    v_add_u32_e32 v1, vcc, Lds at abs32@lo, v1
+; VI-NEXT:    s_mov_b32 m0, -1
+; VI-NEXT:    s_mov_b64 s[0:1], exec
+; VI-NEXT:  .LBB19_1: ; =>This Inner Loop Header: Depth=1
+; VI-NEXT:    v_readfirstlane_b32 s0, v1
+; VI-NEXT:    v_cmp_eq_u32_e64 s[0:1], s0, v1
+; VI-NEXT:    s_and_saveexec_b64 s[0:1], s[0:1]
+; VI-NEXT:    ds_write_b8 v1, v0
+; VI-NEXT:    ; implicit-def: $vgpr1
+; VI-NEXT:    ; implicit-def: $vgpr0
+; VI-NEXT:    s_xor_b64 exec, exec, s[0:1]
+; VI-NEXT:    s_cbranch_execnz .LBB19_1
+; VI-NEXT:  ; %bb.2:
+; VI-NEXT:    s_endpgm
+;
+; GFX9-SDAG-LABEL: ds_write_8:
+; GFX9-SDAG:       ; %bb.0: ; %.entry
+; GFX9-SDAG-NEXT:    s_mov_b32 s0, Lds at abs32@lo
+; GFX9-SDAG-NEXT:    v_lshl_add_u32 v1, v1, 2, s0
+; GFX9-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX9-SDAG-NEXT:  .LBB19_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s0, v1
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[0:1], s0, v1
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[0:1], s[0:1]
+; GFX9-SDAG-NEXT:    ds_write_b8 v1, v0
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr1
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr0
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[0:1]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB19_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_endpgm
+;
+; GFX9-GISEL-LABEL: ds_write_8:
+; GFX9-GISEL:       ; %bb.0: ; %.entry
+; GFX9-GISEL-NEXT:    v_lshlrev_b32_e32 v1, 2, v1
+; GFX9-GISEL-NEXT:    v_add_u32_e32 v1, Lds at abs32@lo, v1
+; GFX9-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX9-GISEL-NEXT:  .LBB19_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s0, v1
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[0:1], s0, v1
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[0:1], s[0:1]
+; GFX9-GISEL-NEXT:    ds_write_b8 v1, v0
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr1
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr0
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[0:1]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB19_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_endpgm
+;
+; GFX10-32-SDAG-LABEL: ds_write_8:
+; GFX10-32-SDAG:       ; %bb.0: ; %.entry
+; GFX10-32-SDAG-NEXT:    v_lshl_add_u32 v1, v1, 2, Lds at abs32@lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-SDAG-NEXT:  .LBB19_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s1, v1
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s1, v1
+; GFX10-32-SDAG-NEXT:    ds_write_b8 v1, v0
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s0, s0
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr1
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr0
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB19_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_endpgm
+;
+; GFX10-32-GISEL-LABEL: ds_write_8:
+; GFX10-32-GISEL:       ; %bb.0: ; %.entry
+; GFX10-32-GISEL-NEXT:    v_lshlrev_b32_e32 v1, 2, v1
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-GISEL-NEXT:    v_add_nc_u32_e32 v1, Lds at abs32@lo, v1
+; GFX10-32-GISEL-NEXT:  .LBB19_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s1, v1
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s1, v1
+; GFX10-32-GISEL-NEXT:    ds_write_b8 v1, v0
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s0, s0
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr1
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr0
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB19_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_endpgm
+;
+; GFX10-64-SDAG-LABEL: ds_write_8:
+; GFX10-64-SDAG:       ; %bb.0: ; %.entry
+; GFX10-64-SDAG-NEXT:    v_lshl_add_u32 v1, v1, 2, Lds at abs32@lo
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-SDAG-NEXT:  .LBB19_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s2, v1
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s2, v1
+; GFX10-64-SDAG-NEXT:    ds_write_b8 v1, v0
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[0:1], s[0:1]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr1
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr0
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB19_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_endpgm
+;
+; GFX10-64-GISEL-LABEL: ds_write_8:
+; GFX10-64-GISEL:       ; %bb.0: ; %.entry
+; GFX10-64-GISEL-NEXT:    v_lshlrev_b32_e32 v1, 2, v1
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-GISEL-NEXT:    v_add_nc_u32_e32 v1, Lds at abs32@lo, v1
+; GFX10-64-GISEL-NEXT:  .LBB19_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s2, v1
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s2, v1
+; GFX10-64-GISEL-NEXT:    ds_write_b8 v1, v0
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[0:1], s[0:1]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr1
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr0
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB19_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_endpgm
+;
+; GFX1150-SDAG-LABEL: ds_write_8:
+; GFX1150-SDAG:       ; %bb.0: ; %.entry
+; GFX1150-SDAG-NEXT:    v_lshl_add_u32 v1, v1, 2, Lds at abs32@lo
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-SDAG-NEXT:  .LBB19_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s2, v1
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s2, v1
+; GFX1150-SDAG-NEXT:    ds_store_b8 v1, v0
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[0:1], s[0:1]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr1
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr0
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB19_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_endpgm
+;
+; GFX1150-GISEL-LABEL: ds_write_8:
+; GFX1150-GISEL:       ; %bb.0: ; %.entry
+; GFX1150-GISEL-NEXT:    v_lshlrev_b32_e32 v1, 2, v1
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-GISEL-NEXT:    v_add_nc_u32_e32 v1, Lds at abs32@lo, v1
+; GFX1150-GISEL-NEXT:  .LBB19_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s2, v1
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s2, v1
+; GFX1150-GISEL-NEXT:    ds_store_b8 v1, v0
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[0:1], s[0:1]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr1
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr0
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB19_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_endpgm
+;
+; GFX12-SDAG-LABEL: ds_write_8:
+; GFX12-SDAG:       ; %bb.0: ; %.entry
+; GFX12-SDAG-NEXT:    v_lshl_add_u32 v1, v1, 2, Lds at abs32@lo
+; GFX12-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-SDAG-NEXT:  .LBB19_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_1) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s2, v1
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s2, v1
+; GFX12-SDAG-NEXT:    ds_store_b8 v1, v0
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[0:1], s[0:1]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr1
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr0
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB19_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_endpgm
+;
+; GFX12-GISEL-LABEL: ds_write_8:
+; GFX12-GISEL:       ; %bb.0: ; %.entry
+; GFX12-GISEL-NEXT:    v_lshlrev_b32_e32 v1, 2, v1
+; GFX12-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_add_nc_u32_e32 v1, Lds at abs32@lo, v1
+; GFX12-GISEL-NEXT:  .LBB19_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_1) | instid1(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s2, v1
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s2, v1
+; GFX12-GISEL-NEXT:    ds_store_b8 v1, v0
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[0:1], s[0:1]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr1
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr0
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB19_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_endpgm
+.entry:
+  %gep = getelementptr [16384 x i32], ptr addrspace(3) @Lds, i32 0, i32 %index
+  %0 = call i32 @llvm.amdgcn.waterfall.begin.p3(i32 0, ptr addrspace(3) %gep)
+  %1 = call ptr addrspace(3) @llvm.amdgcn.waterfall.last.use.vgpr.p3(i32 %0, ptr addrspace(3) %gep)
+  store i8 %value, ptr addrspace(3) %1, align 1
+  ret void
+}
+
+define amdgpu_ps {<4 x float>,<4 x float>} @test_waterfall_multi_end_1loop(
+; VI-SDAG-LABEL: test_waterfall_multi_end_1loop:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; VI-SDAG-NEXT:    s_wqm_b64 exec, exec
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, v2
+; VI-SDAG-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; VI-SDAG-NEXT:    v_mov_b32_e32 v8, v1
+; VI-SDAG-NEXT:    v_mov_b32_e32 v9, v0
+; VI-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[4:5]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v2, s1
+; VI-SDAG-NEXT:    v_add_u32_e32 v0, vcc, s0, v0
+; VI-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; VI-SDAG-NEXT:    v_addc_u32_e32 v1, vcc, v2, v1, vcc
+; VI-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[3:4]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, s3
+; VI-SDAG-NEXT:    v_add_u32_e32 v2, vcc, s2, v2
+; VI-SDAG-NEXT:    v_addc_u32_e32 v3, vcc, v4, v3, vcc
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[10:13], v[0:1]
+; VI-SDAG-NEXT:    v_add_u32_e32 v0, vcc, 16, v0
+; VI-SDAG-NEXT:    v_addc_u32_e32 v1, vcc, 0, v1, vcc
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[14:17], v[0:1]
+; VI-SDAG-NEXT:    flat_load_dwordx4 v[18:21], v[2:3]
+; VI-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; VI-SDAG-NEXT:  .LBB20_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s2, v9
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s6, v8
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v9
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[6:7], s6, v8
+; VI-SDAG-NEXT:    s_and_b64 s[2:3], s[2:3], s[6:7]
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; VI-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, 0.5
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(2)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s8, v10
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s9, v11
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s10, v12
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s11, v13
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s12, v14
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s13, v15
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s14, v16
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s15, v17
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s16, v18
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s17, v19
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s18, v20
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s19, v21
+; VI-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; VI-SDAG-NEXT:    v_mov_b32_e32 v5, v4
+; VI-SDAG-NEXT:    s_nop 2
+; VI-SDAG-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[16:19] dmask:0xf
+; VI-SDAG-NEXT:    image_sample v[4:7], v[4:5], s[8:15], s[16:19] dmask:0xf
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr9
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr8
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15_vgpr16_vgpr17
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr18_vgpr19_vgpr20_vgpr21
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB20_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; VI-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    ; return to shader part epilog
+;
+; VI-GISEL-LABEL: test_waterfall_multi_end_1loop:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; VI-GISEL-NEXT:    s_wqm_b64 exec, exec
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, v3
+; VI-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; VI-GISEL-NEXT:    v_mov_b32_e32 v8, v0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v9, v1
+; VI-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[2:3]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; VI-GISEL-NEXT:    v_add_u32_e32 v0, vcc, v2, v0
+; VI-GISEL-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; VI-GISEL-NEXT:    v_addc_u32_e32 v1, vcc, v3, v1, vcc
+; VI-GISEL-NEXT:    v_lshlrev_b64 v[2:3], 4, v[4:5]
+; VI-GISEL-NEXT:    v_mov_b32_e32 v5, s3
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, s2
+; VI-GISEL-NEXT:    v_add_u32_e32 v2, vcc, v4, v2
+; VI-GISEL-NEXT:    v_addc_u32_e32 v3, vcc, v5, v3, vcc
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[10:13], v[0:1]
+; VI-GISEL-NEXT:    v_add_u32_e32 v0, vcc, 16, v0
+; VI-GISEL-NEXT:    v_addc_u32_e32 v1, vcc, 0, v1, vcc
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[14:17], v[0:1]
+; VI-GISEL-NEXT:    flat_load_dwordx4 v[18:21], v[2:3]
+; VI-GISEL-NEXT:    s_mov_b32 s0, 0
+; VI-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; VI-GISEL-NEXT:    s_mov_b32 s20, 0.5
+; VI-GISEL-NEXT:  .LBB20_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s1, v8
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s8, v9
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[6:7], s1, v8
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[8:9], s8, v9
+; VI-GISEL-NEXT:    s_and_b64 s[6:7], s[6:7], s[8:9]
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[6:7], s[6:7]
+; VI-GISEL-NEXT:    s_mov_b32 s1, s0
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; VI-GISEL-NEXT:    v_mov_b32_e32 v0, s0
+; VI-GISEL-NEXT:    s_mov_b32 s21, s20
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, s20
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s8, v10
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s9, v11
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s10, v12
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s11, v13
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s12, v14
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s13, v15
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s14, v16
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s15, v17
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s16, v18
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s17, v19
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s18, v20
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s19, v21
+; VI-GISEL-NEXT:    v_mov_b32_e32 v1, s1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v5, s21
+; VI-GISEL-NEXT:    s_nop 2
+; VI-GISEL-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[16:19] dmask:0xf
+; VI-GISEL-NEXT:    image_sample v[4:7], v[4:5], s[8:15], s[16:19] dmask:0xf
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr8
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr9
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15_vgpr16_vgpr17
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr18_vgpr19_vgpr20_vgpr21
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[6:7]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB20_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; VI-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX9-SDAG-LABEL: test_waterfall_multi_end_1loop:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX9-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v4, v2
+; GFX9-SDAG-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v8, v1
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v9, v0
+; GFX9-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[4:5]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v2, s1
+; GFX9-SDAG-NEXT:    v_add_co_u32_e32 v0, vcc, s0, v0
+; GFX9-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX9-SDAG-NEXT:    v_addc_co_u32_e32 v1, vcc, v2, v1, vcc
+; GFX9-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[3:4]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v4, s3
+; GFX9-SDAG-NEXT:    v_add_co_u32_e32 v2, vcc, s2, v2
+; GFX9-SDAG-NEXT:    v_addc_co_u32_e32 v3, vcc, v4, v3, vcc
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[14:17], v[0:1], off offset:16
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[10:13], v[0:1], off
+; GFX9-SDAG-NEXT:    global_load_dwordx4 v[18:21], v[2:3], off
+; GFX9-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:  .LBB20_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s2, v9
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s6, v8
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v9
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[6:7], s6, v8
+; GFX9-SDAG-NEXT:    s_and_b64 s[2:3], s[2:3], s[6:7]
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v8, 0
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s8, v10
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s9, v11
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s10, v12
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s11, v13
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s12, v14
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s13, v15
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s14, v16
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s15, v17
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s16, v18
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s17, v19
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s18, v20
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s19, v21
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v9, v8
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v10, 0.5
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v11, v10
+; GFX9-SDAG-NEXT:    s_nop 1
+; GFX9-SDAG-NEXT:    image_sample v[0:3], v[8:9], s[8:15], s[16:19] dmask:0xf
+; GFX9-SDAG-NEXT:    image_sample v[4:7], v[10:11], s[8:15], s[16:19] dmask:0xf
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr9
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr8
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15_vgpr16_vgpr17
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr18_vgpr19_vgpr20_vgpr21
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB20_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX9-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX9-GISEL-LABEL: test_waterfall_multi_end_1loop:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX9-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v4, v3
+; GFX9-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v8, v0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v9, v1
+; GFX9-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[2:3]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX9-GISEL-NEXT:    v_add_co_u32_e32 v0, vcc, v2, v0
+; GFX9-GISEL-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; GFX9-GISEL-NEXT:    v_addc_co_u32_e32 v1, vcc, v3, v1, vcc
+; GFX9-GISEL-NEXT:    v_lshlrev_b64 v[2:3], 4, v[4:5]
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v5, s3
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v4, s2
+; GFX9-GISEL-NEXT:    v_add_co_u32_e32 v2, vcc, v4, v2
+; GFX9-GISEL-NEXT:    v_addc_co_u32_e32 v3, vcc, v5, v3, vcc
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[10:13], v[0:1], off
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[14:17], v[0:1], off offset:16
+; GFX9-GISEL-NEXT:    global_load_dwordx4 v[18:21], v[2:3], off
+; GFX9-GISEL-NEXT:    s_mov_b32 s0, 0
+; GFX9-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX9-GISEL-NEXT:    s_mov_b32 s20, 0.5
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:  .LBB20_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s1, v8
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s8, v9
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[6:7], s1, v8
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[8:9], s8, v9
+; GFX9-GISEL-NEXT:    s_and_b64 s[6:7], s[6:7], s[8:9]
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[6:7], s[6:7]
+; GFX9-GISEL-NEXT:    s_mov_b32 s1, s0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v9, s1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s8, v10
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s9, v11
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s10, v12
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s11, v13
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s12, v14
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s13, v15
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s14, v16
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s15, v17
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s16, v18
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s17, v19
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s18, v20
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s19, v21
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v8, s0
+; GFX9-GISEL-NEXT:    s_mov_b32 s21, s20
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v10, s20
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v11, s21
+; GFX9-GISEL-NEXT:    s_nop 0
+; GFX9-GISEL-NEXT:    image_sample v[0:3], v[8:9], s[8:15], s[16:19] dmask:0xf
+; GFX9-GISEL-NEXT:    image_sample v[4:7], v[10:11], s[8:15], s[16:19] dmask:0xf
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr8
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15_vgpr16_vgpr17
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr18_vgpr19_vgpr20_vgpr21
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[6:7]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB20_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[2:3]
+; GFX9-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-SDAG-LABEL: test_waterfall_multi_end_1loop:
+; GFX10-32-SDAG:       ; %bb.0:
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-SDAG-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v5, v2
+; GFX10-32-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v8, v1
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v9, v0
+; GFX10-32-SDAG-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX10-32-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[3:4]
+; GFX10-32-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[5:6]
+; GFX10-32-SDAG-NEXT:    v_add_co_u32 v0, vcc_lo, s0, v0
+; GFX10-32-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, s1, v1, vcc_lo
+; GFX10-32-SDAG-NEXT:    v_add_co_u32 v2, vcc_lo, s2, v2
+; GFX10-32-SDAG-NEXT:    v_add_co_ci_u32_e32 v3, vcc_lo, s3, v3, vcc_lo
+; GFX10-32-SDAG-NEXT:    s_clause 0x1
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[14:17], v[0:1], off offset:16
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[10:13], v[0:1], off
+; GFX10-32-SDAG-NEXT:    global_load_dwordx4 v[18:21], v[2:3], off
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:  .LBB20_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s2, v9
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s3, v8
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s2, v9
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s3, v8
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s8, v10
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s9, v11
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s10, v12
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s11, v13
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s12, v14
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s13, v15
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s14, v16
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s15, v17
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s16, v18
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s17, v19
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s18, v20
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s19, v21
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v8, 0
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v9, 0.5
+; GFX10-32-SDAG-NEXT:    image_sample v[0:3], [v8, v8], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-SDAG-NEXT:    image_sample v[4:7], [v9, v9], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr9
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr8
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15_vgpr16_vgpr17
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr18_vgpr19_vgpr20_vgpr21
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB20_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-SDAG-NEXT:    s_and_b32 exec_lo, exec_lo, s4
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-GISEL-LABEL: test_waterfall_multi_end_1loop:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-GISEL-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v4, v3
+; GFX10-32-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v8, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v9, v1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v7, s3
+; GFX10-32-GISEL-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; GFX10-32-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[2:3]
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v6, s2
+; GFX10-32-GISEL-NEXT:    v_lshlrev_b64 v[4:5], 4, v[4:5]
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-GISEL-NEXT:    v_add_co_u32 v0, vcc_lo, v2, v0
+; GFX10-32-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc_lo, v3, v1, vcc_lo
+; GFX10-32-GISEL-NEXT:    v_add_co_u32 v2, vcc_lo, v6, v4
+; GFX10-32-GISEL-NEXT:    v_add_co_ci_u32_e32 v3, vcc_lo, v7, v5, vcc_lo
+; GFX10-32-GISEL-NEXT:    s_clause 0x1
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[10:13], v[0:1], off
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[14:17], v[0:1], off offset:16
+; GFX10-32-GISEL-NEXT:    global_load_dwordx4 v[18:21], v[2:3], off
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:  .LBB20_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s2, v8
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s3, v9
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s2, v8
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s3, v9
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s8, v10
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s9, v11
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s10, v12
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s11, v13
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s12, v14
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s13, v15
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s14, v16
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s15, v17
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s16, v18
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s17, v19
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s18, v20
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s19, v21
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v8, 0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v9, 0.5
+; GFX10-32-GISEL-NEXT:    image_sample v[0:3], [v8, v8], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-GISEL-NEXT:    image_sample v[4:7], [v9, v9], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr8
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15_vgpr16_vgpr17
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr18_vgpr19_vgpr20_vgpr21
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB20_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-GISEL-NEXT:    s_and_b32 exec_lo, exec_lo, s4
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-SDAG-LABEL: test_waterfall_multi_end_1loop:
+; GFX10-64-SDAG:       ; %bb.0:
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v5, v2
+; GFX10-64-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v8, v1
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v9, v0
+; GFX10-64-SDAG-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX10-64-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[3:4]
+; GFX10-64-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[5:6]
+; GFX10-64-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX10-64-SDAG-NEXT:    v_add_co_ci_u32_e32 v1, vcc, s1, v1, vcc
+; GFX10-64-SDAG-NEXT:    v_add_co_u32 v2, vcc, s2, v2
+; GFX10-64-SDAG-NEXT:    v_add_co_ci_u32_e32 v3, vcc, s3, v3, vcc
+; GFX10-64-SDAG-NEXT:    s_clause 0x1
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[14:17], v[0:1], off offset:16
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[10:13], v[0:1], off
+; GFX10-64-SDAG-NEXT:    global_load_dwordx4 v[18:21], v[2:3], off
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:  .LBB20_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s6, v9
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s7, v8
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v9
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s7, v8
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s8, v10
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s9, v11
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s10, v12
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s11, v13
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s12, v14
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s13, v15
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s14, v16
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s15, v17
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s16, v18
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s17, v19
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s18, v20
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s19, v21
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v8, 0
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v9, 0.5
+; GFX10-64-SDAG-NEXT:    image_sample v[0:3], [v8, v8], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-SDAG-NEXT:    image_sample v[4:7], [v9, v9], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr9
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr8
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15_vgpr16_vgpr17
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr18_vgpr19_vgpr20_vgpr21
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB20_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-GISEL-LABEL: test_waterfall_multi_end_1loop:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v4, v3
+; GFX10-64-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v8, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v9, v1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v7, s3
+; GFX10-64-GISEL-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; GFX10-64-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[2:3]
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v6, s2
+; GFX10-64-GISEL-NEXT:    v_lshlrev_b64 v[4:5], 4, v[4:5]
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX10-64-GISEL-NEXT:    v_add_co_ci_u32_e32 v1, vcc, v3, v1, vcc
+; GFX10-64-GISEL-NEXT:    v_add_co_u32 v2, vcc, v6, v4
+; GFX10-64-GISEL-NEXT:    v_add_co_ci_u32_e32 v3, vcc, v7, v5, vcc
+; GFX10-64-GISEL-NEXT:    s_clause 0x1
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[10:13], v[0:1], off
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[14:17], v[0:1], off offset:16
+; GFX10-64-GISEL-NEXT:    global_load_dwordx4 v[18:21], v[2:3], off
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:  .LBB20_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s6, v8
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s7, v9
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s6, v8
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s7, v9
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s8, v10
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s9, v11
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s10, v12
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s11, v13
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s12, v14
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s13, v15
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s14, v16
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s15, v17
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s16, v18
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s17, v19
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s18, v20
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s19, v21
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v8, 0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v9, 0.5
+; GFX10-64-GISEL-NEXT:    image_sample v[0:3], [v8, v8], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-GISEL-NEXT:    image_sample v[4:7], [v9, v9], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr8
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15_vgpr16_vgpr17
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr18_vgpr19_vgpr20_vgpr21
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB20_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX1150-SDAG-LABEL: test_waterfall_multi_end_1loop:
+; GFX1150-SDAG:       ; %bb.0:
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v5, v2
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v8, v1
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v9, v0
+; GFX1150-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_4) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-SDAG-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX1150-SDAG-NEXT:    v_lshlrev_b64 v[2:3], 4, v[3:4]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_lshlrev_b64 v[0:1], 5, v[5:6]
+; GFX1150-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_4)
+; GFX1150-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s1, v1, vcc
+; GFX1150-SDAG-NEXT:    v_add_co_u32 v2, vcc, s2, v2
+; GFX1150-SDAG-NEXT:    s_waitcnt_depctr depctr_va_vcc(0)
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-SDAG-NEXT:    v_add_co_ci_u32_e64 v3, null, s3, v3, vcc
+; GFX1150-SDAG-NEXT:    s_clause 0x1
+; GFX1150-SDAG-NEXT:    global_load_b128 v[14:17], v[0:1], off offset:16
+; GFX1150-SDAG-NEXT:    global_load_b128 v[10:13], v[0:1], off
+; GFX1150-SDAG-NEXT:    global_load_b128 v[18:21], v[2:3], off
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-SDAG-NEXT:  .LBB20_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s6, v9
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s7, v8
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v9
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s7, v8
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s8, v10
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s9, v11
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s10, v12
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s11, v13
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s12, v14
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s13, v15
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s14, v16
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s15, v17
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s16, v18
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s17, v19
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s18, v20
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s19, v21
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v4, 0.5
+; GFX1150-SDAG-NEXT:    s_clause 0x1
+; GFX1150-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-SDAG-NEXT:    image_sample v[4:7], [v4, v4], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr9
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr8
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15_vgpr16_vgpr17
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr18_vgpr19_vgpr20_vgpr21
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB20_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX1150-GISEL-LABEL: test_waterfall_multi_end_1loop:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v4, v3
+; GFX1150-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v8, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v9, v1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v7, s3
+; GFX1150-GISEL-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; GFX1150-GISEL-NEXT:    v_lshlrev_b64 v[0:1], 5, v[2:3]
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v6, s2
+; GFX1150-GISEL-NEXT:    v_lshlrev_b64 v[4:5], 4, v[4:5]
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_3)
+; GFX1150-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v3, v1, vcc
+; GFX1150-GISEL-NEXT:    v_add_co_u32 v2, vcc, v6, v4
+; GFX1150-GISEL-NEXT:    s_waitcnt_depctr depctr_va_vcc(0)
+; GFX1150-GISEL-NEXT:    v_add_co_ci_u32_e64 v3, null, v7, v5, vcc
+; GFX1150-GISEL-NEXT:    s_clause 0x1
+; GFX1150-GISEL-NEXT:    global_load_b128 v[10:13], v[0:1], off
+; GFX1150-GISEL-NEXT:    global_load_b128 v[14:17], v[0:1], off offset:16
+; GFX1150-GISEL-NEXT:    global_load_b128 v[18:21], v[2:3], off
+; GFX1150-GISEL-NEXT:  .LBB20_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s6, v8
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s7, v9
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s6, v8
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s7, v9
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(2)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s8, v10
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s9, v11
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s10, v12
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s11, v13
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s12, v14
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s13, v15
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s14, v16
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s15, v17
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s16, v18
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s17, v19
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s18, v20
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s19, v21
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v4, 0.5
+; GFX1150-GISEL-NEXT:    s_clause 0x1
+; GFX1150-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-GISEL-NEXT:    image_sample v[4:7], [v4, v4], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr8
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15_vgpr16_vgpr17
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr18_vgpr19_vgpr20_vgpr21
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB20_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX12-SDAG-LABEL: test_waterfall_multi_end_1loop:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v5, v2
+; GFX12-SDAG-NEXT:    v_ashrrev_i32_e32 v4, 31, v3
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v8, v1
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v9, v0
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_4) | instskip(NEXT) | instid1(VALU_DEP_4)
+; GFX12-SDAG-NEXT:    v_ashrrev_i32_e32 v6, 31, v5
+; GFX12-SDAG-NEXT:    v_lshlrev_b64_e32 v[2:3], 4, v[3:4]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_lshlrev_b64_e32 v[0:1], 5, v[5:6]
+; GFX12-SDAG-NEXT:    v_add_co_u32 v0, vcc, s0, v0
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_4)
+; GFX12-SDAG-NEXT:    v_add_co_ci_u32_e64 v1, null, s1, v1, vcc
+; GFX12-SDAG-NEXT:    v_add_co_u32 v2, vcc, s2, v2
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_vcc(0)
+; GFX12-SDAG-NEXT:    v_add_co_ci_u32_e64 v3, null, s3, v3, vcc
+; GFX12-SDAG-NEXT:    s_clause 0x1
+; GFX12-SDAG-NEXT:    global_load_b128 v[14:17], v[0:1], off offset:16
+; GFX12-SDAG-NEXT:    global_load_b128 v[10:13], v[0:1], off
+; GFX12-SDAG-NEXT:    global_load_b128 v[18:21], v[2:3], off
+; GFX12-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-SDAG-NEXT:  .LBB20_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s6, v9
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s7, v8
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s6, v9
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s7, v8
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s8, v10
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s9, v11
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s10, v12
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s11, v13
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s12, v14
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s13, v15
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s14, v16
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s15, v17
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x0
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s16, v18
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s17, v19
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s18, v20
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s19, v21
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x1
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v4, 0.5
+; GFX12-SDAG-NEXT:    s_clause 0x1
+; GFX12-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-SDAG-NEXT:    image_sample v[4:7], [v4, v4], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr9
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr8
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15_vgpr16_vgpr17
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr18_vgpr19_vgpr20_vgpr21
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB20_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX12-GISEL-LABEL: test_waterfall_multi_end_1loop:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v4, v3
+; GFX12-GISEL-NEXT:    v_ashrrev_i32_e32 v3, 31, v2
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v8, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v9, v1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v7, s3
+; GFX12-GISEL-NEXT:    v_ashrrev_i32_e32 v5, 31, v4
+; GFX12-GISEL-NEXT:    v_lshlrev_b64_e32 v[0:1], 5, v[2:3]
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v3, s1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v2, s0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v6, s2
+; GFX12-GISEL-NEXT:    v_lshlrev_b64_e32 v[4:5], 4, v[4:5]
+; GFX12-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-GISEL-NEXT:    v_add_co_u32 v0, vcc, v2, v0
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_3)
+; GFX12-GISEL-NEXT:    v_add_co_ci_u32_e64 v1, null, v3, v1, vcc
+; GFX12-GISEL-NEXT:    v_add_co_u32 v2, vcc, v6, v4
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_vcc(0)
+; GFX12-GISEL-NEXT:    v_add_co_ci_u32_e64 v3, null, v7, v5, vcc
+; GFX12-GISEL-NEXT:    s_clause 0x1
+; GFX12-GISEL-NEXT:    global_load_b128 v[10:13], v[0:1], off
+; GFX12-GISEL-NEXT:    global_load_b128 v[14:17], v[0:1], off offset:16
+; GFX12-GISEL-NEXT:    global_load_b128 v[18:21], v[2:3], off
+; GFX12-GISEL-NEXT:  .LBB20_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s6, v8
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s7, v9
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s6, v8
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s7, v9
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x2
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s8, v10
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s9, v11
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s10, v12
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s11, v13
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s12, v14
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s13, v15
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s14, v16
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s15, v17
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x0
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s16, v18
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s17, v19
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s18, v20
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s19, v21
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v4, 0.5
+; GFX12-GISEL-NEXT:    s_clause 0x1
+; GFX12-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-GISEL-NEXT:    image_sample v[4:7], [v4, v4], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr8
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr10_vgpr11_vgpr12_vgpr13_vgpr14_vgpr15_vgpr16_vgpr17
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr18_vgpr19_vgpr20_vgpr21
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB20_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    ; return to shader part epilog
+       <8 x i32> addrspace(4)* inreg %in, <4 x i32> addrspace(4)* inreg %s_in,
+       i32 %idx1, i32 %idx2, i32 %s_idx, i32 %s_idx2) #1 {
+  %rptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx
+  %sptr = getelementptr <4 x i32>, <4 x i32> addrspace(4)* %s_in, i32 %s_idx2
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4)* %rptr, align 16
+  %srsrc = load <4 x i32>, <4 x i32> addrspace(4)* %sptr, align 16
+  %tok = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %idx1)
+  %tok1 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 %tok, i32 %idx2)
+  %s_rsrc = call <8 x i32> @llvm.amdgcn.waterfall.readfirstlane.v8i32.v8i32(i32 %tok1, <8 x i32> %rsrc)
+  %s_srsrc = call <4 x i32> @llvm.amdgcn.waterfall.readfirstlane.v4i32.v4i32(i32 %tok1, <4 x i32> %srsrc)
+  %r = call <4 x float> @llvm.amdgcn.image.sample.2d.v4f32.f32(i32 15, float 0.000000e+00, float 0.000000e+00, <8 x i32> %s_rsrc, <4 x i32> %s_srsrc, i1 false, i32 0, i32 0)
+  %r2 = call <4 x float> @llvm.amdgcn.image.sample.2d.v4f32.f32(i32 15, float 0.500000e+00, float 0.500000e+00, <8 x i32> %s_rsrc, <4 x i32> %s_srsrc, i1 false, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok1, <4 x float> %r)
+  %r3 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok1, <4 x float> %r2)
+
+  %insert = insertvalue { <4 x float>, <4 x float> } undef, <4 x float> %r1, 0
+  %insert1 = insertvalue { <4 x float>, <4 x float> } %insert, <4 x float> %r3, 1
+  ret {<4 x float> , <4 x float>} %insert1
+}
+
+define amdgpu_ps {<4 x float>,<4 x float>} @test_waterfall_multi_end_1loop_rsrc_load_inside(
+; VI-SDAG-LABEL: test_waterfall_multi_end_1loop_rsrc_load_inside:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; VI-SDAG-NEXT:    s_wqm_b64 exec, exec
+; VI-SDAG-NEXT:    v_mov_b32_e32 v8, v3
+; VI-SDAG-NEXT:    v_mov_b32_e32 v9, v2
+; VI-SDAG-NEXT:    v_mov_b32_e32 v10, v1
+; VI-SDAG-NEXT:    v_mov_b32_e32 v11, v0
+; VI-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; VI-SDAG-NEXT:  .LBB21_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s8, v11
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s10, v10
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[8:9], s8, v11
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[10:11], s10, v10
+; VI-SDAG-NEXT:    s_and_b64 s[8:9], s[8:9], s[10:11]
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[20:21], s[8:9]
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s8, v9
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s10, v8
+; VI-SDAG-NEXT:    s_ashr_i32 s9, s8, 31
+; VI-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; VI-SDAG-NEXT:    s_lshl_b64 s[8:9], s[8:9], 5
+; VI-SDAG-NEXT:    s_add_u32 s8, s0, s8
+; VI-SDAG-NEXT:    s_addc_u32 s9, s1, s9
+; VI-SDAG-NEXT:    s_lshl_b64 s[10:11], s[10:11], 4
+; VI-SDAG-NEXT:    s_add_u32 s16, s2, s10
+; VI-SDAG-NEXT:    s_addc_u32 s17, s3, s11
+; VI-SDAG-NEXT:    s_load_dwordx8 s[8:15], s[8:9], 0x0
+; VI-SDAG-NEXT:    s_load_dwordx4 s[16:19], s[16:17], 0x0
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; VI-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, 0.5
+; VI-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; VI-SDAG-NEXT:    v_mov_b32_e32 v5, v4
+; VI-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; VI-SDAG-NEXT:    image_sample v[0:3], v[0:1], s[8:15], s[16:19] dmask:0xf
+; VI-SDAG-NEXT:    image_sample v[4:7], v[4:5], s[8:15], s[16:19] dmask:0xf
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr11
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr10
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr9
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr8
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[20:21]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB21_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[6:7]
+; VI-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    ; return to shader part epilog
+;
+; VI-GISEL-LABEL: test_waterfall_multi_end_1loop_rsrc_load_inside:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; VI-GISEL-NEXT:    s_wqm_b64 exec, exec
+; VI-GISEL-NEXT:    s_mov_b32 s6, 0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v8, v0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v11, v1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v9, v2
+; VI-GISEL-NEXT:    v_mov_b32_e32 v10, v3
+; VI-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; VI-GISEL-NEXT:    s_mov_b32 s24, 0.5
+; VI-GISEL-NEXT:  .LBB21_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s7, v8
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s12, v11
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[10:11], s7, v8
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[12:13], s12, v11
+; VI-GISEL-NEXT:    s_and_b64 s[10:11], s[10:11], s[12:13]
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[10:11], s[10:11]
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s12, v9
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s14, v10
+; VI-GISEL-NEXT:    s_ashr_i32 s13, s12, 31
+; VI-GISEL-NEXT:    s_ashr_i32 s15, s14, 31
+; VI-GISEL-NEXT:    s_lshl_b64 s[12:13], s[12:13], 5
+; VI-GISEL-NEXT:    s_add_u32 s12, s0, s12
+; VI-GISEL-NEXT:    s_addc_u32 s13, s1, s13
+; VI-GISEL-NEXT:    s_lshl_b64 s[14:15], s[14:15], 4
+; VI-GISEL-NEXT:    s_add_u32 s20, s2, s14
+; VI-GISEL-NEXT:    s_addc_u32 s21, s3, s15
+; VI-GISEL-NEXT:    s_load_dwordx8 s[12:19], s[12:13], 0x0
+; VI-GISEL-NEXT:    s_load_dwordx4 s[20:23], s[20:21], 0x0
+; VI-GISEL-NEXT:    s_mov_b32 s7, s6
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; VI-GISEL-NEXT:    v_mov_b32_e32 v0, s6
+; VI-GISEL-NEXT:    s_mov_b32 s25, s24
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, s24
+; VI-GISEL-NEXT:    v_mov_b32_e32 v1, s7
+; VI-GISEL-NEXT:    v_mov_b32_e32 v5, s25
+; VI-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; VI-GISEL-NEXT:    image_sample v[0:3], v[0:1], s[12:19], s[20:23] dmask:0xf
+; VI-GISEL-NEXT:    image_sample v[4:7], v[4:5], s[12:19], s[20:23] dmask:0xf
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr8
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr11
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr9
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr10
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[10:11]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB21_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[8:9]
+; VI-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX9-SDAG-LABEL: test_waterfall_multi_end_1loop_rsrc_load_inside:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX9-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v8, v3
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v9, v2
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v10, v1
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v11, v0
+; GFX9-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX9-SDAG-NEXT:  .LBB21_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s8, v11
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s10, v10
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[8:9], s8, v11
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[10:11], s10, v10
+; GFX9-SDAG-NEXT:    s_and_b64 s[8:9], s[8:9], s[10:11]
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[20:21], s[8:9]
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s8, v9
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s10, v8
+; GFX9-SDAG-NEXT:    s_ashr_i32 s9, s8, 31
+; GFX9-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX9-SDAG-NEXT:    s_lshl_b64 s[8:9], s[8:9], 5
+; GFX9-SDAG-NEXT:    s_add_u32 s22, s0, s8
+; GFX9-SDAG-NEXT:    s_addc_u32 s23, s1, s9
+; GFX9-SDAG-NEXT:    s_lshl_b64 s[8:9], s[10:11], 4
+; GFX9-SDAG-NEXT:    s_add_u32 s24, s2, s8
+; GFX9-SDAG-NEXT:    s_addc_u32 s25, s3, s9
+; GFX9-SDAG-NEXT:    s_load_dwordx8 s[8:15], s[22:23], 0x0
+; GFX9-SDAG-NEXT:    s_load_dwordx4 s[16:19], s[24:25], 0x0
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v8, 0
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v9, v8
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v10, 0.5
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v11, v10
+; GFX9-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX9-SDAG-NEXT:    image_sample v[0:3], v[8:9], s[8:15], s[16:19] dmask:0xf
+; GFX9-SDAG-NEXT:    image_sample v[4:7], v[10:11], s[8:15], s[16:19] dmask:0xf
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr11
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr10
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr9
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr8
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[20:21]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB21_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX9-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX9-GISEL-LABEL: test_waterfall_multi_end_1loop_rsrc_load_inside:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX9-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX9-GISEL-NEXT:    s_mov_b32 s6, 0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v8, v0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v11, v1
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v9, v2
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v10, v3
+; GFX9-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX9-GISEL-NEXT:  .LBB21_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s7, v8
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[10:11], s7, v8
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[12:13], s12, v11
+; GFX9-GISEL-NEXT:    s_and_b64 s[10:11], s[10:11], s[12:13]
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[10:11], s[10:11]
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s12, v9
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s14, v10
+; GFX9-GISEL-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX9-GISEL-NEXT:    s_ashr_i32 s15, s14, 31
+; GFX9-GISEL-NEXT:    s_lshl_b64 s[12:13], s[12:13], 5
+; GFX9-GISEL-NEXT:    s_add_u32 s24, s0, s12
+; GFX9-GISEL-NEXT:    s_addc_u32 s25, s1, s13
+; GFX9-GISEL-NEXT:    s_lshl_b64 s[12:13], s[14:15], 4
+; GFX9-GISEL-NEXT:    s_add_u32 s26, s2, s12
+; GFX9-GISEL-NEXT:    s_addc_u32 s27, s3, s13
+; GFX9-GISEL-NEXT:    s_load_dwordx8 s[12:19], s[24:25], 0x0
+; GFX9-GISEL-NEXT:    s_load_dwordx4 s[20:23], s[26:27], 0x0
+; GFX9-GISEL-NEXT:    s_mov_b32 s7, s6
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v9, s7
+; GFX9-GISEL-NEXT:    s_mov_b32 s24, 0.5
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v8, s6
+; GFX9-GISEL-NEXT:    s_mov_b32 s25, s24
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v10, s24
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v11, s25
+; GFX9-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX9-GISEL-NEXT:    image_sample v[0:3], v[8:9], s[12:19], s[20:23] dmask:0xf
+; GFX9-GISEL-NEXT:    image_sample v[4:7], v[10:11], s[12:19], s[20:23] dmask:0xf
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr8
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr11
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr10
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[10:11]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB21_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[8:9]
+; GFX9-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-SDAG-LABEL: test_waterfall_multi_end_1loop_rsrc_load_inside:
+; GFX10-32-SDAG:       ; %bb.0:
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-SDAG-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v8, v3
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v9, v2
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v10, v1
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v11, v0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s6, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s5, exec_lo
+; GFX10-32-SDAG-NEXT:  .LBB21_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s7, v11
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s8, v10
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s7, v11
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s8, v10
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s8, v9
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s10, v8
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v8, 0
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v9, 0.5
+; GFX10-32-SDAG-NEXT:    s_ashr_i32 s9, s8, 31
+; GFX10-32-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX10-32-SDAG-NEXT:    s_lshl_b64 s[8:9], s[8:9], 5
+; GFX10-32-SDAG-NEXT:    s_add_u32 s20, s0, s8
+; GFX10-32-SDAG-NEXT:    s_addc_u32 s21, s1, s9
+; GFX10-32-SDAG-NEXT:    s_lshl_b64 s[8:9], s[10:11], 4
+; GFX10-32-SDAG-NEXT:    s_add_u32 s22, s2, s8
+; GFX10-32-SDAG-NEXT:    s_addc_u32 s23, s3, s9
+; GFX10-32-SDAG-NEXT:    s_load_dwordx8 s[8:15], s[20:21], 0x0
+; GFX10-32-SDAG-NEXT:    s_load_dwordx4 s[16:19], s[22:23], 0x0
+; GFX10-32-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-32-SDAG-NEXT:    image_sample v[0:3], [v8, v8], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-SDAG-NEXT:    image_sample v[4:7], [v9, v9], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s6, s6
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr11
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr10
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr9
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr8
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB21_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s5
+; GFX10-32-SDAG-NEXT:    s_and_b32 exec_lo, exec_lo, s4
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-GISEL-LABEL: test_waterfall_multi_end_1loop_rsrc_load_inside:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v8, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v9, v1
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s4, exec_lo
+; GFX10-32-GISEL-NEXT:    s_wqm_b32 exec_lo, exec_lo
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v10, v2
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v11, v3
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s6, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s5, exec_lo
+; GFX10-32-GISEL-NEXT:  .LBB21_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s7, v8
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s8, v9
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s7, v8
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s8, v9
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s8, v10
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s10, v11
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v8, 0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v9, 0.5
+; GFX10-32-GISEL-NEXT:    s_ashr_i32 s9, s8, 31
+; GFX10-32-GISEL-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX10-32-GISEL-NEXT:    s_lshl_b64 s[8:9], s[8:9], 5
+; GFX10-32-GISEL-NEXT:    s_add_u32 s20, s0, s8
+; GFX10-32-GISEL-NEXT:    s_addc_u32 s21, s1, s9
+; GFX10-32-GISEL-NEXT:    s_lshl_b64 s[8:9], s[10:11], 4
+; GFX10-32-GISEL-NEXT:    s_add_u32 s22, s2, s8
+; GFX10-32-GISEL-NEXT:    s_addc_u32 s23, s3, s9
+; GFX10-32-GISEL-NEXT:    s_load_dwordx8 s[8:15], s[20:21], 0x0
+; GFX10-32-GISEL-NEXT:    s_load_dwordx4 s[16:19], s[22:23], 0x0
+; GFX10-32-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-32-GISEL-NEXT:    image_sample v[0:3], [v8, v8], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-GISEL-NEXT:    image_sample v[4:7], [v9, v9], s[8:15], s[16:19] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s6, s6
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr8
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr10
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr11
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB21_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s5
+; GFX10-32-GISEL-NEXT:    s_and_b32 exec_lo, exec_lo, s4
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-SDAG-LABEL: test_waterfall_multi_end_1loop_rsrc_load_inside:
+; GFX10-64-SDAG:       ; %bb.0:
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v8, v3
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v9, v2
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v10, v1
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v11, v0
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-SDAG-NEXT:  .LBB21_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s10, v11
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s11, v10
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s10, v11
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s11, v10
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s10, v9
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s12, v8
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v8, 0
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v9, 0.5
+; GFX10-64-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX10-64-SDAG-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX10-64-SDAG-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX10-64-SDAG-NEXT:    s_add_u32 s10, s0, s10
+; GFX10-64-SDAG-NEXT:    s_addc_u32 s11, s1, s11
+; GFX10-64-SDAG-NEXT:    s_lshl_b64 s[12:13], s[12:13], 4
+; GFX10-64-SDAG-NEXT:    s_add_u32 s24, s2, s12
+; GFX10-64-SDAG-NEXT:    s_addc_u32 s25, s3, s13
+; GFX10-64-SDAG-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
+; GFX10-64-SDAG-NEXT:    s_load_dwordx4 s[20:23], s[24:25], 0x0
+; GFX10-64-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-64-SDAG-NEXT:    image_sample v[0:3], [v8, v8], s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-SDAG-NEXT:    image_sample v[4:7], [v9, v9], s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[8:9], s[8:9]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr11
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr10
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr9
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr8
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB21_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX10-64-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-GISEL-LABEL: test_waterfall_multi_end_1loop_rsrc_load_inside:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v8, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v9, v1
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX10-64-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v10, v2
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v11, v3
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX10-64-GISEL-NEXT:  .LBB21_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s10, v8
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s11, v9
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s10, v8
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s11, v9
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s10, v10
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v8, 0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v9, 0.5
+; GFX10-64-GISEL-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX10-64-GISEL-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX10-64-GISEL-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX10-64-GISEL-NEXT:    s_add_u32 s10, s0, s10
+; GFX10-64-GISEL-NEXT:    s_addc_u32 s11, s1, s11
+; GFX10-64-GISEL-NEXT:    s_lshl_b64 s[12:13], s[12:13], 4
+; GFX10-64-GISEL-NEXT:    s_add_u32 s24, s2, s12
+; GFX10-64-GISEL-NEXT:    s_addc_u32 s25, s3, s13
+; GFX10-64-GISEL-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
+; GFX10-64-GISEL-NEXT:    s_load_dwordx4 s[20:23], s[24:25], 0x0
+; GFX10-64-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-64-GISEL-NEXT:    image_sample v[0:3], [v8, v8], s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-GISEL-NEXT:    image_sample v[4:7], [v9, v9], s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[8:9], s[8:9]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr8
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr10
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr11
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB21_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX10-64-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX1150-SDAG-LABEL: test_waterfall_multi_end_1loop_rsrc_load_inside:
+; GFX1150-SDAG:       ; %bb.0:
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v8, v3
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v9, v2
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v10, v1
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v11, v0
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-SDAG-NEXT:  .LBB21_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_1) | instid1(VALU_DEP_2)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s10, v11
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s11, v10
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s10, v11
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2)
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s11, v10
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s10, v9
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s12, v8
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(1)
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v4, 0.5
+; GFX1150-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX1150-SDAG-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX1150-SDAG-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(SKIP_2) | instid1(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_add_u32 s10, s0, s10
+; GFX1150-SDAG-NEXT:    s_addc_u32 s11, s1, s11
+; GFX1150-SDAG-NEXT:    s_lshl_b64 s[12:13], s[12:13], 4
+; GFX1150-SDAG-NEXT:    s_add_u32 s20, s2, s12
+; GFX1150-SDAG-NEXT:    s_addc_u32 s21, s3, s13
+; GFX1150-SDAG-NEXT:    s_load_b256 s[12:19], s[10:11], 0x0
+; GFX1150-SDAG-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX1150-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX1150-SDAG-NEXT:    s_clause 0x1
+; GFX1150-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-SDAG-NEXT:    image_sample v[4:7], [v4, v4], s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[8:9], s[8:9]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr11
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr10
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr9
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr8
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB21_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX1150-GISEL-LABEL: test_waterfall_multi_end_1loop_rsrc_load_inside:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v8, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v9, v1
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX1150-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v10, v2
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v11, v3
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX1150-GISEL-NEXT:  .LBB21_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s10, v8
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s11, v9
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s10, v8
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s11, v9
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s10, v10
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(1)
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v4, 0.5
+; GFX1150-GISEL-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX1150-GISEL-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX1150-GISEL-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1) | instskip(SKIP_2) | instid1(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_add_u32 s10, s0, s10
+; GFX1150-GISEL-NEXT:    s_addc_u32 s11, s1, s11
+; GFX1150-GISEL-NEXT:    s_lshl_b64 s[12:13], s[12:13], 4
+; GFX1150-GISEL-NEXT:    s_add_u32 s20, s2, s12
+; GFX1150-GISEL-NEXT:    s_addc_u32 s21, s3, s13
+; GFX1150-GISEL-NEXT:    s_load_b256 s[12:19], s[10:11], 0x0
+; GFX1150-GISEL-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX1150-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX1150-GISEL-NEXT:    s_clause 0x1
+; GFX1150-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-GISEL-NEXT:    image_sample v[4:7], [v4, v4], s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[8:9], s[8:9]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr8
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr10
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr11
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB21_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX12-SDAG-LABEL: test_waterfall_multi_end_1loop_rsrc_load_inside:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-SDAG-NEXT:    s_wqm_b64 exec, exec
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v8, v3
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v9, v2
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v10, v1
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v11, v0
+; GFX12-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX12-SDAG-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-SDAG-NEXT:  .LBB21_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_2) | instid1(VALU_DEP_2)
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s10, v11
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s11, v10
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s10, v11
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_2)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s11, v10
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s10, v9
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s12, v8
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x1
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v4, 0.5
+; GFX12-SDAG-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX12-SDAG-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX12-SDAG-NEXT:    s_lshl_b64 s[12:13], s[12:13], 4
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    s_add_nc_u64 s[10:11], s[0:1], s[10:11]
+; GFX12-SDAG-NEXT:    s_add_nc_u64 s[20:21], s[2:3], s[12:13]
+; GFX12-SDAG-NEXT:    s_load_b256 s[12:19], s[10:11], 0x0
+; GFX12-SDAG-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX12-SDAG-NEXT:    s_wait_kmcnt 0x0
+; GFX12-SDAG-NEXT:    s_clause 0x1
+; GFX12-SDAG-NEXT:    image_sample v[0:3], [v0, v0], s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-SDAG-NEXT:    image_sample v[4:7], [v4, v4], s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[8:9], s[8:9]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr11
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr10
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr9
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr8
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB21_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-SDAG-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX12-GISEL-LABEL: test_waterfall_multi_end_1loop_rsrc_load_inside:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v8, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v9, v1
+; GFX12-GISEL-NEXT:    s_mov_b64 s[4:5], exec
+; GFX12-GISEL-NEXT:    s_wqm_b64 exec, exec
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v10, v2
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v11, v3
+; GFX12-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[6:7], exec
+; GFX12-GISEL-NEXT:  .LBB21_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s10, v8
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s11, v9
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_2) | instskip(NEXT) | instid1(VALU_DEP_2)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s10, v8
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s11, v9
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s10, v10
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s12, v11
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v4, 0.5
+; GFX12-GISEL-NEXT:    s_ashr_i32 s11, s10, 31
+; GFX12-GISEL-NEXT:    s_ashr_i32 s13, s12, 31
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_lshl_b64 s[10:11], s[10:11], 5
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_add_co_u32 s10, s0, s10
+; GFX12-GISEL-NEXT:    s_add_co_ci_u32 s11, s1, s11
+; GFX12-GISEL-NEXT:    s_lshl_b64 s[12:13], s[12:13], 4
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_add_co_u32 s20, s2, s12
+; GFX12-GISEL-NEXT:    s_add_co_ci_u32 s21, s3, s13
+; GFX12-GISEL-NEXT:    s_load_b256 s[12:19], s[10:11], 0x0
+; GFX12-GISEL-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
+; GFX12-GISEL-NEXT:    s_wait_kmcnt 0x0
+; GFX12-GISEL-NEXT:    s_clause 0x1
+; GFX12-GISEL-NEXT:    image_sample v[0:3], [v0, v0], s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-GISEL-NEXT:    image_sample v[4:7], [v4, v4], s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_2D
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[8:9], s[8:9]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr8
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr10
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr11
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB21_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[6:7]
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX12-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    ; return to shader part epilog
+       <8 x i32> addrspace(4)* inreg %in, <4 x i32> addrspace(4)* inreg %s_in,
+       i32 %idx1, i32 %idx2, i32 %s_idx, i32 %s_idx2) #1 {
+  %tok = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %idx1)
+  %tok1 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 %tok, i32 %idx2)
+  %widx0 = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %tok1, i32 %s_idx)
+  %widx1 = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %tok1, i32 %s_idx2)
+  %widx0a = sext i32 %widx0 to i64
+  %widx1a = sext i32 %widx1 to i64
+  %rptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i64 %widx0a
+  %sptr = getelementptr <4 x i32>, <4 x i32> addrspace(4)* %s_in, i64 %widx1a
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4)* %rptr, align 32
+  %srsrc = load <4 x i32>, <4 x i32> addrspace(4)* %sptr, align 16
+
+  %r = call <4 x float> @llvm.amdgcn.image.sample.2d.v4f32.f32(i32 15, float 0.000000e+00, float 0.000000e+00, <8 x i32> %rsrc, <4 x i32> %srsrc, i1 false, i32 0, i32 0)
+  %r2 = call <4 x float> @llvm.amdgcn.image.sample.2d.v4f32.f32(i32 15, float 0.500000e+00, float 0.500000e+00, <8 x i32> %rsrc, <4 x i32> %srsrc, i1 false, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok1, <4 x float> %r)
+  %r3 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok1, <4 x float> %r2)
+
+  %insert = insertvalue { <4 x float>, <4 x float> } undef, <4 x float> %r1, 0
+  %insert1 = insertvalue { <4 x float>, <4 x float> } %insert, <4 x float> %r3, 1
+  ret {<4 x float> , <4 x float>} %insert1
+}
+
+define amdgpu_ps {<4 x float>,float} @test_waterfall_multi_end_struct(
+; VI-SDAG-LABEL: test_waterfall_multi_end_struct:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    v_mov_b32_e32 v9, v4
+; VI-SDAG-NEXT:    v_mov_b32_e32 v8, v3
+; VI-SDAG-NEXT:    v_mov_b32_e32 v7, v2
+; VI-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; VI-SDAG-NEXT:    v_mov_b32_e32 v10, v0
+; VI-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; VI-SDAG-NEXT:  .LBB22_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s2, v10
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v10
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s4, v6
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s5, v7
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s6, v8
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s7, v9
+; VI-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; VI-SDAG-NEXT:    v_mov_b32_e32 v2, v0
+; VI-SDAG-NEXT:    v_mov_b32_e32 v3, v0
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr10
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr6_vgpr7_vgpr8_vgpr9
+; VI-SDAG-NEXT:    s_nop 0
+; VI-SDAG-NEXT:    buffer_load_format_xyzw v[0:4], v5, s[4:7], 0 idxen tfe
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr5
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB22_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    ; return to shader part epilog
+;
+; VI-GISEL-LABEL: test_waterfall_multi_end_struct:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    v_mov_b32_e32 v10, v0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v7, v2
+; VI-GISEL-NEXT:    v_mov_b32_e32 v8, v3
+; VI-GISEL-NEXT:    v_mov_b32_e32 v9, v4
+; VI-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; VI-GISEL-NEXT:  .LBB22_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s2, v10
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v10
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s4, v6
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s5, v7
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s6, v8
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s7, v9
+; VI-GISEL-NEXT:    v_mov_b32_e32 v1, v0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v2, v0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v3, v0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr10
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr6_vgpr7_vgpr8_vgpr9
+; VI-GISEL-NEXT:    s_nop 0
+; VI-GISEL-NEXT:    buffer_load_format_xyzw v[0:4], v5, s[4:7], 0 idxen tfe
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr5
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB22_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX9-SDAG-LABEL: test_waterfall_multi_end_struct:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v9, v4
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v8, v3
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v7, v2
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v10, v0
+; GFX9-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX9-SDAG-NEXT:  .LBB22_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s2, v10
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v10
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s4, v6
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s5, v7
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s6, v8
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s7, v9
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v2, v0
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v3, v0
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr10
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr6_vgpr7_vgpr8_vgpr9
+; GFX9-SDAG-NEXT:    s_nop 0
+; GFX9-SDAG-NEXT:    buffer_load_format_xyzw v[0:4], v5, s[4:7], 0 idxen tfe
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB22_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX9-GISEL-LABEL: test_waterfall_multi_end_struct:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v10, v0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v7, v2
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v8, v3
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v9, v4
+; GFX9-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX9-GISEL-NEXT:  .LBB22_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s2, v10
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v10
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s4, v6
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s5, v7
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s6, v8
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s7, v9
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v1, v0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v2, v0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v3, v0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr10
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr6_vgpr7_vgpr8_vgpr9
+; GFX9-GISEL-NEXT:    s_nop 0
+; GFX9-GISEL-NEXT:    buffer_load_format_xyzw v[0:4], v5, s[4:7], 0 idxen tfe
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB22_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-SDAG-LABEL: test_waterfall_multi_end_struct:
+; GFX10-32-SDAG:       ; %bb.0:
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v9, v4
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v8, v3
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v7, v2
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v10, v0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-SDAG-NEXT:  .LBB22_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s2, v10
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s2, v10
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s4, v6
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s5, v7
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s6, v8
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s7, v9
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v2, v0
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v3, v0
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-32-SDAG-NEXT:    buffer_load_format_xyzw v[0:4], v5, s[4:7], 0 idxen tfe
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr10
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr6_vgpr7_vgpr8_vgpr9
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB22_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-GISEL-LABEL: test_waterfall_multi_end_struct:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v10, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v7, v2
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v8, v3
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v9, v4
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-GISEL-NEXT:  .LBB22_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s2, v10
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s2, v10
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s4, v6
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s5, v7
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s6, v8
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s7, v9
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v1, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v2, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v3, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-32-GISEL-NEXT:    buffer_load_format_xyzw v[0:4], v5, s[4:7], 0 idxen tfe
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr10
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr6_vgpr7_vgpr8_vgpr9
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB22_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-SDAG-LABEL: test_waterfall_multi_end_struct:
+; GFX10-64-SDAG:       ; %bb.0:
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v9, v4
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v8, v3
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v7, v2
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v10, v0
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-SDAG-NEXT:  .LBB22_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s4, v10
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s4, v10
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s4, v6
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s5, v7
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s6, v8
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s7, v9
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v2, v0
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v3, v0
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-64-SDAG-NEXT:    buffer_load_format_xyzw v[0:4], v5, s[4:7], 0 idxen tfe
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr10
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr6_vgpr7_vgpr8_vgpr9
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB22_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-GISEL-LABEL: test_waterfall_multi_end_struct:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v10, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v7, v2
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v8, v3
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v9, v4
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-GISEL-NEXT:  .LBB22_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s4, v10
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v10
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s4, v6
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s5, v7
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s6, v8
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s7, v9
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v1, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v2, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v3, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-64-GISEL-NEXT:    buffer_load_format_xyzw v[0:4], v5, s[4:7], 0 idxen tfe
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr10
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr6_vgpr7_vgpr8_vgpr9
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB22_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX1150-SDAG-LABEL: test_waterfall_multi_end_struct:
+; GFX1150-SDAG:       ; %bb.0:
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v9, v4
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v8, v3
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v7, v2
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v10, v0
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-SDAG-NEXT:  .LBB22_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s4, v10
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s4, v10
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s4, v6
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s5, v7
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s6, v8
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s7, v9
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v2, v0
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v3, v0
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX1150-SDAG-NEXT:    buffer_load_format_xyzw v[0:4], v5, s[4:7], 0 idxen tfe
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr10
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr6_vgpr7_vgpr8_vgpr9
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB22_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX1150-GISEL-LABEL: test_waterfall_multi_end_struct:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v10, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v7, v2
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v8, v3
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v9, v4
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-GISEL-NEXT:  .LBB22_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s4, v10
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v10
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s4, v6
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s5, v7
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s6, v8
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s7, v9
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v1, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v2, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v3, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX1150-GISEL-NEXT:    buffer_load_format_xyzw v[0:4], v5, s[4:7], 0 idxen tfe
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr10
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr6_vgpr7_vgpr8_vgpr9
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB22_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX12-SDAG-LABEL: test_waterfall_multi_end_struct:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v9, v4
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v8, v3
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v7, v2
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v6, v1
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v10, v0
+; GFX12-SDAG-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-SDAG-NEXT:  .LBB22_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_1) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s4, v10
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s4, v10
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x0
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s4, v6
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s5, v7
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s6, v8
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s7, v9
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v2, v0
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v3, v0
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-SDAG-NEXT:    buffer_load_format_xyzw v[0:4], v5, s[4:7], null idxen tfe
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr10
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr6_vgpr7_vgpr8_vgpr9
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr5
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB22_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x0
+; GFX12-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX12-GISEL-LABEL: test_waterfall_multi_end_struct:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v10, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v7, v2
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v8, v3
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v9, v4
+; GFX12-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-GISEL-NEXT:  .LBB22_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s4, v10
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v10
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s4, v6
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s5, v7
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s6, v8
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s7, v9
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v1, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v2, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v3, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-GISEL-NEXT:    buffer_load_format_xyzw v[0:4], v5, s[4:7], null idxen tfe
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr10
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr6_vgpr7_vgpr8_vgpr9
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr5
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB22_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x0
+; GFX12-GISEL-NEXT:    ; return to shader part epilog
+              i32 %idx, <4 x i32> %s_idx, i32 %v_inp) #1 {
+  %tok = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %idx)
+  %widx0 = call <4 x i32> @llvm.amdgcn.waterfall.readfirstlane.v4i32.v4i32(i32 %tok, <4 x i32> %s_idx)
+  %val = call { <4 x float>, i32 } @llvm.amdgcn.struct.buffer.load.format.sl_v4f32i32s(<4 x i32> %widx0, i32 %v_inp, i32 0, i32 0, i32 0)
+  %payl = extractvalue { <4 x float>, i32 } %val, 0
+  %tfe = extractvalue { <4 x float>, i32 } %val, 1
+  %r0 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok, <4 x float> %payl)
+  %r1 = call i32 @llvm.amdgcn.waterfall.end.i32(i32 %tok, i32 %tfe)
+  %r1.float = bitcast i32 %r1 to float
+
+  %insert = insertvalue { <4 x float>, float } undef, <4 x float> %r0, 0
+  %insert1 = insertvalue { <4 x float>, float } %insert, float %r1.float, 1
+  ret {<4 x float> , float} %insert1
+}
+
+define amdgpu_ps {<4 x float>,float} @test_waterfall_multi_end_struct_uniform(
+; VI-SDAG-LABEL: test_waterfall_multi_end_struct_uniform:
+; VI-SDAG:       ; %bb.0:
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s0, v0
+; VI-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; VI-SDAG-NEXT:    v_mov_b32_e32 v5, v4
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s1, v1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s2, v2
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s3, v3
+; VI-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; VI-SDAG-NEXT:    v_mov_b32_e32 v2, v0
+; VI-SDAG-NEXT:    v_mov_b32_e32 v3, v0
+; VI-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; VI-SDAG-NEXT:    s_nop 0
+; VI-SDAG-NEXT:    buffer_load_format_xyzw v[0:4], v5, s[0:3], 0 idxen tfe
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; VI-SDAG-NEXT:    ; return to shader part epilog
+;
+; VI-GISEL-LABEL: test_waterfall_multi_end_struct_uniform:
+; VI-GISEL:       ; %bb.0:
+; VI-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; VI-GISEL-NEXT:    v_mov_b32_e32 v7, v2
+; VI-GISEL-NEXT:    v_mov_b32_e32 v8, v3
+; VI-GISEL-NEXT:    v_mov_b32_e32 v9, v4
+; VI-GISEL-NEXT:    v_mov_b32_e32 v10, s0
+; VI-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; VI-GISEL-NEXT:  .LBB23_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s2, v10
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v10
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s4, v5
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s5, v6
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s6, v7
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s7, v8
+; VI-GISEL-NEXT:    v_mov_b32_e32 v1, v0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v2, v0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v3, v0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr10
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8
+; VI-GISEL-NEXT:    s_nop 0
+; VI-GISEL-NEXT:    buffer_load_format_xyzw v[0:4], v9, s[4:7], 0 idxen tfe
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr9
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB23_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; VI-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX9-SDAG-LABEL: test_waterfall_multi_end_struct_uniform:
+; GFX9-SDAG:       ; %bb.0:
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s0, v0
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v5, v4
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s1, v1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s2, v2
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s3, v3
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v2, v0
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v3, v0
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX9-SDAG-NEXT:    s_nop 0
+; GFX9-SDAG-NEXT:    buffer_load_format_xyzw v[0:4], v5, s[0:3], 0 idxen tfe
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX9-GISEL-LABEL: test_waterfall_multi_end_struct_uniform:
+; GFX9-GISEL:       ; %bb.0:
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v7, v2
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v8, v3
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v9, v4
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v10, s0
+; GFX9-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX9-GISEL-NEXT:  .LBB23_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s2, v10
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[2:3], s2, v10
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[2:3], s[2:3]
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s4, v5
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s5, v6
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s6, v7
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s7, v8
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v1, v0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v2, v0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v3, v0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr10
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8
+; GFX9-GISEL-NEXT:    s_nop 0
+; GFX9-GISEL-NEXT:    buffer_load_format_xyzw v[0:4], v9, s[4:7], 0 idxen tfe
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[2:3]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB23_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX9-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10GFX11-SDAG-LABEL: test_waterfall_multi_end_struct_uniform:
+; GFX10GFX11-SDAG:       ; %bb.0:
+; GFX10GFX11-SDAG-NEXT:    v_mov_b32_e32 v5, v4
+; GFX10GFX11-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10GFX11-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10GFX11-SDAG-NEXT:    v_readfirstlane_b32 s1, v1
+; GFX10GFX11-SDAG-NEXT:    v_readfirstlane_b32 s2, v2
+; GFX10GFX11-SDAG-NEXT:    v_readfirstlane_b32 s3, v3
+; GFX10GFX11-SDAG-NEXT:    v_readfirstlane_b32 s0, v4
+; GFX10GFX11-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; GFX10GFX11-SDAG-NEXT:    v_mov_b32_e32 v2, v0
+; GFX10GFX11-SDAG-NEXT:    v_mov_b32_e32 v3, v0
+; GFX10GFX11-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10GFX11-SDAG-NEXT:    buffer_load_format_xyzw v[0:4], v5, s[0:3], 0 idxen tfe
+; GFX10GFX11-SDAG-NEXT:    s_waitcnt vmcnt(0)
+; GFX10GFX11-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX10-32-GISEL-LABEL: test_waterfall_multi_end_struct_uniform:
+; GFX10-32-GISEL:       ; %bb.0:
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v7, v2
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v8, v3
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v9, v4
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v10, s0
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s1, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s0, exec_lo
+; GFX10-32-GISEL-NEXT:  .LBB23_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s2, v10
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s2, v10
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s4, v5
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s5, v6
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s6, v7
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s7, v8
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v1, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v2, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v3, v0
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-32-GISEL-NEXT:    buffer_load_format_xyzw v[0:4], v9, s[4:7], 0 idxen tfe
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s1, s1
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr10
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB23_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s0
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-32-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX10-64-GISEL-LABEL: test_waterfall_multi_end_struct_uniform:
+; GFX10-64-GISEL:       ; %bb.0:
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v7, v2
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v8, v3
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v9, v4
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v10, s0
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX10-64-GISEL-NEXT:  .LBB23_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s4, v10
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v10
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s4, v5
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s5, v6
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s6, v7
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s7, v8
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v1, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v2, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v3, v0
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX10-64-GISEL-NEXT:    buffer_load_format_xyzw v[0:4], v9, s[4:7], 0 idxen tfe
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[2:3], s[2:3]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr10
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB23_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX10-64-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX1150-GISEL-LABEL: test_waterfall_multi_end_struct_uniform:
+; GFX1150-GISEL:       ; %bb.0:
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v7, v2
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v8, v3
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v9, v4
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v10, s0
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-GISEL-NEXT:  .LBB23_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(NEXT) | instid1(VALU_DEP_1)
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s4, v10
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v10
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s4, v5
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s5, v6
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s6, v7
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s7, v8
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v1, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v2, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v3, v0
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX1150-GISEL-NEXT:    buffer_load_format_xyzw v[0:4], v9, s[4:7], 0 idxen tfe
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr10
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB23_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0)
+; GFX1150-GISEL-NEXT:    ; return to shader part epilog
+;
+; GFX12-SDAG-LABEL: test_waterfall_multi_end_struct_uniform:
+; GFX12-SDAG:       ; %bb.0:
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v5, v4
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s1, v1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s2, v2
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s3, v3
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s0, v4
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v2, v0
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v3, v0
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-SDAG-NEXT:    buffer_load_format_xyzw v[0:4], v5, s[0:3], null idxen tfe
+; GFX12-SDAG-NEXT:    s_wait_loadcnt 0x0
+; GFX12-SDAG-NEXT:    ; return to shader part epilog
+;
+; GFX12-GISEL-LABEL: test_waterfall_multi_end_struct_uniform:
+; GFX12-GISEL:       ; %bb.0:
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v5, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v6, v1
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v7, v2
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v8, v3
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v9, v4
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v10, s0
+; GFX12-GISEL-NEXT:    s_mov_b64 s[2:3], exec
+; GFX12-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-GISEL-NEXT:  .LBB23_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1) | instskip(SKIP_1) | instid1(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s4, v10
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s4, v10
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s4, v5
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s5, v6
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s6, v7
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s7, v8
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v1, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v2, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v3, v0
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v4, v0
+; GFX12-GISEL-NEXT:    buffer_load_format_xyzw v[0:4], v9, s[4:7], null idxen tfe
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[2:3], s[2:3]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr10
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr5_vgpr6_vgpr7_vgpr8
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr9
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB23_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
+; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x0
+; GFX12-GISEL-NEXT:    ; return to shader part epilog
+              i32 inreg %idx, <4 x i32> %s_idx, i32 %v_inp) #1 {
+  %tok = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %idx)
+  %widx0 = call <4 x i32> @llvm.amdgcn.waterfall.readfirstlane.v4i32.v4i32(i32 %tok, <4 x i32> %s_idx)
+  %val = call { <4 x float>, i32 } @llvm.amdgcn.struct.buffer.load.format.sl_v4f32i32s(<4 x i32> %widx0, i32 %v_inp, i32 0, i32 0, i32 0)
+  %payl = extractvalue { <4 x float>, i32 } %val, 0
+  %tfe = extractvalue { <4 x float>, i32 } %val, 1
+  %r0 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok, <4 x float> %payl)
+  %r1 = call i32 @llvm.amdgcn.waterfall.end.i32(i32 %tok, i32 %tfe)
+  %r1.float = bitcast i32 %r1 to float
+
+  %insert = insertvalue { <4 x float>, float } undef, <4 x float> %r0, 0
+  %insert1 = insertvalue { <4 x float>, float } %insert, float %r1.float, 1
+  ret {<4 x float> , float} %insert1
+}
+
+%struct.wobble = type { float, float }
+
+define amdgpu_cs_chain void @wf_scc_check(i32 %arg, %struct.wobble %arg1, float %arg2) {
+; VI-SDAG-LABEL: wf_scc_check:
+; VI-SDAG:       ; %bb.0: ; %bb
+; VI-SDAG-NEXT:    s_waitcnt vmcnt(0) expcnt(0) lgkmcnt(0)
+; VI-SDAG-NEXT:    s_mov_b32 s0, 0
+; VI-SDAG-NEXT:    s_mov_b32 s1, s0
+; VI-SDAG-NEXT:    s_mov_b32 s2, s0
+; VI-SDAG-NEXT:    s_mov_b32 s3, s0
+; VI-SDAG-NEXT:    s_buffer_load_dword s1, s[0:3], 0x0
+; VI-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; VI-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; VI-SDAG-NEXT:    s_cmp_eq_u32 s1, 0
+; VI-SDAG-NEXT:    s_cselect_b64 s[8:9], -1, 0
+; VI-SDAG-NEXT:  .LBB24_1: ; =>This Inner Loop Header: Depth=1
+; VI-SDAG-NEXT:    v_readfirstlane_b32 s1, v8
+; VI-SDAG-NEXT:    v_cmp_eq_u32_e64 s[2:3], s1, v8
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[12:13], s[2:3]
+; VI-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; VI-SDAG-NEXT:    s_mov_b32 s1, s0
+; VI-SDAG-NEXT:    s_mov_b32 s2, s0
+; VI-SDAG-NEXT:    s_mov_b32 s3, s0
+; VI-SDAG-NEXT:    s_mov_b32 s4, s0
+; VI-SDAG-NEXT:    s_mov_b32 s5, s0
+; VI-SDAG-NEXT:    s_mov_b32 s6, s0
+; VI-SDAG-NEXT:    s_mov_b32 s7, s0
+; VI-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; VI-SDAG-NEXT:    image_store v0, v[0:1], s[0:7] unorm
+; VI-SDAG-NEXT:    ; implicit-def: $vgpr8
+; VI-SDAG-NEXT:    s_xor_b64 exec, exec, s[12:13]
+; VI-SDAG-NEXT:    s_cbranch_execnz .LBB24_1
+; VI-SDAG-NEXT:  ; %bb.2:
+; VI-SDAG-NEXT:    s_mov_b64 exec, s[10:11]
+; VI-SDAG-NEXT:    v_rcp_f32_e32 v0, v11
+; VI-SDAG-NEXT:    v_cndmask_b32_e64 v1, 0, 1.0, s[8:9]
+; VI-SDAG-NEXT:    v_mul_f32_e32 v0, v1, v0
+; VI-SDAG-NEXT:    v_cmp_le_f32_e32 vcc, 0, v0
+; VI-SDAG-NEXT:    s_and_saveexec_b64 s[0:1], vcc
+; VI-SDAG-NEXT:    s_endpgm
+;
+; VI-GISEL-LABEL: wf_scc_check:
+; VI-GISEL:       ; %bb.0: ; %bb
+; VI-GISEL-NEXT:    s_waitcnt vmcnt(0) expcnt(0) lgkmcnt(0)
+; VI-GISEL-NEXT:    s_mov_b32 s8, 0
+; VI-GISEL-NEXT:    s_mov_b32 s9, s8
+; VI-GISEL-NEXT:    s_mov_b32 s10, s8
+; VI-GISEL-NEXT:    s_mov_b32 s11, s8
+; VI-GISEL-NEXT:    s_buffer_load_dword s9, s[8:11], 0x0
+; VI-GISEL-NEXT:    s_mov_b32 s0, 0
+; VI-GISEL-NEXT:    s_mov_b32 s1, s8
+; VI-GISEL-NEXT:    s_mov_b32 s2, s8
+; VI-GISEL-NEXT:    s_mov_b32 s3, s8
+; VI-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; VI-GISEL-NEXT:    s_cmp_eq_u32 s9, 0
+; VI-GISEL-NEXT:    s_mov_b32 s4, s8
+; VI-GISEL-NEXT:    s_mov_b32 s5, s8
+; VI-GISEL-NEXT:    s_mov_b32 s6, s8
+; VI-GISEL-NEXT:    s_mov_b32 s7, s8
+; VI-GISEL-NEXT:    s_cselect_b32 s12, 1, 0
+; VI-GISEL-NEXT:    s_mov_b64 s[10:11], exec
+; VI-GISEL-NEXT:  .LBB24_1: ; =>This Inner Loop Header: Depth=1
+; VI-GISEL-NEXT:    v_readfirstlane_b32 s9, v8
+; VI-GISEL-NEXT:    v_cmp_eq_u32_e64 s[14:15], s9, v8
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[14:15], s[14:15]
+; VI-GISEL-NEXT:    s_mov_b32 s9, s8
+; VI-GISEL-NEXT:    v_mov_b32_e32 v0, s8
+; VI-GISEL-NEXT:    v_mov_b32_e32 v2, 0
+; VI-GISEL-NEXT:    v_mov_b32_e32 v1, s9
+; VI-GISEL-NEXT:    image_store v2, v[0:1], s[0:7] unorm
+; VI-GISEL-NEXT:    ; implicit-def: $vgpr8
+; VI-GISEL-NEXT:    s_xor_b64 exec, exec, s[14:15]
+; VI-GISEL-NEXT:    s_cbranch_execnz .LBB24_1
+; VI-GISEL-NEXT:  ; %bb.2:
+; VI-GISEL-NEXT:    s_mov_b64 exec, s[10:11]
+; VI-GISEL-NEXT:    v_rcp_f32_e32 v0, v11
+; VI-GISEL-NEXT:    s_cmp_lg_u32 s12, 0
+; VI-GISEL-NEXT:    s_cselect_b32 s0, 1.0, 0
+; VI-GISEL-NEXT:    v_mul_f32_e32 v0, s0, v0
+; VI-GISEL-NEXT:    v_cmp_le_f32_e32 vcc, 0, v0
+; VI-GISEL-NEXT:    s_and_saveexec_b64 s[0:1], vcc
+; VI-GISEL-NEXT:    s_endpgm
+;
+; GFX9-SDAG-LABEL: wf_scc_check:
+; GFX9-SDAG:       ; %bb.0: ; %bb
+; GFX9-SDAG-NEXT:    s_waitcnt vmcnt(0) expcnt(0) lgkmcnt(0)
+; GFX9-SDAG-NEXT:    s_mov_b32 s0, 0
+; GFX9-SDAG-NEXT:    s_mov_b32 s1, s0
+; GFX9-SDAG-NEXT:    s_mov_b32 s2, s0
+; GFX9-SDAG-NEXT:    s_mov_b32 s3, s0
+; GFX9-SDAG-NEXT:    s_buffer_load_dword s1, s[0:3], 0x0
+; GFX9-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; GFX9-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX9-SDAG-NEXT:    s_cmp_eq_u32 s1, 0
+; GFX9-SDAG-NEXT:    s_cselect_b64 s[8:9], -1, 0
+; GFX9-SDAG-NEXT:  .LBB24_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-SDAG-NEXT:    v_readfirstlane_b32 s1, v8
+; GFX9-SDAG-NEXT:    v_cmp_eq_u32_e64 s[2:3], s1, v8
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[12:13], s[2:3]
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX9-SDAG-NEXT:    s_mov_b32 s1, s0
+; GFX9-SDAG-NEXT:    s_mov_b32 s2, s0
+; GFX9-SDAG-NEXT:    s_mov_b32 s3, s0
+; GFX9-SDAG-NEXT:    s_mov_b32 s4, s0
+; GFX9-SDAG-NEXT:    s_mov_b32 s5, s0
+; GFX9-SDAG-NEXT:    s_mov_b32 s6, s0
+; GFX9-SDAG-NEXT:    s_mov_b32 s7, s0
+; GFX9-SDAG-NEXT:    v_mov_b32_e32 v1, v0
+; GFX9-SDAG-NEXT:    image_store v0, v[0:1], s[0:7] unorm
+; GFX9-SDAG-NEXT:    ; implicit-def: $vgpr8
+; GFX9-SDAG-NEXT:    s_xor_b64 exec, exec, s[12:13]
+; GFX9-SDAG-NEXT:    s_cbranch_execnz .LBB24_1
+; GFX9-SDAG-NEXT:  ; %bb.2:
+; GFX9-SDAG-NEXT:    s_mov_b64 exec, s[10:11]
+; GFX9-SDAG-NEXT:    v_rcp_f32_e32 v0, v11
+; GFX9-SDAG-NEXT:    v_cndmask_b32_e64 v1, 0, 1.0, s[8:9]
+; GFX9-SDAG-NEXT:    v_mul_f32_e32 v0, v1, v0
+; GFX9-SDAG-NEXT:    v_cmp_le_f32_e32 vcc, 0, v0
+; GFX9-SDAG-NEXT:    s_and_saveexec_b64 s[0:1], vcc
+; GFX9-SDAG-NEXT:    s_endpgm
+;
+; GFX9-GISEL-LABEL: wf_scc_check:
+; GFX9-GISEL:       ; %bb.0: ; %bb
+; GFX9-GISEL-NEXT:    s_waitcnt vmcnt(0) expcnt(0) lgkmcnt(0)
+; GFX9-GISEL-NEXT:    s_mov_b32 s8, 0
+; GFX9-GISEL-NEXT:    s_mov_b32 s9, s8
+; GFX9-GISEL-NEXT:    s_mov_b32 s10, s8
+; GFX9-GISEL-NEXT:    s_mov_b32 s11, s8
+; GFX9-GISEL-NEXT:    s_buffer_load_dword s9, s[8:11], 0x0
+; GFX9-GISEL-NEXT:    s_mov_b32 s0, 0
+; GFX9-GISEL-NEXT:    s_mov_b32 s1, s8
+; GFX9-GISEL-NEXT:    s_mov_b32 s2, s8
+; GFX9-GISEL-NEXT:    s_mov_b32 s3, s8
+; GFX9-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX9-GISEL-NEXT:    s_cmp_eq_u32 s9, 0
+; GFX9-GISEL-NEXT:    s_mov_b32 s4, s8
+; GFX9-GISEL-NEXT:    s_mov_b32 s5, s8
+; GFX9-GISEL-NEXT:    s_mov_b32 s6, s8
+; GFX9-GISEL-NEXT:    s_mov_b32 s7, s8
+; GFX9-GISEL-NEXT:    s_cselect_b32 s12, 1, 0
+; GFX9-GISEL-NEXT:    s_mov_b64 s[10:11], exec
+; GFX9-GISEL-NEXT:  .LBB24_1: ; =>This Inner Loop Header: Depth=1
+; GFX9-GISEL-NEXT:    v_readfirstlane_b32 s9, v8
+; GFX9-GISEL-NEXT:    v_cmp_eq_u32_e64 s[14:15], s9, v8
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[14:15], s[14:15]
+; GFX9-GISEL-NEXT:    s_mov_b32 s9, s8
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v0, s8
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v2, 0
+; GFX9-GISEL-NEXT:    v_mov_b32_e32 v1, s9
+; GFX9-GISEL-NEXT:    image_store v2, v[0:1], s[0:7] unorm
+; GFX9-GISEL-NEXT:    ; implicit-def: $vgpr8
+; GFX9-GISEL-NEXT:    s_xor_b64 exec, exec, s[14:15]
+; GFX9-GISEL-NEXT:    s_cbranch_execnz .LBB24_1
+; GFX9-GISEL-NEXT:  ; %bb.2:
+; GFX9-GISEL-NEXT:    s_mov_b64 exec, s[10:11]
+; GFX9-GISEL-NEXT:    v_rcp_f32_e32 v0, v11
+; GFX9-GISEL-NEXT:    s_cmp_lg_u32 s12, 0
+; GFX9-GISEL-NEXT:    s_cselect_b32 s0, 1.0, 0
+; GFX9-GISEL-NEXT:    v_mul_f32_e32 v0, s0, v0
+; GFX9-GISEL-NEXT:    v_cmp_le_f32_e32 vcc, 0, v0
+; GFX9-GISEL-NEXT:    s_and_saveexec_b64 s[0:1], vcc
+; GFX9-GISEL-NEXT:    s_endpgm
+;
+; GFX10-32-SDAG-LABEL: wf_scc_check:
+; GFX10-32-SDAG:       ; %bb.0: ; %bb
+; GFX10-32-SDAG-NEXT:    s_waitcnt vmcnt(0) expcnt(0) lgkmcnt(0)
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s0, 0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s10, exec_lo
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s1, s0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s2, s0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s3, s0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s9, exec_lo
+; GFX10-32-SDAG-NEXT:    s_buffer_load_dword s1, s[0:3], 0x0
+; GFX10-32-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-32-SDAG-NEXT:    s_cmp_eq_u32 s1, 0
+; GFX10-32-SDAG-NEXT:    s_cselect_b32 s8, -1, 0
+; GFX10-32-SDAG-NEXT:  .LBB24_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-SDAG-NEXT:    v_readfirstlane_b32 s1, v8
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-SDAG-NEXT:    v_cmpx_eq_u32_e32 s1, v8
+; GFX10-32-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s1, s0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s2, s0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s3, s0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s4, s0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s5, s0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s6, s0
+; GFX10-32-SDAG-NEXT:    s_mov_b32 s7, s0
+; GFX10-32-SDAG-NEXT:    image_store v0, [v0, v0], s[0:7] dim:SQ_RSRC_IMG_2D unorm
+; GFX10-32-SDAG-NEXT:    s_andn2_wrexec_b32 s10, s10
+; GFX10-32-SDAG-NEXT:    ; implicit-def: $vgpr8
+; GFX10-32-SDAG-NEXT:    s_cbranch_execnz .LBB24_1
+; GFX10-32-SDAG-NEXT:  ; %bb.2:
+; GFX10-32-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-SDAG-NEXT:    s_mov_b32 exec_lo, s9
+; GFX10-32-SDAG-NEXT:    v_rcp_f32_e32 v0, v11
+; GFX10-32-SDAG-NEXT:    v_cndmask_b32_e64 v1, 0, 1.0, s8
+; GFX10-32-SDAG-NEXT:    v_mul_f32_e32 v0, v1, v0
+; GFX10-32-SDAG-NEXT:    v_cmp_le_f32_e32 vcc_lo, 0, v0
+; GFX10-32-SDAG-NEXT:    s_and_saveexec_b32 s0, vcc_lo
+; GFX10-32-SDAG-NEXT:    s_endpgm
+;
+; GFX10-32-GISEL-LABEL: wf_scc_check:
+; GFX10-32-GISEL:       ; %bb.0: ; %bb
+; GFX10-32-GISEL-NEXT:    s_waitcnt vmcnt(0) expcnt(0) lgkmcnt(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s12, 0
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s10, exec_lo
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s13, s12
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s14, s12
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s15, s12
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s0, s12
+; GFX10-32-GISEL-NEXT:    s_buffer_load_dword s7, s[12:15], 0x0
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s1, s12
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s2, s12
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s3, s12
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s4, s12
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s5, s12
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s6, s12
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s9, exec_lo
+; GFX10-32-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-32-GISEL-NEXT:    s_cmp_eq_u32 s7, 0
+; GFX10-32-GISEL-NEXT:    s_mov_b32 s7, s12
+; GFX10-32-GISEL-NEXT:    s_cselect_b32 s8, 1, 0
+; GFX10-32-GISEL-NEXT:  .LBB24_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-32-GISEL-NEXT:    v_readfirstlane_b32 s11, v8
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-32-GISEL-NEXT:    v_cmpx_eq_u32_e32 s11, v8
+; GFX10-32-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-32-GISEL-NEXT:    image_store v0, [v0, v0], s[0:7] dim:SQ_RSRC_IMG_2D unorm
+; GFX10-32-GISEL-NEXT:    s_andn2_wrexec_b32 s10, s10
+; GFX10-32-GISEL-NEXT:    ; implicit-def: $vgpr8
+; GFX10-32-GISEL-NEXT:    s_cbranch_execnz .LBB24_1
+; GFX10-32-GISEL-NEXT:  ; %bb.2:
+; GFX10-32-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-32-GISEL-NEXT:    s_mov_b32 exec_lo, s9
+; GFX10-32-GISEL-NEXT:    v_rcp_f32_e32 v0, v11
+; GFX10-32-GISEL-NEXT:    s_cmp_lg_u32 s8, 0
+; GFX10-32-GISEL-NEXT:    s_cselect_b32 s0, 1.0, 0
+; GFX10-32-GISEL-NEXT:    v_mul_f32_e32 v0, s0, v0
+; GFX10-32-GISEL-NEXT:    v_cmp_le_f32_e32 vcc_lo, 0, v0
+; GFX10-32-GISEL-NEXT:    s_and_saveexec_b32 s0, vcc_lo
+; GFX10-32-GISEL-NEXT:    s_endpgm
+;
+; GFX10-64-SDAG-LABEL: wf_scc_check:
+; GFX10-64-SDAG:       ; %bb.0: ; %bb
+; GFX10-64-SDAG-NEXT:    s_waitcnt vmcnt(0) expcnt(0) lgkmcnt(0)
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s0, 0
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[12:13], exec
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s1, s0
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s2, s0
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s3, s0
+; GFX10-64-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; GFX10-64-SDAG-NEXT:    s_buffer_load_dword s1, s[0:3], 0x0
+; GFX10-64-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-64-SDAG-NEXT:    s_cmp_eq_u32 s1, 0
+; GFX10-64-SDAG-NEXT:    s_cselect_b64 s[8:9], -1, 0
+; GFX10-64-SDAG-NEXT:  .LBB24_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-SDAG-NEXT:    v_readfirstlane_b32 s1, v8
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-SDAG-NEXT:    v_cmpx_eq_u32_e64 s1, v8
+; GFX10-64-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s1, s0
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s2, s0
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s3, s0
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s4, s0
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s5, s0
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s6, s0
+; GFX10-64-SDAG-NEXT:    s_mov_b32 s7, s0
+; GFX10-64-SDAG-NEXT:    image_store v0, [v0, v0], s[0:7] dim:SQ_RSRC_IMG_2D unorm
+; GFX10-64-SDAG-NEXT:    s_andn2_wrexec_b64 s[12:13], s[12:13]
+; GFX10-64-SDAG-NEXT:    ; implicit-def: $vgpr8
+; GFX10-64-SDAG-NEXT:    s_cbranch_execnz .LBB24_1
+; GFX10-64-SDAG-NEXT:  ; %bb.2:
+; GFX10-64-SDAG-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-SDAG-NEXT:    s_mov_b64 exec, s[10:11]
+; GFX10-64-SDAG-NEXT:    v_rcp_f32_e32 v0, v11
+; GFX10-64-SDAG-NEXT:    v_cndmask_b32_e64 v1, 0, 1.0, s[8:9]
+; GFX10-64-SDAG-NEXT:    v_mul_f32_e32 v0, v1, v0
+; GFX10-64-SDAG-NEXT:    v_cmp_le_f32_e32 vcc, 0, v0
+; GFX10-64-SDAG-NEXT:    s_and_saveexec_b64 s[0:1], vcc
+; GFX10-64-SDAG-NEXT:    s_endpgm
+;
+; GFX10-64-GISEL-LABEL: wf_scc_check:
+; GFX10-64-GISEL:       ; %bb.0: ; %bb
+; GFX10-64-GISEL-NEXT:    s_waitcnt vmcnt(0) expcnt(0) lgkmcnt(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s8, 0
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s9, s8
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s10, s8
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s11, s8
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s0, s8
+; GFX10-64-GISEL-NEXT:    s_buffer_load_dword s7, s[8:11], 0x0
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s1, s8
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s2, s8
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s3, s8
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s4, s8
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s5, s8
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s6, s8
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[10:11], exec
+; GFX10-64-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX10-64-GISEL-NEXT:    s_cmp_eq_u32 s7, 0
+; GFX10-64-GISEL-NEXT:    s_mov_b32 s7, s8
+; GFX10-64-GISEL-NEXT:    s_cselect_b32 s12, 1, 0
+; GFX10-64-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX10-64-GISEL-NEXT:  .LBB24_1: ; =>This Inner Loop Header: Depth=1
+; GFX10-64-GISEL-NEXT:    v_readfirstlane_b32 s13, v8
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_sa_sdst(0)
+; GFX10-64-GISEL-NEXT:    v_cmpx_eq_u32_e64 s13, v8
+; GFX10-64-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX10-64-GISEL-NEXT:    image_store v0, [v0, v0], s[0:7] dim:SQ_RSRC_IMG_2D unorm
+; GFX10-64-GISEL-NEXT:    s_andn2_wrexec_b64 s[10:11], s[10:11]
+; GFX10-64-GISEL-NEXT:    ; implicit-def: $vgpr8
+; GFX10-64-GISEL-NEXT:    s_cbranch_execnz .LBB24_1
+; GFX10-64-GISEL-NEXT:  ; %bb.2:
+; GFX10-64-GISEL-NEXT:    s_waitcnt_depctr depctr_vm_vsrc(0)
+; GFX10-64-GISEL-NEXT:    s_mov_b64 exec, s[8:9]
+; GFX10-64-GISEL-NEXT:    v_rcp_f32_e32 v0, v11
+; GFX10-64-GISEL-NEXT:    s_cmp_lg_u32 s12, 0
+; GFX10-64-GISEL-NEXT:    s_cselect_b32 s0, 1.0, 0
+; GFX10-64-GISEL-NEXT:    v_mul_f32_e32 v0, s0, v0
+; GFX10-64-GISEL-NEXT:    v_cmp_le_f32_e32 vcc, 0, v0
+; GFX10-64-GISEL-NEXT:    s_and_saveexec_b64 s[0:1], vcc
+; GFX10-64-GISEL-NEXT:    s_endpgm
+;
+; GFX1150-SDAG-LABEL: wf_scc_check:
+; GFX1150-SDAG:       ; %bb.0: ; %bb
+; GFX1150-SDAG-NEXT:    s_waitcnt vmcnt(0) expcnt(0) lgkmcnt(0)
+; GFX1150-SDAG-NEXT:    s_mov_b32 s0, 0
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[12:13], exec
+; GFX1150-SDAG-NEXT:    s_mov_b32 s1, s0
+; GFX1150-SDAG-NEXT:    s_mov_b32 s2, s0
+; GFX1150-SDAG-NEXT:    s_mov_b32 s3, s0
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; GFX1150-SDAG-NEXT:    s_buffer_load_b32 s1, s[0:3], 0x0
+; GFX1150-SDAG-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX1150-SDAG-NEXT:    s_cmp_eq_u32 s1, 0
+; GFX1150-SDAG-NEXT:    s_cselect_b64 s[8:9], -1, 0
+; GFX1150-SDAG-NEXT:  .LBB24_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-SDAG-NEXT:    v_readfirstlane_b32 s1, v8
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_cmpx_eq_u32_e64 s1, v8
+; GFX1150-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-SDAG-NEXT:    s_mov_b32 s1, s0
+; GFX1150-SDAG-NEXT:    s_mov_b32 s2, s0
+; GFX1150-SDAG-NEXT:    s_mov_b32 s3, s0
+; GFX1150-SDAG-NEXT:    s_mov_b32 s4, s0
+; GFX1150-SDAG-NEXT:    s_mov_b32 s5, s0
+; GFX1150-SDAG-NEXT:    s_mov_b32 s6, s0
+; GFX1150-SDAG-NEXT:    s_mov_b32 s7, s0
+; GFX1150-SDAG-NEXT:    image_store v0, [v0, v0], s[0:7] dim:SQ_RSRC_IMG_2D unorm
+; GFX1150-SDAG-NEXT:    s_and_not1_wrexec_b64 s[12:13], s[12:13]
+; GFX1150-SDAG-NEXT:    ; implicit-def: $vgpr8
+; GFX1150-SDAG-NEXT:    s_cbranch_execnz .LBB24_1
+; GFX1150-SDAG-NEXT:  ; %bb.2:
+; GFX1150-SDAG-NEXT:    s_mov_b64 exec, s[10:11]
+; GFX1150-SDAG-NEXT:    v_rcp_f32_e32 v0, v11
+; GFX1150-SDAG-NEXT:    v_cndmask_b32_e64 v1, 0, 1.0, s[8:9]
+; GFX1150-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(TRANS32_DEP_1) | instid1(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_mul_f32_e32 v0, v1, v0
+; GFX1150-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-SDAG-NEXT:    v_cmpx_le_f32_e32 0, v0
+; GFX1150-SDAG-NEXT:    s_endpgm
+;
+; GFX1150-GISEL-LABEL: wf_scc_check:
+; GFX1150-GISEL:       ; %bb.0: ; %bb
+; GFX1150-GISEL-NEXT:    s_waitcnt vmcnt(0) expcnt(0) lgkmcnt(0)
+; GFX1150-GISEL-NEXT:    s_mov_b32 s8, 0
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    s_mov_b32 s9, s8
+; GFX1150-GISEL-NEXT:    s_mov_b32 s10, s8
+; GFX1150-GISEL-NEXT:    s_mov_b32 s11, s8
+; GFX1150-GISEL-NEXT:    s_mov_b32 s0, s8
+; GFX1150-GISEL-NEXT:    s_buffer_load_b32 s7, s[8:11], 0x0
+; GFX1150-GISEL-NEXT:    s_mov_b32 s1, s8
+; GFX1150-GISEL-NEXT:    s_mov_b32 s2, s8
+; GFX1150-GISEL-NEXT:    s_mov_b32 s3, s8
+; GFX1150-GISEL-NEXT:    s_mov_b32 s4, s8
+; GFX1150-GISEL-NEXT:    s_mov_b32 s5, s8
+; GFX1150-GISEL-NEXT:    s_mov_b32 s6, s8
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[10:11], exec
+; GFX1150-GISEL-NEXT:    s_waitcnt lgkmcnt(0)
+; GFX1150-GISEL-NEXT:    s_cmp_eq_u32 s7, 0
+; GFX1150-GISEL-NEXT:    s_mov_b32 s7, s8
+; GFX1150-GISEL-NEXT:    s_cselect_b32 s12, 1, 0
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX1150-GISEL-NEXT:  .LBB24_1: ; =>This Inner Loop Header: Depth=1
+; GFX1150-GISEL-NEXT:    v_readfirstlane_b32 s13, v8
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-GISEL-NEXT:    v_cmpx_eq_u32_e64 s13, v8
+; GFX1150-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX1150-GISEL-NEXT:    image_store v0, [v0, v0], s[0:7] dim:SQ_RSRC_IMG_2D unorm
+; GFX1150-GISEL-NEXT:    s_and_not1_wrexec_b64 s[10:11], s[10:11]
+; GFX1150-GISEL-NEXT:    ; implicit-def: $vgpr8
+; GFX1150-GISEL-NEXT:    s_cbranch_execnz .LBB24_1
+; GFX1150-GISEL-NEXT:  ; %bb.2:
+; GFX1150-GISEL-NEXT:    s_mov_b64 exec, s[8:9]
+; GFX1150-GISEL-NEXT:    v_rcp_f32_e32 v0, v11
+; GFX1150-GISEL-NEXT:    s_cmp_lg_u32 s12, 0
+; GFX1150-GISEL-NEXT:    s_cselect_b32 s0, 1.0, 0
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(TRANS32_DEP_1) | instid1(SALU_CYCLE_1)
+; GFX1150-GISEL-NEXT:    v_mul_f32_e32 v0, s0, v0
+; GFX1150-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX1150-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX1150-GISEL-NEXT:    v_cmpx_le_f32_e32 0, v0
+; GFX1150-GISEL-NEXT:    s_endpgm
+;
+; GFX12-SDAG-LABEL: wf_scc_check:
+; GFX12-SDAG:       ; %bb.0: ; %bb
+; GFX12-SDAG-NEXT:    s_wait_loadcnt_dscnt 0x0
+; GFX12-SDAG-NEXT:    s_wait_expcnt 0x0
+; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
+; GFX12-SDAG-NEXT:    s_wait_bvhcnt 0x0
+; GFX12-SDAG-NEXT:    s_wait_kmcnt 0x0
+; GFX12-SDAG-NEXT:    s_mov_b32 s0, 0
+; GFX12-SDAG-NEXT:    s_mov_b64 s[10:11], exec
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    s_mov_b32 s1, s0
+; GFX12-SDAG-NEXT:    s_mov_b32 s2, s0
+; GFX12-SDAG-NEXT:    s_mov_b32 s3, s0
+; GFX12-SDAG-NEXT:    s_mov_b64 s[8:9], exec
+; GFX12-SDAG-NEXT:    s_buffer_load_b32 s12, s[0:3], 0x0
+; GFX12-SDAG-NEXT:  .LBB24_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-SDAG-NEXT:    v_readfirstlane_b32 s1, v8
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_cmpx_eq_u32_e64 s1, v8
+; GFX12-SDAG-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-SDAG-NEXT:    s_mov_b32 s1, s0
+; GFX12-SDAG-NEXT:    s_mov_b32 s2, s0
+; GFX12-SDAG-NEXT:    s_mov_b32 s3, s0
+; GFX12-SDAG-NEXT:    s_mov_b32 s4, s0
+; GFX12-SDAG-NEXT:    s_mov_b32 s5, s0
+; GFX12-SDAG-NEXT:    s_mov_b32 s6, s0
+; GFX12-SDAG-NEXT:    s_mov_b32 s7, s0
+; GFX12-SDAG-NEXT:    image_store v0, [v0, v0], s[0:7] dim:SQ_RSRC_IMG_2D
+; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[10:11], s[10:11]
+; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr8
+; GFX12-SDAG-NEXT:    s_cbranch_execnz .LBB24_1
+; GFX12-SDAG-NEXT:  ; %bb.2:
+; GFX12-SDAG-NEXT:    s_mov_b64 exec, s[8:9]
+; GFX12-SDAG-NEXT:    v_rcp_f32_e32 v0, v11
+; GFX12-SDAG-NEXT:    s_wait_kmcnt 0x0
+; GFX12-SDAG-NEXT:    s_cmp_eq_u32 s12, 0
+; GFX12-SDAG-NEXT:    s_cselect_b32 s0, 1.0, 0
+; GFX12-SDAG-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-SDAG-NEXT:    s_delay_alu instid0(TRANS32_DEP_1) | instskip(SKIP_1) | instid1(VALU_DEP_1)
+; GFX12-SDAG-NEXT:    v_mul_f32_e32 v0, s0, v0
+; GFX12-SDAG-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-SDAG-NEXT:    v_cmpx_le_f32_e32 0, v0
+; GFX12-SDAG-NEXT:    s_endpgm
+;
+; GFX12-GISEL-LABEL: wf_scc_check:
+; GFX12-GISEL:       ; %bb.0: ; %bb
+; GFX12-GISEL-NEXT:    s_wait_loadcnt_dscnt 0x0
+; GFX12-GISEL-NEXT:    s_wait_expcnt 0x0
+; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
+; GFX12-GISEL-NEXT:    s_wait_bvhcnt 0x0
+; GFX12-GISEL-NEXT:    s_wait_kmcnt 0x0
+; GFX12-GISEL-NEXT:    s_mov_b32 s8, 0
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_mov_b32 s9, s8
+; GFX12-GISEL-NEXT:    s_mov_b32 s10, s8
+; GFX12-GISEL-NEXT:    s_mov_b32 s11, s8
+; GFX12-GISEL-NEXT:    s_mov_b32 s0, s8
+; GFX12-GISEL-NEXT:    s_buffer_load_b32 s7, s[8:11], 0x0
+; GFX12-GISEL-NEXT:    s_mov_b32 s1, s8
+; GFX12-GISEL-NEXT:    s_mov_b32 s2, s8
+; GFX12-GISEL-NEXT:    s_mov_b32 s3, s8
+; GFX12-GISEL-NEXT:    s_mov_b32 s4, s8
+; GFX12-GISEL-NEXT:    s_mov_b32 s5, s8
+; GFX12-GISEL-NEXT:    s_mov_b32 s6, s8
+; GFX12-GISEL-NEXT:    s_mov_b64 s[10:11], exec
+; GFX12-GISEL-NEXT:    s_wait_kmcnt 0x0
+; GFX12-GISEL-NEXT:    s_cmp_eq_u32 s7, 0
+; GFX12-GISEL-NEXT:    s_mov_b32 s7, s8
+; GFX12-GISEL-NEXT:    s_cselect_b32 s12, 1, 0
+; GFX12-GISEL-NEXT:    s_mov_b64 s[8:9], exec
+; GFX12-GISEL-NEXT:  .LBB24_1: ; =>This Inner Loop Header: Depth=1
+; GFX12-GISEL-NEXT:    v_readfirstlane_b32 s13, v8
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_va_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_cmpx_eq_u32_e64 s13, v8
+; GFX12-GISEL-NEXT:    v_mov_b32_e32 v0, 0
+; GFX12-GISEL-NEXT:    image_store v0, [v0, v0], s[0:7] dim:SQ_RSRC_IMG_2D
+; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[10:11], s[10:11]
+; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr8
+; GFX12-GISEL-NEXT:    s_cbranch_execnz .LBB24_1
+; GFX12-GISEL-NEXT:  ; %bb.2:
+; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[8:9]
+; GFX12-GISEL-NEXT:    v_rcp_f32_e32 v0, v11
+; GFX12-GISEL-NEXT:    s_cmp_lg_u32 s12, 0
+; GFX12-GISEL-NEXT:    s_cselect_b32 s0, 1.0, 0
+; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
+; GFX12-GISEL-NEXT:    s_delay_alu instid0(TRANS32_DEP_1) | instskip(SKIP_1) | instid1(VALU_DEP_1)
+; GFX12-GISEL-NEXT:    v_mul_f32_e32 v0, s0, v0
+; GFX12-GISEL-NEXT:    s_mov_b64 s[0:1], exec
+; GFX12-GISEL-NEXT:    v_cmpx_le_f32_e32 0, v0
+; GFX12-GISEL-NEXT:    s_endpgm
+bb:
+  %call = call i32 @llvm.amdgcn.s.buffer.load.i32(<4 x i32> zeroinitializer, i32 0, i32 0)
+  %icmp = icmp eq i32 %call, 0
+  %call3 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %arg)
+  %call4 = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %call3, i32 %arg)
+  %call5 = call <8 x i32> @llvm.amdgcn.waterfall.last.use.v8i32(i32 %call3, <8 x i32> zeroinitializer)
+  call void @llvm.amdgcn.image.store.2d.f32.i32.v8i32(float 0.000000e+00, i32 0, i32 0, i32 0, <8 x i32> %call5, i32 0, i32 0)
+  %select = select i1 %icmp, float 1.000000e+00, float 0.000000e+00
+  %fdiv = fdiv afn float %select, %arg2
+  %fcmp = fcmp ult float %fdiv, 0.000000e+00
+  br i1 %fcmp, label %bb6, label %bb7
+
+bb6:                                              ; preds = %bb7, %bb
+  ret void
+
+bb7:                                              ; preds = %bb
+  %extractvalue = extractvalue %struct.wobble %arg1, 0
+  %fsub = fsub float 0.000000e+00, %extractvalue
+  %fcmp8 = fcmp olt float %fsub, 0.000000e+00
+  %or = or i1 false, %fcmp8
+  br label %bb6
+}
+
+declare i32 @llvm.amdgcn.waterfall.begin.i32(i32, i32) #6
+declare i32 @llvm.amdgcn.waterfall.begin.v2i32(i32, <2 x i32>) #6
+declare i32 @llvm.amdgcn.waterfall.begin.v4i32(i32, <4 x i32>) #6
+declare i32 @llvm.amdgcn.waterfall.begin.v8i32(i32, <8 x i32>) #1
+declare i32 @llvm.amdgcn.waterfall.begin.p3(i32 , ptr addrspace(3))
+declare i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32, i32) #6
+declare <2 x i32> @llvm.amdgcn.waterfall.readfirstlane.v2i32.v2i32(i32, <2 x i32>) #6
+declare <8 x i32> @llvm.amdgcn.waterfall.readfirstlane.v8i32.v8i32(i32, <8 x i32>) #6
+declare <4 x i32> @llvm.amdgcn.waterfall.readfirstlane.v4i32.v4i32(i32, <4 x i32>) #1
+declare i16 @llvm.amdgcn.waterfall.end.i16(i32, i16) #6
+declare i32 @llvm.amdgcn.waterfall.end.i32(i32, i32) #6
+declare <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32, <4 x float>) #6
+declare <8 x i32> @llvm.amdgcn.waterfall.end.v8i32(i32, <8 x i32>) #6
+declare <8 x i32> @llvm.amdgcn.waterfall.last.use.v8i32(i32, <8 x i32>) #6
+declare ptr addrspace(3) @llvm.amdgcn.waterfall.last.use.vgpr.p3(i32, ptr addrspace(3))
+declare i32 @llvm.amdgcn.readlane(i32, i32) #0
+declare <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32(i32, float, <8 x i32>, <4 x i32>, i1, i32, i32)
+declare <4 x float> @llvm.amdgcn.image.sample.2d.v4f32.f32(i32, float, float, <8 x i32>, <4 x i32>, i1, i32, i32)
+declare void @llvm.amdgcn.image.store.1d.v4f32.i32(<4 x float>, i32, i32, <8 x i32>, i32, i32)
+declare i64 @llvm.amdgcn.s.getpc() #3
+declare float @llvm.amdgcn.buffer.load.ushort(<4 x i32>, i32, i32, i1, i1) #4
+declare float @llvm.amdgcn.buffer.load.f32(<4 x i32>, i32, i32, i1, i1) #4
+declare <4 x float> @llvm.amdgcn.buffer.load.v4f32(<4 x i32>, i32, i32, i1, i1) #4
+declare i32 @llvm.amdgcn.s.buffer.load.i32(<4 x i32>, i32, i32) #2
+declare { <4 x float>, i32 } @llvm.amdgcn.struct.buffer.load.format.sl_v4f32i32s(<4 x i32>, i32, i32, i32, i32) #4
+declare <4 x i32> @llvm.amdgcn.s.buffer.load.v4i32(<4 x i32>, i32, i1) #2
+declare void @llvm.amdgcn.buffer.store.short(float, <4 x i32>, i32, i32, i1, i1) #5
+declare void @llvm.amdgcn.buffer.store.f32(float, <4 x i32>, i32, i32, i1, i1) #5
+declare void @llvm.amdgcn.buffer.store.v4f32(<4 x float>, <4 x i32>, i32, i32, i1, i1) #5
+declare void @llvm.amdgcn.kill(i1) #1
+declare void @llvm.amdgcn.exp.f32(i32 immarg, i32 immarg, float, float, float, float, i1 immarg, i1 immarg) #7
+
+attributes #0 = { nounwind readnone convergent }
+attributes #1 = { nounwind }
+attributes #2 = { nounwind readnone }
+attributes #3 = { nounwind readnone speculatable }
+attributes #4 = { nounwind readonly }
+attributes #5 = { nounwind writeonly }
+attributes #6 = { nounwind }
+attributes #7 = { inaccessiblememonly nounwind writeonly }
diff --git a/llvm/test/CodeGen/AMDGPU/si-insert-waterfall-licm.ll b/llvm/test/CodeGen/AMDGPU/si-insert-waterfall-licm.ll
new file mode 100644
index 0000000000000..3c9aa3f80a32c
--- /dev/null
+++ b/llvm/test/CodeGen/AMDGPU/si-insert-waterfall-licm.ll
@@ -0,0 +1,43 @@
+; RUN: opt -mtriple=amdgcn -mcpu=gfx1100 -amdgpu-codegenprepare -S -o - < %s | FileCheck %s
+
+; Check that llvm.licm.disable is removed for loop forming divergent pointer.
+; CHECK-NOT: !{!"llvm.licm.disable"}
+
+define amdgpu_ps <4 x float> @needs_licm(i32 inreg %userdata, i32 %index, float %s, <4 x i32> inreg %samp, i32 inreg %a, i32 inreg %b) {
+entry:
+  %pc.0 = call i64 @llvm.amdgcn.s.getpc()
+  %pc.1 = and i64 %pc.0, -4294967296
+  %lo = call i32 @llvm.amdgcn.mbcnt.lo(i32 -1, i32 0)
+  %hi = call i32 @llvm.amdgcn.mbcnt.hi(i32 -1, i32 %lo)
+  %cc = icmp uge i32 %hi, 16
+  br i1 %cc, label %alt.entry, label %preheader
+
+alt.entry:
+  br label %preheader
+
+preheader:
+  %limit = phi i32 [ %a, %entry ], [ %b, %alt.entry ]
+  br label %loop.body
+
+loop.body:
+  %i.0 = phi i32 [ 0, %preheader ], [ %inc, %loop.body ]
+  %inc = add nsw i32 %i.0, 1
+  %cmp = icmp slt i32 %inc, %limit
+  %ptr.0 = zext i32 %userdata to i64
+  %ptr.1 = or disjoint i64 %pc.1, %ptr.0
+  %loop.ptr = inttoptr i64 %ptr.1 to ptr addrspace(4)
+  br i1 %cmp, label %loop.body, label %loop.exit, !llvm.loop !1
+
+loop.exit:
+  %wf_token = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %index)
+  %s_idx = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %wf_token, i32 %index)
+  %ptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %loop.ptr, i32 %s_idx
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4)* %ptr, align 32
+  %r = call <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32(i32 15, float %s, <8 x i32> %rsrc, <4 x i32> %samp, i1 0, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %wf_token, <4 x float> %r)
+
+  ret <4 x float> %r1
+}
+
+!1 = !{!1, !2}
+!2 = !{!"llvm.licm.disable"}
diff --git a/llvm/test/CodeGen/AMDGPU/si-insert-waterfall.mir b/llvm/test/CodeGen/AMDGPU/si-insert-waterfall.mir
new file mode 100644
index 0000000000000..30dedf0a89f18
--- /dev/null
+++ b/llvm/test/CodeGen/AMDGPU/si-insert-waterfall.mir
@@ -0,0 +1,54 @@
+# NOTE: Assertions have been autogenerated by utils/update_mir_test_checks.py
+# RUN: llc -mtriple=amdgcn--amdpal -march=amdgcn -mcpu=gfx1030 -run-pass=si-insert-waterfall -o - %s | FileCheck %s
+
+# The test demonstrates a waterfall loop expansion in si-insert-waterfall pass.
+---
+name:             waterfall
+tracksRegLiveness: true
+body:             |
+  bb.0:
+    liveins: $sgpr0, $vgpr0, $vgpr1
+    ; CHECK-LABEL: name: waterfall
+    ; CHECK: successors: %bb.1(0x80000000)
+    ; CHECK-NEXT: liveins: $sgpr0, $vgpr0, $vgpr1
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: [[COPY:%[0-9]+]]:sgpr_32 = COPY $sgpr0
+    ; CHECK-NEXT: [[COPY1:%[0-9]+]]:vgpr_32 = COPY $vgpr0
+    ; CHECK-NEXT: [[COPY2:%[0-9]+]]:vgpr_32 = COPY $vgpr1
+    ; CHECK-NEXT: [[REG_SEQUENCE:%[0-9]+]]:sgpr_128 = REG_SEQUENCE [[COPY]], %subreg.sub0, [[COPY]], %subreg.sub1, [[COPY]], %subreg.sub2, [[COPY]], %subreg.sub3
+    ; CHECK-NEXT: [[S_MOV_B32_:%[0-9]+]]:sreg_32_xm0 = S_MOV_B32 0
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: .1:
+    ; CHECK-NEXT: successors: %bb.2(0x80000000)
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: [[S_MOV_B32_1:%[0-9]+]]:sreg_32_xm0_xexec = S_MOV_B32 $exec_lo
+    ; CHECK-NEXT: [[S_MOV_B32_2:%[0-9]+]]:sreg_32_xm0_xexec = S_MOV_B32 $exec_lo
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: .2:
+    ; CHECK-NEXT: successors: %bb.2(0x40000000), %bb.3(0x40000000)
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: [[PHI:%[0-9]+]]:sreg_32_xm0_xexec = PHI [[S_MOV_B32_1]], %bb.1, %14, %bb.2
+    ; CHECK-NEXT: [[V_READFIRSTLANE_B32_:%[0-9]+]]:sreg_32_xm0 = V_READFIRSTLANE_B32 [[COPY2]], implicit $exec
+    ; CHECK-NEXT: V_CMPX_EQ_U32_nosdst_e32 [[V_READFIRSTLANE_B32_]], [[COPY2]], implicit-def $exec, implicit $exec
+    ; CHECK-NEXT: [[REG_SEQUENCE1:%[0-9]+]]:sreg_64 = REG_SEQUENCE [[V_READFIRSTLANE_B32_]], %subreg.sub0, [[V_READFIRSTLANE_B32_]], %subreg.sub1
+    ; CHECK-NEXT: [[S_LOAD_DWORDX8_IMM:%[0-9]+]]:sgpr_256 = S_LOAD_DWORDX8_IMM [[REG_SEQUENCE1]], 0, 0
+    ; CHECK-NEXT: [[IMAGE_SAMPLE_V4_V1_gfx10_:%[0-9]+]]:vreg_128 = IMAGE_SAMPLE_V4_V1_gfx10 [[COPY1]], [[S_LOAD_DWORDX8_IMM]], [[REG_SEQUENCE]], 15, 0, 0, 0, 0, 0, 0, 0, 0, implicit $exec :: (dereferenceable load (s128), addrspace 8)
+    ; CHECK-NEXT: [[COPY3:%[0-9]+]]:vreg_128 = COPY [[IMAGE_SAMPLE_V4_V1_gfx10_]]
+    ; CHECK-NEXT: [[S_ANDN2_WREXEC_B32_:%[0-9]+]]:sreg_32_xm0_xexec = S_ANDN2_WREXEC_B32 [[PHI]], implicit-def $exec, implicit-def $scc, implicit $exec
+    ; CHECK-NEXT: SI_WATERFALL_LOOP %bb.2, implicit $exec
+    ; CHECK-NEXT: {{  $}}
+    ; CHECK-NEXT: .3:
+    ; CHECK-NEXT: $exec_lo = S_MOV_B32 [[S_MOV_B32_2]]
+  %0:sgpr_32 = COPY $sgpr0
+  %1:vgpr_32 = COPY $vgpr0
+  %2:vgpr_32 = COPY $vgpr1
+  %3:sgpr_128 = REG_SEQUENCE %0:sgpr_32, %subreg.sub0, %0:sgpr_32, %subreg.sub1, %0:sgpr_32, %subreg.sub2, %0:sgpr_32, %subreg.sub3
+  %4:sreg_32_xm0 = S_MOV_B32 0
+  %5:sreg_32 = SI_WATERFALL_BEGIN_V1 killed %4:sreg_32_xm0, %2:vgpr_32, implicit-def $scc
+  %6:sreg_32_xm0 = SI_WATERFALL_READFIRSTLANE_V1 %5:sreg_32, %2:vgpr_32, implicit-def $scc
+  %7:sreg_64 = REG_SEQUENCE %6:sreg_32_xm0, %subreg.sub0, %6:sreg_32_xm0, %subreg.sub1
+  %8:sgpr_256 = S_LOAD_DWORDX8_IMM killed %7:sreg_64, 0, 0
+  %9:vreg_128 = IMAGE_SAMPLE_V4_V1_gfx10 %1:vgpr_32, killed %8:sgpr_256, killed %3:sgpr_128, 15, 0, 0, 0, 0, 0, 0, 0, 0, implicit $exec :: (dereferenceable load (s128), addrspace 8)
+  %10:vreg_128 = SI_WATERFALL_END_V4 %5:sreg_32, killed %9:vreg_128, implicit-def $scc
+...
+
diff --git a/llvm/test/Transforms/InstCombine/AMDGPU/waterfall.ll b/llvm/test/Transforms/InstCombine/AMDGPU/waterfall.ll
new file mode 100644
index 0000000000000..df57d55e2f674
--- /dev/null
+++ b/llvm/test/Transforms/InstCombine/AMDGPU/waterfall.ll
@@ -0,0 +1,73 @@
+; NOTE: Assertions have been autogenerated by utils/update_test_checks.py
+; RUN: opt -S -passes=instcombine -mtriple=amdgcn-amd-amdhsa %s | FileCheck %s
+
+define amdgpu_ps <4 x float> @test_waterfall_same_index2(<8 x i32> addrspace(4)* inreg %in, i32 %index, float %s, <4 x i32> inreg %samp) {
+; CHECK-LABEL: @test_waterfall_same_index2(
+; CHECK-NEXT:    [[WF_TOKEN1:%.*]] = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 [[INDEX:%.*]])
+; CHECK-NEXT:    [[S_IDX:%.*]] = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 [[WF_TOKEN1]], i32 [[INDEX]])
+; CHECK-NEXT:    [[TMP1:%.*]] = sext i32 [[S_IDX]] to i64
+; CHECK-NEXT:    [[PTR:%.*]] = getelementptr <8 x i32>, ptr addrspace(4) [[IN:%.*]], i64 [[TMP1]]
+; CHECK-NEXT:    [[RSRC:%.*]] = load <8 x i32>, ptr addrspace(4) [[PTR]], align 32
+; CHECK-NEXT:    [[R:%.*]] = call <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32.v8i32.v4i32(i32 15, float [[S:%.*]], <8 x i32> [[RSRC]], <4 x i32> [[SAMP:%.*]], i1 false, i32 0, i32 0)
+; CHECK-NEXT:    [[R1:%.*]] = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 [[WF_TOKEN1]], <4 x float> [[R]])
+; CHECK-NEXT:    ret <4 x float> [[R1]]
+;
+  %wf_token1 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %index)
+  %wf_token2 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 %wf_token1, i32 %index)
+  %s_idx = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %wf_token2, i32 %index)
+  %ptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4) * %ptr, align 32
+  %r = call <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32(i32 15, float %s, <8 x i32> %rsrc, <4 x i32> %samp, i1 0, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %wf_token2, <4 x float> %r)
+  ret <4 x float> %r1
+}
+
+define amdgpu_ps <4 x float> @test_waterfall_same_index3(<8 x i32> addrspace(4)* inreg %in, i32 %index, float %s, <4 x i32> inreg %samp) {
+; CHECK-LABEL: @test_waterfall_same_index3(
+; CHECK-NEXT:    [[WF_TOKEN1:%.*]] = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 [[INDEX:%.*]])
+; CHECK-NEXT:    [[S_IDX:%.*]] = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 [[WF_TOKEN1]], i32 [[INDEX]])
+; CHECK-NEXT:    [[TMP1:%.*]] = sext i32 [[S_IDX]] to i64
+; CHECK-NEXT:    [[PTR:%.*]] = getelementptr <8 x i32>, ptr addrspace(4) [[IN:%.*]], i64 [[TMP1]]
+; CHECK-NEXT:    [[RSRC:%.*]] = load <8 x i32>, ptr addrspace(4) [[PTR]], align 32
+; CHECK-NEXT:    [[R:%.*]] = call <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32.v8i32.v4i32(i32 15, float [[S:%.*]], <8 x i32> [[RSRC]], <4 x i32> [[SAMP:%.*]], i1 false, i32 0, i32 0)
+; CHECK-NEXT:    [[R1:%.*]] = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 [[WF_TOKEN1]], <4 x float> [[R]])
+; CHECK-NEXT:    ret <4 x float> [[R1]]
+;
+  %wf_token1 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %index)
+  %wf_token2 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 %wf_token1, i32 %index)
+  %wf_token3 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 %wf_token2, i32 %index)
+  %s_idx = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %wf_token3, i32 %index)
+  %ptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4) * %ptr, align 32
+  %r = call <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32(i32 15, float %s, <8 x i32> %rsrc, <4 x i32> %samp, i1 0, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %wf_token3, <4 x float> %r)
+  ret <4 x float> %r1
+}
+
+define amdgpu_ps <4 x float> @test_waterfall_same_index_aba(<8 x i32> addrspace(4)* inreg %in, i32 %index1, i32 %index2, float %s, <4 x i32> inreg %samp) {
+; CHECK-LABEL: @test_waterfall_same_index_aba(
+; CHECK-NEXT:    [[WF_TOKEN1:%.*]] = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 [[INDEX1:%.*]])
+; CHECK-NEXT:    [[WF_TOKEN2:%.*]] = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 [[WF_TOKEN1]], i32 [[INDEX2:%.*]])
+; CHECK-NEXT:    [[S_IDX:%.*]] = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 [[WF_TOKEN2]], i32 [[INDEX2]])
+; CHECK-NEXT:    [[TMP1:%.*]] = sext i32 [[S_IDX]] to i64
+; CHECK-NEXT:    [[PTR:%.*]] = getelementptr <8 x i32>, ptr addrspace(4) [[IN:%.*]], i64 [[TMP1]]
+; CHECK-NEXT:    [[RSRC:%.*]] = load <8 x i32>, ptr addrspace(4) [[PTR]], align 32
+; CHECK-NEXT:    [[R:%.*]] = call <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32.v8i32.v4i32(i32 15, float [[S:%.*]], <8 x i32> [[RSRC]], <4 x i32> [[SAMP:%.*]], i1 false, i32 0, i32 0)
+; CHECK-NEXT:    [[R1:%.*]] = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 [[WF_TOKEN2]], <4 x float> [[R]])
+; CHECK-NEXT:    ret <4 x float> [[R1]]
+;
+  %wf_token1 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %index1)
+  %wf_token2 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 %wf_token1, i32 %index2)
+  %wf_token3 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 %wf_token2, i32 %index1)
+  %s_idx = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %wf_token3, i32 %index2)
+  %ptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx
+  %rsrc = load <8 x i32>, <8 x i32> addrspace(4) * %ptr, align 32
+  %r = call <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32(i32 15, float %s, <8 x i32> %rsrc, <4 x i32> %samp, i1 0, i32 0, i32 0)
+  %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %wf_token3, <4 x float> %r)
+  ret <4 x float> %r1
+}
+
+declare i32 @llvm.amdgcn.waterfall.begin.i32(i32, i32)
+declare i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32, i32)
+declare <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32, <4 x float>)
+declare <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32(i32, float, <8 x i32>, <4 x i32>, i1, i32, i32)

>From b353b0bf9bfbf67d515bb06c82ca907b12d03d87 Mon Sep 17 00:00:00 2001
From: Greta Yorsh <Greta.Yorsh at amd.com>
Date: Thu, 16 Apr 2026 14:05:14 +0100
Subject: [PATCH 2/5] [AMDGPU] Rename SIInsertWaterfall pass to
 AMDGPUInsertWaterfall

Co-Authored-By: Claude Sonnet 4.6 <noreply at anthropic.com>
---
 llvm/lib/Target/AMDGPU/AMDGPU.h               |  6 +--
 ...aterfall.cpp => AMDGPUInsertWaterfall.cpp} | 38 +++++++++----------
 ...ertWaterfall.h => AMDGPUInsertWaterfall.h} | 11 +++---
 llvm/lib/Target/AMDGPU/AMDGPUPassRegistry.def |  2 +-
 .../Target/AMDGPU/AMDGPURegBankCombiner.cpp   |  2 +-
 .../lib/Target/AMDGPU/AMDGPUTargetMachine.cpp |  8 ++--
 llvm/lib/Target/AMDGPU/CMakeLists.txt         |  2 +-
 7 files changed, 35 insertions(+), 34 deletions(-)
 rename llvm/lib/Target/AMDGPU/{SIInsertWaterfall.cpp => AMDGPUInsertWaterfall.cpp} (96%)
 rename llvm/lib/Target/AMDGPU/{SIInsertWaterfall.h => AMDGPUInsertWaterfall.h} (61%)

diff --git a/llvm/lib/Target/AMDGPU/AMDGPU.h b/llvm/lib/Target/AMDGPU/AMDGPU.h
index 4edc25a391669..ff6e9226b60b8 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPU.h
+++ b/llvm/lib/Target/AMDGPU/AMDGPU.h
@@ -51,7 +51,7 @@ FunctionPass *createSIFixSGPRCopiesLegacyPass();
 FunctionPass *createLowerWWMCopiesPass();
 FunctionPass *createSIMemoryLegalizerPass();
 FunctionPass *createSIInsertWaitcntsPass();
-FunctionPass *createSIInsertWaterfallPass();
+FunctionPass *createAMDGPUInsertWaterfallPass();
 FunctionPass *createSIPreAllocateWWMRegsLegacyPass();
 FunctionPass *createSIFormMemoryClausesLegacyPass();
 
@@ -233,8 +233,8 @@ extern char &SILateBranchLoweringPassID;
 void initializeSIOptimizeExecMaskingLegacyPass(PassRegistry &);
 extern char &SIOptimizeExecMaskingLegacyID;
 
-void initializeSIInsertWaterfallPass(PassRegistry &);
-extern char &SIInsertWaterfallID;
+void initializeAMDGPUInsertWaterfallPass(PassRegistry &);
+extern char &AMDGPUInsertWaterfallID;
 
 void initializeSIPreAllocateWWMRegsLegacyPass(PassRegistry &);
 extern char &SIPreAllocateWWMRegsLegacyID;
diff --git a/llvm/lib/Target/AMDGPU/SIInsertWaterfall.cpp b/llvm/lib/Target/AMDGPU/AMDGPUInsertWaterfall.cpp
similarity index 96%
rename from llvm/lib/Target/AMDGPU/SIInsertWaterfall.cpp
rename to llvm/lib/Target/AMDGPU/AMDGPUInsertWaterfall.cpp
index f0f0662c1b28b..a7e1d97c36b7f 100644
--- a/llvm/lib/Target/AMDGPU/SIInsertWaterfall.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUInsertWaterfall.cpp
@@ -1,4 +1,4 @@
-//===- SIInsertWaterfall.cpp - insert waterall loops at intrinsic markers -===//
+//===- AMDGPUInsertWaterfall.cpp - insert waterfall loops at intrinsic markers -===//
 //
 //                     The LLVM Compiler Infrastructure
 //
@@ -14,7 +14,7 @@
 ///
 //===----------------------------------------------------------------------===//
 
-#include "SIInsertWaterfall.h"
+#include "AMDGPUInsertWaterfall.h"
 #include "AMDGPU.h"
 #include "GCNSubtarget.h"
 #include "MCTargetDesc/AMDGPUMCTargetDesc.h"
@@ -23,7 +23,7 @@
 
 using namespace llvm;
 
-#define DEBUG_TYPE "si-insert-waterfall"
+#define DEBUG_TYPE "amdgpu-insert-waterfall"
 
 namespace {
 
@@ -272,7 +272,7 @@ static void replaceRegIncSubReg(const MachineRegisterInfo *MRI,
     O.substVirtReg(To->getReg(), To->getSubReg(), *TRI);
 }
 
-class SIInsertWaterfall : public MachineFunctionPass {
+class AMDGPUInsertWaterfall : public MachineFunctionPass {
 private:
   struct WaterfallWorkitem {
     const SIInstrInfo *TII;
@@ -435,8 +435,8 @@ class SIInsertWaterfall : public MachineFunctionPass {
 public:
   static char ID;
 
-  SIInsertWaterfall() : MachineFunctionPass(ID) {
-    initializeSIInsertWaterfallPass(*PassRegistry::getPassRegistry());
+  AMDGPUInsertWaterfall() : MachineFunctionPass(ID) {
+    initializeAMDGPUInsertWaterfallPass(*PassRegistry::getPassRegistry());
   }
 
   void getAnalysisUsage(AnalysisUsage &AU) const override {
@@ -453,27 +453,27 @@ class SIInsertWaterfall : public MachineFunctionPass {
 
 } // End anonymous namespace.
 
-INITIALIZE_PASS(SIInsertWaterfall, DEBUG_TYPE, "SI Insert waterfalls", false,
-                false)
+INITIALIZE_PASS(AMDGPUInsertWaterfall, DEBUG_TYPE, "AMDGPU Insert waterfalls",
+                false, false)
 
-char SIInsertWaterfall::ID = 0;
+char AMDGPUInsertWaterfall::ID = 0;
 
-char &llvm::SIInsertWaterfallID = SIInsertWaterfall::ID;
+char &llvm::AMDGPUInsertWaterfallID = AMDGPUInsertWaterfall::ID;
 
-FunctionPass *llvm::createSIInsertWaterfallPass() {
-  return new SIInsertWaterfall;
+FunctionPass *llvm::createAMDGPUInsertWaterfallPass() {
+  return new AMDGPUInsertWaterfall;
 }
 
 PreservedAnalyses
-SIInsertWaterfallPass::run(MachineFunction &MF,
-                           MachineFunctionAnalysisManager &MFAM) {
-  SIInsertWaterfall Impl;
+AMDGPUInsertWaterfallPass::run(MachineFunction &MF,
+                               MachineFunctionAnalysisManager &MFAM) {
+  AMDGPUInsertWaterfall Impl;
   if (!Impl.runOnMachineFunction(MF))
     return PreservedAnalyses::all();
   return PreservedAnalyses::none();
 }
 
-bool SIInsertWaterfall::removeRedundantWaterfall(WaterfallWorkitem &Item) {
+bool AMDGPUInsertWaterfall::removeRedundantWaterfall(WaterfallWorkitem &Item) {
   // In some cases, the waterfall is actually redundant
   // If all the readfirstlane intrinsics are actually for uniform values and
   // the token used in the begin/end isn't used in anything else the waterfall
@@ -576,7 +576,7 @@ bool SIInsertWaterfall::removeRedundantWaterfall(WaterfallWorkitem &Item) {
   return LoopRemoved;
 }
 
-bool SIInsertWaterfall::processWaterfall(MachineBasicBlock &MBB) {
+bool AMDGPUInsertWaterfall::processWaterfall(MachineBasicBlock &MBB) {
   bool Changed = false;
   MachineFunction &MF = *MBB.getParent();
   MachineBasicBlock *CurrMBB = &MBB;
@@ -837,12 +837,12 @@ bool SIInsertWaterfall::processWaterfall(MachineBasicBlock &MBB) {
   return Changed;
 }
 
-Register SIInsertWaterfall::getToken(MachineInstr *MI) {
+Register AMDGPUInsertWaterfall::getToken(MachineInstr *MI) {
   auto CandTokMO = TII->getNamedOperand(*MI, AMDGPU::OpName::tok);
   return CandTokMO->isReg() ? CandTokMO->getReg() : AMDGPU::NoRegister;
 }
 
-bool SIInsertWaterfall::runOnMachineFunction(MachineFunction &MF) {
+bool AMDGPUInsertWaterfall::runOnMachineFunction(MachineFunction &MF) {
   bool Changed = false;
 
   ST = &MF.getSubtarget<GCNSubtarget>();
diff --git a/llvm/lib/Target/AMDGPU/SIInsertWaterfall.h b/llvm/lib/Target/AMDGPU/AMDGPUInsertWaterfall.h
similarity index 61%
rename from llvm/lib/Target/AMDGPU/SIInsertWaterfall.h
rename to llvm/lib/Target/AMDGPU/AMDGPUInsertWaterfall.h
index 406df5a8b3b67..1c337b34ab2c2 100644
--- a/llvm/lib/Target/AMDGPU/SIInsertWaterfall.h
+++ b/llvm/lib/Target/AMDGPU/AMDGPUInsertWaterfall.h
@@ -1,4 +1,4 @@
-//===- SIInsertWaterfall.h --------------------------------------*- C++ -*-===//
+//===- AMDGPUInsertWaterfall.h ----------------------------------*- C++ -*-===//
 //
 // Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
 // See https://llvm.org/LICENSE.txt for license information.
@@ -6,17 +6,18 @@
 //
 //===----------------------------------------------------------------------===//
 
-#ifndef LLVM_LIB_TARGET_AMDGPU_SIINSERTWATERFALL_H
-#define LLVM_LIB_TARGET_AMDGPU_SIINSERTWATERFALL_H
+#ifndef LLVM_LIB_TARGET_AMDGPU_AMDGPUINSERTWATERFALL_H
+#define LLVM_LIB_TARGET_AMDGPU_AMDGPUINSERTWATERFALL_H
 
 #include "llvm/CodeGen/MachinePassManager.h"
 
 namespace llvm {
-class SIInsertWaterfallPass : public PassInfoMixin<SIInsertWaterfallPass> {
+class AMDGPUInsertWaterfallPass
+    : public PassInfoMixin<AMDGPUInsertWaterfallPass> {
 public:
   PreservedAnalyses run(MachineFunction &MF,
                         MachineFunctionAnalysisManager &MFAM);
 };
 } // namespace llvm
 
-#endif // LLVM_LIB_TARGET_AMDGPU_SIINSERTWATERFALL_H
+#endif // LLVM_LIB_TARGET_AMDGPU_AMDGPUINSERTWATERFALL_H
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUPassRegistry.def b/llvm/lib/Target/AMDGPU/AMDGPUPassRegistry.def
index 6f95ef64e1f22..259be87446206 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUPassRegistry.def
+++ b/llvm/lib/Target/AMDGPU/AMDGPUPassRegistry.def
@@ -135,7 +135,7 @@ MACHINE_FUNCTION_PASS("si-form-memory-clauses", SIFormMemoryClausesPass())
 MACHINE_FUNCTION_PASS("si-i1-copies", SILowerI1CopiesPass())
 MACHINE_FUNCTION_PASS("si-insert-hard-clauses", SIInsertHardClausesPass())
 MACHINE_FUNCTION_PASS("si-insert-waitcnts", SIInsertWaitcntsPass())
-MACHINE_FUNCTION_PASS("si-insert-waterfall", SIInsertWaterfallPass())
+MACHINE_FUNCTION_PASS("si-insert-waterfall", AMDGPUInsertWaterfallPass())
 MACHINE_FUNCTION_PASS("si-late-branch-lowering", SILateBranchLoweringPass())
 MACHINE_FUNCTION_PASS("si-load-store-opt", SILoadStoreOptimizerPass())
 MACHINE_FUNCTION_PASS("si-lower-control-flow", SILowerControlFlowPass())
diff --git a/llvm/lib/Target/AMDGPU/AMDGPURegBankCombiner.cpp b/llvm/lib/Target/AMDGPU/AMDGPURegBankCombiner.cpp
index e2df70b3964ca..3cabd3adc089b 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPURegBankCombiner.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPURegBankCombiner.cpp
@@ -98,7 +98,7 @@ class AMDGPURegBankCombinerImpl : public Combiner {
 
   // Combine support to remove amdgcn_waterfall_readfirstlane or
   // amdgcn_waterfall_begin intrinsics if the index is determined to be
-  // uniform. The SIInsertWaterfall pass can handle their removal and in some
+  // uniform. The AMDGPUInsertWaterfall pass can handle their removal and in some
   // cases remove the waterfall altogether
   bool matchRmUniformWF(MachineInstr &MI,
                         RmUniformWFMatchInfo &MatchInfo) const;
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUTargetMachine.cpp b/llvm/lib/Target/AMDGPU/AMDGPUTargetMachine.cpp
index 73201faaf31b9..7cd3b37ff3ac4 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUTargetMachine.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUTargetMachine.cpp
@@ -52,7 +52,7 @@
 #include "SIFixVGPRCopies.h"
 #include "SIFoldOperands.h"
 #include "SIFormMemoryClauses.h"
-#include "SIInsertWaterfall.h"
+#include "AMDGPUInsertWaterfall.h"
 #include "SILoadStoreOptimizer.h"
 #include "SILowerControlFlow.h"
 #include "SILowerSGPRSpills.h"
@@ -674,7 +674,7 @@ extern "C" LLVM_ABI LLVM_EXTERNAL_VISIBILITY void LLVMInitializeAMDGPUTarget() {
   initializeSIFixSGPRCopiesLegacyPass(*PR);
   initializeSIFixVGPRCopiesLegacyPass(*PR);
   initializeSIFoldOperandsLegacyPass(*PR);
-  initializeSIInsertWaterfallPass(*PR);
+  initializeAMDGPUInsertWaterfallPass(*PR);
   initializeSIPeepholeSDWALegacyPass(*PR);
   initializeSIShrinkInstructionsLegacyPass(*PR);
   initializeSIOptimizeExecMaskingPreRALegacyPass(*PR);
@@ -1711,7 +1711,7 @@ void GCNPassConfig::addFastRegAlloc() {
 }
 
 void GCNPassConfig::addPreRegAlloc() {
-  addPass(createSIInsertWaterfallPass());
+  addPass(createAMDGPUInsertWaterfallPass());
   if (getOptLevel() != CodeGenOptLevel::None)
     addPass(&AMDGPUPrepareAGPRAllocLegacyID);
 }
@@ -2526,7 +2526,7 @@ Error AMDGPUCodeGenPassBuilder::addOptimizedRegAlloc(
 }
 
 void AMDGPUCodeGenPassBuilder::addPreRegAlloc(PassManagerWrapper &PMW) const {
-  addMachineFunctionPass(SIInsertWaterfallPass(), PMW);
+  addMachineFunctionPass(AMDGPUInsertWaterfallPass(), PMW);
   if (getOptLevel() != CodeGenOptLevel::None)
     addMachineFunctionPass(AMDGPUPrepareAGPRAllocPass(), PMW);
 }
diff --git a/llvm/lib/Target/AMDGPU/CMakeLists.txt b/llvm/lib/Target/AMDGPU/CMakeLists.txt
index 28f85cee5a67f..54b49ca6792ca 100644
--- a/llvm/lib/Target/AMDGPU/CMakeLists.txt
+++ b/llvm/lib/Target/AMDGPU/CMakeLists.txt
@@ -62,6 +62,7 @@ add_llvm_target(AMDGPUCodeGen
   AMDGPUHazardLatency.cpp
   AMDGPUHSAMetadataStreamer.cpp
   AMDGPUInsertDelayAlu.cpp
+  AMDGPUInsertWaterfall.cpp
   AMDGPUInstCombineIntrinsic.cpp
   AMDGPUUniformIntrinsicCombine.cpp
   AMDGPUInstrInfo.cpp
@@ -164,7 +165,6 @@ add_llvm_target(AMDGPUCodeGen
   SIFrameLowering.cpp
   SIInsertHardClauses.cpp
   SIInsertWaitcnts.cpp
-  SIInsertWaterfall.cpp
   SIInstrInfo.cpp
   SIISelLowering.cpp
   SILateBranchLowering.cpp

>From 07830ec80993af2fe9f3f8d03d0aed223f7b3c59 Mon Sep 17 00:00:00 2001
From: Greta Yorsh <Greta.Yorsh at amd.com>
Date: Thu, 16 Apr 2026 14:38:00 +0100
Subject: [PATCH 3/5] Fix test 
 llvm/test/Transforms/InstCombine/AMDGPU/waterfall.ll

---
 llvm/test/Transforms/InstCombine/AMDGPU/waterfall.ll | 6 +++---
 1 file changed, 3 insertions(+), 3 deletions(-)

diff --git a/llvm/test/Transforms/InstCombine/AMDGPU/waterfall.ll b/llvm/test/Transforms/InstCombine/AMDGPU/waterfall.ll
index df57d55e2f674..367e65fb68045 100644
--- a/llvm/test/Transforms/InstCombine/AMDGPU/waterfall.ll
+++ b/llvm/test/Transforms/InstCombine/AMDGPU/waterfall.ll
@@ -6,7 +6,7 @@ define amdgpu_ps <4 x float> @test_waterfall_same_index2(<8 x i32> addrspace(4)*
 ; CHECK-NEXT:    [[WF_TOKEN1:%.*]] = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 [[INDEX:%.*]])
 ; CHECK-NEXT:    [[S_IDX:%.*]] = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 [[WF_TOKEN1]], i32 [[INDEX]])
 ; CHECK-NEXT:    [[TMP1:%.*]] = sext i32 [[S_IDX]] to i64
-; CHECK-NEXT:    [[PTR:%.*]] = getelementptr <8 x i32>, ptr addrspace(4) [[IN:%.*]], i64 [[TMP1]]
+; CHECK-NEXT:    [[PTR:%.*]] = getelementptr [32 x i8], ptr addrspace(4) [[IN:%.*]], i64 [[TMP1]]
 ; CHECK-NEXT:    [[RSRC:%.*]] = load <8 x i32>, ptr addrspace(4) [[PTR]], align 32
 ; CHECK-NEXT:    [[R:%.*]] = call <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32.v8i32.v4i32(i32 15, float [[S:%.*]], <8 x i32> [[RSRC]], <4 x i32> [[SAMP:%.*]], i1 false, i32 0, i32 0)
 ; CHECK-NEXT:    [[R1:%.*]] = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 [[WF_TOKEN1]], <4 x float> [[R]])
@@ -27,7 +27,7 @@ define amdgpu_ps <4 x float> @test_waterfall_same_index3(<8 x i32> addrspace(4)*
 ; CHECK-NEXT:    [[WF_TOKEN1:%.*]] = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 [[INDEX:%.*]])
 ; CHECK-NEXT:    [[S_IDX:%.*]] = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 [[WF_TOKEN1]], i32 [[INDEX]])
 ; CHECK-NEXT:    [[TMP1:%.*]] = sext i32 [[S_IDX]] to i64
-; CHECK-NEXT:    [[PTR:%.*]] = getelementptr <8 x i32>, ptr addrspace(4) [[IN:%.*]], i64 [[TMP1]]
+; CHECK-NEXT:    [[PTR:%.*]] = getelementptr [32 x i8], ptr addrspace(4) [[IN:%.*]], i64 [[TMP1]]
 ; CHECK-NEXT:    [[RSRC:%.*]] = load <8 x i32>, ptr addrspace(4) [[PTR]], align 32
 ; CHECK-NEXT:    [[R:%.*]] = call <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32.v8i32.v4i32(i32 15, float [[S:%.*]], <8 x i32> [[RSRC]], <4 x i32> [[SAMP:%.*]], i1 false, i32 0, i32 0)
 ; CHECK-NEXT:    [[R1:%.*]] = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 [[WF_TOKEN1]], <4 x float> [[R]])
@@ -50,7 +50,7 @@ define amdgpu_ps <4 x float> @test_waterfall_same_index_aba(<8 x i32> addrspace(
 ; CHECK-NEXT:    [[WF_TOKEN2:%.*]] = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 [[WF_TOKEN1]], i32 [[INDEX2:%.*]])
 ; CHECK-NEXT:    [[S_IDX:%.*]] = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 [[WF_TOKEN2]], i32 [[INDEX2]])
 ; CHECK-NEXT:    [[TMP1:%.*]] = sext i32 [[S_IDX]] to i64
-; CHECK-NEXT:    [[PTR:%.*]] = getelementptr <8 x i32>, ptr addrspace(4) [[IN:%.*]], i64 [[TMP1]]
+; CHECK-NEXT:    [[PTR:%.*]] = getelementptr [32 x i8], ptr addrspace(4) [[IN:%.*]], i64 [[TMP1]]
 ; CHECK-NEXT:    [[RSRC:%.*]] = load <8 x i32>, ptr addrspace(4) [[PTR]], align 32
 ; CHECK-NEXT:    [[R:%.*]] = call <4 x float> @llvm.amdgcn.image.sample.1d.v4f32.f32.v8i32.v4i32(i32 15, float [[S:%.*]], <8 x i32> [[RSRC]], <4 x i32> [[SAMP:%.*]], i1 false, i32 0, i32 0)
 ; CHECK-NEXT:    [[R1:%.*]] = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 [[WF_TOKEN2]], <4 x float> [[R]])

>From 7f177cc6a6c04efc637a6b254cee1c1113c182fd Mon Sep 17 00:00:00 2001
From: Greta Yorsh <Greta.Yorsh at amd.com>
Date: Fri, 17 Apr 2026 08:41:33 +0100
Subject: [PATCH 4/5] Replace "undef" with "poison" or "dummy" arguments in a
 new test.

---
 .../CodeGen/AMDGPU/llvm.amdgcn.waterfall.ll   | 58 +++++++++----------
 1 file changed, 29 insertions(+), 29 deletions(-)

diff --git a/llvm/test/CodeGen/AMDGPU/llvm.amdgcn.waterfall.ll b/llvm/test/CodeGen/AMDGPU/llvm.amdgcn.waterfall.ll
index 4ceb4245c77e9..755a730edab65 100644
--- a/llvm/test/CodeGen/AMDGPU/llvm.amdgcn.waterfall.ll
+++ b/llvm/test/CodeGen/AMDGPU/llvm.amdgcn.waterfall.ll
@@ -12,7 +12,7 @@
 ; RUN: llc -global-isel=0 -march=amdgcn -mcpu=gfx1200 -mattr=-wavefrontsize32,+wavefrontsize64 -verify-machineinstrs < %s | FileCheck -check-prefixes=GFX12,GFX12-SDAG %s
 ; RUN: llc -global-isel=1 -march=amdgcn -mcpu=gfx1200 -mattr=-wavefrontsize32,+wavefrontsize64 -verify-machineinstrs < %s | FileCheck -check-prefixes=GFX12,GFX12-GISEL %s
 
- at Lds = addrspace(3) global [16384 x i32] undef
+ at Lds = addrspace(3) global [16384 x i32] poison
 
 define amdgpu_ps void @test_waterfall_readlane(i32 addrspace(1)* inreg %out, <2 x i32> addrspace(1)* inreg %in, i32 %tid) #1 {
 ; VI-SDAG-LABEL: test_waterfall_readlane:
@@ -2265,7 +2265,7 @@ define amdgpu_ps <4 x float> @test_waterfall_non_uniform_img_multi_rl(<8 x i32>
   ret <4 x float> %r1
 }
 
-define amdgpu_ps <4 x float> @test_waterfall_non_uni_img_2_idx(<8 x i32> addrspace(4)* inreg %in, <4 x i32> addrspace(4)* inreg %samp_in, i32 %index1, i32 %index2, float %s) #1 {
+define amdgpu_ps <4 x float> @test_waterfall_non_uni_img_2_idx(<8 x i32> addrspace(4)* inreg %in, <4 x i32> addrspace(4)* inreg %samp_in, i32 %index1, i32 %index2, float %s, <2 x i32> %dummy) #1 {
 ; VI-SDAG-LABEL: test_waterfall_non_uni_img_2_idx:
 ; VI-SDAG:       ; %bb.0:
 ; VI-SDAG-NEXT:    s_mov_b64 s[4:5], exec
@@ -2743,7 +2743,7 @@ define amdgpu_ps <4 x float> @test_waterfall_non_uni_img_2_idx(<8 x i32> addrspa
 ; GFX12-GISEL-NEXT:    s_and_b64 exec, exec, s[4:5]
 ; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
 ; GFX12-GISEL-NEXT:    ; return to shader part epilog
-  %t_idx = insertelement <2 x i32> undef, i32 %index1, i32 0
+  %t_idx = insertelement <2 x i32> %dummy, i32 %index1, i32 0
   %combined_idx = insertelement <2 x i32> %t_idx, i32 %index2, i32 1
   %wf_token = call i32 @llvm.amdgcn.waterfall.begin.v2i32(i32 0, <2 x i32> %combined_idx)
   %s_c_idx = call <2 x i32> @llvm.amdgcn.waterfall.readfirstlane.v2i32.v2i32(i32 %wf_token, <2 x i32> %combined_idx)
@@ -3850,7 +3850,7 @@ define amdgpu_ps <4 x float> @test_keep_waterfall_multi_rl(<8 x i32> addrspace(4
   ret <4 x float> %r1
 }
 
-define amdgpu_ps void @test_waterfall_sample_with_kill(<8 x i32> addrspace(4)* inreg %in, <4 x i32> addrspace(4)* inreg %samp_in, i32 %index, float %s, i32 inreg %val) #1 {
+define amdgpu_ps void @test_waterfall_sample_with_kill(<8 x i32> addrspace(4)* inreg %in, <4 x i32> addrspace(4)* inreg %samp_in, i32 %index, float %s, i32 inreg %val, float %dummy_f0, float %dummy_f1, float %dummy_f2) #1 {
 ; VI-LABEL: test_waterfall_sample_with_kill:
 ; VI:       ; %bb.0:
 ; VI-NEXT:    s_mov_b64 s[6:7], exec
@@ -3872,7 +3872,7 @@ define amdgpu_ps void @test_waterfall_sample_with_kill(<8 x i32> addrspace(4)* i
 ; VI-NEXT:    s_load_dwordx4 s[20:23], s[20:21], 0x0
 ; VI-NEXT:    ; implicit-def: $vgpr0
 ; VI-NEXT:    s_waitcnt lgkmcnt(0)
-; VI-NEXT:    image_sample v[2:5], v1, s[12:19], s[20:23] dmask:0xf
+; VI-NEXT:    image_sample v[5:8], v1, s[12:19], s[20:23] dmask:0xf
 ; VI-NEXT:    ; implicit-def: $vgpr1
 ; VI-NEXT:    s_xor_b64 exec, exec, s[24:25]
 ; VI-NEXT:    s_cbranch_execnz .LBB10_1
@@ -3880,7 +3880,7 @@ define amdgpu_ps void @test_waterfall_sample_with_kill(<8 x i32> addrspace(4)* i
 ; VI-NEXT:    s_mov_b64 exec, s[8:9]
 ; VI-NEXT:    s_and_b64 exec, exec, s[6:7]
 ; VI-NEXT:    s_waitcnt vmcnt(0)
-; VI-NEXT:    v_cmp_gt_f32_e32 vcc, 0, v2
+; VI-NEXT:    v_cmp_gt_f32_e32 vcc, 0, v5
 ; VI-NEXT:    s_and_saveexec_b64 s[0:1], vcc
 ; VI-NEXT:    s_xor_b64 s[0:1], exec, s[0:1]
 ; VI-NEXT:    s_cbranch_execz .LBB10_5
@@ -3920,7 +3920,7 @@ define amdgpu_ps void @test_waterfall_sample_with_kill(<8 x i32> addrspace(4)* i
 ; GFX9-NEXT:    s_load_dwordx4 s[20:23], s[26:27], 0x0
 ; GFX9-NEXT:    ; implicit-def: $vgpr0
 ; GFX9-NEXT:    s_waitcnt lgkmcnt(0)
-; GFX9-NEXT:    image_sample v[2:5], v1, s[12:19], s[20:23] dmask:0xf
+; GFX9-NEXT:    image_sample v[5:8], v1, s[12:19], s[20:23] dmask:0xf
 ; GFX9-NEXT:    ; implicit-def: $vgpr1
 ; GFX9-NEXT:    s_xor_b64 exec, exec, s[24:25]
 ; GFX9-NEXT:    s_cbranch_execnz .LBB10_1
@@ -3928,7 +3928,7 @@ define amdgpu_ps void @test_waterfall_sample_with_kill(<8 x i32> addrspace(4)* i
 ; GFX9-NEXT:    s_mov_b64 exec, s[8:9]
 ; GFX9-NEXT:    s_and_b64 exec, exec, s[6:7]
 ; GFX9-NEXT:    s_waitcnt vmcnt(0)
-; GFX9-NEXT:    v_cmp_gt_f32_e32 vcc, 0, v2
+; GFX9-NEXT:    v_cmp_gt_f32_e32 vcc, 0, v5
 ; GFX9-NEXT:    s_and_saveexec_b64 s[0:1], vcc
 ; GFX9-NEXT:    s_xor_b64 s[0:1], exec, s[0:1]
 ; GFX9-NEXT:    s_cbranch_execz .LBB10_5
@@ -3968,7 +3968,7 @@ define amdgpu_ps void @test_waterfall_sample_with_kill(<8 x i32> addrspace(4)* i
 ; GFX10-32-NEXT:    s_load_dwordx8 s[12:19], s[10:11], 0x0
 ; GFX10-32-NEXT:    s_load_dwordx4 s[20:23], s[24:25], 0x0
 ; GFX10-32-NEXT:    s_waitcnt lgkmcnt(0)
-; GFX10-32-NEXT:    image_sample v[2:5], v1, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-32-NEXT:    image_sample v[5:8], v1, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
 ; GFX10-32-NEXT:    s_andn2_wrexec_b32 s8, s8
 ; GFX10-32-NEXT:    ; implicit-def: $vgpr0
 ; GFX10-32-NEXT:    ; implicit-def: $vgpr1
@@ -3978,7 +3978,7 @@ define amdgpu_ps void @test_waterfall_sample_with_kill(<8 x i32> addrspace(4)* i
 ; GFX10-32-NEXT:    s_mov_b32 exec_lo, s7
 ; GFX10-32-NEXT:    s_and_b32 exec_lo, exec_lo, s6
 ; GFX10-32-NEXT:    s_waitcnt vmcnt(0)
-; GFX10-32-NEXT:    v_cmp_gt_f32_e32 vcc_lo, 0, v2
+; GFX10-32-NEXT:    v_cmp_gt_f32_e32 vcc_lo, 0, v5
 ; GFX10-32-NEXT:    s_and_saveexec_b32 s0, vcc_lo
 ; GFX10-32-NEXT:    s_xor_b32 s0, exec_lo, s0
 ; GFX10-32-NEXT:    s_cbranch_execz .LBB10_5
@@ -4018,7 +4018,7 @@ define amdgpu_ps void @test_waterfall_sample_with_kill(<8 x i32> addrspace(4)* i
 ; GFX10-64-NEXT:    s_load_dwordx8 s[12:19], s[24:25], 0x0
 ; GFX10-64-NEXT:    s_load_dwordx4 s[20:23], s[26:27], 0x0
 ; GFX10-64-NEXT:    s_waitcnt lgkmcnt(0)
-; GFX10-64-NEXT:    image_sample v[2:5], v1, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX10-64-NEXT:    image_sample v[5:8], v1, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
 ; GFX10-64-NEXT:    s_andn2_wrexec_b64 s[10:11], s[10:11]
 ; GFX10-64-NEXT:    ; implicit-def: $vgpr0
 ; GFX10-64-NEXT:    ; implicit-def: $vgpr1
@@ -4028,7 +4028,7 @@ define amdgpu_ps void @test_waterfall_sample_with_kill(<8 x i32> addrspace(4)* i
 ; GFX10-64-NEXT:    s_mov_b64 exec, s[8:9]
 ; GFX10-64-NEXT:    s_and_b64 exec, exec, s[6:7]
 ; GFX10-64-NEXT:    s_waitcnt vmcnt(0)
-; GFX10-64-NEXT:    v_cmp_gt_f32_e32 vcc, 0, v2
+; GFX10-64-NEXT:    v_cmp_gt_f32_e32 vcc, 0, v5
 ; GFX10-64-NEXT:    s_and_saveexec_b64 s[0:1], vcc
 ; GFX10-64-NEXT:    s_xor_b64 s[0:1], exec, s[0:1]
 ; GFX10-64-NEXT:    s_cbranch_execz .LBB10_5
@@ -4072,7 +4072,7 @@ define amdgpu_ps void @test_waterfall_sample_with_kill(<8 x i32> addrspace(4)* i
 ; GFX1150-NEXT:    s_load_b256 s[12:19], s[12:13], 0x0
 ; GFX1150-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
 ; GFX1150-NEXT:    s_waitcnt vmcnt(0) lgkmcnt(0)
-; GFX1150-NEXT:    image_sample v[2:5], v1, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX1150-NEXT:    image_sample v[5:8], v1, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
 ; GFX1150-NEXT:    s_and_not1_wrexec_b64 s[10:11], s[10:11]
 ; GFX1150-NEXT:    ; implicit-def: $vgpr0
 ; GFX1150-NEXT:    ; implicit-def: $vgpr1
@@ -4083,7 +4083,7 @@ define amdgpu_ps void @test_waterfall_sample_with_kill(<8 x i32> addrspace(4)* i
 ; GFX1150-NEXT:    s_and_b64 exec, exec, s[6:7]
 ; GFX1150-NEXT:    s_mov_b64 s[0:1], exec
 ; GFX1150-NEXT:    s_waitcnt vmcnt(0)
-; GFX1150-NEXT:    v_cmpx_gt_f32_e32 0, v2
+; GFX1150-NEXT:    v_cmpx_gt_f32_e32 0, v5
 ; GFX1150-NEXT:    s_xor_b64 s[0:1], exec, s[0:1]
 ; GFX1150-NEXT:    s_cbranch_execz .LBB10_5
 ; GFX1150-NEXT:  ; %bb.3: ; %.kill
@@ -4132,7 +4132,7 @@ define amdgpu_ps void @test_waterfall_sample_with_kill(<8 x i32> addrspace(4)* i
 ; GFX12-SDAG-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
 ; GFX12-SDAG-NEXT:    s_wait_kmcnt 0x0
 ; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
-; GFX12-SDAG-NEXT:    image_sample v[2:5], v1, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-SDAG-NEXT:    image_sample v[5:8], v1, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
 ; GFX12-SDAG-NEXT:    s_and_not1_wrexec_b64 s[10:11], s[10:11]
 ; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr0
 ; GFX12-SDAG-NEXT:    ; implicit-def: $vgpr1
@@ -4143,7 +4143,7 @@ define amdgpu_ps void @test_waterfall_sample_with_kill(<8 x i32> addrspace(4)* i
 ; GFX12-SDAG-NEXT:    s_and_b64 exec, exec, s[6:7]
 ; GFX12-SDAG-NEXT:    s_mov_b64 s[0:1], exec
 ; GFX12-SDAG-NEXT:    s_wait_samplecnt 0x0
-; GFX12-SDAG-NEXT:    v_cmpx_gt_f32_e32 0, v2
+; GFX12-SDAG-NEXT:    v_cmpx_gt_f32_e32 0, v5
 ; GFX12-SDAG-NEXT:    s_xor_b64 s[0:1], exec, s[0:1]
 ; GFX12-SDAG-NEXT:    s_cbranch_execz .LBB10_5
 ; GFX12-SDAG-NEXT:  ; %bb.3: ; %.kill
@@ -4189,7 +4189,7 @@ define amdgpu_ps void @test_waterfall_sample_with_kill(<8 x i32> addrspace(4)* i
 ; GFX12-GISEL-NEXT:    s_load_b128 s[20:23], s[20:21], 0x0
 ; GFX12-GISEL-NEXT:    s_wait_kmcnt 0x0
 ; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
-; GFX12-GISEL-NEXT:    image_sample v[2:5], v1, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
+; GFX12-GISEL-NEXT:    image_sample v[5:8], v1, s[12:19], s[20:23] dmask:0xf dim:SQ_RSRC_IMG_1D
 ; GFX12-GISEL-NEXT:    s_and_not1_wrexec_b64 s[10:11], s[10:11]
 ; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr0
 ; GFX12-GISEL-NEXT:    ; implicit-def: $vgpr1
@@ -4200,7 +4200,7 @@ define amdgpu_ps void @test_waterfall_sample_with_kill(<8 x i32> addrspace(4)* i
 ; GFX12-GISEL-NEXT:    s_and_b64 exec, exec, s[6:7]
 ; GFX12-GISEL-NEXT:    s_mov_b64 s[0:1], exec
 ; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
-; GFX12-GISEL-NEXT:    v_cmpx_gt_f32_e32 0, v2
+; GFX12-GISEL-NEXT:    v_cmpx_gt_f32_e32 0, v5
 ; GFX12-GISEL-NEXT:    s_xor_b64 s[0:1], exec, s[0:1]
 ; GFX12-GISEL-NEXT:    s_cbranch_execz .LBB10_5
 ; GFX12-GISEL-NEXT:  ; %bb.3: ; %.kill
@@ -4236,7 +4236,7 @@ define amdgpu_ps void @test_waterfall_sample_with_kill(<8 x i32> addrspace(4)* i
   br label %.exit
 
 .exit:
-  call void @llvm.amdgcn.exp.f32(i32 immarg 0, i32 immarg 1, float 0.000000e+00, float undef, float undef, float undef, i1 immarg true, i1 immarg true)
+  call void @llvm.amdgcn.exp.f32(i32 immarg 0, i32 immarg 1, float 0.000000e+00, float %dummy_f0, float %dummy_f1, float %dummy_f2, i1 immarg true, i1 immarg true)
   ret void
 }
 
@@ -9468,7 +9468,7 @@ define amdgpu_ps {<4 x float>,<4 x float>} @test_waterfall_multi_begin_uniform_i
 ; GFX12-GISEL-NEXT:    ; return to shader part epilog
        <8 x i32> addrspace(4)* inreg %in, <4 x i32> addrspace(4)* inreg %s_in,
        i32 %idx1, i32 %idx2, i32 inreg %idx3, i32 inreg%idx4,
-       i32 %s_idx, i32 %s_idx2) #1 {
+       i32 %s_idx, i32 %s_idx2, { <4 x float>, <4 x float> } %dummy) #1 {
   %rptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx
   %sptr = getelementptr <4 x i32>, <4 x i32> addrspace(4)* %s_in, i32 %s_idx2
   %rsrc = load <8 x i32>, <8 x i32> addrspace(4)* %rptr, align 16
@@ -9487,7 +9487,7 @@ define amdgpu_ps {<4 x float>,<4 x float>} @test_waterfall_multi_begin_uniform_i
   %r2 = call <4 x float> @llvm.amdgcn.image.sample.2d.v4f32.f32(i32 15, float 0.000000e+00, float 0.000000e+00, <8 x i32> %s_rsrc1, <4 x i32> %s_srsrc1, i1 false, i32 0, i32 0)
   %r3 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok3, <4 x float> %r2)
 
-  %insert = insertvalue { <4 x float>, <4 x float> } undef, <4 x float> %r1, 0
+  %insert = insertvalue { <4 x float>, <4 x float> } %dummy, <4 x float> %r1, 0
   %insert1 = insertvalue { <4 x float>, <4 x float> } %insert, <4 x float> %r3, 1
   ret {<4 x float> , <4 x float>} %insert1
 }
@@ -10737,7 +10737,7 @@ define amdgpu_ps {<4 x float>,<4 x float>} @test_waterfall_multi_end_1loop(
 ; GFX12-GISEL-NEXT:    s_wait_alu depctr_sa_sdst(0)
 ; GFX12-GISEL-NEXT:    ; return to shader part epilog
        <8 x i32> addrspace(4)* inreg %in, <4 x i32> addrspace(4)* inreg %s_in,
-       i32 %idx1, i32 %idx2, i32 %s_idx, i32 %s_idx2) #1 {
+       i32 %idx1, i32 %idx2, i32 %s_idx, i32 %s_idx2, { <4 x float>, <4 x float> } %dummy) #1 {
   %rptr = getelementptr <8 x i32>, <8 x i32> addrspace(4)* %in, i32 %s_idx
   %sptr = getelementptr <4 x i32>, <4 x i32> addrspace(4)* %s_in, i32 %s_idx2
   %rsrc = load <8 x i32>, <8 x i32> addrspace(4)* %rptr, align 16
@@ -10751,7 +10751,7 @@ define amdgpu_ps {<4 x float>,<4 x float>} @test_waterfall_multi_end_1loop(
   %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok1, <4 x float> %r)
   %r3 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok1, <4 x float> %r2)
 
-  %insert = insertvalue { <4 x float>, <4 x float> } undef, <4 x float> %r1, 0
+  %insert = insertvalue { <4 x float>, <4 x float> } %dummy, <4 x float> %r1, 0
   %insert1 = insertvalue { <4 x float>, <4 x float> } %insert, <4 x float> %r3, 1
   ret {<4 x float> , <4 x float>} %insert1
 }
@@ -11346,7 +11346,7 @@ define amdgpu_ps {<4 x float>,<4 x float>} @test_waterfall_multi_end_1loop_rsrc_
 ; GFX12-GISEL-NEXT:    s_wait_samplecnt 0x0
 ; GFX12-GISEL-NEXT:    ; return to shader part epilog
        <8 x i32> addrspace(4)* inreg %in, <4 x i32> addrspace(4)* inreg %s_in,
-       i32 %idx1, i32 %idx2, i32 %s_idx, i32 %s_idx2) #1 {
+       i32 %idx1, i32 %idx2, i32 %s_idx, i32 %s_idx2, { <4 x float>, <4 x float> } %dummy) #1 {
   %tok = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %idx1)
   %tok1 = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 %tok, i32 %idx2)
   %widx0 = call i32 @llvm.amdgcn.waterfall.readfirstlane.i32.i32(i32 %tok1, i32 %s_idx)
@@ -11363,7 +11363,7 @@ define amdgpu_ps {<4 x float>,<4 x float>} @test_waterfall_multi_end_1loop_rsrc_
   %r1 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok1, <4 x float> %r)
   %r3 = call <4 x float> @llvm.amdgcn.waterfall.end.v4f32(i32 %tok1, <4 x float> %r2)
 
-  %insert = insertvalue { <4 x float>, <4 x float> } undef, <4 x float> %r1, 0
+  %insert = insertvalue { <4 x float>, <4 x float> } %dummy, <4 x float> %r1, 0
   %insert1 = insertvalue { <4 x float>, <4 x float> } %insert, <4 x float> %r3, 1
   ret {<4 x float> , <4 x float>} %insert1
 }
@@ -11782,7 +11782,7 @@ define amdgpu_ps {<4 x float>,float} @test_waterfall_multi_end_struct(
 ; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
 ; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x0
 ; GFX12-GISEL-NEXT:    ; return to shader part epilog
-              i32 %idx, <4 x i32> %s_idx, i32 %v_inp) #1 {
+              i32 %idx, <4 x i32> %s_idx, i32 %v_inp, { <4 x float>, float } %dummy) #1 {
   %tok = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %idx)
   %widx0 = call <4 x i32> @llvm.amdgcn.waterfall.readfirstlane.v4i32.v4i32(i32 %tok, <4 x i32> %s_idx)
   %val = call { <4 x float>, i32 } @llvm.amdgcn.struct.buffer.load.format.sl_v4f32i32s(<4 x i32> %widx0, i32 %v_inp, i32 0, i32 0, i32 0)
@@ -11792,7 +11792,7 @@ define amdgpu_ps {<4 x float>,float} @test_waterfall_multi_end_struct(
   %r1 = call i32 @llvm.amdgcn.waterfall.end.i32(i32 %tok, i32 %tfe)
   %r1.float = bitcast i32 %r1 to float
 
-  %insert = insertvalue { <4 x float>, float } undef, <4 x float> %r0, 0
+  %insert = insertvalue { <4 x float>, float } %dummy, <4 x float> %r0, 0
   %insert1 = insertvalue { <4 x float>, float } %insert, float %r1.float, 1
   ret {<4 x float> , float} %insert1
 }
@@ -12078,7 +12078,7 @@ define amdgpu_ps {<4 x float>,float} @test_waterfall_multi_end_struct_uniform(
 ; GFX12-GISEL-NEXT:    s_mov_b64 exec, s[0:1]
 ; GFX12-GISEL-NEXT:    s_wait_loadcnt 0x0
 ; GFX12-GISEL-NEXT:    ; return to shader part epilog
-              i32 inreg %idx, <4 x i32> %s_idx, i32 %v_inp) #1 {
+              i32 inreg %idx, <4 x i32> %s_idx, i32 %v_inp, { <4 x float>, float } %dummy) #1 {
   %tok = call i32 @llvm.amdgcn.waterfall.begin.i32(i32 0, i32 %idx)
   %widx0 = call <4 x i32> @llvm.amdgcn.waterfall.readfirstlane.v4i32.v4i32(i32 %tok, <4 x i32> %s_idx)
   %val = call { <4 x float>, i32 } @llvm.amdgcn.struct.buffer.load.format.sl_v4f32i32s(<4 x i32> %widx0, i32 %v_inp, i32 0, i32 0, i32 0)
@@ -12088,7 +12088,7 @@ define amdgpu_ps {<4 x float>,float} @test_waterfall_multi_end_struct_uniform(
   %r1 = call i32 @llvm.amdgcn.waterfall.end.i32(i32 %tok, i32 %tfe)
   %r1.float = bitcast i32 %r1 to float
 
-  %insert = insertvalue { <4 x float>, float } undef, <4 x float> %r0, 0
+  %insert = insertvalue { <4 x float>, float } %dummy, <4 x float> %r0, 0
   %insert1 = insertvalue { <4 x float>, float } %insert, float %r1.float, 1
   ret {<4 x float> , float} %insert1
 }

>From b819e094dae63dbce8e03c2342dfd5086147f7f1 Mon Sep 17 00:00:00 2001
From: Greta Yorsh <Greta.Yorsh at amd.com>
Date: Fri, 17 Apr 2026 16:42:37 +0100
Subject: [PATCH 5/5] Improve documentation and comments

---
 llvm/docs/AMDGPUUsage.rst                | 146 ++++++++++++++++-------
 llvm/include/llvm/IR/IntrinsicsAMDGPU.td |   4 +-
 2 files changed, 106 insertions(+), 44 deletions(-)

diff --git a/llvm/docs/AMDGPUUsage.rst b/llvm/docs/AMDGPUUsage.rst
index 96ad9c0b755f6..9301255196362 100644
--- a/llvm/docs/AMDGPUUsage.rst
+++ b/llvm/docs/AMDGPUUsage.rst
@@ -1776,64 +1776,126 @@ The AMDGPU backend implements the following LLVM IR intrinsics.
                                                    * :ref:`Synchronization Scope<amdgpu-intrinsics-syncscope-metadata-operand>`.
                                                      Note that the scope used must ensure that the L2 cache will be hit.
 
-  llvm.amdgcn.waterfall.begin                      Marks the beginning of a waterfall region of code.
-
-                                                   The compiler generates a waterfall loop around the region.
-                                                   A waterfall loop handles the case where an operation that requires
-                                                   a uniform operand (e.g., held in an SGPR) is applied to a divergent operand
-                                                   (held in a VGPR, with values varying per lane).
-                                                   Each iteration of the waterfall loop activates a subset of lanes that
-                                                   share the same value of the VGPR (the value in the first active lane).
-                                                   The operation is then executed using that value as the uniform operand.
-                                                   If the VGPR is already uniform, the waterfall loop executes only once.
-                                                   The worst case for a waterfall loop is one iteration per lane (all lanes
-                                                   have different values of the VGPR), but it is not common in practice.
-
-                                                   The intrinsic takes a previous token
-                                                   (``i32``; use a null/zero value if this is the first ``waterfall.begin`` in
-                                                   a waterfall group) and a VGPR.
+  ==============================================   ==========================================================
 
-                                                   The intrinsic returns a new token that must be threaded through the
-                                                   corresponding ``waterfall.readfirstlane`` and ``waterfall.end`` or
-                                                   ``waterfall.last_use`` intrinsics, forming a waterfall group of intrinsics
-                                                   that together define a waterfall region.
+.. TODO::
+
+   List AMDGPU intrinsics.
+
+'``llvm.amdgcn.waterfall``' Intrinsics
+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+The ``llvm.amdgcn.waterfall`` :ref:`family of intrinsics<amdgpu-waterfall-intrinsics-table>`
+describe how to generate a waterfall loop around a region of code.
+
+:Background:
+
+A waterfall loop handles the case where an operation that requires a uniform
+operand (e.g., in an SGPR) is applied to a non-uniform operand (held in a
+VGPR, with values varying per lane).
+
+Each iteration of the waterfall loop activates a subset of lanes that share
+the same value of the non-uniform operand (the value in the first active
+lane). The operation is then executed using that value as the uniform
+operand.
+
+If the operand is already uniform, the waterfall loop executes only
+once. The worst case for a waterfall loop is one iteration per lane (all
+lanes have different values of the operand), but this is not common in
+practice.
+
+:Motivation:
+
+The ``llvm.amdgc.waterfall.*`` intrinsics provide a practical way for a
+frontend to combine several operations into a single waterfall loop with an
+efficient iteration strategy over multiple non-uniform operands.
+
+In particular, a frontend can specify a subset of non-uniform operands that
+is sufficient to use as the "index" to loop over, i.e., lanes that have the
+same values for the specified operands must have the same value for all
+other non-uniform operands that need to be uniform. It is hard for the
+backend to find a fitting subset of non-uniform operands, but it can be easy
+for a frontend.
 
-                                                   All intrinsics in a waterfall group must reside in the same basic block.
+:Implementation:
 
-                                                   Multiple ``waterfall.begin`` intrinsics can be chained by
-                                                   passing the token of the preceding ``waterfall.begin`` as the first argument.
-                                                   This allows a front-end to create one waterfall loop for
-                                                   multiple non-uniform values.
-                                                   Later compiler passes may remove values determined as uniform.
-                                                   The final token is used for other waterfall intrinsics in the same group.
+A group of waterfall intrinsics that depend on the same token defines a
+single waterfall loop. The group identifies a region of code, the
+corresponding non-uniform operands that need to be uniform, and a subset of
+these operands that is sufficient to use as the "index" of the loop.
 
-  llvm.amdgcn.waterfall.readfirstlane              Reads the first active lane's value of the VGPR and returns it as
-                                                   an SGPR for use within a waterfall region.
+A waterfall group must contain at least one
+``waterfall.begin``, at least one ``waterfall.readfirstlane``, and at least one
+of ``waterfall.end`` or ``waterfall.loop_end`` intrinsic.
+A group can contain more than one of each of the waterfall intrinsics.
+The token on the final
+``waterfall.begin`` must be used for other waterfall intrinsics in the same
+group. The waterfall loop will enclose all instructions from the earliest
+``waterfall.begin`` to the latest ``waterfall.end`` or ``waterfall.loop_end``
+intrinsic in the group.
 
-                                                   Takes the ``i32`` token from the final ``waterfall.begin`` in the waterfall
-                                                   group and the VGPR. Returns the uniform (SGPR) result.
+Each waterfall group must be contained within a single basic block.
+A single basic block can contain more than one waterfall group.
 
-                                                   If the VGPR is determined to be uniform at compile time, this intrinsic
-                                                   is optimized away (the input VGPR value is used directly).
+Later compiler passes can remove operands determined as uniform.
+If all operands are uniform, the compiler will not insert the waterfall loop.
 
-  llvm.amdgcn.waterfall.end                        Marks the end of a waterfall region.
-                                                   Takes the ``i32`` token from the final ``waterfall.begin``
+.. table:: AMDGPU Waterfall Intrinsics
+  :name: amdgpu-waterfall-intrinsics-table
 
-  llvm.amdgcn.waterfall.last_use                   Variant of ``waterfall.end`` for values whose last use is in a
+  ==============================================   =========================================================================
+  LLVM Intrinsic                                   Description
+  ==============================================   =========================================================================
+  ``llvm.amdgcn.waterfall.begin``                  Marks the beginning of a waterfall region of code.
+                                                   Specifies a non-uniform operand that needs to be uniform in the region
+                                                   and should be used as part of the index for the waterfall loop to iterate
+                                                   over.
+
+                                                   Multiple non-uniform operands of different types can be used together as
+                                                   an index by threading tokens through multiple ``waterfall.begin`` intrinsics::
+
+                                                     %tok0 = llvm.amdgcn.waterfall.begin(i32 0, i32 %idx0)
+                                                     %tok1 = llvm.amdgcn.waterfall.begin(i32 %tok0, i32 %idx1)
+                                                     ...
+
+                                                   The intrinsic takes a token
+                                                   (use a null/zero value if this is the first ``waterfall.begin`` in a
+                                                   waterfall group) and an operand.
+
+                                                   The intrinsic returns a new token that must be threaded through the
+                                                   corresponding ``waterfall.readfirstlane``, ``waterfall.end`` or
+                                                   ``waterfall.last.use**`` intrinsics.
+
+  ``llvm.amdgcn.waterfall.readfirstlane``          Reads the first active lane's value of the non-uniform operand and
+                                                   returns it for use within a waterfall region.
+
+                                                   Takes a token from the final ``waterfall.begin`` in the waterfall
+                                                   group and a non-uniform operand. Returns the uniform result.
+
+                                                   Every non-inform operand from ``waterfall.begin`` must have a
+                                                   corresponding ``waterfall.readfirstlane`` intrinsic in the group.
+                                                   Additional ``waterfall.readfirstlane`` intrinsics in the group can be used
+                                                   to specify non-uniform operands that are not part of the index,
+                                                   but need to be uniform in the waterfall loop, as long as they satisfy
+                                                   the correctness requirement (i.e., they have the same values in lanes
+                                                   where the index has the same values).
+
+  ``llvm.amdgcn.waterfall.end``                    Marks the end of a waterfall region.
+                                                   Takes the token from the final ``waterfall.begin`` in the group.
+
+  ``llvm.amdgcn.waterfall.last.use``               Variant of ``waterfall.end`` for values whose last use is in a
                                                    non-defining operation such as a store. Marks that the use of the value
                                                    constitutes the end of the waterfall region.
 
-  llvm.amdgcn.waterfall.last_use_vgpr              Variant of ``waterfall.last_use`` for values that remain in a VGPR.
+  ``llvm.amdgcn.waterfall.last.use.vgpr``          Variant of ``waterfall.last_use`` for values that remain in a VGPR.
 
-  llvm.amdgcn.waterfall.loop_end                   Inserted later by the compiler to be used with ```waterfall.last_use*``
+  ``llvm.amdgcn.waterfall.loop.end``               Inserted later by the compiler to be used with ``waterfall.last_use*``
                                                    to mark the loop-end point for special
                                                    handling such as SCC clobber tracking.
+                                                   This intrinsic should not be generated by a frontend.
 
-  ==============================================   ==========================================================
 
-.. TODO::
-
-   List AMDGPU intrinsics.
+  ==============================================   =========================================================================
 
 '``llvm.amdgcn.cooperative.atomic``' Intrinsics
 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
diff --git a/llvm/include/llvm/IR/IntrinsicsAMDGPU.td b/llvm/include/llvm/IR/IntrinsicsAMDGPU.td
index e4aee0ad0310f..3a7d2e0e2a810 100644
--- a/llvm/include/llvm/IR/IntrinsicsAMDGPU.td
+++ b/llvm/include/llvm/IR/IntrinsicsAMDGPU.td
@@ -2696,8 +2696,8 @@ def int_amdgcn_cs_chain:
 // Waterfall intrinsics are used to mark a region as requiring waterfall loops to
 // activate and deactivate lanes in a loop (and sometimes transform VGPR values
 // into SGPR values).
-// Best case (uniform index in VGPR) the loop will execute once, worst case (all
-// values in the index are different) the loop will execute wave-size times.
+// Best case (all lanes have the same values) the loop will execute once, worst case (all lanes
+// have different values) the loop will execute wave-size times.
 // The waterfall.begin intrinsic returns a new token that must be threaded through the
 // corresponding waterfall.readfirstlane and waterfall.end or
 // waterfall.last_use intrinsics, forming a waterfall group of intrinsics that together



More information about the llvm-commits mailing list