[llvm] [SystemZ] XPLINK64: emit narrow sign/zero-extend instructions for sub-i32 formal args (PR #206833)

Zibi Sarbinowski via llvm-commits llvm-commits at lists.llvm.org
Fri Aug 28 09:33:39 PDT 2026


https://github.com/zibi2 updated https://github.com/llvm/llvm-project/pull/206833

>From 379b9039d14405c28f1bacc98b6ff797563c98ac Mon Sep 17 00:00:00 2001
From: Zibi Sarbinowski <zibi at ca.ibm.com>
Date: Tue, 11 Aug 2026 17:56:46 -0400
Subject: [PATCH 1/6] [SystemZ] XPLINK64: emit narrow sign/zero-extend
 instructions for sub-i32 formal args

Previously, i8/i16 integer formal arguments with signext/zeroext attributes
were promoted to i64 via CC_XPLINK_Promote_i32 using LocVT=i64, causing the
DAG to select LGFR/LLGFR (word-sized) for all sub-i64 types instead of the
narrower byte/halfword forms required for XL compiler binary compatibility.

Root cause: LLVM type-legalizes i8/i16 to i32 before the CC table runs, so
ValVT is always i32 by the time CC_XPLINK_Promote_i32 is called. The pre-
legalization original type (Ins[i].ArgVT) is the only way to distinguish
i8/i16 formals from i32 formals at that point.

Fix:
- Add SystemZCCState (subclass of CCState) with an IsFormalArgLowering flag,
  an ArgOrigVTs vector stashed in AnalyzeFormalArguments from Ins[i].ArgVT
  (using MVT::Other as sentinel for non-simple EVTs), and a safe
  getArgOrigVT(ValNo) accessor (returns MVT::Other when out of range, e.g.
  when called from AnalyzeCallOperands where ArgOrigVTs is never populated).
- In CC_XPLINK_Promote_i32 with isFormalArgLowering=true: for original i8/i16,
  keep LocVT=i32 (GR32 live-in) only when a GPR is still available (probed via
  getFirstUnallocated([R1D,R2D,R3D])). When all GPRs are taken, the arg goes
  to an 8-byte stack slot where the caller stores a sign-extended i64 with the
  actual value in bytes 6-7; LocVT=i64 is used in that case so the full 8
  bytes are loaded correctly.
- In LowerFormalArguments: after convertLocVTToValVT, truncate from the
  legalized i32 down to the true i8/i16 type so the subsequent sign/zero-
  extend to i64 selects the correct narrow instruction (LGBR/LGHR/LLGCR/
  LLGHR for register args, LGB/LGH/LLGC/LLGH for stack args).
- Extend CCIfType from [i32] to [i8, i16, i32] in CC_SystemZ_XPLINK64.
- Update call-zos-01.ll and call-zos-vararg.ll to expect the new narrow
  instructions.

Result:
  jbyte    (i8  signext) -> LGBR/LGB   (was LGFR)
  jboolean (i8  zeroext) -> LLGCR/LLGC (was LLGFR)
  jshort   (i16 signext) -> LGHR/LGH   (was LGFR)
  jchar    (i16 zeroext) -> LLGHR/LLGH (was LLGFR)
  jint     (i32 signext) -> LGFR        (unchanged)

Adds lit test zos-xplink64-formal-sub64-extend.ll covering all five integer
types plus optnone spill path and small struct coercion cases.

This commit was assisted by IBM Bob (AI).
---
 llvm/lib/Target/SystemZ/SystemZCallingConv.h  |  79 ++++++
 llvm/lib/Target/SystemZ/SystemZCallingConv.td |   7 +-
 .../Target/SystemZ/SystemZISelLowering.cpp    |  27 +-
 llvm/test/CodeGen/SystemZ/call-zos-01.ll      |   9 +-
 llvm/test/CodeGen/SystemZ/call-zos-vararg.ll  |   8 +-
 .../zos-xplink64-formal-sub64-extend.ll       | 251 ++++++++++++++++++
 6 files changed, 367 insertions(+), 14 deletions(-)
 create mode 100644 llvm/test/CodeGen/SystemZ/zos-xplink64-formal-sub64-extend.ll

diff --git a/llvm/lib/Target/SystemZ/SystemZCallingConv.h b/llvm/lib/Target/SystemZ/SystemZCallingConv.h
index f5ffbf5c04d60..8b8afd4b61e94 100644
--- a/llvm/lib/Target/SystemZ/SystemZCallingConv.h
+++ b/llvm/lib/Target/SystemZ/SystemZCallingConv.h
@@ -85,6 +85,45 @@ inline bool CC_SystemZ_I128Indirect(unsigned &ValNo, MVT &ValVT,
   return true;
 }
 
+// Extends CCState with:
+//   - a flag to distinguish formal-argument lowering from outgoing-call
+//     lowering (used by CC_XPLINK_Promote_i32)
+//   - the pre-legalization original MVT for each argument (needed because
+//     i8/i16 are both legalized to i32 before the CC table runs, so ValVT
+//     alone cannot distinguish them)
+class SystemZCCState : public CCState {
+  bool IsFormalArgLowering = false;
+  SmallVector<MVT, 8> ArgOrigVTs;
+
+public:
+  using CCState::CCState;
+
+  void AnalyzeFormalArguments(const SmallVectorImpl<ISD::InputArg> &Ins,
+                              CCAssignFn Fn) {
+    // Record the pre-legalization original type for each argument.
+    // Use MVT::Other as a sentinel when ArgVT is not a simple type (e.g.
+    // split or extended EVTs that occur in 32-bit mode).
+    ArgOrigVTs.clear();
+    for (const auto &In : Ins)
+      ArgOrigVTs.push_back(In.ArgVT.isSimple() ? In.ArgVT.getSimpleVT()
+                                               : MVT::Other);
+    CCState::AnalyzeFormalArguments(Ins, Fn);
+  }
+
+  // Returns the pre-legalization MVT for argument ValNo.
+  // Only valid after AnalyzeFormalArguments; returns MVT::Other if out of
+  // range (e.g. called from AnalyzeCallOperands context where ArgOrigVTs
+  // was never populated).
+  MVT getArgOrigVT(unsigned ValNo) const {
+    if (ValNo >= ArgOrigVTs.size())
+      return MVT::Other;
+    return ArgOrigVTs[ValNo];
+  }
+
+  bool isFormalArgLowering() const { return IsFormalArgLowering; }
+  void setIsFormalArgLowering() { IsFormalArgLowering = true; }
+};
+
 // A pointer in 64bit mode is always passed as 64bit.
 inline bool CC_XPLINK64_Pointer(unsigned &ValNo, MVT &ValVT, MVT &LocVT,
                                 CCValAssign::LocInfo &LocInfo,
@@ -96,6 +135,46 @@ inline bool CC_XPLINK64_Pointer(unsigned &ValNo, MVT &ValVT, MVT &LocVT,
   return false;
 }
 
+inline bool CC_XPLINK_Promote_i32(unsigned &ValNo, MVT &ValVT, MVT &LocVT,
+                                  CCValAssign::LocInfo &LocInfo,
+                                  ISD::ArgFlagsTy &ArgFlags, CCState &State) {
+  SystemZCCState *SZState = static_cast<SystemZCCState *>(&State);
+  if (SZState->isFormalArgLowering()) {
+    // For formal arguments, use the pre-legalization original MVT to
+    // distinguish i8/i16 from i32 — all three are legalized to i32 (ValVT)
+    // before the CC table runs, so ValVT alone cannot tell them apart.
+    MVT OrigVT = SZState->getArgOrigVT(ValNo);
+    if (OrigVT == MVT::i8 || OrigVT == MVT::i16) {
+      // Keep LocVT=i32 (GR32 live-in) only if a GR32 register is still free.
+      // The CC table assigns i32 args to R1L/R2L/R3L in XPLINK64; if all
+      // three are taken, this arg goes to an 8-byte stack slot.  The caller
+      // stores a sign-extended i64 in that slot (value in bytes 6-7), so we
+      // must use LocVT=i64 to load all 8 bytes correctly.
+      // We probe the 64-bit register aliases: AllocateReg marks all aliases,
+      // so if R1D is allocated (by a prior i64/ptr arg) R1L is also marked.
+      static const MCPhysReg GPR64s[] = {SystemZ::R1D, SystemZ::R2D,
+                                         SystemZ::R3D};
+      if (State.getFirstUnallocated(GPR64s) < 3) {
+        LocInfo = CCValAssign::AExt;
+        return false;
+      }
+      // All GR64s (and their GR32 aliases) taken — stack path: promote to i64.
+    }
+    // i32 formal (or stack-bound i8/i16): promote to GR64 via AExt.
+    LocVT = MVT::i64;
+    LocInfo = CCValAssign::AExt;
+  } else {
+    LocVT = MVT::i64;
+    if (ArgFlags.isSExt())
+      LocInfo = CCValAssign::SExt;
+    else if (ArgFlags.isZExt())
+      LocInfo = CCValAssign::ZExt;
+    else
+      LocInfo = CCValAssign::AExt;
+  }
+  return false;
+}
+
 inline bool CC_XPLINK64_Shadow_Reg(unsigned &ValNo, MVT &ValVT, MVT &LocVT,
                                    CCValAssign::LocInfo &LocInfo,
                                    ISD::ArgFlagsTy &ArgFlags, CCState &State) {
diff --git a/llvm/lib/Target/SystemZ/SystemZCallingConv.td b/llvm/lib/Target/SystemZ/SystemZCallingConv.td
index 69202e3fcbc57..e080463a853a7 100644
--- a/llvm/lib/Target/SystemZ/SystemZCallingConv.td
+++ b/llvm/lib/Target/SystemZ/SystemZCallingConv.td
@@ -213,9 +213,10 @@ def RetCC_SystemZ_XPLINK64 : CallingConv<[
 // examples.
 
 def CC_SystemZ_XPLINK64 : CallingConv<[
-  // XPLINK64 ABI compliant code widens integral types smaller than i64
-  // to i64 before placing the parameters either on the stack or in registers.
-  CCIfType<[i32], CCIfExtend<CCPromoteToType<i64>>>,
+  // Promote sub-i64 integers to i64 if they have an explicit extension type.
+  // The convention is that true integer arguments smaller than 64 bits should
+  // be marked as extended, but structures coerced to an integer type shouldn't.
+  CCIfType<[i8, i16, i32], CCIfExtend<CCCustom<"CC_XPLINK_Promote_i32">>>,
   // Promote f32 to f64 and bitcast to i64, if it needs to be passed in GPRs.
   // Although we assign the f32 vararg to be bitcast, it will first be promoted
   // to an f64 within convertValVTToLocVT().
diff --git a/llvm/lib/Target/SystemZ/SystemZISelLowering.cpp b/llvm/lib/Target/SystemZ/SystemZISelLowering.cpp
index 9c81005d6bd59..8fbfab795ed9d 100644
--- a/llvm/lib/Target/SystemZ/SystemZISelLowering.cpp
+++ b/llvm/lib/Target/SystemZ/SystemZISelLowering.cpp
@@ -2047,7 +2047,8 @@ SDValue SystemZTargetLowering::LowerFormalArguments(
 
   // Assign locations to all of the incoming arguments.
   SmallVector<CCValAssign, 16> ArgLocs;
-  CCState CCInfo(CallConv, IsVarArg, MF, ArgLocs, *DAG.getContext());
+  SystemZCCState CCInfo(CallConv, IsVarArg, MF, ArgLocs, *DAG.getContext());
+  CCInfo.setIsFormalArgLowering();
   CCInfo.AnalyzeFormalArguments(Ins, CC_SystemZ);
   FuncInfo->setSizeOfFnParams(CCInfo.getStackSize());
 
@@ -2152,8 +2153,24 @@ SDValue SystemZTargetLowering::LowerFormalArguments(
           assert(PartOffset && "Offset should be non-zero.");
         }
       }
-    } else
-      InVals.push_back(convertLocVTToValVT(DAG, DL, VA, Chain, ArgValue));
+    } else {
+      SDValue Val = convertLocVTToValVT(DAG, DL, VA, Chain, ArgValue);
+      // For i8/i16 formals under XPLINK64, CC_XPLINK_Promote_i32 either keeps
+      // LocVT=i32 (register arg) or uses LocVT=i64 (stack arg, 8-byte slot).
+      // Either way convertLocVTToValVT yields ValVT=i32 (legalized).  Truncate
+      // down to the true original i8/i16 so the DAG folds the truncate+extend
+      // into a narrow instruction (LGBR/LGHR/LGB/LGH etc).
+      // Only for XPLINK64: ELF receives i8/i16 already correctly widened.
+      // Guard with isSimple() since non-simple EVTs must not reach
+      // getSimpleVT().
+      if (Subtarget.isTargetXPLINK64() && Ins[I].ArgVT.isSimple()) {
+        MVT OrigVT = Ins[I].ArgVT.getSimpleVT();
+        if (OrigVT != VA.getValVT() && OrigVT.isScalarInteger() &&
+            OrigVT.getSizeInBits() < VA.getValVT().getSizeInBits())
+          Val = DAG.getNode(ISD::TRUNCATE, DL, OrigVT, Val);
+      }
+      InVals.push_back(Val);
+    }
   }
 
   if (IsVarArg && Subtarget.isTargetXPLINK64()) {
@@ -2360,8 +2377,10 @@ SystemZTargetLowering::LowerCall(CallLoweringInfo &CLI,
   verifyNarrowIntegerArgs_Call(Outs, &MF.getFunction(), Callee);
 
   // Analyze the operands of the call, assigning locations to each operand.
+  // Use SystemZCCState so CC_XPLINK_Promote_i32 can safely cast State to
+  // SystemZCCState& (isFormalArgLowering() returns false by default).
   SmallVector<CCValAssign, 16> ArgLocs;
-  CCState ArgCCInfo(CallConv, IsVarArg, MF, ArgLocs, Ctx);
+  SystemZCCState ArgCCInfo(CallConv, IsVarArg, MF, ArgLocs, Ctx);
   ArgCCInfo.AnalyzeCallOperands(Outs, CC_SystemZ);
 
   // We don't support GuaranteedTailCallOpt, only automatically-detected
diff --git a/llvm/test/CodeGen/SystemZ/call-zos-01.ll b/llvm/test/CodeGen/SystemZ/call-zos-01.ll
index a6006035dcaa1..d1a79b10bfe9e 100644
--- a/llvm/test/CodeGen/SystemZ/call-zos-01.ll
+++ b/llvm/test/CodeGen/SystemZ/call-zos-01.ll
@@ -129,7 +129,7 @@ entry:
 
 define signext i8 @pass_char(i8 signext %arg) {
 ; CHECK-LABEL: pass_char DS 0H
-; CHECK:         lgr 3,1
+; CHECK:         lgb 3,
 ; CHECK-NEXT:    b 2(7)
 entry:
   ret i8 %arg
@@ -137,7 +137,7 @@ entry:
 
 define signext i16 @pass_short(i16 signext %arg) {
 ; CHECK-LABEL: pass_short DS 0H
-; CHECK:         lgr 3,1
+; CHECK:         lgh 3,
 ; CHECK-NEXT:    b 2(7)
 entry:
   ret i16 %arg
@@ -145,7 +145,7 @@ entry:
 
 define signext i32 @pass_int(i32 signext %arg0, i32 signext %arg1) {
 ; CHECK-LABEL: pass_int DS 0H
-; CHECK:         lgr 3,2
+; CHECK:         lgfr 3,2
 ; CHECK-NEXT:    b 2(7)
 entry:
   ret i32 %arg1
@@ -164,8 +164,7 @@ entry:
 
 define signext i64 @pass_integrals0(i64 signext %arg0, i32 signext %arg1, i16 signext %arg2, i64 signext %arg3) {
 ; CHECK-LABEL: pass_integrals0 DS 0H
-; CHECK:         ag 2,2200(4)
-; CHECK-NEXT:    lgr 3,2
+; CHECK:         agfr 3,2
 ; CHECK-NEXT:    b 2(7)
 entry:
   %N = sext i32 %arg1 to i64
diff --git a/llvm/test/CodeGen/SystemZ/call-zos-vararg.ll b/llvm/test/CodeGen/SystemZ/call-zos-vararg.ll
index 3bcc583adec7f..5ecc2bb4b152a 100644
--- a/llvm/test/CodeGen/SystemZ/call-zos-vararg.ll
+++ b/llvm/test/CodeGen/SystemZ/call-zos-vararg.ll
@@ -287,10 +287,14 @@ define void @call_vec_double_vararg_straddle(<2 x double> %v) {
 ; CHECK-NEXT:    aghi 4,-192
 ; CHECK-NEXT:    *FENCE
 ; CHECK-NEXT:    L#end_of_prologue{{[0-9]+}} DS 0H
-; CHECK-NEXT:    lg 0,2392(4)
+; CHECK-NEXT:    lgh 0,
+; CHECK-NEXT:    lgb 7,
 ; CHECK-NEXT:    lg 6,40(5)
 ; CHECK-NEXT:    lg 5,32(5)
-; CHECK-NEXT:    stg 0,2200(4)
+; CHECK-NEXT:    lgr 3,2
+; CHECK-NEXT:    lgfr 1,1
+; CHECK-NEXT:    stg 7,2200(4)
+; CHECK-NEXT:    lgr 2,0
 ; CHECK-NEXT:    basr 7,6
 ; CHECK-NEXT:    bcr 0,0
 ; CHECK-NEXT:    lg 7,2072(4)
diff --git a/llvm/test/CodeGen/SystemZ/zos-xplink64-formal-sub64-extend.ll b/llvm/test/CodeGen/SystemZ/zos-xplink64-formal-sub64-extend.ll
new file mode 100644
index 0000000000000..61534a6f04d56
--- /dev/null
+++ b/llvm/test/CodeGen/SystemZ/zos-xplink64-formal-sub64-extend.ll
@@ -0,0 +1,251 @@
+; Test that sub-64-bit integer formal arguments and small struct formal
+; arguments are correctly sign/zero-extended by the callee in the z/OS
+; XPLINK64 calling convention.
+;
+; The XPLINK64 ABI does not require the caller to have extended the upper bits
+; of a GPR for named (non-variadic) integer arguments.  The callee therefore
+; must perform its own extension when it needs the full 64-bit value.
+;
+; Integer formals (CCIfExtend -> CC_XPLINK_Promote_i32 with isFormalArgLowering=true):
+;   For i8/i16 formals, CC_XPLINK_Promote_i32 keeps LocVT=i32 (GR32 live-in).
+;   LowerFormalArguments truncates the GR32 to the true i8/i16 type, so the
+;   subsequent sign/zero-extend to i64 selects the narrow register-extend
+;   instructions (LGBR/LGHR/LLGCR/LLGHR) matching XL compiler output.
+;   For i32 formals, LocVT=i64 (GR64), so LGFR is used as before.
+;
+;   jbyte    (i8  signext)  lgbr  -- sign-extend byte register to 64 bits
+;   jboolean (i8  zeroext)  llgcr -- zero-extend char register to 64 bits
+;   jshort   (i16 signext)  lghr  -- sign-extend halfword register to 64 bits
+;   jchar    (i16 zeroext)  llghr -- zero-extend halfword register to 64 bits
+;   jint     (i32 signext)  lgfr  -- sign-extend word register to 64 bits
+;
+; With optnone (alloca spill path), the reload uses the natural-width
+; sign/zero-extend load instruction (lgb/stc+lgb).
+;
+; Small struct formals (coerced to i64, no CCIfExtend):
+;   Struct member occupies the high bits of the GPR (big-endian).  The callee
+;   must shift down to recover the signed value via srag (arithmetic shift).
+;   S8  { signed char  }  srag 1,1,56
+;   S16 { short        }  srag 1,1,48
+;   S32 { int          }  srag 1,1,32
+;
+; Source:
+;   void print(long long);
+;   typedef signed char jbyte;  typedef unsigned char jboolean;
+;   typedef short jshort;       typedef unsigned short jchar;
+;   typedef int   jint;
+;   struct S8 { signed char x; };  struct S16 { short x; };  struct S32 { int x; };
+;   void callee_jbyte   (jbyte    c) { print((long long)c); }
+;   void callee_jboolean(jboolean c) { print((long long)c); }
+;   void callee_jshort  (jshort   c) { print((long long)c); }
+;   void callee_jchar   (jchar    c) { print((long long)c); }
+;   void callee_jint    (jint     c) { print((long long)c); }
+;   void callee_jbyte_optnone(jbyte c) { print((long long)c); }  // optnone
+;   void callee_s8 (struct S8  s)   { print((long long)s.x); }
+;   void callee_s16(struct S16 s)   { print((long long)s.x); }
+;   void callee_s32(struct S32 s)   { print((long long)s.x); }
+;
+; RUN: llc -mtriple=s390x-ibm-zos -O3 < %s | FileCheck %s --check-prefix=OPT
+; RUN: llc -mtriple=s390x-ibm-zos -O0 < %s | FileCheck %s --check-prefix=NOOPT
+
+;----------------------------------------------------------------------------
+; jbyte (i8 signext): callee sign-extends incoming R1 (GR32) to 64 bits.
+; At -O3: lgbr 1,1 (register form) or lgb (memory form, optimizer may spill).
+; At -O0: lr copies R1 to R0, then lgbr 1,0.
+;----------------------------------------------------------------------------
+; OPT-LABEL: @callee_jbyte
+; OPT:       L#end_of_prologue0
+; OPT:        lgb
+; OPT:        basr 7,6
+
+; NOOPT-LABEL: @callee_jbyte
+; NOOPT:       L#end_of_prologue0
+; NOOPT:        lgbr 1,
+; NOOPT:        basr 7,6
+
+define hidden void @callee_jbyte(i8 noundef signext %c) local_unnamed_addr {
+entry:
+  %conv = sext i8 %c to i64
+  tail call void @print(i64 noundef %conv)
+  ret void
+}
+
+;----------------------------------------------------------------------------
+; jboolean (i8 zeroext): callee zero-extends incoming R1 (GR32) to 64 bits.
+; At -O3: llgcr 1,1 or llgc (memory form).
+; At -O0: lr + llgcr 1,0.
+;----------------------------------------------------------------------------
+; OPT-LABEL: @callee_jboolean
+; OPT:       L#end_of_prologue1
+; OPT:        llgc
+; OPT:        basr 7,6
+
+; NOOPT-LABEL: @callee_jboolean
+; NOOPT:       L#end_of_prologue1
+; NOOPT:        llgcr 1,
+; NOOPT:        basr 7,6
+
+define hidden void @callee_jboolean(i8 noundef zeroext %c) local_unnamed_addr {
+entry:
+  %conv = zext i8 %c to i64
+  tail call void @print(i64 noundef %conv)
+  ret void
+}
+
+;----------------------------------------------------------------------------
+; jshort (i16 signext): callee sign-extends incoming R1 (GR32) to 64 bits.
+; At -O3: lghr 1,1 or lgh (memory form).
+; At -O0: lr + lghr 1,0.
+;----------------------------------------------------------------------------
+; OPT-LABEL: @callee_jshort
+; OPT:       L#end_of_prologue2
+; OPT:        lgh
+; OPT:        basr 7,6
+
+; NOOPT-LABEL: @callee_jshort
+; NOOPT:       L#end_of_prologue2
+; NOOPT:        lghr 1,
+; NOOPT:        basr 7,6
+
+define hidden void @callee_jshort(i16 noundef signext %c) local_unnamed_addr {
+entry:
+  %conv = sext i16 %c to i64
+  tail call void @print(i64 noundef %conv)
+  ret void
+}
+
+;----------------------------------------------------------------------------
+; jchar (i16 zeroext): callee zero-extends incoming R1 (GR32) to 64 bits.
+; At -O3: llghr 1,1 or llgh (memory form).
+; At -O0: lr + llghr 1,0.
+;----------------------------------------------------------------------------
+; OPT-LABEL: @callee_jchar
+; OPT:       L#end_of_prologue3
+; OPT:        llgh
+; OPT:        basr 7,6
+
+; NOOPT-LABEL: @callee_jchar
+; NOOPT:       L#end_of_prologue3
+; NOOPT:        llghr 1,
+; NOOPT:        basr 7,6
+
+define hidden void @callee_jchar(i16 noundef zeroext %c) local_unnamed_addr {
+entry:
+  %conv = zext i16 %c to i64
+  tail call void @print(i64 noundef %conv)
+  ret void
+}
+
+;----------------------------------------------------------------------------
+; jint (i32 signext): callee sign-extends incoming R1 (GR64) to 64 bits.
+; At -O3: lgfr 1,1.
+; At -O0: lr + lgfr 1,0.
+;----------------------------------------------------------------------------
+; OPT-LABEL: @callee_jint
+; OPT:       L#end_of_prologue4
+; OPT:        lgfr 1,1
+; OPT:        basr 7,6
+
+; NOOPT-LABEL: @callee_jint
+; NOOPT:       L#end_of_prologue4
+; NOOPT:        lgfr 1,
+; NOOPT:        basr 7,6
+
+define hidden void @callee_jint(i32 noundef signext %c) local_unnamed_addr {
+entry:
+  %conv = sext i32 %c to i64
+  tail call void @print(i64 noundef %conv)
+  ret void
+}
+
+;----------------------------------------------------------------------------
+; jbyte optnone: value spilled to alloca.  The reload uses lgb (sign-extend
+; byte load) rather than lgbr, because the value passes through memory.
+; At both -O3 and -O0: stc+lgb.  The stc source register may be 0 or 1.
+;----------------------------------------------------------------------------
+; OPT-LABEL: @callee_jbyte_optnone
+; OPT:       L#end_of_prologue5
+; OPT:        stc {{[01]}},{{[0-9]+}}(4)
+; OPT:        lgb 1,{{[0-9]+}}(4)
+; OPT:        basr 7,6
+
+; NOOPT-LABEL: @callee_jbyte_optnone
+; NOOPT:       L#end_of_prologue5
+; NOOPT:        stc {{[01]}},{{[0-9]+}}(4)
+; NOOPT:        lgb 1,{{[0-9]+}}(4)
+; NOOPT:        basr 7,6
+
+define hidden void @callee_jbyte_optnone(i8 noundef signext %c) noinline optnone {
+entry:
+  %c.addr = alloca i8, align 1
+  store i8 %c, ptr %c.addr, align 1
+  %0 = load i8, ptr %c.addr, align 1
+  %conv = sext i8 %0 to i64
+  call void @print(i64 noundef %conv)
+  ret void
+}
+
+;----------------------------------------------------------------------------
+; struct S8 { signed char x } coerced to i64 (struct, no CCIfExtend).
+; The byte occupies the high bits; srag 1,1,56 extracts and sign-extends.
+;----------------------------------------------------------------------------
+; OPT-LABEL: @callee_s8
+; OPT:       L#end_of_prologue6
+; OPT:        srag 1,1,56
+; OPT:        basr 7,6
+
+; NOOPT-LABEL: @callee_s8
+; NOOPT:       L#end_of_prologue6
+; NOOPT:        srag 1,1,56
+; NOOPT:        basr 7,6
+
+define hidden void @callee_s8(i64 %s.coerce) local_unnamed_addr {
+entry:
+  %conv = ashr i64 %s.coerce, 56
+  tail call void @print(i64 noundef %conv)
+  ret void
+}
+
+;----------------------------------------------------------------------------
+; struct S16 { short x } coerced to i64.
+; The halfword occupies the high bits; srag 1,1,48 extracts and sign-extends.
+;----------------------------------------------------------------------------
+; OPT-LABEL: @callee_s16
+; OPT:       L#end_of_prologue7
+; OPT:        srag 1,1,48
+; OPT:        basr 7,6
+
+; NOOPT-LABEL: @callee_s16
+; NOOPT:       L#end_of_prologue7
+; NOOPT:        srag 1,1,48
+; NOOPT:        basr 7,6
+
+define hidden void @callee_s16(i64 %s.coerce) local_unnamed_addr {
+entry:
+  %conv = ashr i64 %s.coerce, 48
+  tail call void @print(i64 noundef %conv)
+  ret void
+}
+
+;----------------------------------------------------------------------------
+; struct S32 { int x } coerced to i64.
+; The word occupies the high bits; srag 1,1,32 extracts and sign-extends.
+;----------------------------------------------------------------------------
+; OPT-LABEL: @callee_s32
+; OPT:       L#end_of_prologue8
+; OPT:        srag 1,1,32
+; OPT:        basr 7,6
+
+; NOOPT-LABEL: @callee_s32
+; NOOPT:       L#end_of_prologue8
+; NOOPT:        srag 1,1,32
+; NOOPT:        basr 7,6
+
+define hidden void @callee_s32(i64 %s.coerce) local_unnamed_addr {
+entry:
+  %conv = ashr i64 %s.coerce, 32
+  tail call void @print(i64 noundef %conv)
+  ret void
+}
+
+declare void @print(i64 noundef)

>From e0763edba3814d5b4afec8d48b2d32773747ce44 Mon Sep 17 00:00:00 2001
From: Zibi Sarbinowski <zibi at ca.ibm.com>
Date: Tue, 25 Aug 2026 12:15:32 -0400
Subject: [PATCH 2/6] [SystemZ] XPLINK64: probe GR32 aliases in
 CC_XPLINK_Promote_i32; add i32 CC table rule
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit

The previous implementation probed R1D/R2D/R3D (GR64) to decide whether
an i8/i16 formal argument should take the register path (LocVT=i32) or
the stack path (LocVT=i64), but the CC table had no CCAssignToRegAndStack
rule for i32 — only for i64. This meant that even when a GR32 slot was
free the i32 LocVT fell through to the stack, generating LGB/LGH (memory
forms) instead of LGBR/LGHR (register forms) for the first argument.

Fix:
- Probe R1L/R2L/R3L (GR32) instead of R1D/R2D/R3D in CC_XPLINK_Promote_i32,
  consistent with the register set that the CC table actually assigns i32 to.
- Add CCIfType<[i32], CCAssignToRegAndStack<[R1L, R2L, R3L], 8, 8>> to
  CC_SystemZ_XPLINK64 so that i32 LocVT args are assigned to GR32 registers
  (matching the Woz branch behaviour).

Result: single-argument i8/i16 signext/zeroext formals now land in R1L
and are sign/zero-extended with the register form (LGBR/LGHR/LLGCR/LLGHR)
rather than the memory form.

Update call-zos-01.ll and call-zos-vararg.ll to match the new codegen.
---
 llvm/lib/Target/SystemZ/SystemZCallingConv.h  | 15 +++++++++------
 llvm/lib/Target/SystemZ/SystemZCallingConv.td |  6 +++---
 llvm/test/CodeGen/SystemZ/call-zos-01.ll      |  7 ++++---
 llvm/test/CodeGen/SystemZ/call-zos-vararg.ll  |  9 ++++-----
 4 files changed, 20 insertions(+), 17 deletions(-)

diff --git a/llvm/lib/Target/SystemZ/SystemZCallingConv.h b/llvm/lib/Target/SystemZ/SystemZCallingConv.h
index 8b8afd4b61e94..bfabfccc8547a 100644
--- a/llvm/lib/Target/SystemZ/SystemZCallingConv.h
+++ b/llvm/lib/Target/SystemZ/SystemZCallingConv.h
@@ -150,15 +150,18 @@ inline bool CC_XPLINK_Promote_i32(unsigned &ValNo, MVT &ValVT, MVT &LocVT,
       // three are taken, this arg goes to an 8-byte stack slot.  The caller
       // stores a sign-extended i64 in that slot (value in bytes 6-7), so we
       // must use LocVT=i64 to load all 8 bytes correctly.
-      // We probe the 64-bit register aliases: AllocateReg marks all aliases,
-      // so if R1D is allocated (by a prior i64/ptr arg) R1L is also marked.
-      static const MCPhysReg GPR64s[] = {SystemZ::R1D, SystemZ::R2D,
-                                         SystemZ::R3D};
-      if (State.getFirstUnallocated(GPR64s) < 3) {
+      // Keep LocVT=i32 (GR32 live-in) only if a GR32 register is still free.
+      // XPLINK64 assigns i32 to R1L/R2L/R3L; if all three are taken this arg
+      // goes to an 8-byte stack slot where the value lives in bytes 6-7 (the
+      // caller stores a sign-extended i64).  In that case we must use
+      // LocVT=i64 so the 8-byte slot is loaded correctly.
+      static const MCPhysReg GPR32s[] = {SystemZ::R1L, SystemZ::R2L,
+                                         SystemZ::R3L};
+      if (State.getFirstUnallocated(GPR32s) < 3) {
         LocInfo = CCValAssign::AExt;
         return false;
       }
-      // All GR64s (and their GR32 aliases) taken — stack path: promote to i64.
+      // All GR32s taken — stack path: promote to i64.
     }
     // i32 formal (or stack-bound i8/i16): promote to GR64 via AExt.
     LocVT = MVT::i64;
diff --git a/llvm/lib/Target/SystemZ/SystemZCallingConv.td b/llvm/lib/Target/SystemZ/SystemZCallingConv.td
index e080463a853a7..1fdad8c627a90 100644
--- a/llvm/lib/Target/SystemZ/SystemZCallingConv.td
+++ b/llvm/lib/Target/SystemZ/SystemZCallingConv.td
@@ -243,9 +243,9 @@ def CC_SystemZ_XPLINK64 : CallingConv<[
   // If i128 is not legal, such values are already split into two i64 here,
   // so we have to use a custom handler.
   CCIfType<[i64], CCCustom<"CC_SystemZ_I128Indirect">>,
-  // The first 3 integer arguments are passed in registers R1D-R3D.
-  // The rest will be passed in the user area. The address offset of the user
-  // area can be found in register R4D.
+  // The first 3 integer arguments are passed in registers R1-R3.
+  // The rest will be passed in the user area.
+  CCIfType<[i32], CCAssignToRegAndStack<[R1L, R2L, R3L], 8, 8>>,
   CCIfType<[i64], CCAssignToRegAndStack<[R1D, R2D, R3D], 8, 8>>,
 
   // The first 8 named vector arguments are passed in V24-V31. Sub-128 vectors
diff --git a/llvm/test/CodeGen/SystemZ/call-zos-01.ll b/llvm/test/CodeGen/SystemZ/call-zos-01.ll
index d1a79b10bfe9e..19ed55f5a5670 100644
--- a/llvm/test/CodeGen/SystemZ/call-zos-01.ll
+++ b/llvm/test/CodeGen/SystemZ/call-zos-01.ll
@@ -129,7 +129,7 @@ entry:
 
 define signext i8 @pass_char(i8 signext %arg) {
 ; CHECK-LABEL: pass_char DS 0H
-; CHECK:         lgb 3,
+; CHECK:         lgbr 3,1
 ; CHECK-NEXT:    b 2(7)
 entry:
   ret i8 %arg
@@ -137,7 +137,7 @@ entry:
 
 define signext i16 @pass_short(i16 signext %arg) {
 ; CHECK-LABEL: pass_short DS 0H
-; CHECK:         lgh 3,
+; CHECK:         lghr 3,1
 ; CHECK-NEXT:    b 2(7)
 entry:
   ret i16 %arg
@@ -164,7 +164,8 @@ entry:
 
 define signext i64 @pass_integrals0(i64 signext %arg0, i32 signext %arg1, i16 signext %arg2, i64 signext %arg3) {
 ; CHECK-LABEL: pass_integrals0 DS 0H
-; CHECK:         agfr 3,2
+; CHECK:         lgfr 3,2
+; CHECK-NEXT:    ag 3,2200(4)
 ; CHECK-NEXT:    b 2(7)
 entry:
   %N = sext i32 %arg1 to i64
diff --git a/llvm/test/CodeGen/SystemZ/call-zos-vararg.ll b/llvm/test/CodeGen/SystemZ/call-zos-vararg.ll
index 5ecc2bb4b152a..ca654394ee150 100644
--- a/llvm/test/CodeGen/SystemZ/call-zos-vararg.ll
+++ b/llvm/test/CodeGen/SystemZ/call-zos-vararg.ll
@@ -287,14 +287,13 @@ define void @call_vec_double_vararg_straddle(<2 x double> %v) {
 ; CHECK-NEXT:    aghi 4,-192
 ; CHECK-NEXT:    *FENCE
 ; CHECK-NEXT:    L#end_of_prologue{{[0-9]+}} DS 0H
-; CHECK-NEXT:    lgh 0,
-; CHECK-NEXT:    lgb 7,
+; CHECK-NEXT:    lgb 0,
 ; CHECK-NEXT:    lg 6,40(5)
 ; CHECK-NEXT:    lg 5,32(5)
-; CHECK-NEXT:    lgr 3,2
+; CHECK-NEXT:    * kill: def $r2l killed $r2l def $r2d
+; CHECK-NEXT:    lghr 2,2
 ; CHECK-NEXT:    lgfr 1,1
-; CHECK-NEXT:    stg 7,2200(4)
-; CHECK-NEXT:    lgr 2,0
+; CHECK-NEXT:    stg 0,2200(4)
 ; CHECK-NEXT:    basr 7,6
 ; CHECK-NEXT:    bcr 0,0
 ; CHECK-NEXT:    lg 7,2072(4)

>From f62b19cd5658b7a953b0984967f8f24706ee8588 Mon Sep 17 00:00:00 2001
From: Zibi Sarbinowski <zibi at ca.ibm.com>
Date: Wed, 26 Aug 2026 14:10:02 -0400
Subject: [PATCH 3/6] [SystemZ] XPLINK64: simplify CC_XPLINK_Promote_i32 per
 uweigand review
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit

Remove the ArgOrigVTs tracking, AnalyzeFormalArguments override, and
getArgOrigVT() accessor from SystemZCCState — they were used to
distinguish i8/i16 from i32 in CC_XPLINK_Promote_i32 in order to
decide between LocVT=i32 (GR32 register path) and LocVT=i64 (stack
path).  As uweigand pointed out, AExt-to-i32 and AExt-to-i64 have
exactly the same semantics (only the low bytes are valid; high bytes
are undefined), so the distinction is unnecessary.

CC_XPLINK_Promote_i32 now always uses LocVT=i64:
- Formal args: AExt (callee cannot rely on caller having extended)
- Call operands: preserve SExt/ZExt/AExt from ArgFlags (caller extends)

The TRUNCATE insertion in LowerFormalArguments (Ins[I].ArgVT vs
VA.getValVT()) is kept: it truncates the AExt i64 back to the true
i8/i16 type so the subsequent DAG sign/zero-extend selects the correct
narrow instruction (LGBR/LGHR/LLGCR/LLGHR for register args,
LGB/LGH/LLGC/LLGH for stack args), reading from the natural-width
bits rather than from bit 31.

Also revert CCIfType<[i8, i16, i32]> to CCIfType<[i32]>: i8/i16 are
always legalized to i32 before the CC table runs and are never seen as
i8/i16 types there.

Remove the stale '* kill: def $r2l' annotation from call-zos-vararg.ll
that was an artifact of the old GR32 live-in path.
---
 llvm/lib/Target/SystemZ/SystemZCallingConv.h  | 77 ++++---------------
 llvm/lib/Target/SystemZ/SystemZCallingConv.td |  2 +-
 .../Target/SystemZ/SystemZISelLowering.cpp    | 11 ++-
 llvm/test/CodeGen/SystemZ/call-zos-vararg.ll  |  1 -
 4 files changed, 20 insertions(+), 71 deletions(-)

diff --git a/llvm/lib/Target/SystemZ/SystemZCallingConv.h b/llvm/lib/Target/SystemZ/SystemZCallingConv.h
index bfabfccc8547a..9ac0f798f6071 100644
--- a/llvm/lib/Target/SystemZ/SystemZCallingConv.h
+++ b/llvm/lib/Target/SystemZ/SystemZCallingConv.h
@@ -85,41 +85,16 @@ inline bool CC_SystemZ_I128Indirect(unsigned &ValNo, MVT &ValVT,
   return true;
 }
 
-// Extends CCState with:
-//   - a flag to distinguish formal-argument lowering from outgoing-call
-//     lowering (used by CC_XPLINK_Promote_i32)
-//   - the pre-legalization original MVT for each argument (needed because
-//     i8/i16 are both legalized to i32 before the CC table runs, so ValVT
-//     alone cannot distinguish them)
+// Extends CCState to track whether we are lowering formal arguments or
+// outgoing call operands.  CC_XPLINK_Promote_i32 uses this to decide
+// whether to use AExt (formal args — callee cannot rely on caller extension)
+// or to preserve the SExt/ZExt flags (call operands — caller must extend).
 class SystemZCCState : public CCState {
   bool IsFormalArgLowering = false;
-  SmallVector<MVT, 8> ArgOrigVTs;
 
 public:
   using CCState::CCState;
 
-  void AnalyzeFormalArguments(const SmallVectorImpl<ISD::InputArg> &Ins,
-                              CCAssignFn Fn) {
-    // Record the pre-legalization original type for each argument.
-    // Use MVT::Other as a sentinel when ArgVT is not a simple type (e.g.
-    // split or extended EVTs that occur in 32-bit mode).
-    ArgOrigVTs.clear();
-    for (const auto &In : Ins)
-      ArgOrigVTs.push_back(In.ArgVT.isSimple() ? In.ArgVT.getSimpleVT()
-                                               : MVT::Other);
-    CCState::AnalyzeFormalArguments(Ins, Fn);
-  }
-
-  // Returns the pre-legalization MVT for argument ValNo.
-  // Only valid after AnalyzeFormalArguments; returns MVT::Other if out of
-  // range (e.g. called from AnalyzeCallOperands context where ArgOrigVTs
-  // was never populated).
-  MVT getArgOrigVT(unsigned ValNo) const {
-    if (ValNo >= ArgOrigVTs.size())
-      return MVT::Other;
-    return ArgOrigVTs[ValNo];
-  }
-
   bool isFormalArgLowering() const { return IsFormalArgLowering; }
   void setIsFormalArgLowering() { IsFormalArgLowering = true; }
 };
@@ -138,42 +113,18 @@ inline bool CC_XPLINK64_Pointer(unsigned &ValNo, MVT &ValVT, MVT &LocVT,
 inline bool CC_XPLINK_Promote_i32(unsigned &ValNo, MVT &ValVT, MVT &LocVT,
                                   CCValAssign::LocInfo &LocInfo,
                                   ISD::ArgFlagsTy &ArgFlags, CCState &State) {
-  SystemZCCState *SZState = static_cast<SystemZCCState *>(&State);
-  if (SZState->isFormalArgLowering()) {
-    // For formal arguments, use the pre-legalization original MVT to
-    // distinguish i8/i16 from i32 — all three are legalized to i32 (ValVT)
-    // before the CC table runs, so ValVT alone cannot tell them apart.
-    MVT OrigVT = SZState->getArgOrigVT(ValNo);
-    if (OrigVT == MVT::i8 || OrigVT == MVT::i16) {
-      // Keep LocVT=i32 (GR32 live-in) only if a GR32 register is still free.
-      // The CC table assigns i32 args to R1L/R2L/R3L in XPLINK64; if all
-      // three are taken, this arg goes to an 8-byte stack slot.  The caller
-      // stores a sign-extended i64 in that slot (value in bytes 6-7), so we
-      // must use LocVT=i64 to load all 8 bytes correctly.
-      // Keep LocVT=i32 (GR32 live-in) only if a GR32 register is still free.
-      // XPLINK64 assigns i32 to R1L/R2L/R3L; if all three are taken this arg
-      // goes to an 8-byte stack slot where the value lives in bytes 6-7 (the
-      // caller stores a sign-extended i64).  In that case we must use
-      // LocVT=i64 so the 8-byte slot is loaded correctly.
-      static const MCPhysReg GPR32s[] = {SystemZ::R1L, SystemZ::R2L,
-                                         SystemZ::R3L};
-      if (State.getFirstUnallocated(GPR32s) < 3) {
-        LocInfo = CCValAssign::AExt;
-        return false;
-      }
-      // All GR32s taken — stack path: promote to i64.
-    }
-    // i32 formal (or stack-bound i8/i16): promote to GR64 via AExt.
-    LocVT = MVT::i64;
+  LocVT = MVT::i64;
+  if (static_cast<SystemZCCState &>(State).isFormalArgLowering()) {
+    // Formal arguments: the callee cannot rely on the caller having extended
+    // the value, so treat the incoming register as any-extended (upper bits
+    // undefined).  The callee will re-extend from the natural width if needed.
     LocInfo = CCValAssign::AExt;
+  } else if (ArgFlags.isSExt()) {
+    LocInfo = CCValAssign::SExt;
+  } else if (ArgFlags.isZExt()) {
+    LocInfo = CCValAssign::ZExt;
   } else {
-    LocVT = MVT::i64;
-    if (ArgFlags.isSExt())
-      LocInfo = CCValAssign::SExt;
-    else if (ArgFlags.isZExt())
-      LocInfo = CCValAssign::ZExt;
-    else
-      LocInfo = CCValAssign::AExt;
+    LocInfo = CCValAssign::AExt;
   }
   return false;
 }
diff --git a/llvm/lib/Target/SystemZ/SystemZCallingConv.td b/llvm/lib/Target/SystemZ/SystemZCallingConv.td
index 1fdad8c627a90..2968c47e0090f 100644
--- a/llvm/lib/Target/SystemZ/SystemZCallingConv.td
+++ b/llvm/lib/Target/SystemZ/SystemZCallingConv.td
@@ -216,7 +216,7 @@ def CC_SystemZ_XPLINK64 : CallingConv<[
   // Promote sub-i64 integers to i64 if they have an explicit extension type.
   // The convention is that true integer arguments smaller than 64 bits should
   // be marked as extended, but structures coerced to an integer type shouldn't.
-  CCIfType<[i8, i16, i32], CCIfExtend<CCCustom<"CC_XPLINK_Promote_i32">>>,
+  CCIfType<[i32], CCIfExtend<CCCustom<"CC_XPLINK_Promote_i32">>>,
   // Promote f32 to f64 and bitcast to i64, if it needs to be passed in GPRs.
   // Although we assign the f32 vararg to be bitcast, it will first be promoted
   // to an f64 within convertValVTToLocVT().
diff --git a/llvm/lib/Target/SystemZ/SystemZISelLowering.cpp b/llvm/lib/Target/SystemZ/SystemZISelLowering.cpp
index 8fbfab795ed9d..d52815b3b5b8a 100644
--- a/llvm/lib/Target/SystemZ/SystemZISelLowering.cpp
+++ b/llvm/lib/Target/SystemZ/SystemZISelLowering.cpp
@@ -2155,12 +2155,11 @@ SDValue SystemZTargetLowering::LowerFormalArguments(
       }
     } else {
       SDValue Val = convertLocVTToValVT(DAG, DL, VA, Chain, ArgValue);
-      // For i8/i16 formals under XPLINK64, CC_XPLINK_Promote_i32 either keeps
-      // LocVT=i32 (register arg) or uses LocVT=i64 (stack arg, 8-byte slot).
-      // Either way convertLocVTToValVT yields ValVT=i32 (legalized).  Truncate
-      // down to the true original i8/i16 so the DAG folds the truncate+extend
-      // into a narrow instruction (LGBR/LGHR/LGB/LGH etc).
-      // Only for XPLINK64: ELF receives i8/i16 already correctly widened.
+      // For i8/i16 formals under XPLINK64, LocVT=i64/AExt means the DAG
+      // would otherwise select LGFR (sign-extend from bit 31).  Truncate
+      // down to the true original i8/i16 so the subsequent sign/zero-extend
+      // to i64 selects the correct narrow instruction (LGBR/LGHR/LLGCR/LLGHR
+      // for register args, LGB/LGH/LLGC/LLGH for stack args).
       // Guard with isSimple() since non-simple EVTs must not reach
       // getSimpleVT().
       if (Subtarget.isTargetXPLINK64() && Ins[I].ArgVT.isSimple()) {
diff --git a/llvm/test/CodeGen/SystemZ/call-zos-vararg.ll b/llvm/test/CodeGen/SystemZ/call-zos-vararg.ll
index ca654394ee150..fe25f7a03eba4 100644
--- a/llvm/test/CodeGen/SystemZ/call-zos-vararg.ll
+++ b/llvm/test/CodeGen/SystemZ/call-zos-vararg.ll
@@ -290,7 +290,6 @@ define void @call_vec_double_vararg_straddle(<2 x double> %v) {
 ; CHECK-NEXT:    lgb 0,
 ; CHECK-NEXT:    lg 6,40(5)
 ; CHECK-NEXT:    lg 5,32(5)
-; CHECK-NEXT:    * kill: def $r2l killed $r2l def $r2d
 ; CHECK-NEXT:    lghr 2,2
 ; CHECK-NEXT:    lgfr 1,1
 ; CHECK-NEXT:    stg 0,2200(4)

>From 4ab0db330ba1cdf90af8cdfde1ec6ce06d3c221c Mon Sep 17 00:00:00 2001
From: Zibi Sarbinowski <zibi at ca.ibm.com>
Date: Thu, 27 Aug 2026 16:36:00 -0400
Subject: [PATCH 4/6] [SystemZ] Remove dead i32 CCAssignToRegAndStack rule in
 CC_SystemZ_XPLINK64

All i32 arguments with extend flags are already promoted to i64 by
CC_XPLINK_Promote_i32 before reaching this point, so the i32
CCAssignToRegAndStack rule can never fire. Remove it.

Suggested-by: Ulrich Weigand <uweigand at de.ibm.com>
---
 llvm/lib/Target/SystemZ/SystemZCallingConv.td | 1 -
 1 file changed, 1 deletion(-)

diff --git a/llvm/lib/Target/SystemZ/SystemZCallingConv.td b/llvm/lib/Target/SystemZ/SystemZCallingConv.td
index 2968c47e0090f..fd46686205741 100644
--- a/llvm/lib/Target/SystemZ/SystemZCallingConv.td
+++ b/llvm/lib/Target/SystemZ/SystemZCallingConv.td
@@ -245,7 +245,6 @@ def CC_SystemZ_XPLINK64 : CallingConv<[
   CCIfType<[i64], CCCustom<"CC_SystemZ_I128Indirect">>,
   // The first 3 integer arguments are passed in registers R1-R3.
   // The rest will be passed in the user area.
-  CCIfType<[i32], CCAssignToRegAndStack<[R1L, R2L, R3L], 8, 8>>,
   CCIfType<[i64], CCAssignToRegAndStack<[R1D, R2D, R3D], 8, 8>>,
 
   // The first 8 named vector arguments are passed in V24-V31. Sub-128 vectors

>From 53b9545d17a01f680282205fa4c0c6c1afb23def Mon Sep 17 00:00:00 2001
From: Zibi Sarbinowski <zibi at ca.ibm.com>
Date: Fri, 28 Aug 2026 12:32:57 -0400
Subject: [PATCH 5/6] [SystemZ] Add test: bare i32 formal args must land in
 R1L/R2L/R3L under XPLINK64

Bare i32 arguments (no signext/zeroext, e.g. sitofp sources or
struct-coerced i32s) bypass CC_XPLINK_Promote_i32 and must be
assigned to R1L/R2L/R3L by the CCAssignToRegAndStack rule.
Without that rule they fall through to the stack fallback,
producing wrong codegen (stack loads instead of register references).

This test demonstrates the regression caused by removing the
CCIfType<[i32], CCAssignToRegAndStack<[R1L,R2L,R3L],8,8>> rule
from CC_SystemZ_XPLINK64 and serves as justification for keeping it.
---
 .../SystemZ/zos-xplink64-bare-i32-arg.ll      | 38 +++++++++++++++++++
 1 file changed, 38 insertions(+)
 create mode 100644 llvm/test/CodeGen/SystemZ/zos-xplink64-bare-i32-arg.ll

diff --git a/llvm/test/CodeGen/SystemZ/zos-xplink64-bare-i32-arg.ll b/llvm/test/CodeGen/SystemZ/zos-xplink64-bare-i32-arg.ll
new file mode 100644
index 0000000000000..60523cf706643
--- /dev/null
+++ b/llvm/test/CodeGen/SystemZ/zos-xplink64-bare-i32-arg.ll
@@ -0,0 +1,38 @@
+; Test that bare i32 formal arguments (without signext/zeroext, e.g. struct
+; coercions or sitofp sources) are correctly assigned to R1L/R2L/R3L under
+; the XPLINK64 calling convention.
+;
+; The CCIfExtend-guarded CC_XPLINK_Promote_i32 handler only fires for i32
+; arguments that carry an explicit extension type.  Bare i32 arguments (no
+; extend flags) must be assigned to R1L/R2L/R3L by the separate
+; CCAssignToRegAndStack<[R1L,R2L,R3L]> rule; without that rule they fall
+; through to the stack fallback, producing a load from the parameter area
+; instead of a register reference.
+;
+; RUN: llc < %s -mtriple=s390x-ibm-zos | FileCheck %s
+
+; Bare i32 passed to sitofp.  Correct: cefbr 0,1 (R1 is the source).
+; Without the i32 rule: l 0,<offset>(4) then cefbr 0,0 (stack load).
+define float @sitofp_bare_i32(i32 %x) {
+; CHECK-LABEL: sitofp_bare_i32 DS 0H
+; CHECK: cefbr 0,1
+  %r = sitofp i32 %x to float
+  ret float %r
+}
+
+; Bare i32 passed to sitofp (double).  Correct: cdfbr 0,1.
+define double @sitofp_bare_i32_double(i32 %x) {
+; CHECK-LABEL: sitofp_bare_i32_double DS 0H
+; CHECK: cdfbr 0,1
+  %r = sitofp i32 %x to double
+  ret double %r
+}
+
+; Two bare i32 arguments compared.  Correct: cr 1,2 (register compare).
+; Without the i32 rule: two stack loads then a memory/register compare.
+define i1 @cmp_bare_i32(i32 %a, i32 %b) {
+; CHECK-LABEL: cmp_bare_i32 DS 0H
+; CHECK: cr 1,2
+  %r = icmp eq i32 %a, %b
+  ret i1 %r
+}

>From 5314b3d87a00a6a9b3b3493105acbbe92a42cb13 Mon Sep 17 00:00:00 2001
From: Zibi Sarbinowski <zibi at ca.ibm.com>
Date: Fri, 28 Aug 2026 12:33:10 -0400
Subject: [PATCH 6/6] Revert "[SystemZ] Remove dead i32 CCAssignToRegAndStack
 rule in CC_SystemZ_XPLINK64"

This reverts commit 4ab0db330ba1cdf90af8cdfde1ec6ce06d3c221c.
---
 llvm/lib/Target/SystemZ/SystemZCallingConv.td | 1 +
 1 file changed, 1 insertion(+)

diff --git a/llvm/lib/Target/SystemZ/SystemZCallingConv.td b/llvm/lib/Target/SystemZ/SystemZCallingConv.td
index fd46686205741..2968c47e0090f 100644
--- a/llvm/lib/Target/SystemZ/SystemZCallingConv.td
+++ b/llvm/lib/Target/SystemZ/SystemZCallingConv.td
@@ -245,6 +245,7 @@ def CC_SystemZ_XPLINK64 : CallingConv<[
   CCIfType<[i64], CCCustom<"CC_SystemZ_I128Indirect">>,
   // The first 3 integer arguments are passed in registers R1-R3.
   // The rest will be passed in the user area.
+  CCIfType<[i32], CCAssignToRegAndStack<[R1L, R2L, R3L], 8, 8>>,
   CCIfType<[i64], CCAssignToRegAndStack<[R1D, R2D, R3D], 8, 8>>,
 
   // The first 8 named vector arguments are passed in V24-V31. Sub-128 vectors



More information about the llvm-commits mailing list