[llvm] [NVPTX] Add support for f32x2 mixed-precision add/sub (PR #221957)

Durgadoss R via llvm-commits llvm-commits at lists.llvm.org
Wed Sep 9 05:22:52 PDT 2026


================
@@ -2518,6 +2549,40 @@ let Predicates = [SM100, doNoF32FTZ] in {
             (INT_NVVM_MIXED_ADD_rn_f32_bf16 B16:$a, B32:$b)>;
 }
 
+let Predicates = [SM107f, doNoF32FTZ] in
+  foreach t = [F16X2RT, BF16X2RT] in {
+    defvar Inst = !cast<Instruction>("INT_NVVM_MIXED_ADD_rn_f32x2_" # t.PtxType);
+
+    def : Pat<(v2f32 (fadd (fpextend_v2f32_from_packed t.Ty:$a), v2f32:$b)),
+              (Inst $a, $b)>;
+
+    def : Pat<(v2f32 (fadd (fpextend_v2f32_from_lanes t.Ty.ElementType:$a0,
+                                                      t.Ty.ElementType:$a1),
+                           v2f32:$b)),
+              (Inst (t.Ty (V2I16toI32 $a0, $a1)), $b)>;
+  }
+
+
+// mixed precision - downconverting
+
+// f16x2 - rz rounding mode, only with ftz
+// bf16x2 - rz rounding mode, without ftz
+foreach t = [F16X2RT, BF16X2RT] in {
+  defvar ftz = !if(!eq(t, F16X2RT), "ftz", "");
+  defvar AddOp = !cast<Intrinsic>(StrJoin<"_", ["int_nvvm_fadd", ftz]>.ret);
+  defvar CvtOp = !cast<Intrinsic>("int_nvvm_ff2" # t.PtxType # "_rz");
+
+  def INT_NVVM_MIXED_ADD_rz_ # t.PtxType # _f32x2 :
+    BasicNVPTXInst<(outs t.RC:$dst), (ins B64:$a, B64:$b),
+      StrJoin<".", ["add.rz", ftz, t.PtxType, "f32x2", "f32x2"]>.ret,
+      // The cvt packs its first argument into the high half of the result.
----------------
durga4github wrote:

nit: upper half (or) high bits 

https://github.com/llvm/llvm-project/pull/221957


More information about the llvm-commits mailing list