[llvm] [NVPTX] Add support for f32x2 mixed-precision add/sub (PR #221957)
Srinivasa Ravi via llvm-commits
llvm-commits at lists.llvm.org
Wed Sep 9 22:47:20 PDT 2026
================
@@ -2508,6 +2537,24 @@ foreach rnd = FPRoundingModes in {
(f32 (fpextend type:$a)),
f32:$b, rnd_imm))]>,
Requires<[SM100]>;
+
+ let Predicates = [hasRubinFamilySupport] in {
+ foreach t = [F16X2RT, BF16X2RT] in {
+ def INT_NVVM_MIXED_ADD_ # rnd # _f32x2_ # t.PtxType :
+ BasicNVPTXInst<(outs B64:$dst), (ins B32:$a, B64:$b),
+ "add." # rnd # ".f32x2." # t.PtxType # ".f32x2",
+ [(set v2f32:$dst,
+ (int_nvvm_fadd (fpextend_v2f32_from_packed t.Ty:$a), v2f32:$b,
+ rnd_imm))]>;
+
+ def : Pat<(v2f32 (int_nvvm_fadd
+ (fpextend_v2f32_from_lanes t.Ty.ElementType:$a0,
+ t.Ty.ElementType:$a1),
+ v2f32:$b, rnd_imm)),
+ (!cast<Instruction>("INT_NVVM_MIXED_ADD_" # rnd # "_f32x2_" #
----------------
Wolfram70 wrote:
Makes sense, done!
https://github.com/llvm/llvm-project/pull/221957
More information about the llvm-commits
mailing list