[llvm] [SelectionDAG] Widen vector math libcalls when no routine is available (PR #218948)
David Sherwood via llvm-commits
llvm-commits at lists.llvm.org
Mon Sep 7 04:08:28 PDT 2026
=?utf-8?q?Mattéo?= Rizza Murgier,=?utf-8?q?Mattéo?= Rizza Murgier,
=?utf-8?q?Mattéo?= Rizza Murgier,=?utf-8?q?Mattéo?= Rizza Murgier,
=?utf-8?q?Mattéo?= Rizza Murgier,=?utf-8?q?Mattéo?= Rizza Murgier,
=?utf-8?q?Mattéo?= Rizza Murgier,=?utf-8?q?Mattéo?= Rizza Murgier
Message-ID:
In-Reply-To: <llvm.org/llvm/llvm-project/pull/218948 at github.com>
================
@@ -112,5 +112,208 @@ define <vscale x 2 x double> @frem_strict_nxv2f64(<vscale x 2 x double> %unused,
ret <vscale x 2 x double> %res
}
+; Expected to be widened.
+define <2 x float> @frem_v2f32(<2 x float> %unused, <2 x float> %a, <2 x float> %b) #0 {
+; ARMPL-LABEL: frem_v2f32:
+; ARMPL: // %bb.0:
+; ARMPL-NEXT: str x30, [sp, #-16]! // 8-byte Folded Spill
+; ARMPL-NEXT: .cfi_def_cfa_offset 16
+; ARMPL-NEXT: .cfi_offset w30, -16
+; ARMPL-NEXT: fmov d0, d1
+; ARMPL-NEXT: // kill: def $d2 killed $d2 def $q2
+; ARMPL-NEXT: mov v1.16b, v2.16b
+; ARMPL-NEXT: bl armpl_vfmodq_f32
----------------
david-arm wrote:
OK thanks for this! The reason I mentioned this is because I have seen real workloads write code like this:
```
float p0 = powf(x[i], y[i]);
float p1 = powf(x[i+1], y[i+1]);
float p2 = powf(x[i+2], y[i+2]);
float p3 = powf(x[i+3], y[i+3]);
float res0 = cond[i] ? p0 : ...;
float res1 = cond[i+1] ? p1 : ...;
float res2 = cond[i+2] ? p2 : ...;
float res3 = cond[i+3] ? p3 : ...;
```
where in the cases where the powf result was unused, e.g. when `cond[i] = 0` the exponents passed in to `powf` were also junk, which took us down the slow path in vector math calls. So by duplicating the first two lanes into the top lanes it's certainly better than poison, but also not a guarantee that the inputs themselves are sane. :) In addition, LLVM aggressively speculatively executes math calls for C code like this:
```
float p0 = 0.0f, p1 = 0.0f, p2 = 0.0f, p3 = 0.0f;
if (cond[i]) {
p0 = powf(x[i], y[i]);
}
...
```
by transforming this into
```
float p0 = powf(x[i], y[i]);
float p1 = powf(x[i+1], y[i+1]);
float p2 = powf(x[i+2], y[i+2]);
float p3 = powf(x[i+3], y[i+3]);
```
when building with -ffast-math. Such aggressive speculation increases the chance of going down the slow path in math calls too. If we see any performance issues due to this in future, we can always revisit this to use explicit values.
https://github.com/llvm/llvm-project/pull/218948
More information about the llvm-commits
mailing list