[llvm] [AMDGPU] Fix uniform fcopysign pattern (PR #218644)
Jay Foad via llvm-commits
llvm-commits at lists.llvm.org
Tue Aug 25 02:41:29 PDT 2026
================
@@ -2476,6 +2476,21 @@ def : GCNPat <
>;
}
let True16Predicate = UseRealTrue16Insts in {
+// Uniform true16 values use SReg_32. Copy them to VGPR_16 before packing so
+// subregister liveness tracks the source lane in the correct coordinates.
+let AddedComplexity = 1 in
+def : GCNPat <
----------------
jayfoad wrote:
Instead of adding a new pattern can you just fix the existing pattern? You should be able to use this pattern for all inputs, uniform and divergent (but you might need to remove the `SReg_32:` from the COPY_TO_REGCLASSs). In the divergent case the COPY_TO_REGCLASS should be a no-op.
Stepping back a bit, we might not want to use V_BFI for uniform inputs. It's usually best to use SALU instructions instead. But that is just an optimization, not a correctness issue.
https://github.com/llvm/llvm-project/pull/218644
More information about the llvm-commits
mailing list