[llvm] [AMDGPU] Convert ds_bpermute/wave_shuffle XOR patterns to DPP row_xmask and permlanex16 (PR #194565)
Jay Foad via llvm-commits
llvm-commits at lists.llvm.org
Tue Apr 28 06:15:03 PDT 2026
jayfoad wrote:
IIUC permlanex16 can handle all XOR values 16..31 and permlane16 could handle all XOR values 0..15.
> The emitted DPP uses bound_ctrl=true, enabling GCNDPPCombine to fold the shuffle into the consuming math op (e.g. v_add_f32_dpp row_xmask:1).
Right but if there is no consuming math op then it's probably better to generate permlane16 instead of a DPP op?
https://github.com/llvm/llvm-project/pull/194565
More information about the llvm-commits
mailing list