[all-commits] [llvm/llvm-project] 90d8ed: [LowerMemIntrinsics][AMDGPU] Optimize memset.patte...

Fabian Ritter via All-commits all-commits at lists.llvm.org
Thu Mar 12 05:52:59 PDT 2026


  Branch: refs/heads/users/ritter-x2a/03-11-_lowermemintrinsics_amdgpu_optimize_memset.pattern_lowering
  Home:   https://github.com/llvm/llvm-project
  Commit: 90d8edfae004fe3a504c99cb055317e960b639f0
      https://github.com/llvm/llvm-project/commit/90d8edfae004fe3a504c99cb055317e960b639f0
  Author: Fabian Ritter <fabian.ritter at amd.com>
  Date:   2026-03-12 (Thu, 12 Mar 2026)

  Changed paths:
    M llvm/include/llvm/Transforms/Utils/LowerMemIntrinsics.h
    M llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp
    M llvm/lib/Target/AMDGPU/AMDGPULowerBufferFatPointers.cpp
    M llvm/lib/Transforms/Utils/LowerMemIntrinsics.cpp
    M llvm/test/CodeGen/AMDGPU/lower-buffer-fat-pointers-mem-transfer.ll
    A llvm/test/CodeGen/AMDGPU/memset-pattern.ll
    M llvm/test/CodeGen/RISCV/memset-pattern.ll
    A llvm/test/Transforms/PreISelIntrinsicLowering/AMDGPU/memset-pattern.ll
    M llvm/test/Transforms/PreISelIntrinsicLowering/PowerPC/memset-pattern.ll
    M llvm/test/Transforms/PreISelIntrinsicLowering/RISCV/memset-pattern.ll
    M llvm/test/Transforms/PreISelIntrinsicLowering/X86/memset-pattern.ll

  Log Message:
  -----------
  [LowerMemIntrinsics][AMDGPU] Optimize memset.pattern lowering

This patch changes the lowering of the [experimental.memset.pattern intrinsic](https://llvm.org/docs/LangRef.html#llvm-experimental-memset-pattern-intrinsic)
to match the optimized memset and memcpy lowering when possible. (The tl;dr of
memset.pattern is that it is like memset, except that you can use it to set
values that are wider than a single byte.)

The memset.pattern lowering now queries `TTI::getMemcpyLoopLoweringType` for a
preferred memory access type. If the size of that type is a multiple of the set
value's type, and if both types have consistent store and alloc sizes (since
memset.pattern behaves in a way that is not well suitable for access widening
if store and alloc size differ), the memset.pattern is lowered into two loops:
a main loop that stores a sufficiently wide vector splat of the SetValue with
the preferred memory access type and a residual loop that covers the remaining
set values individually.

In contrast to the memset lowering, this patch doesn't include a specialized
lowering for residual loops with known constant lengths. Loops that are
statically known to be unreachable will not be emitted.

For backends that don't override `TTI::getMemcpyLoopLoweringType`, the
generated code is mostly unchanged except for more consistent basic block
names, no more `br i1 false` for memset.patterns with known size, and a flipped
loop condition for memset.patterns with known size (see test changes).

This is a follow-up to a similar patch for memset: #169040


  Commit: 3a8c16f7ae600268096c5121ae5545f3a355c3b5
      https://github.com/llvm/llvm-project/commit/3a8c16f7ae600268096c5121ae5545f3a355c3b5
  Author: Fabian Ritter <fabian.ritter at amd.com>
  Date:   2026-03-12 (Thu, 12 Mar 2026)

  Changed paths:
    M llvm/test/CodeGen/AMDGPU/memset-pattern.ll

  Log Message:
  -----------
  Add AS7 tests


Compare: https://github.com/llvm/llvm-project/compare/da546fcd03e5...3a8c16f7ae60

To unsubscribe from these emails, change your notification settings at https://github.com/llvm/llvm-project/settings/notifications


More information about the All-commits mailing list