[all-commits] [llvm/llvm-project] 90d8ed: [LowerMemIntrinsics][AMDGPU] Optimize memset.patte...
Fabian Ritter via All-commits
all-commits at lists.llvm.org
Thu Mar 12 05:52:59 PDT 2026
Branch: refs/heads/users/ritter-x2a/03-11-_lowermemintrinsics_amdgpu_optimize_memset.pattern_lowering
Home: https://github.com/llvm/llvm-project
Commit: 90d8edfae004fe3a504c99cb055317e960b639f0
https://github.com/llvm/llvm-project/commit/90d8edfae004fe3a504c99cb055317e960b639f0
Author: Fabian Ritter <fabian.ritter at amd.com>
Date: 2026-03-12 (Thu, 12 Mar 2026)
Changed paths:
M llvm/include/llvm/Transforms/Utils/LowerMemIntrinsics.h
M llvm/lib/CodeGen/PreISelIntrinsicLowering.cpp
M llvm/lib/Target/AMDGPU/AMDGPULowerBufferFatPointers.cpp
M llvm/lib/Transforms/Utils/LowerMemIntrinsics.cpp
M llvm/test/CodeGen/AMDGPU/lower-buffer-fat-pointers-mem-transfer.ll
A llvm/test/CodeGen/AMDGPU/memset-pattern.ll
M llvm/test/CodeGen/RISCV/memset-pattern.ll
A llvm/test/Transforms/PreISelIntrinsicLowering/AMDGPU/memset-pattern.ll
M llvm/test/Transforms/PreISelIntrinsicLowering/PowerPC/memset-pattern.ll
M llvm/test/Transforms/PreISelIntrinsicLowering/RISCV/memset-pattern.ll
M llvm/test/Transforms/PreISelIntrinsicLowering/X86/memset-pattern.ll
Log Message:
-----------
[LowerMemIntrinsics][AMDGPU] Optimize memset.pattern lowering
This patch changes the lowering of the [experimental.memset.pattern intrinsic](https://llvm.org/docs/LangRef.html#llvm-experimental-memset-pattern-intrinsic)
to match the optimized memset and memcpy lowering when possible. (The tl;dr of
memset.pattern is that it is like memset, except that you can use it to set
values that are wider than a single byte.)
The memset.pattern lowering now queries `TTI::getMemcpyLoopLoweringType` for a
preferred memory access type. If the size of that type is a multiple of the set
value's type, and if both types have consistent store and alloc sizes (since
memset.pattern behaves in a way that is not well suitable for access widening
if store and alloc size differ), the memset.pattern is lowered into two loops:
a main loop that stores a sufficiently wide vector splat of the SetValue with
the preferred memory access type and a residual loop that covers the remaining
set values individually.
In contrast to the memset lowering, this patch doesn't include a specialized
lowering for residual loops with known constant lengths. Loops that are
statically known to be unreachable will not be emitted.
For backends that don't override `TTI::getMemcpyLoopLoweringType`, the
generated code is mostly unchanged except for more consistent basic block
names, no more `br i1 false` for memset.patterns with known size, and a flipped
loop condition for memset.patterns with known size (see test changes).
This is a follow-up to a similar patch for memset: #169040
Commit: 3a8c16f7ae600268096c5121ae5545f3a355c3b5
https://github.com/llvm/llvm-project/commit/3a8c16f7ae600268096c5121ae5545f3a355c3b5
Author: Fabian Ritter <fabian.ritter at amd.com>
Date: 2026-03-12 (Thu, 12 Mar 2026)
Changed paths:
M llvm/test/CodeGen/AMDGPU/memset-pattern.ll
Log Message:
-----------
Add AS7 tests
Compare: https://github.com/llvm/llvm-project/compare/da546fcd03e5...3a8c16f7ae60
To unsubscribe from these emails, change your notification settings at https://github.com/llvm/llvm-project/settings/notifications
More information about the All-commits
mailing list