[llvm] [MemCpyOpt] Extend call slot optimization for non-dereferenceable destinations. (PR #217436)
Nikita Popov via llvm-commits
llvm-commits at lists.llvm.org
Mon Aug 24 00:57:57 PDT 2026
================
@@ -924,8 +924,18 @@ bool MemCpyOptPass::performCallSlotOptzn(Instruction *cpyLoad,
ExplicitlyDereferenceableOnly) ||
!isDereferenceablePointer(cpyDest, APInt(64, cpySize),
SimplifyQuery(DL, DT, AC, C))) {
- LLVM_DEBUG(dbgs() << "Call Slot: Dest pointer not dereferenceable\n");
- return false;
+ // If the call is guaranteed to return normally (willreturn + nounwind),
+ // and there are no instructions between the call and the store that might
+ // trap or throw, execution will reach the store. Since the store would
+ // trap anyway if the pointer was not dereferenceable, we can forward the
+ // pointer to the call. Perform optimization only for non-memcpy/memset
+ // calls, as those are special cased later.
----------------
nikic wrote:
> For callslot_deref.ll:must_not_remove_memcpy() The first memset gets removed which I think is correct.
Yeah, this is correct, though it might make sense to adjust the test to insert a non-willreturn operation to preserve the previous test intent.
> For the memcpy inline cases, it's going to keep the "kind" - inline or not - of the first one, not the second though, as currently tested. I'm not sure this change is correct.
If previously two different memcpy kinds were present, I think keeping either one would be correct.
> But if MemCpyOpt turns 2 memcpys into one, DSE doesn't remove anything, so the alloca isn't removed.
So we just leave behind an alloca without users? I think this is fine, it will get dropped by InstCombine. Whether alloca gets dropped by MemCpyOpt depends on which transform you hit (e.g. stack-move will drop it, but most others not).
> is this the place to handle it?
I think if there are no clear regressions in the optimizations results, we probably shouldn't go out of the way to *not* handle it.
After all, all of those optimizations already take place for the dereferenceable case, right? I think if we wanted to exclude them, we would have to exclude them from the entire call-slot optimization, not just this specific special case.
https://github.com/llvm/llvm-project/pull/217436
More information about the llvm-commits
mailing list