[llvm] [MemCpyOpt] Extend call slot optimization for non-dereferenceable destinations. (PR #217436)

Nikita Popov via llvm-commits llvm-commits at lists.llvm.org
Mon Aug 24 00:57:57 PDT 2026


================
@@ -924,8 +924,18 @@ bool MemCpyOptPass::performCallSlotOptzn(Instruction *cpyLoad,
                         ExplicitlyDereferenceableOnly) ||
       !isDereferenceablePointer(cpyDest, APInt(64, cpySize),
                                 SimplifyQuery(DL, DT, AC, C))) {
-    LLVM_DEBUG(dbgs() << "Call Slot: Dest pointer not dereferenceable\n");
-    return false;
+    // If the call is guaranteed to return normally (willreturn + nounwind),
+    // and there are no instructions between the call and the store that might
+    // trap or throw, execution will reach the store. Since the store would
+    // trap anyway if the pointer was not dereferenceable, we can forward the
+    // pointer to the call. Perform optimization only for non-memcpy/memset
+    // calls, as those are special cased later.
----------------
nikic wrote:

> For callslot_deref.ll:must_not_remove_memcpy() The first memset gets removed which I think is correct.

Yeah, this is correct, though it might make sense to adjust the test to insert a non-willreturn operation to preserve the previous test intent.

> For the memcpy inline cases, it's going to keep the "kind" - inline or not - of the first one, not the second though, as currently tested. I'm not sure this change is correct.

If previously two different memcpy kinds were present, I think keeping either one would be correct.

> But if MemCpyOpt turns 2 memcpys into one, DSE doesn't remove anything, so the alloca isn't removed.

So we just leave behind an alloca without users? I think this is fine, it will get dropped by InstCombine. Whether alloca gets dropped by MemCpyOpt depends on which transform you hit (e.g. stack-move will drop it, but most others not).

> is this the place to handle it?

I think if there are no clear regressions in the optimizations results, we probably shouldn't go out of the way to *not* handle it.

After all, all of those optimizations already take place for the dereferenceable case, right? I think if we wanted to exclude them, we would have to exclude them from the entire call-slot optimization, not just this specific special case.


https://github.com/llvm/llvm-project/pull/217436


More information about the llvm-commits mailing list