[llvm] [NVPTX] Lower allocas to the local address space (PR #204346)
Drew Kersnar via llvm-commits
llvm-commits at lists.llvm.org
Fri Jul 10 09:34:48 PDT 2026
================
@@ -51,82 +56,228 @@ char NVPTXLowerAlloca::ID = 1;
INITIALIZE_PASS(NVPTXLowerAlloca, "nvptx-lower-alloca", "Lower Alloca", false,
false)
+static Value *getOrCreateGenericPtr(Value *LocalPtr,
+ DenseMap<Value *, Value *> &GenericPtrs) {
+ auto It = GenericPtrs.find(LocalPtr);
+ if (It != GenericPtrs.end())
+ return It->second;
+
+ auto *LocalInst = cast<Instruction>(LocalPtr);
+ auto *GenericPtr = new AddrSpaceCastInst(
+ LocalPtr, PointerType::get(LocalPtr->getContext(), ADDRESS_SPACE_GENERIC),
+ "");
+ GenericPtr->insertAfter(LocalInst->getIterator());
+ GenericPtrs[LocalPtr] = GenericPtr;
+ return GenericPtr;
+}
+
+static void updateMemIntrinsicDeclaration(MemIntrinsic *MI) {
+ SmallVector<Type *, 3> Tys;
+ if (auto *MTI = dyn_cast<MemTransferInst>(MI)) {
+ Tys.push_back(MTI->getRawDest()->getType());
+ Tys.push_back(MTI->getRawSource()->getType());
+ Tys.push_back(MTI->getLength()->getType());
+ } else {
+ auto *MSI = cast<MemSetInst>(MI);
+ Tys.push_back(MSI->getRawDest()->getType());
+ Tys.push_back(MSI->getLength()->getType());
+ }
+
+ Function *Decl = Intrinsic::getOrInsertDeclaration(MI->getModule(),
+ MI->getIntrinsicID(), Tys);
+ MI->setCalledFunction(Decl);
+}
+
+static void convertPointerUsersToLocal(Value *OldPtr, Value *LocalPtr,
+ DenseMap<Value *, Value *> &GenericPtrs,
+ SmallPtrSetImpl<Value *> &Visited) {
+ if (!Visited.insert(OldPtr).second)
+ return;
+
+ // Debug records aren't on the use-list visited below, so retarget them here;
----------------
dakersnar wrote:
Did we determine that not having InferAddressSpaces at O0 was acceptable? Just making sure I'm interpreting correctly: it looks like you were worried about correctness concerns but thought it over and decided there were none. I assume the only remaining tradeoff that comes from not running InferAddressSpaces at O0 is that the motivating example still expands byte-by-byte, right? I assume we find that acceptable since it is O0?
https://github.com/llvm/llvm-project/pull/204346
More information about the llvm-commits
mailing list