[all-commits] [llvm/llvm-project] 94c1a8: [NVPTX] Fold symbol addresses into memory operands...

Tim Besard via All-commits all-commits at lists.llvm.org
Thu Jun 25 08:22:54 PDT 2026


  Branch: refs/heads/main
  Home:   https://github.com/llvm/llvm-project
  Commit: 94c1a8bc3e6b790ccc5723803129b3b956f11d13
      https://github.com/llvm/llvm-project/commit/94c1a8bc3e6b790ccc5723803129b3b956f11d13
  Author: Tim Besard <tim.besard at gmail.com>
  Date:   2026-06-25 (Thu, 25 Jun 2026)

  Changed paths:
    M llvm/lib/Target/NVPTX/CMakeLists.txt
    M llvm/lib/Target/NVPTX/MCTargetDesc/NVPTXMCTargetDesc.cpp
    M llvm/lib/Target/NVPTX/MCTargetDesc/NVPTXMCTargetDesc.h
    M llvm/lib/Target/NVPTX/NVPTX.h
    A llvm/lib/Target/NVPTX/NVPTXAddressFolder.cpp
    M llvm/lib/Target/NVPTX/NVPTXInstrFormats.td
    M llvm/lib/Target/NVPTX/NVPTXInstrInfo.td
    M llvm/lib/Target/NVPTX/NVPTXTargetMachine.cpp
    A llvm/test/CodeGen/NVPTX/address-folder.ll
    A llvm/test/CodeGen/NVPTX/address-folder.mir

  Log Message:
  -----------
  [NVPTX] Fold symbol addresses into memory operands (#202379)

SelectionDAG can fold a symbol address (a kernel parameter, global
variable, or external symbol) directly into a memory instruction's
address operand, but only within a single basic block. When the address
crosses a block boundary, ISel materializes it with `MOV_B{32,64}_sym`
and the memory instruction becomes register-relative:

```ptx
mov.b64       %rd1, kernel_param_0;
ld.param.b64  %rd2, [%rd1];
ld.param.b64  %rd3, [%rd1+8];
```

instead of:

```
ld.param.b64  %rd2, [kernel_param_0];
ld.param.b64  %rd3, [kernel_param_0+8];
```

This patch adds NVPTXAddressFolder, a pre-regalloc pass that looks for
loads and stores whose address operand is defined by `MOV_B{32,64}_sym`,
then folds the symbol back into the memory operand. The mov is erased
once it has no remaining uses; if the address also feeds arithmetic or
escapes, it is kept.

To make this generic over NVPTX memory instructions, the patch enables
named operand tables for NVPTX instructions and uses them to find `addr`
and `addsp` operands instead of hardcoding opcode-specific operand
indices.

This was motivated by a CUDA.jl performance regression where `byval`
kernel parameters stopped being pre-lowered into simple loads and
exposed the missing cross-block fold in the backend.

Disclaimer: LLMs (GPT 5.5, Opus 4.8) were used to develop this PR.

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply at anthropic.com>



To unsubscribe from these emails, change your notification settings at https://github.com/llvm/llvm-project/settings/notifications


More information about the All-commits mailing list