[all-commits] [llvm/llvm-project] 188aa8: [AMDGPU] Add ptr.s.buffer.load intrinsic, use it f...

Krzysztof Drewniak via All-commits all-commits at lists.llvm.org
Thu Jul 30 09:27:50 PDT 2026


  Branch: refs/heads/main
  Home:   https://github.com/llvm/llvm-project
  Commit: 188aa82fd8d0b9f322f6d5b3812f8cb4a970b60c
      https://github.com/llvm/llvm-project/commit/188aa82fd8d0b9f322f6d5b3812f8cb4a970b60c
  Author: Krzysztof Drewniak <Krzysztof.Drewniak at amd.com>
  Date:   2026-07-30 (Thu, 30 Jul 2026)

  Changed paths:
    M clang/include/clang/Basic/BuiltinsAMDGPUDocs.td
    M clang/lib/CodeGen/TargetBuiltins/AMDGPU.cpp
    M clang/test/CodeGenOpenCL/builtins-amdgcn-s-buffer-load.cl
    M llvm/include/llvm/IR/IntrinsicsAMDGPU.td
    M llvm/lib/Target/AMDGPU/AMDGPUAsanInstrumentation.cpp
    M llvm/lib/Target/AMDGPU/AMDGPUInstCombineIntrinsic.cpp
    M llvm/lib/Target/AMDGPU/AMDGPULegalizerInfo.cpp
    M llvm/lib/Target/AMDGPU/AMDGPULowerIntrinsics.cpp
    M llvm/lib/Target/AMDGPU/AMDGPURegBankLegalizeHelper.cpp
    M llvm/lib/Target/AMDGPU/AMDGPURegisterBankInfo.cpp
    M llvm/lib/Target/AMDGPU/SIISelLowering.cpp
    M llvm/lib/Target/AMDGPU/SIISelLowering.h
    M llvm/lib/Target/AMDGPU/SIInstrInfo.td
    M llvm/test/CodeGen/AMDGPU/GlobalISel/is-safe-to-sink-bug.ll
    A llvm/test/CodeGen/AMDGPU/GlobalISel/llvm.amdgcn.ptr.s.buffer.load.ll
    M llvm/test/CodeGen/AMDGPU/GlobalISel/llvm.amdgcn.set.inactive.ll
    A llvm/test/CodeGen/AMDGPU/GlobalISel/regbanklegalize-amdgcn.ptr.s.buffer.load.ll
    A llvm/test/CodeGen/AMDGPU/GlobalISel/regbanklegalize-amdgcn.ptr.s.buffer.load.subdword.ll
    A llvm/test/CodeGen/AMDGPU/GlobalISel/regbankselect-amdgcn.ptr.s.buffer.load.ll
    M llvm/test/CodeGen/AMDGPU/GlobalISel/regbankselect-amdgcn.s.buffer.load.mir
    M llvm/test/CodeGen/AMDGPU/amdgpu-uniform-intrinsic-wwm-single-lane.ll
    M llvm/test/CodeGen/AMDGPU/bug-deadlanes.ll
    M llvm/test/CodeGen/AMDGPU/bug-vopc-commute.ll
    M llvm/test/CodeGen/AMDGPU/dagcombine-fma-fmad.ll
    M llvm/test/CodeGen/AMDGPU/fneg-modifier-casting.ll
    A llvm/test/CodeGen/AMDGPU/gfx12_scalar_subword_ptr_s_buffer_loads.ll
    M llvm/test/CodeGen/AMDGPU/group-image-instructions.ll
    M llvm/test/CodeGen/AMDGPU/insert_waitcnt_for_precise_memory.ll
    A llvm/test/CodeGen/AMDGPU/llvm.amdgcn.ptr.s.buffer.load-gfx12.ll
    A llvm/test/CodeGen/AMDGPU/llvm.amdgcn.ptr.s.buffer.load-mmo.ll
    A llvm/test/CodeGen/AMDGPU/llvm.amdgcn.ptr.s.buffer.load.ll
    M llvm/test/CodeGen/AMDGPU/llvm.amdgcn.set.inactive.ll
    A llvm/test/CodeGen/AMDGPU/ptr-s-buffer-load-mmo-offsets.ll
    M llvm/test/CodeGen/AMDGPU/scalar-float-sop2.ll
    M llvm/test/CodeGen/AMDGPU/scheduler-subrange-crash.ll
    M llvm/test/CodeGen/AMDGPU/set_kill_i1_for_floation_point_comparison.ll
    M llvm/test/CodeGen/AMDGPU/sgpr-copy.ll
    M llvm/test/CodeGen/AMDGPU/si-scheduler.ll
    M llvm/test/CodeGen/AMDGPU/si-sgpr-spill.ll
    M llvm/test/CodeGen/AMDGPU/si-spill-cf.ll
    M llvm/test/CodeGen/AMDGPU/splitkit-getsubrangeformask.ll
    M llvm/test/CodeGen/AMDGPU/vgpr-spill-emergency-stack-slot.ll
    M llvm/test/CodeGen/AMDGPU/wqm.ll
    M llvm/test/CodeGen/AMDGPU/wwm-reserved-spill.ll
    M llvm/test/CodeGen/AMDGPU/wwm-reserved.ll
    M llvm/test/Transforms/EarlyCSE/AMDGPU/intrinsics.ll
    M llvm/test/Transforms/InstCombine/AMDGPU/amdgcn-demanded-vector-elts-inseltpoison.ll
    M llvm/test/Transforms/InstCombine/AMDGPU/amdgcn-demanded-vector-elts.ll
    M mlir/include/mlir/Dialect/LLVMIR/ROCDLOps.td
    M mlir/test/Dialect/LLVMIR/rocdl.mlir
    M mlir/test/Target/LLVMIR/rocdl.mlir

  Log Message:
  -----------
  [AMDGPU] Add ptr.s.buffer.load intrinsic, use it from Clang (#209243)

This commit adds a version of the existing s_buffer_load intrinsic that
more accurately models the memory semantics of the s_buffer_load
instruction, namely that it is, in fact, a memory load.

To preserve the existing behavior that the "nomem" s.buffer.load
intrinsic was using, Clang and MLIR add !invariant.load metadata when
constructing the intrinsic (matching documented requirements on
scalarazable buffer loads) and a late codegen pass adds the metadata
just to be safe.

Tests that were "about" s.buffer.load have been copied to create
versions that use the new intrinsic, as was done for the other
*.ptr.buffer.* operations.

Other tests have been upgraded to use the new intrinsic. This has mainly
resulted in minor instruction ordering changes in prologues, if any
change at all. However, CodeGen/AMDGPU/dagcombine-fma-fmad.ll has seen a
v_fma => v_mad pattern fail to match in one case, I don't know if this
is a real regression.

While I was here, the s_buffer_load => buffer_load fallback has been
updated to preserve cachepolity.

AI disclosure: Primarily AI-written code, but I have at least looked at
and tried to find the worst of the silliness.

---------

Co-authored-by: Codex <codex at openai.com>



To unsubscribe from these emails, change your notification settings at https://github.com/llvm/llvm-project/settings/notifications


More information about the All-commits mailing list