[llvm] [mlir] [LLVM][NVPTX] Add async bulk copy global to shared extensions (PR #222323)
Durgadoss R via llvm-commits
llvm-commits at lists.llvm.org
Fri Sep 11 01:50:08 PDT 2026
================
@@ -1822,10 +1877,75 @@ The '`@llvm.nvvm.cp.async.bulk.global.to.shared.cta`' intrinsic corresponds to
the `cp.async.bulk.shared::cta.global.*` family of PTX instructions. These
instructions initiate an asynchronous copy of bulk data from global memory to
shared::cta memory. The 32-bit operand `%size` specifies the amount of memory
-to be copied and it must be a multiple of 16. The last argument (denoted by
-`i1 %flag_ch`) is a compile-time constant. When set, it indicates a valid
-cache_hint (`i64 %ch`) and generates the `.L2::cache_hint` variant of the
-PTX instruction.
+to be copied and it must be a multiple of 16.
+
+- The trailing `%flag_ch`, `%flag_oob`, and `%validate_pattern` arguments
+ control the cache-hint modifier, the out-of-bounds handling, and the
+ data-validity reporting pattern. They must be compile-time constants.
+- The argument denoted by `i1 %flag_ch`, when set, indicates a valid cache_hint
+ (`i64 %ch`) and generates the `.L2::cache_hint` variant of the PTX
+ instruction.
+- The argument denoted by `i1 %flag_oob`, when set, generates the
+ `.ignore_oob` variant of the PTX instruction. In that case, the `i32 %ibl`
+ and `i32 %ibr` operands specify the numbers of bytes to ignore on the left
+ and right sides of the source, respectively. Both operands must be in the
+ range \[0, 15]; otherwise, the behavior is undefined. The `.ignore_oob`
----------------
durga4github wrote:
I looked at this range and assumed that the ignore bytes are immediates!
https://github.com/llvm/llvm-project/pull/222323
More information about the llvm-commits
mailing list