[llvm] 9247e89 - [IR] Add llvm.structured.gep instruction (#176145)
via llvm-commits
llvm-commits at lists.llvm.org
Wed Jan 21 07:45:13 PST 2026
Author: Nathan Gauër
Date: 2026-01-21T15:45:08Z
New Revision: 9247e897060846154f17e515994de0de5eb5c830
URL: https://github.com/llvm/llvm-project/commit/9247e897060846154f17e515994de0de5eb5c830
DIFF: https://github.com/llvm/llvm-project/commit/9247e897060846154f17e515994de0de5eb5c830.diff
LOG: [IR] Add llvm.structured.gep instruction (#176145)
This commit adds initial support for `@llvm.structured.gep` instruction
in Clang. This intrinsic is supposed to be used as an alternative to
ptrdiff/GEP when pointers arithmetic is invalid and only structured
access is possible.
Link to the RFC:
https://discourse.llvm.org/t/rfc-adding-instructions-to-to-carry-gep-type-traversal-information/
Previous discussion around the documentation:
https://github.com/llvm/llvm-project/pull/167883
Added:
llvm/test/Verifier/structured-gep-indices-bad.ll
llvm/test/Verifier/structured-gep-indices.ll
Modified:
llvm/docs/LangRef.rst
llvm/include/llvm/IR/IRBuilder.h
llvm/include/llvm/IR/IntrinsicInst.h
llvm/include/llvm/IR/Intrinsics.td
llvm/lib/IR/Verifier.cpp
Removed:
################################################################################
diff --git a/llvm/docs/LangRef.rst b/llvm/docs/LangRef.rst
index 811a878bb92df..8633453ec7928 100644
--- a/llvm/docs/LangRef.rst
+++ b/llvm/docs/LangRef.rst
@@ -14990,6 +14990,155 @@ Semantics:
See the description for :ref:`llvm.stacksave <int_stacksave>`.
+.. _i_structured_gep:
+
+'``llvm.structured.gep``' Intrinsic
+^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
+
+Syntax:
+"""""""
+
+::
+
+ declare <ret_type>
+ @llvm.structured.gep(ptr elementtype(<basetype>) <source>
+ {, [i32/i64] <index> }*)
+
+Overview:
+"""""""""
+
+The '``llvm.structured.gep``' intrinsic (structured
+**G**\ et\ **E**\ lement\ **P**\ tr) computes a new pointer address resulting
+from a logical indexing into the ``<source>`` pointer. The returned address
+depends on the indices and may depend on the layout of ``<basetype>``
+at runtime.
+
+Arguments:
+""""""""""
+
+``ptr elementtype(<basetype>) <source>``:
+A pointer to the memory location used as base for the address computation.
+
+The ``source`` argument must be annotated with an :ref:`elementtype
+<attr_elementtype>` attribute at the call-site. This attribute specifies the
+type of the element pointed to by the pointer source. This type will be
+used along with the provided indices and source operand to compute a new
+pointer representing the result of a logical indexing into the basetype
+pointed by source.
+
+The ``basetype`` is only associated with ``<source>`` for this particular
+call. A frontend could possibly emit multiple structured
+GEP with the same source pointer but a
diff erent ``basetype``.
+
+``[i32/i64] index, ...``:
+Indices used to traverse into the ``basetype`` and compute a pointer to the
+target element. Indices can be 32-bit or 64-bit unsigned integers. Indices being
+handled one by one, both sizes can be mixed in the same instruction. The
+precision used to compute the resulting pointer is target-dependent.
+When used to index into a struct, only integer constants are allowed.
+
+Semantics:
+""""""""""
+
+The ``llvm.structured.gep`` performs a logical traversal of the type
+``basetype`` using the list of provided indices, computing the pointer
+addressing the targeted element/field assuming ``source`` points to a
+physically laid out ``basetype``. The physical layout of the source depends
+on the target and does not necessarily match the one described by the
+datalayout.
+
+The first index determines which element/field of ``basetype`` is selected,
+computes the pointer to access this element/field assuming ``source`` points
+to the start of ``basetype``.
+This pointer becomes the new ``source``, the current type the new
+``basetype``, and the next indices is consumed until a scalar type is
+reached or all indices are consumed.
+
+All indices must be consumed, and it is illegal to index into a scalar type.
+Meaning the maximum number of indices depends on the depth of the basetype.
+
+Because this instruction performs a logical addressing, all indices are
+assumed to be inbounds. This means it is not possible to access the next
+element in the logical layout by overflowing:
+
+- If the indexed type is a struct with N fields, the index must be an
+ immediate/constant value in the range ``[0; N[``.
+- If indexing into an array or vector, the index can be a variable, but
+ is assumed to be inbounds with regards to the current basetype logical layout.
+- If the traversed type is an array or vector of N elements with ``N > 0``,
+ the index is assumed to belong to ``[0; N[``.
+- If the traversed type is an array of size ``0``, the array size is assumed
+ to be known at runtime, and the instruction assumes the index is always
+ inbounds.
+
+In all cases **except** when the accessed type is a 0-sized array, indexing
+out of bounds yields `poison`. When the index value is unknown, optimizations
+can use the type bounds to determine the range of values the index can have.
+If the source pointer is poison, the instruction returns poison.
+The resulting pointer belongs to the same address space as ``source``.
+This instruction does not dereference the pointer.
+
+Example:
+""""""""
+
+**Simple case: logical access of a struct field**
+
+.. code-block:: cpp
+
+ struct A { int a, int b, int c, int d };
+ int val = my_struct->b;
+
+Could be translated to:
+
+.. code-block:: llvm
+
+ %A = type { i32, i32, i32, i32 }
+ %src = call ptr @llvm.structured.gep(ptr elementtype(%A) %my_struct, i32 1)
+ %val = load i32, ptr %src
+
+**A more complex case**
+
+This instruction can also be used on the same pointer with
diff erent
+basetypes, as long as codegen knows how those are physically laid out.
+Let’s consider the following code:
+
+.. code-block:: cpp
+
+ struct S {
+ uint a;
+ uint b;
+ uint c;
+ uint d;
+ }
+
+ int val = my_struct->b;
+
+
+In this example, the frontend doesn't know the exact physical layout, but
+knows those logical layouts are lowered to the same physical layout:
+
+ - `{ i32, i32, i32, i32 }`
+ - `[ i32 x 4 ]`
+
+This means is is valid to lower the following code to either:
+
+.. code-block:: llvm
+
+ %S = type { i32, i32, i32, i32 }
+ %src = call ptr @llvm.structured.gep(ptr elementtype(%S) %my_struct, i32 1)
+ load i32, ptr %src
+
+Or:
+
+.. code-block:: llvm
+
+ %src = call ptr @llvm.structured.gep(ptr elementtype([ 4 x i32 ]) %my_struct, i32 1)
+ load i32, ptr %src
+
+This is, however, dependent on context that codegen has an insight on. The
+fact that `[ i32 x 4 ]` and `%S` are equivalent depends on the target.
+
+
.. _int_get_dynamic_area_offset:
'``llvm.get.dynamic.area.offset``' Intrinsic
diff --git a/llvm/include/llvm/IR/IRBuilder.h b/llvm/include/llvm/IR/IRBuilder.h
index bc4e909269b80..a17fc1348f0b9 100644
--- a/llvm/include/llvm/IR/IRBuilder.h
+++ b/llvm/include/llvm/IR/IRBuilder.h
@@ -1923,6 +1923,20 @@ class IRBuilderBase {
return Insert(new AtomicRMWInst(Op, Ptr, Val, *Align, Ordering, SSID));
}
+ CallInst *CreateStructuredGEP(Type *BaseType, Value *PtrBase,
+ ArrayRef<Value *> Indices,
+ const Twine &Name = "") {
+ SmallVector<Value *> Args;
+ Args.push_back(PtrBase);
+ llvm::append_range(Args, Indices);
+
+ CallInst *Output = CreateIntrinsic(Intrinsic::structured_gep,
+ {PtrBase->getType()}, Args, {}, Name);
+ Output->addParamAttr(
+ 0, Attribute::get(getContext(), Attribute::ElementType, BaseType));
+ return Output;
+ }
+
Value *CreateGEP(Type *Ty, Value *Ptr, ArrayRef<Value *> IdxList,
const Twine &Name = "",
GEPNoWrapFlags NW = GEPNoWrapFlags::none()) {
diff --git a/llvm/include/llvm/IR/IntrinsicInst.h b/llvm/include/llvm/IR/IntrinsicInst.h
index 0b25baa465a71..ba42601085448 100644
--- a/llvm/include/llvm/IR/IntrinsicInst.h
+++ b/llvm/include/llvm/IR/IntrinsicInst.h
@@ -1804,6 +1804,48 @@ class ConvergenceControlInst : public IntrinsicInst {
CreateLoop(BasicBlock &BB, ConvergenceControlInst *Parent);
};
+class StructuredGEPInst : public IntrinsicInst {
+public:
+ static bool classof(const IntrinsicInst *I) {
+ return I->getIntrinsicID() == Intrinsic::structured_gep;
+ }
+
+ static bool classof(const Value *V) {
+ return isa<IntrinsicInst>(V) && classof(cast<IntrinsicInst>(V));
+ }
+
+ Type *getBaseType() const {
+ return getParamAttr(0, Attribute::ElementType).getValueAsType();
+ }
+
+ unsigned getNumIndices() const { return arg_size() - 1; }
+
+ Value *getIndexOperand(size_t Index) const {
+ assert(Index < getNumIndices());
+ return getOperand(Index + 1);
+ }
+
+ Type *getResultElementType() const {
+ Type *CurrentType = getBaseType();
+ for (unsigned I = 0; I < getNumIndices(); I++) {
+ if (ArrayType *AT = dyn_cast<ArrayType>(CurrentType)) {
+ CurrentType = AT->getElementType();
+ } else if (VectorType *VT = dyn_cast<VectorType>(CurrentType)) {
+ CurrentType = VT->getElementType();
+ } else if (StructType *ST = dyn_cast<StructType>(CurrentType)) {
+ ConstantInt *CI = cast<ConstantInt>(getIndexOperand(I));
+ CurrentType = ST->getElementType(CI->getZExtValue());
+ } else {
+ // FIXME(Keenuts): add testing reaching those places once initial
+ // implementation has landed.
+ llvm_unreachable("unimplemented");
+ }
+ }
+
+ return CurrentType;
+ }
+};
+
} // end namespace llvm
#endif // LLVM_IR_INTRINSICINST_H
diff --git a/llvm/include/llvm/IR/Intrinsics.td b/llvm/include/llvm/IR/Intrinsics.td
index 3bf53ed08380c..ea6bd59c5aeca 100644
--- a/llvm/include/llvm/IR/Intrinsics.td
+++ b/llvm/include/llvm/IR/Intrinsics.td
@@ -1029,6 +1029,11 @@ def int_call_preallocated_teardown : DefaultAttrsIntrinsic<[], [llvm_token_ty]>;
def int_callbr_landingpad : Intrinsic<[llvm_any_ty], [LLVMMatchType<0>],
[IntrNoMerge]>;
+def int_structured_gep
+ : DefaultAttrsIntrinsic<[llvm_anyptr_ty],
+ [LLVMMatchType<0>, llvm_vararg_ty],
+ [IntrNoMem, IntrSpeculatable]>;
+
//===------------------- Standard C Library Intrinsics --------------------===//
//
diff --git a/llvm/lib/IR/Verifier.cpp b/llvm/lib/IR/Verifier.cpp
index 93bf9ea2f0db1..b7cfb4ddbd54c 100644
--- a/llvm/lib/IR/Verifier.cpp
+++ b/llvm/lib/IR/Verifier.cpp
@@ -6868,6 +6868,38 @@ void Verifier::visitIntrinsicCall(Intrinsic::ID ID, CallBase &Call) {
&Call);
break;
}
+ case Intrinsic::structured_gep: {
+ // Parser should refuse those 2 cases.
+ assert(Call.arg_size() >= 1);
+ assert(Call.getOperand(0)->getType()->isPointerTy());
+
+ Check(Call.paramHasAttr(0, Attribute::ElementType),
+ "Intrinsic first parameter is missing an ElementType attribute",
+ &Call);
+
+ Type *T = Call.getParamAttr(0, Attribute::ElementType).getValueAsType();
+ for (unsigned I = 1; I < Call.arg_size(); ++I) {
+ Value *Index = Call.getOperand(I);
+ ConstantInt *CI = dyn_cast<ConstantInt>(Index);
+ Check(Index->getType()->isIntegerTy(),
+ "Index operand type must be an integer", &Call);
+
+ if (ArrayType *AT = dyn_cast<ArrayType>(T)) {
+ T = AT->getElementType();
+ } else if (StructType *ST = dyn_cast<StructType>(T)) {
+ Check(CI, "Indexing into a struct requires a constant int", &Call);
+ Check(CI->getZExtValue() < ST->getNumElements(),
+ "Indexing in a struct should be inbounds", &Call);
+ T = ST->getElementType(CI->getZExtValue());
+ } else if (VectorType *VT = dyn_cast<VectorType>(T)) {
+ T = VT->getElementType();
+ } else {
+ CheckFailed("Reached a non-composite type with more indices to process",
+ &Call);
+ }
+ }
+ break;
+ }
case Intrinsic::amdgcn_cs_chain: {
auto CallerCC = Call.getCaller()->getCallingConv();
switch (CallerCC) {
diff --git a/llvm/test/Verifier/structured-gep-indices-bad.ll b/llvm/test/Verifier/structured-gep-indices-bad.ll
new file mode 100644
index 0000000000000..5cb7c09406496
--- /dev/null
+++ b/llvm/test/Verifier/structured-gep-indices-bad.ll
@@ -0,0 +1,38 @@
+; RUN: not llvm-as -disable-output %s 2>&1 | FileCheck %s
+
+%S = type { i32, i32 }
+
+define void @too_many_indices(ptr %src) {
+entry:
+; CHECK: Reached a non-composite type with more indices to process
+ %ptr = call ptr (ptr, ...) @llvm.structured.gep.p0(ptr elementtype(%S) %src, i32 0, i32 0)
+ ret void
+}
+
+define void @out_of_bounds_struct_access(ptr %src) {
+entry:
+; CHECK: Indexing in a struct should be inbounds
+ %ptr = call ptr (ptr, ...) @llvm.structured.gep.p0(ptr elementtype(%S) %src, i32 2)
+ ret void
+}
+
+define void @dynamic_index_struct(ptr %src, i32 %index) {
+entry:
+; CHECK: Indexing into a struct requires a constant int
+ %ptr = call ptr (ptr, ...) @llvm.structured.gep.p0(ptr elementtype(%S) %src, i32 %index)
+ ret void
+}
+
+define void @non_integer_operand(ptr %src) {
+entry:
+; CHECK: Index operand type must be an integer
+ %ptr = call ptr (ptr, ...) @llvm.structured.gep.p0(ptr elementtype([ 2 x i32 ]) %src, float 1.0)
+ ret void
+}
+
+define void @missing_attribute(ptr %src) {
+entry:
+; CHECK: Intrinsic first parameter is missing an ElementType attribute
+ %ptr = call ptr (ptr, ...) @llvm.structured.gep.p0(ptr %src, i32 0)
+ ret void
+}
diff --git a/llvm/test/Verifier/structured-gep-indices.ll b/llvm/test/Verifier/structured-gep-indices.ll
new file mode 100644
index 0000000000000..0e914f77a875a
--- /dev/null
+++ b/llvm/test/Verifier/structured-gep-indices.ll
@@ -0,0 +1,57 @@
+; RUN: llvm-as -disable-output %s
+
+%S = type { i32, i32 }
+
+define void @runtime_array_nested_access(ptr %src, i32 %index) {
+entry:
+ %ptr = call ptr (ptr, ...) @llvm.structured.gep.p0(ptr elementtype([0 x %S]) %src, i32 %index, i32 1)
+ ret void
+}
+
+define void @normal_array_access(ptr %src, i32 %index) {
+entry:
+ %ptr = call ptr (ptr, ...) @llvm.structured.gep.p0(ptr elementtype([2 x %S]) %src, i32 %index, i32 1)
+ ret void
+}
+
+define void @normal_array_constant_index(ptr %src) {
+entry:
+ %ptr = call ptr (ptr, ...) @llvm.structured.gep.p0(ptr elementtype([2 x %S]) %src, i32 1, i32 1)
+ ret void
+}
+
+define void @struct_access(ptr %src) {
+entry:
+ %ptr = call ptr (ptr, ...) @llvm.structured.gep.p0(ptr elementtype(%S) %src, i32 0)
+ ret void
+}
+
+define void @nested_array_access(ptr %src) {
+entry:
+ %ptr = call ptr (ptr, ...) @llvm.structured.gep.p0(ptr elementtype([ 3 x [ 2 x i32 ] ]) %src, i32 2, i32 1)
+ ret void
+}
+
+define void @runtime_array_index(ptr %src) {
+entry:
+ %ptr = call ptr (ptr, ...) @llvm.structured.gep.p0(ptr elementtype([ 0 x i32 ]) %src, i32 1)
+ ret void
+}
+
+define void @scalar_with_no_index(ptr %src) {
+entry:
+ %ptr = call ptr (ptr, ...) @llvm.structured.gep.p0(ptr elementtype(i32) %src)
+ ret void
+}
+
+define void @access_64bit(ptr %src) {
+entry:
+ %ptr = call ptr (ptr, ...) @llvm.structured.gep.p0(ptr elementtype([ 0 x i32 ]) %src, i64 1)
+ ret void
+}
+
+define void @access_8bit(ptr %src) {
+entry:
+ %ptr = call ptr (ptr, ...) @llvm.structured.gep.p0(ptr elementtype([ 0 x i32 ]) %src, i8 1)
+ ret void
+}
More information about the llvm-commits
mailing list