[all-commits] [llvm/llvm-project] f4a8eb: [offload] add support for aligned allocations (#20...
EuphoricThinking via All-commits
all-commits at lists.llvm.org
Wed Jun 24 10:27:23 PDT 2026
Branch: refs/heads/main
Home: https://github.com/llvm/llvm-project
Commit: f4a8eb0a6ea0dd268a23aa08cf0384c6f8542172
https://github.com/llvm/llvm-project/commit/f4a8eb0a6ea0dd268a23aa08cf0384c6f8542172
Author: EuphoricThinking <agata.momot at intel.com>
Date: 2026-06-24 (Wed, 24 Jun 2026)
Changed paths:
M offload/liboffload/API/Memory.td
M offload/liboffload/src/OffloadImpl.cpp
M offload/plugins-nextgen/amdgpu/src/rtl.cpp
M offload/plugins-nextgen/common/include/MemoryManager.h
M offload/plugins-nextgen/common/include/PluginInterface.h
M offload/plugins-nextgen/common/src/PluginInterface.cpp
M offload/plugins-nextgen/cuda/src/rtl.cpp
M offload/plugins-nextgen/host/src/rtl.cpp
M offload/plugins-nextgen/level_zero/include/L0Device.h
M offload/plugins-nextgen/level_zero/src/L0Device.cpp
M offload/unittests/OffloadAPI/CMakeLists.txt
A offload/unittests/OffloadAPI/memory/olMemAllocAligned.cpp
Log Message:
-----------
[offload] add support for aligned allocations (#203353)
This patch is the first step towards introducing alignment support in
memory allocations using liboffload, in order to enable SYCL
implementation of aligned allocations.
At the level of device allocators, it does not modify the Level Zero
code except for forwarding the alignment parameter, since Level Zero
already allows for specifying the alignment in its device allocator
implementation. For AMD and CUDA, it checks whether the alignment passed
by the caller is supported by the given backend; the reasoning behind
this verification is described in the following paragraphs. At the API
level, it adds a new function olMemAllocAligned, which is expected to
work similarly to olMemAlloc, with the difference that the buffers
returned by olMemAllocAligned should be aligned to the alignment passed
by the user. At the level of the plugin interface internal abstractions,
it adds a new argument Alignment to existing functions and delegates
memory allocation between olMemAllocAligned and olMemAlloc
implementations by using a common helper function.
The goal of the anticipated series of patches is to implement handling
of the alignment in the memory manager at the plugin interface level. At
the first stage, presented in this patch, the information about the
passed alignment is used mainly for checking whether the buffer returned
by the device allocators meets the requirements. In the case of the
requested memory size exceeding the thresholds of allocations handled by
the memory manager, the request is forwarded directly to the device.
Otherwise, the memory manager is responsible for allocating memory in
full pages and pooling it according to the requested chunks. In the
first scenario, the requested size, which is greater than the
aforementioned threshold, is usually a multiple of the page size.
Therefore, any alignment smaller than the page size would be correct.
Neither CUDA nor HSA provides users with the ability to specify the
alignment of the allocated memory. Their APIs include the alignment as
one of the possible arguments only in functions that reserve virtual
address space. CUDA enables users only to check the allocation
granularity, which is usually synonymous with the page size. In the case
of HSA, if the memory is allocated using a pool, the user is able to
check not only the granularity, but also the alignment of the buffers
allocated using the given pool. However, these values - granularity and
alignment - are defined only if
HSA_AMD_MEMORY_POOL_INFO_RUNTIME_ALLOC_ALLOWED is set to true, which
should always be set since the current implementation of the memory pool
for the AMD plugin in liboffload uses only hsa_amd_memory_pool_ts for
allocations. Only Level Zero accepts the alignment parameter during the
memory allocation process, but manages returned pointers internally.
Since pooling is also implemented in the memory manager at the plugin
level, such a design in the Level Zero device allocator duplicates
pooling in the whole application.
In the final version, the memory manager at the plugin interface level
will handle memory pooling and the buffer alignment from different
devices instead of delegating memory management to device allocator
implementations, as this would simplify the Level Zero plugin design and
provide AMD and CUDA plugins with support for the alignment in memory
allocations, which is not natively included in their APIs.
To unsubscribe from these emails, change your notification settings at https://github.com/llvm/llvm-project/settings/notifications
More information about the All-commits
mailing list