[llvm] [Offload][libomptarget] Replace slow `omp_target_memset` implementation by `dataFill` (PR #200202)
Jan André Reuter via llvm-commits
llvm-commits at lists.llvm.org
Fri May 29 00:26:57 PDT 2026
https://github.com/Thyre updated https://github.com/llvm/llvm-project/pull/200202
>From 418c51d5bbbd78417045ad0b281fa40dc8d645c1 Mon Sep 17 00:00:00 2001
From: =?UTF-8?q?Jan=20Andr=C3=A9=20Reuter?= <j.reuter at fz-juelich.de>
Date: Thu, 28 May 2026 12:46:53 +0200
Subject: [PATCH] [Offload] Replace slow `omp_target_memset` implementation by
`dataFill`
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
`omp_target_memset` was initially implemented before the existance of
`offload`. Because of this, a slow path was chosen to implement
`omp_target_memset`, first allocating memory on the host, calling
`memset` on that memory, and then transferring this to the device.
Aside from the inefficient way of setting device memory, this also
causes a data transfer event for the OpenMP Tools Interface, interfering
with the adding memset event in OpenMP v6.0.
Since offload implements setting data via `dataFill` by now, replace the
slow path by just calling `dataFill` on the RTL instead.
This resolves both the inefficiency, and removes the superfluous event
dispatched to a tool.
Signed-off-by: Jan André Reuter <j.reuter at fz-juelich.de>
---
offload/libomptarget/OpenMP/API.cpp | 30 +++++++++++------------------
1 file changed, 11 insertions(+), 19 deletions(-)
diff --git a/offload/libomptarget/OpenMP/API.cpp b/offload/libomptarget/OpenMP/API.cpp
index dc4bccd01dfea..8f46dc30f026a 100644
--- a/offload/libomptarget/OpenMP/API.cpp
+++ b/offload/libomptarget/OpenMP/API.cpp
@@ -482,26 +482,18 @@ EXTERN void *omp_target_memset(void *Ptr, int ByteVal, size_t NumBytes,
nullptr, const_cast<void *>(Ptr), NumBytes,
__builtin_return_address(0)));
- // TODO: replace the omp_target_memset() slow path with the fast path.
- // That will require the ability to execute a kernel from within
- // libomptarget.so (which we do not have at the moment).
-
- // This is a very slow path: create a filled array on the host and upload
- // it to the GPU device.
- int InitialDevice = omp_get_initial_device();
- void *Shadow = omp_target_alloc(NumBytes, InitialDevice);
- if (Shadow) {
- (void)memset(Shadow, ByteVal, NumBytes);
- (void)omp_target_memcpy(Ptr, Shadow, NumBytes, 0, 0, DeviceNum,
- InitialDevice);
- (void)omp_target_free(Shadow, InitialDevice);
- } else {
- // If the omp_target_alloc has failed, let's just not do anything.
- // omp_target_memset does not have any good way to fail, so we
- // simply avoid a catastrophic failure of the process for now.
+ auto DeviceOrErr = PM->getDevice(DeviceNum);
+ if (!DeviceOrErr)
+ FATAL_MESSAGE(DeviceNum, "%s", toString(DeviceOrErr.takeError()).c_str());
+ AsyncInfoTy AsyncInfo(*DeviceOrErr);
+ if (auto Error = DeviceOrErr->RTL->getDevice(DeviceOrErr->RTLDeviceID)
+ .dataFill(Ptr, &ByteVal, 1, NumBytes, AsyncInfo)) {
ODBG(ODT_Interface)
- << __func__
- << " failed to fill memory due to error with omp_target_alloc";
+ << __func__ << " failed to fill memory due to error with dataFill";
+ // If the dataFill failed, let's just not do anything.
+ // omp_target_memset does not have any good way to fail.
+ // Depending on the RTL implementation, the application will
+ // abort anyway.
}
}
More information about the llvm-commits
mailing list