[llvm] [Offload][libomptarget] Replace slow `omp_target_memset` implementation by `dataFill` (PR #200202)

Jan André Reuter via llvm-commits llvm-commits at lists.llvm.org
Fri May 29 00:26:57 PDT 2026


https://github.com/Thyre updated https://github.com/llvm/llvm-project/pull/200202

>From 418c51d5bbbd78417045ad0b281fa40dc8d645c1 Mon Sep 17 00:00:00 2001
From: =?UTF-8?q?Jan=20Andr=C3=A9=20Reuter?= <j.reuter at fz-juelich.de>
Date: Thu, 28 May 2026 12:46:53 +0200
Subject: [PATCH] [Offload] Replace slow `omp_target_memset` implementation by
 `dataFill`
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit

`omp_target_memset` was initially implemented before the existance of
`offload`. Because of this, a slow path was chosen to implement
`omp_target_memset`, first allocating memory on the host, calling
`memset` on that memory, and then transferring this to the device.

Aside from the inefficient way of setting device memory, this also
causes a data transfer event for the OpenMP Tools Interface, interfering
with the adding memset event in OpenMP v6.0.

Since offload implements setting data via `dataFill` by now, replace the
slow path by just calling `dataFill` on the RTL instead.
This resolves both the inefficiency, and removes the superfluous event
dispatched to a tool.

Signed-off-by: Jan André Reuter <j.reuter at fz-juelich.de>
---
 offload/libomptarget/OpenMP/API.cpp | 30 +++++++++++------------------
 1 file changed, 11 insertions(+), 19 deletions(-)

diff --git a/offload/libomptarget/OpenMP/API.cpp b/offload/libomptarget/OpenMP/API.cpp
index dc4bccd01dfea..8f46dc30f026a 100644
--- a/offload/libomptarget/OpenMP/API.cpp
+++ b/offload/libomptarget/OpenMP/API.cpp
@@ -482,26 +482,18 @@ EXTERN void *omp_target_memset(void *Ptr, int ByteVal, size_t NumBytes,
         nullptr, const_cast<void *>(Ptr), NumBytes,
         __builtin_return_address(0)));
 
-    // TODO: replace the omp_target_memset() slow path with the fast path.
-    // That will require the ability to execute a kernel from within
-    // libomptarget.so (which we do not have at the moment).
-
-    // This is a very slow path: create a filled array on the host and upload
-    // it to the GPU device.
-    int InitialDevice = omp_get_initial_device();
-    void *Shadow = omp_target_alloc(NumBytes, InitialDevice);
-    if (Shadow) {
-      (void)memset(Shadow, ByteVal, NumBytes);
-      (void)omp_target_memcpy(Ptr, Shadow, NumBytes, 0, 0, DeviceNum,
-                              InitialDevice);
-      (void)omp_target_free(Shadow, InitialDevice);
-    } else {
-      // If the omp_target_alloc has failed, let's just not do anything.
-      // omp_target_memset does not have any good way to fail, so we
-      // simply avoid a catastrophic failure of the process for now.
+    auto DeviceOrErr = PM->getDevice(DeviceNum);
+    if (!DeviceOrErr)
+      FATAL_MESSAGE(DeviceNum, "%s", toString(DeviceOrErr.takeError()).c_str());
+    AsyncInfoTy AsyncInfo(*DeviceOrErr);
+    if (auto Error = DeviceOrErr->RTL->getDevice(DeviceOrErr->RTLDeviceID)
+                         .dataFill(Ptr, &ByteVal, 1, NumBytes, AsyncInfo)) {
       ODBG(ODT_Interface)
-          << __func__
-          << " failed to fill memory due to error with omp_target_alloc";
+          << __func__ << " failed to fill memory due to error with dataFill";
+      // If the dataFill failed, let's just not do anything.
+      // omp_target_memset does not have any good way to fail.
+      // Depending on the RTL implementation, the application will
+      // abort anyway.
     }
   }
 



More information about the llvm-commits mailing list