[clang] [clang][SPIR-V] Add sse/sse2 for x86-64 MSVC hosts (PR #227665)
Tom Honermann via cfe-commits
cfe-commits at lists.llvm.org
Fri Oct 2 14:53:17 PDT 2026
================
@@ -0,0 +1,48 @@
+/// Check that the x86 intrinsic headers parse and compile for a SPIR-V device
+/// when the auxiliary host target is x86-64 MSVC, which is what the sse/sse2
+/// device features derived from that host are for.
+
+// RUN: %clang_cc1 -triple spirv64-unknown-unknown -aux-triple x86_64-pc-windows-msvc \
+// RUN: -fsycl-is-device -ffreestanding -emit-llvm -o - %s | FileCheck %s
+
+/// With sse/sse2 disabled, the always_inline intrinsics cannot be inlined into
+/// device code; this is the failure the derived features prevent.
+/// Codegen stops at the first failing function, so only add() is diagnosed.
+// RUN: %clang_cc1 -triple spirv64-unknown-unknown -aux-triple x86_64-pc-windows-msvc \
+// RUN: -fsycl-is-device -ffreestanding -target-feature -sse -target-feature -sse2 \
+// RUN: -emit-llvm -verify=no-sse -o - %s
+
+/// Intrinsics that lower to x86 builtins are rejected for the device.
+// RUN: %clang_cc1 -triple spirv64-unknown-unknown -aux-triple x86_64-pc-windows-msvc \
+// RUN: -fsycl-is-device -ffreestanding -DX86_BUILTIN -emit-llvm -verify=x86-builtin \
+// RUN: -o - %s
+
+#include <immintrin.h>
+
+// CHECK-LABEL: define {{.*}}spir_func noundef float @_Z3addff
+[[clang::sycl_external]] float add(float x, float y) {
+ // no-sse-error at +1 {{always_inline function '_mm_set1_ps' requires target feature 'sse', but would be inlined into function 'add' that is compiled without support for 'sse'}}
+ __m128 a = _mm_set1_ps(x);
+ // no-sse-error at +1 {{always_inline function '_mm_set1_ps' requires target feature 'sse', but would be inlined into function 'add' that is compiled without support for 'sse'}}
+ __m128 b = _mm_set1_ps(y);
+ // CHECK: fadd <4 x float>
----------------
tahonermann wrote:
Ah, yes, I missed that. `_mm_set1_ps()` is defined in `clang/lib/Headers/xmmintrin.h` as:
```c++
36 #define __DEFAULT_FN_ATTRS \
37 __attribute__((__always_inline__, __nodebug__, __target__("sse"), \
38 __min_vector_width__(128)))
39 #define __DEFAULT_FN_ATTRS_SSE2 \
40 __attribute__((__always_inline__, __nodebug__, __target__("sse2"), \
41 __min_vector_width__(128)))
42
43 #if defined(__cplusplus) && (__cplusplus >= 201103L)
44 #define __DEFAULT_FN_ATTRS_CONSTEXPR __DEFAULT_FN_ATTRS constexpr
45 #define __DEFAULT_FN_ATTRS_SSE2_CONSTEXPR __DEFAULT_FN_ATTRS_SSE2 constexpr
46 #else
47 #define __DEFAULT_FN_ATTRS_CONSTEXPR __DEFAULT_FN_ATTRS
48 #define __DEFAULT_FN_ATTRS_SSE2_CONSTEXPR __DEFAULT_FN_ATTRS_SSE2
49 #endif
..
1914 static __inline__ __m128 __DEFAULT_FN_ATTRS_CONSTEXPR
1915 _mm_set1_ps(float __w) {
1916 return __extension__ (__m128){ __w, __w, __w, __w };
1917 }
```
The body of the function is just a C99 compound literal, so the only basis we have for rejection is the [`__target__` attribute](https://clang.llvm.org/docs/AttributeReference.html#target) attached to the function. Diagnosing the `arch=`, `cpu=`, `tune=`, and `branch-protection=` cases should be straight forward. Diagnosing the subtarget feature cases looks like it will require differentiating which features are enabled for host compatibility vs which ones are supported for the actual target. I don't have a good sense of how difficult that would be to do.
https://github.com/llvm/llvm-project/pull/227665
More information about the cfe-commits
mailing list