[llvm] [AArch64] Match scalar_to_vector of frozen extended loads (PR #224213)
Cyrus Ding via llvm-commits
llvm-commits at lists.llvm.org
Thu Sep 17 00:33:49 PDT 2026
https://github.com/dingcyrus created https://github.com/llvm/llvm-project/pull/224213
DAGCombiner sinks freeze below SCALAR_TO_VECTOR (freeze(stv(x)) folds to stv(freeze(x)) when only element zero is demanded, and visitFREEZE pushes freeze through single-use poison-propagating ops). A freeze of a load can never be folded away, because the loaded value may be poison in memory, so scalar_to_vector(freeze(extload)) is a common shape reaching ISel.
The ExtLoad8_16AllModes / ExtLoad8_16_32AllModes scalar_to_vector patterns only match plain zextload/extload scalars, so the frozen shape fell back to a GPR load plus fmov instead of the direct ldr b/h/s SIMD-register forms.
Add PatFrags that additionally match a freeze of the extended load and use them in the scalar_to_vector pattern instantiations. Matching through the freeze is safe: the LDRB/LDRH/LDRS forms write a fully defined value into the destination register, which is exactly what freeze promises, and the load stays visible so the memory operand is attached to the selected instruction.
Fixes #224181
>From a76864df680378270c860c0c877eb02578bcdae0 Mon Sep 17 00:00:00 2001
From: Cyrus Ding <785101675 at qq.com>
Date: Thu, 17 Sep 2026 15:08:42 +0800
Subject: [PATCH] [AArch64] Match scalar_to_vector of frozen extended loads
DAGCombiner sinks freeze below SCALAR_TO_VECTOR (freeze(stv(x)) folds
to stv(freeze(x)) when only element zero is demanded, and visitFREEZE
pushes freeze through single-use poison-propagating ops). A freeze of
a load can never be folded away, because the loaded value may be
poison in memory, so scalar_to_vector(freeze(extload)) is a common
shape reaching ISel.
The ExtLoad8_16AllModes / ExtLoad8_16_32AllModes scalar_to_vector
patterns only match plain zextload/extload scalars, so the frozen
shape fell back to a GPR load plus fmov instead of the direct
ldr b/h/s SIMD-register forms.
Add PatFrags that additionally match a freeze of the extended load and
use them in the scalar_to_vector pattern instantiations. Matching
through the freeze is safe: the LDRB/LDRH/LDRS forms write a fully
defined value into the destination register, which is exactly what
freeze promises, and the load stays visible so the memory operand is
attached to the selected instruction.
Fixes #224181
---
llvm/lib/Target/AArch64/AArch64InstrInfo.td | 44 ++++++--
.../AArch64/scalar-to-vector-frozen-load.ll | 104 ++++++++++++++++++
2 files changed, 139 insertions(+), 9 deletions(-)
create mode 100644 llvm/test/CodeGen/AArch64/scalar-to-vector-frozen-load.ll
diff --git a/llvm/lib/Target/AArch64/AArch64InstrInfo.td b/llvm/lib/Target/AArch64/AArch64InstrInfo.td
index f09f9c2a2f61e..6a1e37dea5190 100644
--- a/llvm/lib/Target/AArch64/AArch64InstrInfo.td
+++ b/llvm/lib/Target/AArch64/AArch64InstrInfo.td
@@ -4691,20 +4691,46 @@ multiclass ExtLoad8_16_32AllModes<ValueType OutTy, ValueType InnerTy,
(SUBREG_TO_REG (LDRSroX GPR64sp:$Rn, GPR64:$Rm, ro32.Xext:$extend), ssub)>;
}
+// Load fragments that additionally match a freeze of the load.
+//
+// DAGCombiner sinks freeze below SCALAR_TO_VECTOR (freeze(stv(x)) is folded to
+// stv(freeze(x)) when only element zero is demanded, and visitFREEZE pushes
+// freeze through single-use poison-propagating ops). A freeze of a load can
+// never be folded away, because the loaded value may be poison in memory, so
+// scalar_to_vector(freeze(extload)) is a common shape reaching ISel. Matching
+// it to the LDRB/LDRH/LDRS forms is safe: those instructions write a fully
+// defined value into the destination register, which is exactly what freeze
+// promises, and matching through the freeze keeps the load visible so the
+// memory operand is attached to the selected instruction.
+def zextloadi8_or_frz : PatFrags<(ops node:$ptr),
+ [(zextloadi8 node:$ptr), (freeze (zextloadi8 node:$ptr))]>;
+def zextloadi16_or_frz : PatFrags<(ops node:$ptr),
+ [(zextloadi16 node:$ptr), (freeze (zextloadi16 node:$ptr))]>;
+def zextloadi32_or_frz : PatFrags<(ops node:$ptr),
+ [(zextloadi32 node:$ptr), (freeze (zextloadi32 node:$ptr))]>;
+def extloadi8_or_frz : PatFrags<(ops node:$ptr),
+ [(extloadi8 node:$ptr), (freeze (extloadi8 node:$ptr))]>;
+def extloadi16_or_frz : PatFrags<(ops node:$ptr),
+ [(extloadi16 node:$ptr), (freeze (extloadi16 node:$ptr))]>;
+def extloadi32_or_frz : PatFrags<(ops node:$ptr),
+ [(extloadi32 node:$ptr), (freeze (extloadi32 node:$ptr))]>;
+
// Instantiate bitconvert patterns for floating-point types.
defm : ExtLoad8_16AllModes<f32, i32, bitconvert, zextloadi8, zextloadi16>;
defm : ExtLoad8_16_32AllModes<f64, i64, bitconvert, zextloadi8, zextloadi16, zextloadi32>;
-// Instantiate scalar_to_vector patterns for all vector types.
-defm : ExtLoad8_16AllModes<v16i8, i32, scalar_to_vector, zextloadi8, zextloadi16>;
-defm : ExtLoad8_16AllModes<v16i8, i32, scalar_to_vector, extloadi8, extloadi16>;
-defm : ExtLoad8_16AllModes<v8i16, i32, scalar_to_vector, zextloadi8, zextloadi16>;
-defm : ExtLoad8_16AllModes<v8i16, i32, scalar_to_vector, extloadi8, extloadi16>;
-defm : ExtLoad8_16AllModes<v4i32, i32, scalar_to_vector_v4f32, zextloadi8, zextloadi16>;
-defm : ExtLoad8_16AllModes<v4i32, i32, scalar_to_vector_v4f32, extloadi8, extloadi16>;
-defm : ExtLoad8_16_32AllModes<v2i64, i64, scalar_to_vector_v2f64, zextloadi8, zextloadi16, zextloadi32>;
-defm : ExtLoad8_16_32AllModes<v2i64, i64, scalar_to_vector_v2f64, extloadi8, extloadi16, extloadi32>;
+// Instantiate scalar_to_vector patterns for all vector types. The frozen-load
+// variants keep the direct SIMD-register load forms selected when a freeze was
+// sunk onto the scalar operand (see *_or_frz above).
+defm : ExtLoad8_16AllModes<v16i8, i32, scalar_to_vector, zextloadi8_or_frz, zextloadi16_or_frz>;
+defm : ExtLoad8_16AllModes<v16i8, i32, scalar_to_vector, extloadi8_or_frz, extloadi16_or_frz>;
+defm : ExtLoad8_16AllModes<v8i16, i32, scalar_to_vector, zextloadi8_or_frz, zextloadi16_or_frz>;
+defm : ExtLoad8_16AllModes<v8i16, i32, scalar_to_vector, extloadi8_or_frz, extloadi16_or_frz>;
+defm : ExtLoad8_16AllModes<v4i32, i32, scalar_to_vector_v4f32, zextloadi8_or_frz, zextloadi16_or_frz>;
+defm : ExtLoad8_16AllModes<v4i32, i32, scalar_to_vector_v4f32, extloadi8_or_frz, extloadi16_or_frz>;
+defm : ExtLoad8_16_32AllModes<v2i64, i64, scalar_to_vector_v2f64, zextloadi8_or_frz, zextloadi16_or_frz, zextloadi32_or_frz>;
+defm : ExtLoad8_16_32AllModes<v2i64, i64, scalar_to_vector_v2f64, extloadi8_or_frz, extloadi16_or_frz, extloadi32_or_frz>;
// Pre-fetch.
defm PRFUM : PrefetchUnscaled<0b11, 0, 0b10, "prfum",
diff --git a/llvm/test/CodeGen/AArch64/scalar-to-vector-frozen-load.ll b/llvm/test/CodeGen/AArch64/scalar-to-vector-frozen-load.ll
new file mode 100644
index 0000000000000..f4a8404a13502
--- /dev/null
+++ b/llvm/test/CodeGen/AArch64/scalar-to-vector-frozen-load.ll
@@ -0,0 +1,104 @@
+; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py UTC_ARGS: --version 6
+; RUN: llc -mtriple=aarch64-linux-gnu < %s | FileCheck %s
+
+; DAGCombiner sinks freeze below SCALAR_TO_VECTOR (and visitFREEZE pushes it
+; through single-use poison-propagating ops), and freeze(load) can never be
+; folded away because the loaded value may be poison in memory. The resulting
+; scalar_to_vector(freeze(extload)) shape must still select the direct
+; SIMD-register load forms (ldr b/h) instead of a GPR load plus fmov.
+
+define <16 x i8> @stv_frozen_load_i8_base(ptr %p) {
+; CHECK-LABEL: stv_frozen_load_i8_base:
+; CHECK: // %bb.0:
+; CHECK-NEXT: ldr b0, [x0]
+; CHECK-NEXT: ret
+ %v = load i8, ptr %p
+ %f = freeze i8 %v
+ %s = insertelement <16 x i8> poison, i8 %f, i32 0
+ ret <16 x i8> %s
+}
+
+define <16 x i8> @stv_frozen_load_i8_scaled_offset(ptr %p) {
+; CHECK-LABEL: stv_frozen_load_i8_scaled_offset:
+; CHECK: // %bb.0:
+; CHECK-NEXT: ldr b0, [x0, #5]
+; CHECK-NEXT: ret
+ %q = getelementptr i8, ptr %p, i64 5
+ %v = load i8, ptr %q
+ %f = freeze i8 %v
+ %s = insertelement <16 x i8> poison, i8 %f, i32 0
+ ret <16 x i8> %s
+}
+
+define <16 x i8> @stv_frozen_load_i8_unscaled_offset(ptr %p) {
+; CHECK-LABEL: stv_frozen_load_i8_unscaled_offset:
+; CHECK: // %bb.0:
+; CHECK-NEXT: ldur b0, [x0, #-3]
+; CHECK-NEXT: ret
+ %q = getelementptr i8, ptr %p, i64 -3
+ %v = load i8, ptr %q
+ %f = freeze i8 %v
+ %s = insertelement <16 x i8> poison, i8 %f, i32 0
+ ret <16 x i8> %s
+}
+
+define <16 x i8> @stv_frozen_load_i8_reg_offset_x(ptr %p, i64 %i) {
+; CHECK-LABEL: stv_frozen_load_i8_reg_offset_x:
+; CHECK: // %bb.0:
+; CHECK-NEXT: ldr b0, [x0, x1]
+; CHECK-NEXT: ret
+ %q = getelementptr i8, ptr %p, i64 %i
+ %v = load i8, ptr %q
+ %f = freeze i8 %v
+ %s = insertelement <16 x i8> poison, i8 %f, i32 0
+ ret <16 x i8> %s
+}
+
+define <16 x i8> @stv_frozen_load_i8_reg_offset_w(ptr %p, i32 %i) {
+; CHECK-LABEL: stv_frozen_load_i8_reg_offset_w:
+; CHECK: // %bb.0:
+; CHECK-NEXT: ldr b0, [x0, w1, sxtw]
+; CHECK-NEXT: ret
+ %idx = sext i32 %i to i64
+ %q = getelementptr i8, ptr %p, i64 %idx
+ %v = load i8, ptr %q
+ %f = freeze i8 %v
+ %s = insertelement <16 x i8> poison, i8 %f, i32 0
+ ret <16 x i8> %s
+}
+
+define <8 x i16> @stv_frozen_load_i16_base(ptr %p) {
+; CHECK-LABEL: stv_frozen_load_i16_base:
+; CHECK: // %bb.0:
+; CHECK-NEXT: ldr h0, [x0]
+; CHECK-NEXT: ret
+ %v = load i16, ptr %p
+ %f = freeze i16 %v
+ %s = insertelement <8 x i16> poison, i16 %f, i32 0
+ ret <8 x i16> %s
+}
+
+define <8 x i16> @stv_frozen_load_i16_scaled_offset(ptr %p) {
+; CHECK-LABEL: stv_frozen_load_i16_scaled_offset:
+; CHECK: // %bb.0:
+; CHECK-NEXT: ldr h0, [x0, #10]
+; CHECK-NEXT: ret
+ %q = getelementptr i8, ptr %p, i64 10
+ %v = load i16, ptr %q
+ %f = freeze i16 %v
+ %s = insertelement <8 x i16> poison, i16 %f, i32 0
+ ret <8 x i16> %s
+}
+
+define <16 x i8> @stv_noundef_frozen_load_i8(ptr %p) {
+; CHECK-LABEL: stv_noundef_frozen_load_i8:
+; CHECK: // %bb.0:
+; CHECK-NEXT: ldr b0, [x0]
+; CHECK-NEXT: ret
+ %v = load i8, ptr %p
+ %z = zext i8 %v to i32
+ %f = freeze i32 %z
+ %t = trunc i32 %f to i8
+ %s = insertelement <16 x i8> poison, i8 %t, i32 0
+ ret <16 x i8> %s
+}
More information about the llvm-commits
mailing list