[llvm] [X86][CostModel] Cost scalar integer divide/remainder by a constant (PR #211529)
via llvm-commits
llvm-commits at lists.llvm.org
Thu Jul 23 04:54:33 PDT 2026
llvmorg-github-actions[bot] wrote:
<!--LLVM PR SUMMARY COMMENT-->
@llvm/pr-subscribers-backend-x86
Author: Rito Takeuchi (Licht-T)
<details>
<summary>Changes</summary>
Follow-up to #<!-- -->208491.
### What this does
1. In `getArithmeticInstrCost`, before the vector paths, price a scalar integer div/rem by a constant as its real sequence: **5** for `UDIV`/`SDIV`, **6** for `UREM`/`SREM` (the extra multiply-back and subtract), **+2** for the `CodeSize`/`SizeAndLatency` cost kinds. Power-of-two divisors lower to a shift and are left to the generic handling.
2. With the scalar lane priced correctly, the `9` workaround is no longer needed, so **restore the honest `vXi64` divide/remainder vector costs** that #<!-- -->208491 lowered: uniform `15`, non-uniform `19` (signed) / `22` (unsigned). The cost model now reports the sequence it actually emits, and the vector forms still win: by a comfortable margin rather than by one.
The value is intentionally divisor-independent: `TTI` does not see the constant, and whether the magic number needs the extra "add" fixup only moves the count by one or two, within the tolerance of a cost label.
### Testing
- `Analysis/CostModel/X86/{div,rem}.ll` regenerated: every scalar `udiv`/`sdiv`/`urem`/`srem`-by-constant row moves from `1` to `5`/`6` (larger cost kinds `+2`), and the `vXi64` vector rows move from `9` back to `15`/`19`/`22`.
- `Transforms/SLPVectorizer/X86/idiv-by-const.ll` and four SLP tests that carry constant-divisor div/rem (`non-schedulable-user-different-bb`, `reused-extract-scalar-lanes`, `split-node-reused-in-later-vector`, `split-vector-operand-with-reuses`) are regenerated; each still vectorizes.
Part of #<!-- -->37771.
### AI Usage Disclosure
This PR was prepared with the assistance of Claude Code.
---
Patch is 300.21 KiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/211529.diff
8 Files Affected:
- (modified) llvm/lib/Target/X86/X86TargetTransformInfo.cpp (+23-8)
- (modified) llvm/test/Analysis/CostModel/X86/div.ll (+260-260)
- (modified) llvm/test/Analysis/CostModel/X86/rem.ll (+240-240)
- (modified) llvm/test/Transforms/SLPVectorizer/X86/idiv-by-const.ll (+43-168)
- (modified) llvm/test/Transforms/SLPVectorizer/X86/non-schedulable-user-different-bb.ll (+3-4)
- (modified) llvm/test/Transforms/SLPVectorizer/X86/reused-extract-scalar-lanes.ll (+8-8)
- (modified) llvm/test/Transforms/SLPVectorizer/X86/split-node-reused-in-later-vector.ll (+1-14)
- (modified) llvm/test/Transforms/SLPVectorizer/X86/split-vector-operand-with-reuses.ll (+12-9)
``````````diff
diff --git a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
index 188c8f76b8e38..7b9493ff39cdb 100644
--- a/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
+++ b/llvm/lib/Target/X86/X86TargetTransformInfo.cpp
@@ -278,6 +278,21 @@ InstructionCost X86TTIImpl::getArithmeticInstrCost(
int ISD = TLI->InstructionOpcodeToISD(Opcode);
assert(ISD && "Invalid opcode");
+ // A scalar integer divide/remainder by a constant is not a hardware divide;
+ // it lowers to a magic-number multiply-high plus a few fixup ops. Cost it as
+ // that sequence rather than the generic single-instruction divide, so the
+ // vectorizers do not compare against an artificially cheap scalar lane.
+ // Power-of-two divisors lower to a shift and are left to the generic handling.
+ if (!Ty->isVectorTy() && Op2Info.isConstant() && !Op2Info.isPowerOf2() &&
+ !Op2Info.isNegatedPowerOf2() &&
+ (ISD == ISD::UDIV || ISD == ISD::SDIV || ISD == ISD::UREM ||
+ ISD == ISD::SREM)) {
+ unsigned Cost = ISD == ISD::UREM || ISD == ISD::SREM ? 6 : 5;
+ if (CostKind == TTI::TCK_CodeSize || CostKind == TTI::TCK_SizeAndLatency)
+ Cost += 2;
+ return Cost;
+ }
+
if (ISD == ISD::MUL && Args.size() == 2 && LT.second.isVector() &&
(LT.second.getScalarType() == MVT::i32 ||
LT.second.getScalarType() == MVT::i64)) {
@@ -423,9 +438,9 @@ InstructionCost X86TTIImpl::getArithmeticInstrCost(
return LT.first * *KindCost;
static const CostKindTblEntry AVX512DQUniformConstCostTable[] = {
- { ISD::SDIV, MVT::v4i64, { 9 } }, // vpmullq-based MULHS sequence
+ { ISD::SDIV, MVT::v4i64, { 15 } }, // vpmullq-based MULHS sequence
{ ISD::SREM, MVT::v4i64, { 17 } }, // vpmullq-based MULHS+mul+sub sequence
- { ISD::SDIV, MVT::v8i64, { 9 } }, // vpmullq-based MULHS sequence
+ { ISD::SDIV, MVT::v8i64, { 15 } }, // vpmullq-based MULHS sequence
{ ISD::SREM, MVT::v8i64, { 17 } }, // vpmullq-based MULHS+mul+sub sequence
// The remainder's multiply-back is a single vpmullq with DQ, just like the
// pmulld the vXi32 entries above rely on. Without DQ it is another
@@ -469,7 +484,7 @@ InstructionCost X86TTIImpl::getArithmeticInstrCost(
{ ISD::UDIV, MVT::v16i32, { 5 } }, // pmuludq sequence
{ ISD::UREM, MVT::v16i32, { 7 } }, // pmuludq+mul+sub sequence
- { ISD::UDIV, MVT::v8i64, { 9 } }, // pmuludq-based MULHU sequence
+ { ISD::UDIV, MVT::v8i64, { 15 } }, // pmuludq-based MULHU sequence
{ ISD::UREM, MVT::v8i64, { 21 } }, // pmuludq-based MULHU+mul+sub sequence
};
@@ -513,7 +528,7 @@ InstructionCost X86TTIImpl::getArithmeticInstrCost(
{ ISD::UDIV, MVT::v8i32, { 5 } }, // pmuludq sequence
{ ISD::UREM, MVT::v8i32, { 7 } }, // pmuludq+mul+sub sequence
- { ISD::UDIV, MVT::v4i64, { 9 } }, // pmuludq-based MULHU sequence
+ { ISD::UDIV, MVT::v4i64, { 15 } }, // pmuludq-based MULHU sequence
{ ISD::UREM, MVT::v4i64, { 21 } }, // pmuludq-based MULHU+mul+sub sequence
};
@@ -616,9 +631,9 @@ InstructionCost X86TTIImpl::getArithmeticInstrCost(
return LT.first * *KindCost;
static const CostKindTblEntry AVX512DQConstCostTable[] = {
- { ISD::SDIV, MVT::v4i64, { 9 } }, // vpmullq-based MULHS sequence
+ { ISD::SDIV, MVT::v4i64, { 19 } }, // vpmullq-based MULHS sequence
{ ISD::SREM, MVT::v4i64, { 21 } }, // vpmullq-based MULHS+mul+sub sequence
- { ISD::SDIV, MVT::v8i64, { 9 } }, // vpmullq-based MULHS sequence
+ { ISD::SDIV, MVT::v8i64, { 19 } }, // vpmullq-based MULHS sequence
{ ISD::SREM, MVT::v8i64, { 21 } }, // vpmullq-based MULHS+mul+sub sequence
// The remainder's multiply-back is a single vpmullq with DQ, whereas the
// AVX512/AVX2 tables have to charge for another vpmuludq schoolbook.
@@ -648,7 +663,7 @@ InstructionCost X86TTIImpl::getArithmeticInstrCost(
{ ISD::UDIV, MVT::v16i32, { 15 } }, // vpmuludq sequence
{ ISD::UREM, MVT::v16i32, { 17 } }, // vpmuludq+mul+sub sequence
- { ISD::UDIV, MVT::v8i64, { 9 } }, // vpmuludq-based MULHU sequence
+ { ISD::UDIV, MVT::v8i64, { 22 } }, // vpmuludq-based MULHU sequence
{ ISD::UREM, MVT::v8i64, { 28 } }, // vpmuludq-based MULHU+mul+sub sequence
};
@@ -674,7 +689,7 @@ InstructionCost X86TTIImpl::getArithmeticInstrCost(
{ ISD::UDIV, MVT::v8i32, { 15 } }, // vpmuludq sequence
{ ISD::UREM, MVT::v8i32, { 19 } }, // vpmuludq+mul+sub sequence
- { ISD::UDIV, MVT::v4i64, { 9 } }, // vpmuludq-based MULHU sequence
+ { ISD::UDIV, MVT::v4i64, { 22 } }, // vpmuludq-based MULHU sequence
{ ISD::UREM, MVT::v4i64, { 28 } }, // vpmuludq-based MULHU+mul+sub sequence
};
diff --git a/llvm/test/Analysis/CostModel/X86/div.ll b/llvm/test/Analysis/CostModel/X86/div.ll
index 59d232a2101c8..151ac37681ed7 100644
--- a/llvm/test/Analysis/CostModel/X86/div.ll
+++ b/llvm/test/Analysis/CostModel/X86/div.ll
@@ -100,152 +100,152 @@ define i32 @udiv() {
define i32 @sdiv_const() {
; SSE2-LABEL: 'sdiv_const'
-; SSE2-NEXT: Cost Model: Found costs of RThru:1 CodeSize:4 Lat:4 SizeLat:4 for: %I64 = sdiv i64 undef, 7
-; SSE2-NEXT: Cost Model: Found costs of RThru:40 CodeSize:4 Lat:4 SizeLat:4 for: %V2i64 = sdiv <2 x i64> undef, <i64 6, i64 7>
-; SSE2-NEXT: Cost Model: Found costs of RThru:80 CodeSize:4 Lat:4 SizeLat:4 for: %V4i64 = sdiv <4 x i64> undef, <i64 4, i64 5, i64 6, i64 7>
-; SSE2-NEXT: Cost Model: Found costs of RThru:160 CodeSize:4 Lat:4 SizeLat:4 for: %V8i64 = sdiv <8 x i64> undef, <i64 4, i64 5, i64 6, i64 7, i64 8, i64 9, i64 10, i64 11>
-; SSE2-NEXT: Cost Model: Found costs of RThru:1 CodeSize:4 Lat:4 SizeLat:4 for: %I32 = sdiv i32 undef, 7
+; SSE2-NEXT: Cost Model: Found costs of RThru:5 CodeSize:7 Lat:5 SizeLat:7 for: %I64 = sdiv i64 undef, 7
+; SSE2-NEXT: Cost Model: Found costs of RThru:200 CodeSize:4 Lat:4 SizeLat:4 for: %V2i64 = sdiv <2 x i64> undef, <i64 6, i64 7>
+; SSE2-NEXT: Cost Model: Found costs of RThru:400 CodeSize:4 Lat:4 SizeLat:4 for: %V4i64 = sdiv <4 x i64> undef, <i64 4, i64 5, i64 6, i64 7>
+; SSE2-NEXT: Cost Model: Found costs of RThru:800 CodeSize:4 Lat:4 SizeLat:4 for: %V8i64 = sdiv <8 x i64> undef, <i64 4, i64 5, i64 6, i64 7, i64 8, i64 9, i64 10, i64 11>
+; SSE2-NEXT: Cost Model: Found costs of RThru:5 CodeSize:7 Lat:5 SizeLat:7 for: %I32 = sdiv i32 undef, 7
; SSE2-NEXT: Cost Model: Found costs of RThru:19 CodeSize:4 Lat:4 SizeLat:4 for: %V4i32 = sdiv <4 x i32> undef, <i32 4, i32 5, i32 6, i32 7>
; SSE2-NEXT: Cost Model: Found costs of RThru:38 CodeSize:4 Lat:4 SizeLat:4 for: %V8i32 = sdiv <8 x i32> undef, <i32 4, i32 5, i32 6, i32 7, i32 8, i32 9, i32 10, i32 11>
; SSE2-NEXT: Cost Model: Found costs of RThru:76 CodeSize:4 Lat:4 SizeLat:4 for: %V16i32 = sdiv <16 x i32> undef, <i32 4, i32 5, i32 6, i32 7, i32 8, i32 9, i32 10, i32 11, i32 12, i32 13, i32 14, i32 15, i32 16, i32 17, i32 18, i32 19>
-; SSE2-NEXT: Cost Model: Found costs of RThru:1 CodeSize:4 Lat:4 SizeLat:4 for: %I16 = sdiv i16 undef, 7
+; SSE2-NEXT: Cost Model: Found costs of RThru:5 CodeSize:7 Lat:5 SizeLat:7 for: %I16 = sdiv i16 undef, 7
; SSE2-NEXT: Cost Model: Found costs of RThru:6 CodeSize:4 Lat:4 SizeLat:4 for: %V8i16 = sdiv <8 x i16> undef, <i16 4, i16 5, i16 6, i16 7, i16 8, i16 9, i16 10, i16 11>
; SSE2-NEXT: Cost Model: Found costs of RThru:12 CodeSize:4 Lat:4 SizeLat:4 for: %V16i16 = sdiv <16 x i16> undef, <i16 4, i16 5, i16 6, i16 7, i16 8, i16 9, i16 10, i16 11, i16 12, i16 13, i16 14, i16 15, i16 16, i16 17, i16 18, i16 19>
; SSE2-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:4 SizeLat:4 for: %V32i16 = sdiv <32 x i16> undef, <i16 4, i16 5, i16 6, i16 7, i16 8, i16 9, i16 10, i16 11, i16 12, i16 13, i16 14, i16 15, i16 16, i16 17, i16 18, i16 19, i16 4, i16 5, i16 6, i16 7, i16 8, i16 9, i16 10, i16 11, i16 12, i16 13, i16 14, i16 15, i16 16, i16 17, i16 18, i16 19>
-; SSE2-NEXT: Cost Model: Found costs of RThru:1 CodeSize:4 Lat:4 SizeLat:4 for: %I8 = sdiv i8 undef, 7
+; SSE2-NEXT: Cost Model: Found costs of RThru:5 CodeSize:7 Lat:5 SizeLat:7 for: %I8 = sdiv i8 undef, 7
; SSE2-NEXT: Cost Model: Found costs of RThru:14 CodeSize:4 Lat:4 SizeLat:4 for: %V16i8 = sdiv <16 x i8> undef, <i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19>
; SSE2-NEXT: Cost Model: Found costs of RThru:28 CodeSize:4 Lat:4 SizeLat:4 for: %V32i8 = sdiv <32 x i8> undef, <i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19, i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19>
; SSE2-NEXT: Cost Model: Found costs of RThru:56 CodeSize:4 Lat:4 SizeLat:4 for: %V64i8 = sdiv <64 x i8> undef, <i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19, i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19, i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19, i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19>
; SSE2-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret i32 undef
;
; SSE42-LABEL: 'sdiv_const'
-; SSE42-NEXT: Cost Model: Found costs of RThru:1 CodeSize:4 Lat:4 SizeLat:4 for: %I64 = sdiv i64 undef, 7
-; SSE42-NEXT: Cost Model: Found costs of RThru:40 CodeSize:4 Lat:4 SizeLat:4 for: %V2i64 = sdiv <2 x i64> undef, <i64 6, i64 7>
-; SSE42-NEXT: Cost Model: Found costs of RThru:80 CodeSize:4 Lat:4 SizeLat:4 for: %V4i64 = sdiv <4 x i64> undef, <i64 4, i64 5, i64 6, i64 7>
-; SSE42-NEXT: Cost Model: Found costs of RThru:160 CodeSize:4 Lat:4 SizeLat:4 for: %V8i64 = sdiv <8 x i64> undef, <i64 4, i64 5, i64 6, i64 7, i64 8, i64 9, i64 10, i64 11>
-; SSE42-NEXT: Cost Model: Found costs of RThru:1 CodeSize:4 Lat:4 SizeLat:4 for: %I32 = sdiv i32 undef, 7
+; SSE42-NEXT: Cost Model: Found costs of RThru:5 CodeSize:7 Lat:5 SizeLat:7 for: %I64 = sdiv i64 undef, 7
+; SSE42-NEXT: Cost Model: Found costs of RThru:200 CodeSize:4 Lat:4 SizeLat:4 for: %V2i64 = sdiv <2 x i64> undef, <i64 6, i64 7>
+; SSE42-NEXT: Cost Model: Found costs of RThru:400 CodeSize:4 Lat:4 SizeLat:4 for: %V4i64 = sdiv <4 x i64> undef, <i64 4, i64 5, i64 6, i64 7>
+; SSE42-NEXT: Cost Model: Found costs of RThru:800 CodeSize:4 Lat:4 SizeLat:4 for: %V8i64 = sdiv <8 x i64> undef, <i64 4, i64 5, i64 6, i64 7, i64 8, i64 9, i64 10, i64 11>
+; SSE42-NEXT: Cost Model: Found costs of RThru:5 CodeSize:7 Lat:5 SizeLat:7 for: %I32 = sdiv i32 undef, 7
; SSE42-NEXT: Cost Model: Found costs of RThru:15 CodeSize:4 Lat:4 SizeLat:4 for: %V4i32 = sdiv <4 x i32> undef, <i32 4, i32 5, i32 6, i32 7>
; SSE42-NEXT: Cost Model: Found costs of RThru:30 CodeSize:4 Lat:4 SizeLat:4 for: %V8i32 = sdiv <8 x i32> undef, <i32 4, i32 5, i32 6, i32 7, i32 8, i32 9, i32 10, i32 11>
; SSE42-NEXT: Cost Model: Found costs of RThru:60 CodeSize:4 Lat:4 SizeLat:4 for: %V16i32 = sdiv <16 x i32> undef, <i32 4, i32 5, i32 6, i32 7, i32 8, i32 9, i32 10, i32 11, i32 12, i32 13, i32 14, i32 15, i32 16, i32 17, i32 18, i32 19>
-; SSE42-NEXT: Cost Model: Found costs of RThru:1 CodeSize:4 Lat:4 SizeLat:4 for: %I16 = sdiv i16 undef, 7
+; SSE42-NEXT: Cost Model: Found costs of RThru:5 CodeSize:7 Lat:5 SizeLat:7 for: %I16 = sdiv i16 undef, 7
; SSE42-NEXT: Cost Model: Found costs of RThru:6 CodeSize:4 Lat:4 SizeLat:4 for: %V8i16 = sdiv <8 x i16> undef, <i16 4, i16 5, i16 6, i16 7, i16 8, i16 9, i16 10, i16 11>
; SSE42-NEXT: Cost Model: Found costs of RThru:12 CodeSize:4 Lat:4 SizeLat:4 for: %V16i16 = sdiv <16 x i16> undef, <i16 4, i16 5, i16 6, i16 7, i16 8, i16 9, i16 10, i16 11, i16 12, i16 13, i16 14, i16 15, i16 16, i16 17, i16 18, i16 19>
; SSE42-NEXT: Cost Model: Found costs of RThru:24 CodeSize:4 Lat:4 SizeLat:4 for: %V32i16 = sdiv <32 x i16> undef, <i16 4, i16 5, i16 6, i16 7, i16 8, i16 9, i16 10, i16 11, i16 12, i16 13, i16 14, i16 15, i16 16, i16 17, i16 18, i16 19, i16 4, i16 5, i16 6, i16 7, i16 8, i16 9, i16 10, i16 11, i16 12, i16 13, i16 14, i16 15, i16 16, i16 17, i16 18, i16 19>
-; SSE42-NEXT: Cost Model: Found costs of RThru:1 CodeSize:4 Lat:4 SizeLat:4 for: %I8 = sdiv i8 undef, 7
+; SSE42-NEXT: Cost Model: Found costs of RThru:5 CodeSize:7 Lat:5 SizeLat:7 for: %I8 = sdiv i8 undef, 7
; SSE42-NEXT: Cost Model: Found costs of RThru:14 CodeSize:4 Lat:4 SizeLat:4 for: %V16i8 = sdiv <16 x i8> undef, <i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19>
; SSE42-NEXT: Cost Model: Found costs of RThru:28 CodeSize:4 Lat:4 SizeLat:4 for: %V32i8 = sdiv <32 x i8> undef, <i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19, i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19>
; SSE42-NEXT: Cost Model: Found costs of RThru:56 CodeSize:4 Lat:4 SizeLat:4 for: %V64i8 = sdiv <64 x i8> undef, <i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19, i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19, i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19, i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19>
; SSE42-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret i32 undef
;
; AVX1-LABEL: 'sdiv_const'
-; AVX1-NEXT: Cost Model: Found costs of RThru:1 CodeSize:4 Lat:4 SizeLat:4 for: %I64 = sdiv i64 undef, 7
-; AVX1-NEXT: Cost Model: Found costs of RThru:40 CodeSize:4 Lat:4 SizeLat:4 for: %V2i64 = sdiv <2 x i64> undef, <i64 6, i64 7>
-; AVX1-NEXT: Cost Model: Found costs of RThru:80 CodeSize:4 Lat:4 SizeLat:4 for: %V4i64 = sdiv <4 x i64> undef, <i64 4, i64 5, i64 6, i64 7>
-; AVX1-NEXT: Cost Model: Found costs of RThru:160 CodeSize:4 Lat:4 SizeLat:4 for: %V8i64 = sdiv <8 x i64> undef, <i64 4, i64 5, i64 6, i64 7, i64 8, i64 9, i64 10, i64 11>
-; AVX1-NEXT: Cost Model: Found costs of RThru:1 CodeSize:4 Lat:4 SizeLat:4 for: %I32 = sdiv i32 undef, 7
+; AVX1-NEXT: Cost Model: Found costs of RThru:5 CodeSize:7 Lat:5 SizeLat:7 for: %I64 = sdiv i64 undef, 7
+; AVX1-NEXT: Cost Model: Found costs of RThru:200 CodeSize:4 Lat:4 SizeLat:4 for: %V2i64 = sdiv <2 x i64> undef, <i64 6, i64 7>
+; AVX1-NEXT: Cost Model: Found costs of RThru:400 CodeSize:4 Lat:4 SizeLat:4 for: %V4i64 = sdiv <4 x i64> undef, <i64 4, i64 5, i64 6, i64 7>
+; AVX1-NEXT: Cost Model: Found costs of RThru:800 CodeSize:4 Lat:4 SizeLat:4 for: %V8i64 = sdiv <8 x i64> undef, <i64 4, i64 5, i64 6, i64 7, i64 8, i64 9, i64 10, i64 11>
+; AVX1-NEXT: Cost Model: Found costs of RThru:5 CodeSize:7 Lat:5 SizeLat:7 for: %I32 = sdiv i32 undef, 7
; AVX1-NEXT: Cost Model: Found costs of RThru:15 CodeSize:4 Lat:4 SizeLat:4 for: %V4i32 = sdiv <4 x i32> undef, <i32 4, i32 5, i32 6, i32 7>
; AVX1-NEXT: Cost Model: Found costs of RThru:32 CodeSize:4 Lat:4 SizeLat:4 for: %V8i32 = sdiv <8 x i32> undef, <i32 4, i32 5, i32 6, i32 7, i32 8, i32 9, i32 10, i32 11>
; AVX1-NEXT: Cost Model: Found costs of RThru:64 CodeSize:4 Lat:4 SizeLat:4 for: %V16i32 = sdiv <16 x i32> undef, <i32 4, i32 5, i32 6, i32 7, i32 8, i32 9, i32 10, i32 11, i32 12, i32 13, i32 14, i32 15, i32 16, i32 17, i32 18, i32 19>
-; AVX1-NEXT: Cost Model: Found costs of RThru:1 CodeSize:4 Lat:4 SizeLat:4 for: %I16 = sdiv i16 undef, 7
+; AVX1-NEXT: Cost Model: Found costs of RThru:5 CodeSize:7 Lat:5 SizeLat:7 for: %I16 = sdiv i16 undef, 7
; AVX1-NEXT: Cost Model: Found costs of RThru:6 CodeSize:4 Lat:4 SizeLat:4 for: %V8i16 = sdiv <8 x i16> undef, <i16 4, i16 5, i16 6, i16 7, i16 8, i16 9, i16 10, i16 11>
; AVX1-NEXT: Cost Model: Found costs of RThru:14 CodeSize:4 Lat:4 SizeLat:4 for: %V16i16 = sdiv <16 x i16> undef, <i16 4, i16 5, i16 6, i16 7, i16 8, i16 9, i16 10, i16 11, i16 12, i16 13, i16 14, i16 15, i16 16, i16 17, i16 18, i16 19>
; AVX1-NEXT: Cost Model: Found costs of RThru:28 CodeSize:4 Lat:4 SizeLat:4 for: %V32i16 = sdiv <32 x i16> undef, <i16 4, i16 5, i16 6, i16 7, i16 8, i16 9, i16 10, i16 11, i16 12, i16 13, i16 14, i16 15, i16 16, i16 17, i16 18, i16 19, i16 4, i16 5, i16 6, i16 7, i16 8, i16 9, i16 10, i16 11, i16 12, i16 13, i16 14, i16 15, i16 16, i16 17, i16 18, i16 19>
-; AVX1-NEXT: Cost Model: Found costs of RThru:1 CodeSize:4 Lat:4 SizeLat:4 for: %I8 = sdiv i8 undef, 7
+; AVX1-NEXT: Cost Model: Found costs of RThru:5 CodeSize:7 Lat:5 SizeLat:7 for: %I8 = sdiv i8 undef, 7
; AVX1-NEXT: Cost Model: Found costs of RThru:14 CodeSize:4 Lat:4 SizeLat:4 for: %V16i8 = sdiv <16 x i8> undef, <i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19>
; AVX1-NEXT: Cost Model: Found costs of RThru:30 CodeSize:4 Lat:4 SizeLat:4 for: %V32i8 = sdiv <32 x i8> undef, <i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19, i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19>
; AVX1-NEXT: Cost Model: Found costs of RThru:60 CodeSize:4 Lat:4 SizeLat:4 for: %V64i8 = sdiv <64 x i8> undef, <i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19, i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19, i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19, i8 4, i8 5, i8 6, i8 7, i8 8, i8 9, i8 10, i8 11, i8 12, i8 13, i8 14, i8 15, i8 16, i8 17, i8 18, i8 19>
; AVX1-NEXT: Cost Model: Found costs of RThru:0 CodeSize:1 Lat:1 SizeLat:1 for: ret i32 undef
;
; AVX2-LABEL: 'sdiv_const'
-; AVX2-NEXT: Cost Model: Found costs of RThru:1 CodeSize:4 Lat:4 SizeLat:4 for: %I64 = sdiv i64 undef, 7
-; AVX2-NEXT: Cost Model: Found costs of RThru:40 CodeSize:4 Lat:4 SizeLat:4 for: %V2i64 = sdiv <2 x i64> undef, <i64 6, i64 7>
-; AVX2-NEXT: Cost Model: Found costs of RThru:80 CodeSize:4 Lat:4 SizeLat:4 for: %V4i64 = sdiv <4 x i64> undef, <i64 4, i64 5, i64 6, i64 7>
-; AVX2-NEXT: Cost Model: Found costs of RThru:160 CodeSize:4 Lat:4 SizeLat:4 for: %V8i64 = sdiv <8 x i64> undef, <i64 4, i64 5, i64 6, i64 7, i64 8, i64 9, i64 10, i64 11>
-; AVX2-NEXT: Cost Model: Found costs of RThru:1 CodeSize:4 Lat:4 SizeLat:4 for: %I32 = sdiv i32 undef, 7
+; AVX2-NEXT: Cost Model: Found costs of RThru:5 CodeSize:7 Lat:5 SizeLat:7 for: %I64 = sdiv i64 undef, 7
+; AVX2-NEXT: Cost Model: Found costs of RThru:200 CodeSize:4 Lat:4 SizeLat:4 for: %V2i64 = sdiv <2 x i64> undef, <i64 6, i64 7>
+; AVX2-NEXT: Cost Model: Found costs of RThru:400 CodeSize:4 Lat:4 SizeLat:4 for: %V4i64 = sdiv <4 x i64> undef, <i64 4, i64 5, i64 6, i64 7>
+; AVX2-NEXT: Cost Model: Found costs of RThru:800 CodeSize:4 Lat:4 SizeLat:4 for: %V8i64 = sdiv <8 x i64> undef, <i64 4, i64 5, i64 6, i64 7, i64 8, i64 9, i64 10, i64 11>
+; AVX2-NEXT: Cost Model: Found costs of RThru:5 CodeSize:7 Lat:5 SizeLat:7 for: %I32 = sdiv i32 undef, 7
; AVX2-NEXT: Cost Model: Found costs of RThru:15 CodeSize:4 Lat:4 SizeLat:4 for: %V4i32 = sdiv <4 x i32> undef, <i32 4, i32 5, i32 6, i32 7>
; AVX2-NEXT: Cost Model: Found costs of RThru:15 CodeSize:4 Lat:4 SizeLat:4 for: %V8i32 = sdiv <8 x i32> undef, <i32 4, i32 5, i32 6, i32 7, i32 8, i32 9, i32 10, i32 11>
; AVX2-NEXT: Cost Model: Found costs of RThru:30 CodeSize:4 Lat:4 SizeLat:4 for: %V16i32 = sdiv <16 x i32> undef, <i32 4, i32 5, i32 6, i32 7, i32 8, i32 9, i32 10, i32 11, i32 12, i32 13, i32 14, i32 15, i32 16, i32 17, i32 18, i32 19>
-; AVX2-NEXT: Cost Model: Found costs of RThru:1 CodeSize:4 Lat:4 SizeLat:4 for: %I16 = sdiv i16 undef, 7
+; AVX2-NEXT: Cost Model: Found costs of RThru:5 CodeSize:7 Lat:5 SizeLat:7 for: %I16 = sdiv i16 undef, 7
; AVX2-NEXT: Cost Model: Found costs of RThru:6 CodeSize:4 Lat:4 SizeLat:4 for: %V8i16 = sdiv <8 x i16> undef, <i16 4, i16 5, i16 6, i16 7, i16 8, i16 9, i16 10, i16 11>
; AVX2-NEXT: Cost Model: Found costs of RThru:6 CodeSize:4 Lat:4 SizeLat:4 for: %V16i16 = sdiv <16 x i16> undef, <i16 4, i16 5, i16 6, i16 7, i16 8, i16 9, i16 10, i16 11, i16 12, i16 13, i16 14, i16 15, i16 16, i16 17, i16 18, i16 19>
; AVX2-NEXT: Cost Model: Found costs of RThru:12 CodeSize:4 ...
[truncated]
``````````
</details>
https://github.com/llvm/llvm-project/pull/211529
More information about the llvm-commits
mailing list