[llvm] [LoopIdiomRecognize] Enable clmul optimization for CRC loops (PR #203405)
Ramkumar Ramachandra via llvm-commits
llvm-commits at lists.llvm.org
Sat Jun 27 05:13:23 PDT 2026
================
@@ -1549,7 +1553,145 @@ bool LoopIdiomRecognize::avoidLIRForMultiBlockLoop(bool IsMemset,
return false;
}
-bool LoopIdiomRecognize::optimizeCRCLoop(const PolynomialInfo &Info) {
+bool LoopIdiomRecognize::optimizeCRCLoopUsingClmul(const PolynomialInfo &Info) {
+ Type *CRCTy = Info.LHS->getType();
+ LLVMContext &Ctx = CRCTy->getContext();
+ unsigned CRCBW = CRCTy->getIntegerBitWidth();
+ // The TripCount determines how many bits of data are processed, regardless of
+ // whether the actual data bit width matches (if auxiliary data is even used
+ // at all).
+ unsigned TC = Info.TripCount;
+ // The first clmul uses 2*TC bits, and the second clmul uses CRCBW+TC bits.
+ // For simplicity, have both operate on the same bit width.
+ unsigned ClmulBW = std::max(2 * TC, CRCBW + TC);
+ auto *ClmulTy = IntegerType::get(Ctx, ClmulBW);
+
+ // This optimization should only be applied if clmul for the required width is
+ // a fast operation on the target.
+ // TODO: If TC > CRCBW, then the data could probably be split into multiple
+ // chunks and processed in a loop.
+ if (!TTI->haveFastClmul(ClmulTy))
+ return false;
+
+ // First, generate the constants required for GF(2) Barrett reduction.
+ auto [Mu, FullGenPoly] =
+ HashRecognize::genBarrettConstants(Info.RHS, TC, Info.ByteOrderSwapped);
+ Value *MuExt = ConstantInt::get(Ctx, Mu.zext(ClmulBW));
+ Value *FullGenPolyExt = ConstantInt::get(Ctx, FullGenPoly.zext(ClmulBW));
+
+ IRBuilder<> Builder(CurLoop->getLoopPreheader()->getTerminator());
+
+ auto GetMostSignificantTCBits = [&, TC](Value *Op, unsigned BW,
----------------
artagnon wrote:
```suggestion
auto GetMostSignificantTCBits = [&](Value *Op, unsigned BW,
```
`&` would capture everything in the context, including TC?
https://github.com/llvm/llvm-project/pull/203405
More information about the llvm-commits
mailing list