[llvm] [LoopIdiomRecognize] Enable clmul optimization for CRC loops (PR #203405)
Ramkumar Ramachandra via llvm-commits
llvm-commits at lists.llvm.org
Mon Jun 29 17:48:21 PDT 2026
================
@@ -1549,7 +1554,144 @@ bool LoopIdiomRecognize::avoidLIRForMultiBlockLoop(bool IsMemset,
return false;
}
-bool LoopIdiomRecognize::optimizeCRCLoop(const PolynomialInfo &Info) {
+bool LoopIdiomRecognize::optimizeCRCLoopUsingClmul(const PolynomialInfo &Info) {
+ Type *CRCTy = Info.LHS->getType();
+ LLVMContext &Ctx = CRCTy->getContext();
+ unsigned CRCBW = CRCTy->getIntegerBitWidth();
+ // The TripCount determines how many bits of data are processed, regardless of
+ // whether the actual data bit width matches (if auxiliary data is even used
+ // at all).
+ unsigned TC = Info.TripCount;
+ // The first clmul uses 2*TC bits, and the second clmul uses CRCBW+TC bits.
+ // For simplicity, have both operate on the same bit width.
+ unsigned ClmulBW = std::max(2 * TC, CRCBW + TC);
+ auto *ClmulTy = IntegerType::get(Ctx, ClmulBW);
+
+ // This optimization should only be applied if clmul for the required width is
+ // a fast operation on the target.
+ // TODO: If TC > CRCBW, then the data could probably be split into multiple
+ // chunks and processed in a loop.
+ if (!TTI->haveFastClmul(ClmulTy))
+ return false;
+
+ // First, generate the constants required for GF(2) Barrett reduction.
+ auto [Mu, FullGenPoly] =
+ HashRecognize::genBarrettConstants(Info.RHS, TC, Info.IsBigEndian);
+ Value *MuExt = ConstantInt::get(Ctx, Mu.zext(ClmulBW));
+ Value *FullGenPolyExt = ConstantInt::get(Ctx, FullGenPoly.zext(ClmulBW));
+
+ IRBuilder<> Builder(CurLoop->getLoopPreheader()->getTerminator());
+
+ auto GetMostSignificantTCBits = [&](Value *Op, unsigned BW,
+ const Twine &Name) {
+ unsigned OpBW = Op->getType()->getIntegerBitWidth();
+ assert(BW > TC && OpBW > TC &&
+ "Trip count should not exceed the bit width");
+ if (Info.IsBigEndian)
+ return Builder.CreateLShr(Op, BW - TC, Name + ".be.lshr");
+ auto *Mask = ConstantInt::get(Ctx, APInt::getLowBitsSet(OpBW, TC));
+ return Builder.CreateAnd(Op, Mask, Name + ".le.mask");
+ };
+
+ Value *LHS = Builder.CreateZExt(Info.LHS, ClmulTy, "crc.ext");
+
+ // For the big-endian case, align the leftmost bit of the CRC with the
+ // leftmost bit of the data which is used. For the little-endian case, align
+ // the rightmost bits (nothing to do).
+ Value *CRCAlignTC;
+ if (!Info.IsBigEndian || CRCBW == TC)
+ CRCAlignTC = LHS;
+ else if (CRCBW > TC)
+ CRCAlignTC = Builder.CreateLShr(LHS, CRCBW - TC, "crc.be.tcbits");
+ else
+ CRCAlignTC = Builder.CreateShl(LHS, TC - CRCBW, "crc.be.tcbits");
+
+ // If auxiliary data is present, XOR it in with the CRC.
+ Value *ClmulMuInput;
+ if (Value *Data = Info.LHSAux) {
+ unsigned DataBW = Data->getType()->getIntegerBitWidth();
+ // For big-endian CRC loops where auxiliary data is XORed with the CRC
+ // inside the loop, the bits won't be aligned properly if the bit widths
+ // don't match, and thus the CRC computation is incorrect, but HashRecognize
+ // will still detect the loop. To handle the case where the data is zexted
+ // before XORing with the CRC, just ignore the auxiliary data entirely,
+ // because the extracted bit will always be zero.
+ if (!Info.IsBigEndian || DataBW >= CRCBW) {
+ // For the aforementioned HashRecognize quirk, to handle the case where
+ // the data is truncated before XORing with the CRC, shift the data so the
+ // CRCBW-1 bit becomes the leftmost bit, and then the remaining logic
+ // treats the DataBW-1 bit as the first bit to be processed.
----------------
artagnon wrote:
I think this is due to missing a part of the Barrett Reduction algorithm:
Input: degree-127 polynomial R(x), degree-64 polynomial P(x), µ = floor(x^128 / P(x))
Output: C(x) = R(x) mod P(x)
Step 1: T1(x) = clmul(floor(R(x)/x^64)), µ)
Step 2: T2(x) = clmul(floor(T1(x)/x^64)), P(x))
Step 3: C(x) = R(x) ^ T2(x) mod x^64
I think you missed the floor division by x^64 in steps 1 and 2 entirely. Example: floor(x^128 + x^64 + x^16 + 1/x^32) = floor(0b11010001 / 0b100000) = x^4 + x^2 = 0b110. The remainder is 0b10001, which is required in step 3. In other words, simply right-shift by log2(32) = 5. In all these steps, it would simply be a right-shift by CRCBW.
https://github.com/llvm/llvm-project/pull/203405
More information about the llvm-commits
mailing list