[llvm] [LoopIdiomRecognize] Enable clmul optimization for CRC loops (PR #203405)

Piotr Fusik via llvm-commits llvm-commits at lists.llvm.org
Wed Jul 8 04:03:23 PDT 2026


================
@@ -376,6 +376,74 @@ CRCTable HashRecognize::genSarwateTable(const APInt &GenPoly,
   return Table;
 }
 
+// Perform polynomial (GF(2)) floor division. This is based on the
+// floor_division(S, P) algorithm in
+// https://www.corsix.org/content/barrett-reduction-polynomials. Note that the
+// maximum degree of the returned polynomial is
+// max(0, deg(Dividend) - deg(Divisor)), but the bit width will be the same as
+// that of Dividend.
+static APInt floorDivideGF2(APInt Dividend, APInt Divisor) {
+  assert(!Divisor.isZero() && "Cannot divide by zero");
+
+  // Extend the divisor bit width to match the dividend.
+  Divisor = Divisor.zext(Dividend.getBitWidth());
+
+  // Note that getActiveBits(_) returns deg(_)+1, but the computation below
+  // still holds.
+  unsigned DivisorActiveBits = Divisor.getActiveBits();
+
+  // Q = 0
+  APInt Quotient = APInt::getZero(Dividend.getBitWidth());
+  // S != 0 and deg(S) >= deg(P)
+  // (S != 0 implied by DivisorActiveBits > 0)
+  while (Dividend.getActiveBits() >= DivisorActiveBits) {
+    // T = S[deg(S)] / P[deg(P)]
+    unsigned Shift = Dividend.getActiveBits() - DivisorActiveBits;
+    // Q = Q + T
+    Quotient.setBit(Shift);
+    // S = S - T * P
+    Dividend ^= Divisor.shl(Shift);
+  }
+  return Quotient;
+}
+
+// Generate the constants for performing a Polynomial (GF(2)) Barrett Reduction
+// according to Intel's Fast CRC Computation white paper
+// with some adjustments to account for the fact that bit width and trip count
+// can vary.
+std::pair<APInt, APInt> HashRecognize::genBarrettConstants(const APInt &GenPoly,
+                                                           unsigned TripCount,
+                                                           bool IsBigEndian) {
+  unsigned BW = GenPoly.getBitWidth();
+
+  // Recover the full generating polynomial in normal form by reflecting the LE
+  // case and adding the implied x^BW term.
+  // deg(P(x)) = BW due to the implied term, and thus P(x) must fit in exactly
+  // BW+1 bits.
+  APInt FullGenPoly =
+      (IsBigEndian ? GenPoly : GenPoly.reverseBits()).zext(BW + 1);
+  FullGenPoly.setBit(BW);
+
+  // Calculate mu = floor(x^(BW+TC) / P(x)).
+  // deg(mu) <= deg(x^(BW+TC)) - deg(P(x)) = BW+TC - BW = TC, and thus mu must
+  // fit in at most TC+1 bits.
+  unsigned DivBW = BW + TripCount + 1;
+  APInt Mu =
+      floorDivideGF2(APInt::getOneBitSet(DivBW, BW + TripCount), FullGenPoly)
+          .trunc(TripCount + 1);
+
+  // In the bit-reflected case, mu and P(x) must be bit-reflected across their
+  // respective widths for the corresponding Barrett reduction steps.
+  if (!IsBigEndian) {
+    Mu = Mu.reverseBits();
+    FullGenPoly = FullGenPoly.reverseBits();
+  }
+
+  // The final constants are (mu or mu') and (P(x) or P(x)'), depending on
----------------
pfusik wrote:

It's unclear what `mu'` and `P(x)'` are - I see no other references to these symbols.
If that's just the bit-reversed `mu` and `P(x)`, this comment only adds confusion.

https://github.com/llvm/llvm-project/pull/203405


More information about the llvm-commits mailing list