[llvm] [LoopIdiomRecognize] Enable clmul optimization for CRC loops (PR #203405)
Piotr Fusik via llvm-commits
llvm-commits at lists.llvm.org
Wed Jul 8 04:03:23 PDT 2026
================
@@ -376,6 +376,74 @@ CRCTable HashRecognize::genSarwateTable(const APInt &GenPoly,
return Table;
}
+// Perform polynomial (GF(2)) floor division. This is based on the
+// floor_division(S, P) algorithm in
+// https://www.corsix.org/content/barrett-reduction-polynomials. Note that the
+// maximum degree of the returned polynomial is
+// max(0, deg(Dividend) - deg(Divisor)), but the bit width will be the same as
+// that of Dividend.
+static APInt floorDivideGF2(APInt Dividend, APInt Divisor) {
+ assert(!Divisor.isZero() && "Cannot divide by zero");
+
+ // Extend the divisor bit width to match the dividend.
+ Divisor = Divisor.zext(Dividend.getBitWidth());
+
+ // Note that getActiveBits(_) returns deg(_)+1, but the computation below
+ // still holds.
+ unsigned DivisorActiveBits = Divisor.getActiveBits();
+
+ // Q = 0
+ APInt Quotient = APInt::getZero(Dividend.getBitWidth());
+ // S != 0 and deg(S) >= deg(P)
+ // (S != 0 implied by DivisorActiveBits > 0)
+ while (Dividend.getActiveBits() >= DivisorActiveBits) {
+ // T = S[deg(S)] / P[deg(P)]
+ unsigned Shift = Dividend.getActiveBits() - DivisorActiveBits;
+ // Q = Q + T
+ Quotient.setBit(Shift);
+ // S = S - T * P
+ Dividend ^= Divisor.shl(Shift);
+ }
+ return Quotient;
+}
+
+// Generate the constants for performing a Polynomial (GF(2)) Barrett Reduction
+// according to Intel's Fast CRC Computation white paper
+// with some adjustments to account for the fact that bit width and trip count
+// can vary.
+std::pair<APInt, APInt> HashRecognize::genBarrettConstants(const APInt &GenPoly,
+ unsigned TripCount,
+ bool IsBigEndian) {
+ unsigned BW = GenPoly.getBitWidth();
+
+ // Recover the full generating polynomial in normal form by reflecting the LE
+ // case and adding the implied x^BW term.
+ // deg(P(x)) = BW due to the implied term, and thus P(x) must fit in exactly
+ // BW+1 bits.
+ APInt FullGenPoly =
+ (IsBigEndian ? GenPoly : GenPoly.reverseBits()).zext(BW + 1);
+ FullGenPoly.setBit(BW);
+
+ // Calculate mu = floor(x^(BW+TC) / P(x)).
+ // deg(mu) <= deg(x^(BW+TC)) - deg(P(x)) = BW+TC - BW = TC, and thus mu must
+ // fit in at most TC+1 bits.
+ unsigned DivBW = BW + TripCount + 1;
+ APInt Mu =
+ floorDivideGF2(APInt::getOneBitSet(DivBW, BW + TripCount), FullGenPoly)
+ .trunc(TripCount + 1);
+
+ // In the bit-reflected case, mu and P(x) must be bit-reflected across their
+ // respective widths for the corresponding Barrett reduction steps.
+ if (!IsBigEndian) {
+ Mu = Mu.reverseBits();
+ FullGenPoly = FullGenPoly.reverseBits();
+ }
+
+ // The final constants are (mu or mu') and (P(x) or P(x)'), depending on
----------------
pfusik wrote:
It's unclear what `mu'` and `P(x)'` are - I see no other references to these symbols.
If that's just the bit-reversed `mu` and `P(x)`, this comment only adds confusion.
https://github.com/llvm/llvm-project/pull/203405
More information about the llvm-commits
mailing list