[llvm] 466df93 - [APFloat] Fix sign bit corrupting exponent in unsigned FP8 formats (#223286)

via llvm-commits llvm-commits at lists.llvm.org
Mon Sep 21 04:50:31 PDT 2026


Author: Cheng-Chieh Liu
Date: 2026-09-21T11:50:25Z
New Revision: 466df937833ef60bc9f549a01059869ca7143d32

URL: https://github.com/llvm/llvm-project/commit/466df937833ef60bc9f549a01059869ca7143d32
DIFF: https://github.com/llvm/llvm-project/commit/466df937833ef60bc9f549a01059869ca7143d32.diff

LOG: [APFloat] Fix sign bit corrupting exponent in unsigned FP8 formats (#223286)

`convertIEEEFloatToAPInt` unconditionally ORs the sign bit into the top
bit of the encoded value. `Float8E8M0FNU` and `Float8E5M3FNU` have no
sign bit (`hasSignedRepr = false`), so that top bit actually belongs to
the exponent field. Encoding a value with sign set therefore corrupts
the exponent, in one case colliding with the reserved NaN bit pattern.

Guard the OR with `hasSignedRepr`, matching the existing check in the
decode path (`initFromIEEEAPInt`).

Extend `ConvertLosesUnrepresentableSignAndZero` to also check
`bitcastToAPInt()` after dropping an unrepresentable sign, using `-1.0`
instead of `-2.0` so the corruption is actually visible in the bits.

Added: 
    

Modified: 
    llvm/lib/Support/APFloat.cpp
    llvm/unittests/ADT/APFloatTest.cpp

Removed: 
    


################################################################################
diff  --git a/llvm/lib/Support/APFloat.cpp b/llvm/lib/Support/APFloat.cpp
index 8a1ef696f9350..3d1932454761f 100644
--- a/llvm/lib/Support/APFloat.cpp
+++ b/llvm/lib/Support/APFloat.cpp
@@ -3591,9 +3591,11 @@ APInt IEEEFloat::convertIEEEFloatToAPInt() const {
   }
   std::fill(words_iter, words.end(), uint64_t{0});
   constexpr size_t last_word = words.size() - 1;
-  uint64_t shifted_sign = static_cast<uint64_t>(sign & 1)
-                          << ((S.sizeInBits - 1) % 64);
-  words[last_word] |= shifted_sign;
+  if constexpr (S.hasSignedRepr) {
+    uint64_t shifted_sign = static_cast<uint64_t>(sign & 1)
+                            << ((S.sizeInBits - 1) % 64);
+    words[last_word] |= shifted_sign;
+  }
   uint64_t shifted_exponent = (myexponent & exponent_mask)
                               << (trailing_significand_bits % 64);
   words[last_word] |= shifted_exponent;

diff  --git a/llvm/unittests/ADT/APFloatTest.cpp b/llvm/unittests/ADT/APFloatTest.cpp
index 35e97c6b2d9ea..40a121232fd29 100644
--- a/llvm/unittests/ADT/APFloatTest.cpp
+++ b/llvm/unittests/ADT/APFloatTest.cpp
@@ -2522,23 +2522,27 @@ TEST(APFloatTest, ConvertLosesUnrepresentableSignAndZero) {
 
   for (const fltSemantics *Sem : NoSignSemantics) {
     // The magnitude converts exactly, so the sign is the whole of the loss.
-    APFloat test(-2.0);
+    APFloat test(-1.0);
     bool losesInfo = false;
     APFloat::opStatus status =
         test.convert(*Sem, APFloat::rmNearestTiesToEven, &losesInfo);
     EXPECT_TRUE(losesInfo);
     EXPECT_EQ(status, APFloat::opInexact);
     EXPECT_TRUE(test.isNegative());
-    EXPECT_EQ(-2.0, test.convertToDouble());
+    EXPECT_EQ(-1.0, test.convertToDouble());
+    APInt negBits = test.bitcastToAPInt();
 
     // The same magnitude without the sign has nothing to report.
-    test = APFloat(2.0);
+    test = APFloat(1.0);
     losesInfo = true;
     status = test.convert(*Sem, APFloat::rmNearestTiesToEven, &losesInfo);
     EXPECT_FALSE(losesInfo);
     EXPECT_EQ(status, APFloat::opOK);
     EXPECT_FALSE(test.isNegative());
-    EXPECT_EQ(2.0, test.convertToDouble());
+    EXPECT_EQ(1.0, test.convertToDouble());
+
+    // No sign bit exists, so the bits must match the positive magnitude.
+    EXPECT_EQ(test.bitcastToAPInt(), negBits);
   }
 
   // Float8E8M0FNU has no zero either, and substitutes 2^-127 for one. That


        


More information about the llvm-commits mailing list