[llvm] [APFloat] Fix sign bit corrupting exponent in unsigned FP8 formats (PR #223286)

via llvm-commits llvm-commits at lists.llvm.org
Sun Sep 13 19:20:52 PDT 2026


llvmorg-github-actions[bot] wrote:


<!--LLVM PR SUMMARY COMMENT-->

@llvm/pr-subscribers-llvm-adt

Author: Cheng-Chieh Liu (jay87406)

<details>
<summary>Changes</summary>

`convertIEEEFloatToAPInt` unconditionally ORs the sign bit into the top bit of the encoded value. `Float8E8M0FNU` and `Float8E5M3FNU` have no sign bit (`hasSignedRepr = false`), so that top bit actually belongs to the exponent field. Encoding a value with sign set therefore corrupts the exponent, in one case colliding with the reserved NaN bit pattern.

Guard the OR with `hasSignedRepr`, matching the existing check in the decode path (`initFromIEEEAPInt`).

Extend `ConvertLosesUnrepresentableSignAndZero` to also check `bitcastToAPInt()` after dropping an unrepresentable sign, using `-1.0` instead of `-2.0` so the corruption is actually visible in the bits.

---
Full diff: https://github.com/llvm/llvm-project/pull/223286.diff


2 Files Affected:

- (modified) llvm/lib/Support/APFloat.cpp (+5-3) 
- (modified) llvm/unittests/ADT/APFloatTest.cpp (+8-4) 


``````````diff
diff --git a/llvm/lib/Support/APFloat.cpp b/llvm/lib/Support/APFloat.cpp
index 59a6ba3867d40..e824b00162d33 100644
--- a/llvm/lib/Support/APFloat.cpp
+++ b/llvm/lib/Support/APFloat.cpp
@@ -3569,9 +3569,11 @@ APInt IEEEFloat::convertIEEEFloatToAPInt() const {
   }
   std::fill(words_iter, words.end(), uint64_t{0});
   constexpr size_t last_word = words.size() - 1;
-  uint64_t shifted_sign = static_cast<uint64_t>(sign & 1)
-                          << ((S.sizeInBits - 1) % 64);
-  words[last_word] |= shifted_sign;
+  if constexpr (S.hasSignedRepr) {
+    uint64_t shifted_sign = static_cast<uint64_t>(sign & 1)
+                            << ((S.sizeInBits - 1) % 64);
+    words[last_word] |= shifted_sign;
+  }
   uint64_t shifted_exponent = (myexponent & exponent_mask)
                               << (trailing_significand_bits % 64);
   words[last_word] |= shifted_exponent;
diff --git a/llvm/unittests/ADT/APFloatTest.cpp b/llvm/unittests/ADT/APFloatTest.cpp
index e161071e63af4..66147c5f9b46e 100644
--- a/llvm/unittests/ADT/APFloatTest.cpp
+++ b/llvm/unittests/ADT/APFloatTest.cpp
@@ -2494,23 +2494,27 @@ TEST(APFloatTest, ConvertLosesUnrepresentableSignAndZero) {
 
   for (const fltSemantics *Sem : NoSignSemantics) {
     // The magnitude converts exactly, so the sign is the whole of the loss.
-    APFloat test(-2.0);
+    APFloat test(-1.0);
     bool losesInfo = false;
     APFloat::opStatus status =
         test.convert(*Sem, APFloat::rmNearestTiesToEven, &losesInfo);
     EXPECT_TRUE(losesInfo);
     EXPECT_EQ(status, APFloat::opInexact);
     EXPECT_TRUE(test.isNegative());
-    EXPECT_EQ(-2.0, test.convertToDouble());
+    EXPECT_EQ(-1.0, test.convertToDouble());
+    APInt negBits = test.bitcastToAPInt();
 
     // The same magnitude without the sign has nothing to report.
-    test = APFloat(2.0);
+    test = APFloat(1.0);
     losesInfo = true;
     status = test.convert(*Sem, APFloat::rmNearestTiesToEven, &losesInfo);
     EXPECT_FALSE(losesInfo);
     EXPECT_EQ(status, APFloat::opOK);
     EXPECT_FALSE(test.isNegative());
-    EXPECT_EQ(2.0, test.convertToDouble());
+    EXPECT_EQ(1.0, test.convertToDouble());
+
+    // No sign bit exists, so the bits must match the positive magnitude.
+    EXPECT_EQ(test.bitcastToAPInt(), negBits);
   }
 
   // Float8E8M0FNU has no zero either, and substitutes 2^-127 for one. That

``````````

</details>


https://github.com/llvm/llvm-project/pull/223286


More information about the llvm-commits mailing list