[llvm] [APFloat] Fix sign bit corrupting exponent in unsigned FP8 formats (PR #223286)

Cheng-Chieh Liu via llvm-commits llvm-commits at lists.llvm.org
Mon Sep 21 02:38:44 PDT 2026


https://github.com/jay87406 updated https://github.com/llvm/llvm-project/pull/223286

>From 09c5609798e0513cb1f0e51700b3a56db8b7d96b Mon Sep 17 00:00:00 2001
From: Cheng-Chieh Liu <jay87406 at gmail.com>
Date: Mon, 21 Sep 2026 04:38:17 -0500
Subject: [PATCH] [APFloat] Fix sign bit corrupting exponent in unsigned FP8
 formats

convertIEEEFloatToAPInt unconditionally ORs the sign bit into the top
bit of the encoded value. Float8E8M0FNU and Float8E5M3FNU have no sign
bit (hasSignedRepr = false), so that top bit actually belongs to the
exponent field. Encoding a value with sign set therefore corrupts the
exponent, in one case colliding with the reserved NaN bit pattern.

Guard the OR with hasSignedRepr, matching the existing check in the
decode path (initFromIEEEAPInt).

Extend ConvertLosesUnrepresentableSignAndZero to also check
bitcastToAPInt() after dropping an unrepresentable sign, using -1.0
instead of -2.0 so the corruption is actually visible in the bits.
---
 llvm/lib/Support/APFloat.cpp       |  8 +++++---
 llvm/unittests/ADT/APFloatTest.cpp | 12 ++++++++----
 2 files changed, 13 insertions(+), 7 deletions(-)

diff --git a/llvm/lib/Support/APFloat.cpp b/llvm/lib/Support/APFloat.cpp
index 8a1ef696f9350..3d1932454761f 100644
--- a/llvm/lib/Support/APFloat.cpp
+++ b/llvm/lib/Support/APFloat.cpp
@@ -3591,9 +3591,11 @@ APInt IEEEFloat::convertIEEEFloatToAPInt() const {
   }
   std::fill(words_iter, words.end(), uint64_t{0});
   constexpr size_t last_word = words.size() - 1;
-  uint64_t shifted_sign = static_cast<uint64_t>(sign & 1)
-                          << ((S.sizeInBits - 1) % 64);
-  words[last_word] |= shifted_sign;
+  if constexpr (S.hasSignedRepr) {
+    uint64_t shifted_sign = static_cast<uint64_t>(sign & 1)
+                            << ((S.sizeInBits - 1) % 64);
+    words[last_word] |= shifted_sign;
+  }
   uint64_t shifted_exponent = (myexponent & exponent_mask)
                               << (trailing_significand_bits % 64);
   words[last_word] |= shifted_exponent;
diff --git a/llvm/unittests/ADT/APFloatTest.cpp b/llvm/unittests/ADT/APFloatTest.cpp
index 35e97c6b2d9ea..40a121232fd29 100644
--- a/llvm/unittests/ADT/APFloatTest.cpp
+++ b/llvm/unittests/ADT/APFloatTest.cpp
@@ -2522,23 +2522,27 @@ TEST(APFloatTest, ConvertLosesUnrepresentableSignAndZero) {
 
   for (const fltSemantics *Sem : NoSignSemantics) {
     // The magnitude converts exactly, so the sign is the whole of the loss.
-    APFloat test(-2.0);
+    APFloat test(-1.0);
     bool losesInfo = false;
     APFloat::opStatus status =
         test.convert(*Sem, APFloat::rmNearestTiesToEven, &losesInfo);
     EXPECT_TRUE(losesInfo);
     EXPECT_EQ(status, APFloat::opInexact);
     EXPECT_TRUE(test.isNegative());
-    EXPECT_EQ(-2.0, test.convertToDouble());
+    EXPECT_EQ(-1.0, test.convertToDouble());
+    APInt negBits = test.bitcastToAPInt();
 
     // The same magnitude without the sign has nothing to report.
-    test = APFloat(2.0);
+    test = APFloat(1.0);
     losesInfo = true;
     status = test.convert(*Sem, APFloat::rmNearestTiesToEven, &losesInfo);
     EXPECT_FALSE(losesInfo);
     EXPECT_EQ(status, APFloat::opOK);
     EXPECT_FALSE(test.isNegative());
-    EXPECT_EQ(2.0, test.convertToDouble());
+    EXPECT_EQ(1.0, test.convertToDouble());
+
+    // No sign bit exists, so the bits must match the positive magnitude.
+    EXPECT_EQ(test.bitcastToAPInt(), negBits);
   }
 
   // Float8E8M0FNU has no zero either, and substitutes 2^-127 for one. That



More information about the llvm-commits mailing list