[llvm] [APFloat] Fix sign bit corrupting exponent in unsigned FP8 formats (PR #223286)
via llvm-commits
llvm-commits at lists.llvm.org
Sun Sep 13 19:20:52 PDT 2026
llvmorg-github-actions[bot] wrote:
<!--LLVM PR SUMMARY COMMENT-->
@llvm/pr-subscribers-llvm-adt
Author: Cheng-Chieh Liu (jay87406)
<details>
<summary>Changes</summary>
`convertIEEEFloatToAPInt` unconditionally ORs the sign bit into the top bit of the encoded value. `Float8E8M0FNU` and `Float8E5M3FNU` have no sign bit (`hasSignedRepr = false`), so that top bit actually belongs to the exponent field. Encoding a value with sign set therefore corrupts the exponent, in one case colliding with the reserved NaN bit pattern.
Guard the OR with `hasSignedRepr`, matching the existing check in the decode path (`initFromIEEEAPInt`).
Extend `ConvertLosesUnrepresentableSignAndZero` to also check `bitcastToAPInt()` after dropping an unrepresentable sign, using `-1.0` instead of `-2.0` so the corruption is actually visible in the bits.
---
Full diff: https://github.com/llvm/llvm-project/pull/223286.diff
2 Files Affected:
- (modified) llvm/lib/Support/APFloat.cpp (+5-3)
- (modified) llvm/unittests/ADT/APFloatTest.cpp (+8-4)
``````````diff
diff --git a/llvm/lib/Support/APFloat.cpp b/llvm/lib/Support/APFloat.cpp
index 59a6ba3867d40..e824b00162d33 100644
--- a/llvm/lib/Support/APFloat.cpp
+++ b/llvm/lib/Support/APFloat.cpp
@@ -3569,9 +3569,11 @@ APInt IEEEFloat::convertIEEEFloatToAPInt() const {
}
std::fill(words_iter, words.end(), uint64_t{0});
constexpr size_t last_word = words.size() - 1;
- uint64_t shifted_sign = static_cast<uint64_t>(sign & 1)
- << ((S.sizeInBits - 1) % 64);
- words[last_word] |= shifted_sign;
+ if constexpr (S.hasSignedRepr) {
+ uint64_t shifted_sign = static_cast<uint64_t>(sign & 1)
+ << ((S.sizeInBits - 1) % 64);
+ words[last_word] |= shifted_sign;
+ }
uint64_t shifted_exponent = (myexponent & exponent_mask)
<< (trailing_significand_bits % 64);
words[last_word] |= shifted_exponent;
diff --git a/llvm/unittests/ADT/APFloatTest.cpp b/llvm/unittests/ADT/APFloatTest.cpp
index e161071e63af4..66147c5f9b46e 100644
--- a/llvm/unittests/ADT/APFloatTest.cpp
+++ b/llvm/unittests/ADT/APFloatTest.cpp
@@ -2494,23 +2494,27 @@ TEST(APFloatTest, ConvertLosesUnrepresentableSignAndZero) {
for (const fltSemantics *Sem : NoSignSemantics) {
// The magnitude converts exactly, so the sign is the whole of the loss.
- APFloat test(-2.0);
+ APFloat test(-1.0);
bool losesInfo = false;
APFloat::opStatus status =
test.convert(*Sem, APFloat::rmNearestTiesToEven, &losesInfo);
EXPECT_TRUE(losesInfo);
EXPECT_EQ(status, APFloat::opInexact);
EXPECT_TRUE(test.isNegative());
- EXPECT_EQ(-2.0, test.convertToDouble());
+ EXPECT_EQ(-1.0, test.convertToDouble());
+ APInt negBits = test.bitcastToAPInt();
// The same magnitude without the sign has nothing to report.
- test = APFloat(2.0);
+ test = APFloat(1.0);
losesInfo = true;
status = test.convert(*Sem, APFloat::rmNearestTiesToEven, &losesInfo);
EXPECT_FALSE(losesInfo);
EXPECT_EQ(status, APFloat::opOK);
EXPECT_FALSE(test.isNegative());
- EXPECT_EQ(2.0, test.convertToDouble());
+ EXPECT_EQ(1.0, test.convertToDouble());
+
+ // No sign bit exists, so the bits must match the positive magnitude.
+ EXPECT_EQ(test.bitcastToAPInt(), negBits);
}
// Float8E8M0FNU has no zero either, and substitutes 2^-127 for one. That
``````````
</details>
https://github.com/llvm/llvm-project/pull/223286
More information about the llvm-commits
mailing list