[llvm] [WebAssembly] Fix lowering of (extending) loads from addrspace(1) globals (PR #155937)
Demetrius Kanios via llvm-commits
llvm-commits at lists.llvm.org
Thu Apr 9 13:49:25 PDT 2026
================
@@ -1850,16 +1879,94 @@ SDValue WebAssemblyTargetLowering::LowerLoad(SDValue Op,
LoadSDNode *LN = cast<LoadSDNode>(Op.getNode());
const SDValue &Base = LN->getBasePtr();
const SDValue &Offset = LN->getOffset();
+ ISD::LoadExtType ExtType = LN->getExtensionType();
+ EVT ResultType = LN->getValueType(0);
if (IsWebAssemblyGlobal(Base)) {
if (!Offset->isUndef())
report_fatal_error(
"unexpected offset when loading from webassembly global", false);
- SDVTList Tys = DAG.getVTList(LN->getValueType(0), MVT::Other);
- SDValue Ops[] = {LN->getChain(), Base};
- return DAG.getMemIntrinsicNode(WebAssemblyISD::GLOBAL_GET, DL, Tys, Ops,
- LN->getMemoryVT(), LN->getMemOperand());
+ if (!ResultType.isInteger() && !ResultType.isFloatingPoint()) {
+ SDVTList Tys = DAG.getVTList(ResultType, MVT::Other);
+ SDValue Ops[] = {LN->getChain(), Base};
+ SDValue GlobalGetNode =
+ DAG.getMemIntrinsicNode(WebAssemblyISD::GLOBAL_GET, DL, Tys, Ops,
+ LN->getMemoryVT(), LN->getMemOperand());
+ return GlobalGetNode;
+ }
+
+ EVT GT = MVT::INVALID_SIMPLE_VALUE_TYPE;
+
+ if (auto *GA = dyn_cast<GlobalAddressSDNode>(
+ Base->getOpcode() == WebAssemblyISD::Wrapper ? Base->getOperand(0)
+ : Base))
+ GT = EVT::getEVT(GA->getGlobal()->getValueType());
+
+ if (GT != MVT::i8 && GT != MVT::i16 && GT != MVT::i32 && GT != MVT::i64 &&
+ GT != MVT::f32 && GT != MVT::f64)
+ report_fatal_error("encountered unexpected global type for Base when "
+ "loading from webassembly global",
+ false);
+
+ EVT PromotedGT = getTypeToTransformTo(*DAG.getContext(), GT);
+
+ switch (ExtType) {
+ case ISD::NON_EXTLOAD: {
+ // A normal, non-extending load may try to load more or less than the
+ // underlying global, which is invalid. We lower this to a load of the
+ // global (i32 or i64) then truncate or extend as needed
+
+ // Modify the MMO to load the full global
+ // This is assumed to be safe without copy/dup, as the original load will
+ // be removed
+ MachineMemOperand *MMO = LN->getMemOperand();
+ MMO->setType(LLT(PromotedGT.getSimpleVT()));
+
+ SDVTList Tys = DAG.getVTList(PromotedGT, MVT::Other);
+ SDValue Ops[] = {LN->getChain(), Base};
+ SDValue GlobalGetNode = DAG.getMemIntrinsicNode(
+ WebAssemblyISD::GLOBAL_GET, DL, Tys, Ops, PromotedGT, MMO);
+
+ if (ResultType.bitsEq(PromotedGT)) {
+ return GlobalGetNode;
+ }
+
+ SDValue ValRes;
+ if (ResultType.isFloatingPoint())
+ ValRes = DAG.getFPExtendOrRound(GlobalGetNode, DL, ResultType);
----------------
QuantumSegfault wrote:
Pretty much; you have the general idea.
In regards to need for truncate/extend, sort of, but not quite.
Basic (non-extending loads) treat both bits outside the loaded value, and outside the declared IR global, as undefined/garbage. If I load i16 from i32 IR global, it boils down to a single `global.get` (as something else will eventually extend LLVM i16 => i32 if the upper bits are needed, so this is pretty usual/normal). Loading i16 from i64 loads the Wasm i64 global, and truncates (wraps) to Wasm i32. Loading i16 from i8 IR global results in the same operations as i16 from i32, but results is VISIBLE (to LLVM) undefined behavior because you are loading outside the defined bounds of the IR global, meaning LLVM makes no guarentees that the upper 24-bits of the underlying i32 Wasm global are going to be e.g. zeroed out upon store, and you are read out of bounds garbage.
Loading i64 from i32 fills the undefined upper 32-bits with zero (load the i32 global, extend_u to i64).
Extending loads are pretty literal, since there are no special instructions. Takes the result of the LOAD as described before, and extends it (EXTLOAD/anyext defaults to zext).
-------
Basically, I'm sort of emulating linear memory, but within the constraints of Wasm globals.
With floats the question becomes what to do when loading f32 from f64 (or vice versa). Do we treat it as floating point ops truncate and extend, or do we treat the global as a bag of bits and extract the lower 32-bits (or insert zeros for the upper) and interpret it as a float.
Bag of bits makes more sense to match the expected behavior of the load, even if doing so is ill-advised (why would you ever try to load f32 from f64 global????).
https://github.com/llvm/llvm-project/pull/155937
More information about the llvm-commits
mailing list