[llvm] [X86] Faster truncf and roundf on x86 SSE2 (PR #226513)
Divyansh Yadav via llvm-commits
llvm-commits at lists.llvm.org
Fri Sep 25 08:04:16 PDT 2026
schizophrenicmaniac wrote:
### Performance Improvements
By removing the reliance on soft-float library calls for `truncf` and `roundf` on pre-SSE4.1 hardware, the following observations can be made about the improvements in the compiler's output:
1. **Instruction Overhead Elimination:**
Previously, truncation loops generated a `calll _truncf` per scalar iteration. This forced arguments to the stack, flushed pipelines, and hindered optimizations. The new lowering reduces this to a 5-instruction branchless inline sequence (`cvttps2dq` -> `cvtdq2ps` -> `andps` -> `pcmpgtd` -> `orps`), completely removing call overhead.
2. **Cost Model & Vectorization (SLP/Loop):**
Prior to this PR, `ISD::FTRUNC` and `ISD::FROUND` were considered "Expensive" library calls (Cost: ~10). With `Custom` lowering, the TargetTransformInfo (TTI) now correctly recognizes these operations as "Legal" (Cost: 1-4). As verified by the updated test `Transforms/SLPVectorizer/X86/fround.ll`, this directly unlocks SLP Vectorization for loops containing these operations, allowing 4 scalar `float` truncations to be folded into a single SIMD `v4f32` operation.
https://github.com/llvm/llvm-project/pull/226513
More information about the llvm-commits
mailing list