[llvm] [AArch64] C1-Premium SME definitions (PR #212482)
Asher Dobrescu via llvm-commits
llvm-commits at lists.llvm.org
Wed Aug 26 04:52:36 PDT 2026
https://github.com/Asher8118 requested changes to this pull request.
Thank you for all the changes on this patch. This is close to final now.
>We model loads/stores on the CME as having a dependency on the load/store units on the main processor. This is a design consideration as we want to illustrate that the main processor will offload SME instructions to the CME co-processor.
Thank you for the explanation, that makes sense.
> Duplicate, scalar form is showing incorrect throughput, it should be 3 - this didn't seem possible in tablegen. I can get it to express the throughput as 2, like other instructions in its group, but the definition of the PERMF pipeline makes 3 seemingly impossible. If you're happy with the compromise of 0.5 then that is the current change
That is fine for now. Could you please add a TODO comment to explain this in the scheduler model?
Some leftover discrepancies:
- Conditional extract operations, SIMD&FP scalar and vector forms, it is currently showing with latency 3 when it should be 4 and using pipeline `PERMF`.
- Similarly for Convert to floating point, 64b to float/double, 32b to single or half, 16b to half, it should have latency of 4 and use the `VXALU` pipeline
- Not a discrepancy, but I would like to see an example of duplicate instructions in `C1Premium-sve-instructions.s`, eg: `dup z0.s, w1`.
- Cache Prefetch Prediction Restriction by Context is still showing incorrect information. It shows as having throughput of 1 when it should be 2.
https://github.com/llvm/llvm-project/pull/212482
More information about the llvm-commits
mailing list