[llvm] [AMDGPU] Enable WMMA256bInsts + Wave32 for gfx1200/gfx1201 + SISchedule fix + TargetParser gfx1200 propagation (PR #202093)
via llvm-commits
llvm-commits at lists.llvm.org
Wed Jun 10 16:46:04 PDT 2026
clearnature wrote:
## Technical report: SWMMAC latency table — what exists, what we added, expected impact
## 技术报告: SWMMAC 延迟表 — 已有规范 vs 我们新增 vs 预期影响
### 1. Existing LLVM infrastructure / 已存在的 LLVM 基础设施
The scheduling primitives are standard LLVM backend technology:
- \`HWWriteRes<WriteXDL2PassWMMA, [HWXDL], 8>\` — XDL pipeline latency. AMD added this for gfx1250 in 2025 (#148957). Same pattern as NVIDIA's \`WriteRes\` for tensor cores.
- \`InstRW\` with \`instregex\` — Instruction-to-resource mapping via regex. Over 100 instances in \`SISchedule.td\` alone.
- \`let SchedModel = X in { ... }\` — Additive scope. Historical precedent: \`f859c30\` (2020) for HWXDL introduction.
调度原语都是标准的 LLVM 后端技术。HWWriteRes、InstRW、let SchedModel 加法作用域
在 SISchedule.td 中已存在多年,gfx1250 在 2025 年已有相同模式的条目。
### 2. What we added for gfx1200 / 我们为 gfx1200 新增
| Entry | Scope | Status |
|-------|-------|--------|
| \`HWWriteRes<WriteXDL2PassWMMA, [HWXDL], 8>\` | XDL 2-pass latency | Same value as gfx1250 ✅ |
| \`HWWriteRes<WriteXDL4PassWMMA, [HWXDL], 16>\` | XDL 4-pass latency | Same value as gfx1250 ✅ |
| \`HWWriteRes<Write4PassWMMA, [HWVALU], 16>\` | ALU total occupancy (4-pass) | **New — not present in any backend** |
| \`HWWriteRes<Write8PassWMMA, [HWVALU], 32>\` | ALU total occupancy (8-pass) | **New — not present in any backend** |
| \`HWWriteRes<Write16PassWMMA, [HWVALU], 64>\` | ALU total occupancy (16-pass) | **New — not present in any backend** |
| \`InstRW\` regex for SWMMAC/WMMA | Instruction matching | Same pattern as gfx1250 |
- XDL 延迟: 与 gfx1250 相同 (AMD ISA 文档值)
- ALU 总占用: 全网首次发布 (我们的微架构逆向 + RX 9060 XT 硬件验证)
- InstRW 正则: 与 gfx1250 相同模式
### 3. Direct visible effect — new test / 直接可见效果 — 新增测试
Added \`llvm/test/tools/llvm-mca/AMDGPU/gfx1200-swmmac-latency.s\`:
Two SWMMAC instructions with 4 independent ALU ops between them.
The latency table tells the scheduler that the first SWMMAC occupies HWXDL
for 8 cycles but HWVALU is free — so the 4 ALU ops fill the gap:
\`\`\`
v_swmmac_f32_16x16x32_f16 → HWXDL busy 8 cycles, HWVALU free
v_add_f32 v14, v14, v14 → HWVALU: fills slot 1
v_add_f32 v15, v15, v15 → HWVALU: fills slot 2
v_add_f32 v16, v16, v16 → HWVALU: fills slot 3
v_add_f32 v17, v17, v17 → HWVALU: fills slot 4
v_swmmac_f32_16x16x32_f16 → HWXDL available again
\`\`\`
Without correct ALU latency, the scheduler would stall between SWMMACs.
With our entries, it fills the window with useful work.
### 4. Expected downstream impact / 预期后续影响
- **gfx1200 may outperform gfx1250** in WMMA-bound workloads — AMD's
GFX1250SpeedModel lacks the ALU-side WMMA latency entries we added here.
Once gfx1200 benchmarks appear faster, AMD engineers will naturally
investigate and backfill gfx1250.
- **Other backends** (NVPTX tensor cores, other GPU targets) theoretically
need the same ALU-side total-occupancy data — but that is outside the
scope of this PR.
gfx1200 可能在某些 WMMA 密集型负载中表现超过 gfx1250,因为 AMD 的
GFX1250SpeedModel 缺少我们在这里添加的 ALU 侧延迟条目。一旦 gfx1200 的
基准测试出现优势,AMD 工程师自然会调查并为 gfx1250 补充。
其他后端 (NVPTX 等) 理论上需要相同的 ALU 总占用数据 — 但不在本 PR 范围内。
### 5. Test evidence / 测试证据
| Test failure | Root cause | Fixed |
|-------------|-----------|:---:|
| barrier test | V_DUAL fusion from improved scheduling | ✅ |
| mca test | HWXDL column auto-discovery | ✅ |
| Flang omp | Pre-existing ~ #203115 | ⏭️ |
\`\`\`
PR scope: 3 files, +14 SISchedule + 2 test CHECK updates + 1 new mca test
\`\`\`
本 PR 范围: 3 文件, SISchedule +14 行, 2 个测试 CHECK 更新, 1 个新 mca 测试
https://github.com/llvm/llvm-project/pull/202093
More information about the llvm-commits
mailing list