[llvm-branch-commits] [mlir] [mlir] Migrate AMDGPU/ROCDL to targets, not chipset versions (PR #223563)
via llvm-branch-commits
llvm-branch-commits at lists.llvm.org
Mon Sep 14 16:51:22 PDT 2026
llvmorg-github-actions[bot] wrote:
<!--LLVM PR SUMMARY COMMENT-->
@llvm/pr-subscribers-mlir-llvm
@llvm/pr-subscribers-mlir-gpu
Author: Krzysztof Drewniak (krzysz00)
<details>
<summary>Changes</summary>
**migration tl;dr:** Replace usages of `amdgpu::Chipset` with `ROCDL::TargetInfo`, ideally move from `chipset=` to `arch=`. If you don't use upstream pipelines, call 'TargetInfo::migrateArchFeaturesToModuleFlags` at the appropriate location.
Further note: if you've got a build pipeline that's getting a `gfxXXX` name from something like `rocm_agent_enumerator`, using a full triple name like the ones you get from `rocminfo` is preferred.
`amdgpu::Chipset` was an awkward hack that was hard to keep up to date
with changes in the compiler/new architectures, and didn't properly
support generic targets (and has been strongly disfavored by the
compiler team).
This PR replaces `amdgpu::Chipset` with `ROCDL::TargetInfo`, a
structure that uses LLVM's TargetParser and the underlying LLVM
features tables to get the real nature of the target being compiled
for.
This also helps MLIR move to
new-style (`-mtriple=amdgpuX.YZ-amd-amdhsa`) over "old
style" (`-mtriple=amdgcn-amd-amdhsa -mcpu=gfxXYZ`) triples.
The utility structure is moved from AMDGPU to ROCDL, both because it's
tied to LLVM rather directly and because projects like Triton should
be able to use these feature tests without pulling in the AMDGPU
dialect and its memref dependencies.
This migration also fixes a few correctness issues:
- Atomic emulation was producing floating-point additions that don't
exist on gfx90c (even though it's "after" gfx90a) and gfx908's more
precise about where emulation is needed.
- gfx90c was also being handed a bare `s_barrier`, but it has no
hardware barrier back-off, so it needs the inline asm workaround.
- gfx950 won't allow xf32 MFMAs anymore.
- permlane_swap forms that don't exist on some architectures no
longer lower.
- gfx11.7 is now listed as an OCP FP8-having target.
Some checks still need to check for a generation (ex. when encoding
s_waitcnt or what the semantics of a WMMA are) go through an
`isGeneration(N)` method, which checks for having instructions from
generation N but not N+1.
(The barrier lowering has been reordered to account for gfx13 not
having, but also not needing, BackOffBarrier.)
This migration renames the `chipset` or `chip` options on most passes to `arch`, but keeps the old name for compatibility.
Please note that, because of changes to how LLVM handles xnack and sramecc (they aren't target features anymore), you need to set `rocdl.xnack` and `rocdl.sramecc` on the modules you're translating if you know the values of those modifiers. `TargetInfo::migrateArchFeaturesToModuleFlags(Operation *op)` will do this for you.
Chipset is deprecated instead of being removed so that folks have time
to migrate.
Pre-commit tests in #<!-- -->220104.
AI disclosure: I steered this, Claude wrote the code, I tried to clean
up the docs.
---
<sub>Stack created with <a href="https://github.com/github/gh-stack">GitHub Stacks CLI</a> • <a href="https://gh.io/stacks-feedback">Give Feedback 💬</a></sub>
---
Patch is 219.42 KiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/223563.diff
109 Files Affected:
- (modified) mlir/docs/ReleaseNotes.md (+33)
- (modified) mlir/include/mlir/Conversion/AMDGPUToROCDL/AMDGPUToROCDL.h (+2-2)
- (modified) mlir/include/mlir/Conversion/ArithToAMDGPU/ArithToAMDGPU.h (+2-2)
- (modified) mlir/include/mlir/Conversion/GPUToROCDL/GPUToROCDLPass.h (+2-5)
- (modified) mlir/include/mlir/Conversion/MathToROCDL/MathToROCDL.h (+4-4)
- (modified) mlir/include/mlir/Conversion/Passes.td (+31-15)
- (modified) mlir/include/mlir/Dialect/AMDGPU/Transforms/Passes.h (+2-2)
- (modified) mlir/include/mlir/Dialect/AMDGPU/Transforms/Passes.td (+10-6)
- (modified) mlir/include/mlir/Dialect/AMDGPU/Utils/Chipset.h (+12)
- (modified) mlir/include/mlir/Dialect/GPU/Pipelines/Passes.h (+14-15)
- (modified) mlir/include/mlir/Dialect/GPU/TransformOps/GPUTransformOps.td (+19-8)
- (modified) mlir/include/mlir/Dialect/GPU/Transforms/Passes.h (+7-8)
- (modified) mlir/include/mlir/Dialect/GPU/Transforms/Passes.td (+14-5)
- (modified) mlir/lib/Conversion/AMDGPUToROCDL/AMDGPUToROCDL.cpp (+316-323)
- (modified) mlir/lib/Conversion/ArithToAMDGPU/ArithToAMDGPU.cpp (+24-24)
- (modified) mlir/lib/Conversion/GPUToROCDL/LowerGpuOpsToROCDLOps.cpp (+30-30)
- (modified) mlir/lib/Conversion/MathToROCDL/MathToROCDL.cpp (+16-11)
- (modified) mlir/lib/Dialect/AMDGPU/Transforms/CMakeLists.txt (+1)
- (modified) mlir/lib/Dialect/AMDGPU/Transforms/EmulateAtomics.cpp (+46-40)
- (modified) mlir/lib/Dialect/GPU/Pipelines/GPUToROCDLPipeline.cpp (+8-6)
- (modified) mlir/lib/Dialect/GPU/TransformOps/GPUTransformOps.cpp (+27-18)
- (modified) mlir/lib/Dialect/GPU/Transforms/PromoteShuffleToAMDGPU.cpp (+3-5)
- (modified) mlir/lib/Dialect/GPU/Transforms/ROCDLAttachTarget.cpp (+87-4)
- (modified) mlir/lib/Dialect/GPU/Transforms/SubgroupReduceLowering.cpp (+18-16)
- (removed) mlir/test/Conversion/AMDGPUToROCDL/8-bit-floats-ocp-gfx1170.mlir (-23)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/8-bit-floats-ocp.mlir (+3-2)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/8-bit-floats.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/amdgpu-to-rocdl.mlir (+7-7)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/dot-gfx11.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/dot-gfx12.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/dot-gfx9.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/dot-invalid.mlir (+2-2)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/dpp.mlir (+4-3)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/gfx1250.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/global-prefetch.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/global_transpose_load.mlir (+3-3)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/lds-barrier-gfx90c.mlir (+3-2)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/load_lds-gfx950.mlir (+2-2)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/load_lds.mlir (+2-2)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/memory_counter_wait.mlir (+4-4)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/memory_counter_wait_tensor.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/memory_counter_wait_unsupported.mlir (+3-3)
- (added) mlir/test/Conversion/AMDGPUToROCDL/mfma-fp8-invalid.mlir (+21)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/mfma-gfx950.mlir (+1-12)
- (added) mlir/test/Conversion/AMDGPUToROCDL/mfma-reduce-precision-invalid.mlir (+23)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/mfma.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/packed-ext.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/packed-trunc-invalid.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/packed-trunc.mlir (+1-1)
- (added) mlir/test/Conversion/AMDGPUToROCDL/permlane-gfx1200-invalid.mlir (+21)
- (added) mlir/test/Conversion/AMDGPUToROCDL/permlane-gfx1250-invalid.mlir (+11)
- (added) mlir/test/Conversion/AMDGPUToROCDL/permlane-gfx1250.mlir (+12)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/permlane-var.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/permlane.mlir (+1-2)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/sparse-mfma-gfx950.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/sparse-mfma.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/swizzle.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/swmmac-gfx12.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/swmmac-gfx1250.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/transpose_load.mlir (+2-2)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/transpose_load_gfx1250.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/transpose_load_gfx1250_invalid.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/transpose_load_gfx950_invalid.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/transpose_load_reject.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/wmma-gfx11.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/wmma-gfx12.mlir (+1-1)
- (modified) mlir/test/Conversion/AMDGPUToROCDL/wmma-gfx1250.mlir (+1-1)
- (modified) mlir/test/Conversion/ArithToAMDGPU/16-bit-floats.mlir (+1-1)
- (modified) mlir/test/Conversion/ArithToAMDGPU/8-bit-float-saturation-ocp.mlir (+2-2)
- (modified) mlir/test/Conversion/ArithToAMDGPU/8-bit-float-saturation.mlir (+1-1)
- (modified) mlir/test/Conversion/ArithToAMDGPU/8-bit-floats-ocp.mlir (+3-3)
- (modified) mlir/test/Conversion/ArithToAMDGPU/8-bit-floats.mlir (+1-1)
- (added) mlir/test/Conversion/ArithToAMDGPU/deprecated-chipset-alias.mlir (+27)
- (modified) mlir/test/Conversion/ArithToAMDGPU/scaling-extf.mlir (+2-2)
- (modified) mlir/test/Conversion/ArithToAMDGPU/scaling-truncf-tensor.mlir (+1-1)
- (modified) mlir/test/Conversion/ArithToAMDGPU/scaling-truncf.mlir (+2-2)
- (modified) mlir/test/Conversion/GPUCommon/lower-global-id.mlir (+1-1)
- (modified) mlir/test/Conversion/GPUCommon/lower-memory-space-attrs.mlir (+1-1)
- (modified) mlir/test/Conversion/GPUCommon/memory-attrbution.mlir (+1-1)
- (modified) mlir/test/Conversion/GPUCommon/memref-arg-attrs.mlir (+1-1)
- (modified) mlir/test/Conversion/GPUCommon/memref-arg-noalias-attrs.mlir (+1-1)
- (modified) mlir/test/Conversion/GPUCommon/memref-arg-noalias-warning.mlir (+1-1)
- (modified) mlir/test/Conversion/GPUToROCDL/constant-address-space.mlir (+1-1)
- (modified) mlir/test/Conversion/GPUToROCDL/gpu-to-rocdl-barrier.mlir (+2-2)
- (modified) mlir/test/Conversion/GPUToROCDL/gpu-to-rocdl-barriers-gfx12.mlir (+1-1)
- (modified) mlir/test/Conversion/GPUToROCDL/gpu-to-rocdl-hip.mlir (+3-1)
- (modified) mlir/test/Conversion/GPUToROCDL/gpu-to-rocdl-invalid-ballot.mlir (+1-1)
- (modified) mlir/test/Conversion/GPUToROCDL/gpu-to-rocdl-invalid-dialect.mlir (+1-1)
- (modified) mlir/test/Conversion/GPUToROCDL/gpu-to-rocdl-invalid-named-barrier.mlir (+1-1)
- (modified) mlir/test/Conversion/GPUToROCDL/gpu-to-rocdl-named-barrier-non-const.mlir (+1-1)
- (modified) mlir/test/Conversion/GPUToROCDL/gpu-to-rocdl-opencl.mlir (+1-1)
- (modified) mlir/test/Conversion/GPUToROCDL/gpu-to-rocdl-subgroup-id.mlir (+2-2)
- (modified) mlir/test/Conversion/GPUToROCDL/gpu-to-rocdl.mlir (+3-3)
- (modified) mlir/test/Conversion/GPUToROCDL/memref.mlir (+2-2)
- (modified) mlir/test/Conversion/GPUToROCDL/private-func-no-c-interface.mlir (+1-1)
- (modified) mlir/test/Conversion/MathToROCDL/math-to-rocdl.mlir (+2-2)
- (modified) mlir/test/Dialect/AMDGPU/amdgpu-emulate-atomics.mlir (+49-24)
- (added) mlir/test/Dialect/GPU/promote-shuffle-amdgpu-invalid.mlir (+58)
- (modified) mlir/test/Dialect/GPU/promote-shuffle-amdgpu.mlir (+1-1)
- (modified) mlir/test/Dialect/LLVMIR/attach-targets.mlir (+1-1)
- (added) mlir/test/Dialect/LLVMIR/rocdl-attach-target-arch.mlir (+105)
- (modified) mlir/test/Integration/GPU/ROCM/gpu-lower-to-rocdl-pipeline.mlir (+1-1)
- (modified) mlir/test/Integration/GPU/ROCM/gpu-to-hsaco.mlir (+1-1)
- (modified) mlir/test/Integration/GPU/ROCM/printf.mlir (+1-1)
- (modified) mlir/test/Integration/GPU/ROCM/two-modules.mlir (+1-1)
- (modified) mlir/test/Integration/GPU/ROCM/vecadd.mlir (+1-1)
- (modified) mlir/test/Integration/GPU/ROCM/vector-transferops.mlir (+1-1)
- (modified) mlir/test/lib/Dialect/GPU/TestGpuRewrite.cpp (+8-5)
- (modified) mlir/unittests/Dialect/AMDGPU/AMDGPUUtilsTest.cpp (+8)
``````````diff
diff --git a/mlir/docs/ReleaseNotes.md b/mlir/docs/ReleaseNotes.md
index 16b93c8909670..b4650b82841b3 100644
--- a/mlir/docs/ReleaseNotes.md
+++ b/mlir/docs/ReleaseNotes.md
@@ -8,6 +8,39 @@ specifically, it is a snapshot of the MLIR development at the time of the releas
[TOC]
+## LLVM 24
+
+### GPU/AMDGPU Changes
+
+- `mlir::amdgpu::Chipset` is deprecated in favour of `mlir::ROCDL::TargetInfo`,
+ which describes a target by its triple, subarch, and the resolved set of
+ target features from LLVM's own tables. Lowerings should ask whether a target
+ has a feature rather than inaccurately compare chipset versions.
+ `TargetInfo` also represents generic targets such as `gfx9-4-generic` and, unlike
+ `Chipset`, explicitly stores the wavesize for targets where it is configurable.
+- The `chipset` option in AMDGPU passes is renamed to an `arch` option, which uses
+ Clang target naming syntax. It accepts a GPU name with optional
+ modifiers (`gfx942`, `gfx942:xnack+`, `gfx9-4-generic`), a triple
+ (`amdgpu9.42-amd-amdhsa`), or a full target ID
+ (`amdgpu9.42-amd-amdhsa--gfx90a:sramecc+:xnack-`, which is what `rocminfo` prints
+ for a device's ISA). `chipset` or `chip` remain as compatibility names.
+ The default arch is `invalid`, so a target must be passed
+ explicitly, removing the old "fallback" `gfx000` GPU.
+- Wavefront size is not a target-ID feature, so `convert-gpu-to-rocdl` takes it
+ as a separate `wavesize` option (32, 64, or 0 for the architecture's
+ default). The `wave64` flag on `gpu-lower-to-rocdl-pipeline` and on
+ `rocdl-attach-target` is likewise replaced by the same `wavesize` option,
+ which has the same allowed values.
+- In keeping with broader LLVM changes, `xnack` and `sramecc` are no longer
+ architecture features but module flags. In keeping with Clang, the target
+ specifier still includes these xnack/sramecc flags where they're configurable,
+ but lowering passes now convert these to module flags. Downstream users should
+ call `migrateArchFeaturesToModuleFlags` to lower these attributes in custom
+ pipelines.
+- `rocdl-attach-target` gains `arch` alongside its existing `triple`, `chip` and
+ `features`. When `arch` is given, it overrides `triple` and `chip`, and handles
+ xnack/sramecc modifier migration.
+
## LLVM 21
### GPU/NVVM Changes
diff --git a/mlir/include/mlir/Conversion/AMDGPUToROCDL/AMDGPUToROCDL.h b/mlir/include/mlir/Conversion/AMDGPUToROCDL/AMDGPUToROCDL.h
index 393658652dbac..861bbe0cc8f75 100644
--- a/mlir/include/mlir/Conversion/AMDGPUToROCDL/AMDGPUToROCDL.h
+++ b/mlir/include/mlir/Conversion/AMDGPUToROCDL/AMDGPUToROCDL.h
@@ -8,7 +8,7 @@
#ifndef MLIR_CONVERSION_AMDGPUTOROCDL_AMDGPUTOROCDL_H_
#define MLIR_CONVERSION_AMDGPUTOROCDL_AMDGPUTOROCDL_H_
-#include "mlir/Dialect/AMDGPU/Utils/Chipset.h"
+#include "mlir/Dialect/LLVMIR/ROCDLTargetInfo.h"
#include <memory>
#include <string>
@@ -27,7 +27,7 @@ class Pass;
/// populateAMDGPUTypeAndAttributeConversions().
void populateAMDGPUToROCDLConversionPatterns(LLVMTypeConverter &converter,
RewritePatternSet &patterns,
- amdgpu::Chipset chipset);
+ const ROCDL::TargetInfo &target);
namespace amdgpu {
/// Remap common GPU memory spaces (Workgroup, Private, etc) to LLVM address
diff --git a/mlir/include/mlir/Conversion/ArithToAMDGPU/ArithToAMDGPU.h b/mlir/include/mlir/Conversion/ArithToAMDGPU/ArithToAMDGPU.h
index fd144edf77452..e6e137759a76a 100644
--- a/mlir/include/mlir/Conversion/ArithToAMDGPU/ArithToAMDGPU.h
+++ b/mlir/include/mlir/Conversion/ArithToAMDGPU/ArithToAMDGPU.h
@@ -9,7 +9,7 @@
#ifndef MLIR_CONVERSION_ARITHTOAMDGPU_ARITHTOAMDGPU_H
#define MLIR_CONVERSION_ARITHTOAMDGPU_ARITHTOAMDGPU_H
-#include "mlir/Dialect/AMDGPU/Utils/Chipset.h"
+#include "mlir/Dialect/LLVMIR/ROCDLTargetInfo.h"
#include "mlir/IR/PatternMatch.h"
#include <memory>
#include <string>
@@ -31,7 +31,7 @@ namespace arith {
void populateArithToAMDGPUConversionPatterns(
RewritePatternSet &patterns, bool convertFP8Arithmetic,
bool saturateFP8Truncf, bool allowPackedF16Rtz, bool supportsScaledExtTrunc,
- amdgpu::Chipset chipset, PatternBenefit benefit = 1);
+ const ROCDL::TargetInfo &target, PatternBenefit benefit = 1);
} // namespace arith
} // namespace mlir
diff --git a/mlir/include/mlir/Conversion/GPUToROCDL/GPUToROCDLPass.h b/mlir/include/mlir/Conversion/GPUToROCDL/GPUToROCDLPass.h
index 220da0ad3c08f..494ee8cfa11ab 100644
--- a/mlir/include/mlir/Conversion/GPUToROCDL/GPUToROCDLPass.h
+++ b/mlir/include/mlir/Conversion/GPUToROCDL/GPUToROCDLPass.h
@@ -10,6 +10,7 @@
#include "mlir/Conversion/GPUToROCDL/Runtimes.h"
#include "mlir/Conversion/LLVMCommon/LoweringOptions.h"
+#include "mlir/Dialect/LLVMIR/ROCDLTargetInfo.h"
#include <memory>
namespace mlir {
@@ -21,10 +22,6 @@ class RewritePatternSet;
template <typename OpT>
class OperationPass;
-namespace amdgpu {
-struct Chipset;
-} // namespace amdgpu
-
namespace gpu {
class GPUModuleOp;
} // namespace gpu
@@ -38,7 +35,7 @@ class GPUModuleOp;
void populateGpuToROCDLConversionPatterns(const LLVMTypeConverter &converter,
RewritePatternSet &patterns,
gpu::amd::Runtime runtime,
- amdgpu::Chipset chipset);
+ const ROCDL::TargetInfo &target);
/// Configure target to convert from the GPU dialect to ROCDL.
void configureGpuToROCDLConversionLegality(ConversionTarget &target);
diff --git a/mlir/include/mlir/Conversion/MathToROCDL/MathToROCDL.h b/mlir/include/mlir/Conversion/MathToROCDL/MathToROCDL.h
index 60f1888569362..8ba104972abff 100644
--- a/mlir/include/mlir/Conversion/MathToROCDL/MathToROCDL.h
+++ b/mlir/include/mlir/Conversion/MathToROCDL/MathToROCDL.h
@@ -9,7 +9,7 @@
#define MLIR_CONVERSION_MATHTOROCDL_MATHTOROCDL_H_
#include "mlir/Conversion/LLVMCommon/TypeConverter.h"
-#include "mlir/Dialect/AMDGPU/Utils/Chipset.h"
+#include "mlir/Dialect/LLVMIR/ROCDLTargetInfo.h"
#include "mlir/IR/PatternMatch.h"
#include <memory>
@@ -20,11 +20,11 @@ class Pass;
#include "mlir/Conversion/Passes.h.inc"
/// Populate the given list with patterns that convert from Math to ROCDL calls.
-// `chipset` specifies the AMDGPU chipset to target. If `std::nullopt`,
-// none of the chipset dependent patterns are added.
+// `target` describes the AMDGPU target. If `std::nullopt`, none of the
+// target-dependent patterns are added.
void populateMathToROCDLConversionPatterns(
const LLVMTypeConverter &converter, RewritePatternSet &patterns,
- std::optional<amdgpu::Chipset> chipset);
+ std::optional<ROCDL::TargetInfo> target);
} // namespace mlir
#endif // MLIR_CONVERSION_MATHTOROCDL_MATHTOROCDL_H_
diff --git a/mlir/include/mlir/Conversion/Passes.td b/mlir/include/mlir/Conversion/Passes.td
index c80e1753ea2a9..5291f9d683016 100644
--- a/mlir/include/mlir/Conversion/Passes.td
+++ b/mlir/include/mlir/Conversion/Passes.td
@@ -130,9 +130,13 @@ def ConvertAMDGPUToROCDLPass : Pass<"convert-amdgpu-to-rocdl"> {
"LLVM::LLVMDialect",
"ROCDL::ROCDLDialect",
];
- let options = [Option<"chipset", "chipset", "std::string",
- /*default=*/"\"gfx000\"",
- "Chipset that these operations will run on">];
+ let options = [
+ Option<"arch", "arch", "std::string",
+ /*default=*/"\"invalid\"",
+ "Target architecture, as in Clang, with optional target-ID modifiers. New-style triples such as amdgpu9.42-amd-amdhsa are preferred, and can be extended to a full target ID like amdgpu9.42-amd-amdhsa--gfx942:xnack+:sramecc-. Bare chip names like gfx1250 or gfx942:xnack- are also supported. Defaults to an invalid target so that one must be given explicitly">,
+ Option<"chipset", "chipset", "std::string", /*default=*/"\"\"",
+ "Deprecated alias for 'arch'.">,
+ ];
}
//===----------------------------------------------------------------------===//
@@ -150,9 +154,11 @@ def ArithToAMDGPUConversionPass : Pass<"convert-arith-to-amdgpu"> {
let dependentDialects = ["amdgpu::AMDGPUDialect", "vector::VectorDialect"];
let options = [
- Option<"chipset", "chipset", "std::string",
- /*default=*/"\"gfx000\"",
- "Chipset that these operations will run on">,
+ Option<"arch", "arch", "std::string",
+ /*default=*/"\"invalid\"",
+ "Target architecture, as in Clang, with optional target-ID modifiers. New-style triples such as amdgpu9.42-amd-amdhsa are preferred, and can be extended to a full target ID like amdgpu9.42-amd-amdhsa--gfx942:xnack+:sramecc-. Bare chip names like gfx1250 or gfx942:xnack- are also supported. Defaults to an invalid target so that one must be given explicitly">,
+ Option<"chipset", "chipset", "std::string", /*default=*/"\"\"",
+ "Deprecated alias for 'arch'.">,
Option<"saturateFP8Truncf", "saturate-fp8-truncf", "bool",
/*default=*/"false",
"Use saturating truncation for 8-bit float types">,
@@ -684,9 +690,14 @@ def ConvertGpuOpsToROCDLOps : Pass<"convert-gpu-to-rocdl", "gpu::GPUModuleOp"> {
"memref::MemRefDialect",
];
let options = [
- Option<"chipset", "chipset", "std::string",
- /*default=*/"\"gfx000\"",
- "Chipset that these operations will run on">,
+ Option<"arch", "arch", "std::string",
+ /*default=*/"\"invalid\"",
+ "Target architecture, as in Clang, with optional target-ID modifiers. New-style triples such as amdgpu9.42-amd-amdhsa are preferred, and can be extended to a full target ID like amdgpu9.42-amd-amdhsa--gfx942:xnack+:sramecc-. Bare chip names like gfx1250 or gfx942:xnack- are also supported. Defaults to an invalid target so that one must be given explicitly">,
+ Option<"chipset", "chipset", "std::string", /*default=*/"\"\"",
+ "Deprecated alias for 'arch'.">,
+ Option<"waveSize", "wavesize", "unsigned", /*default=*/"0",
+ "Wavefront size (32 or 64) for targets that run at either, or 0 to "
+ "use the architecture's default">,
Option<"indexBitwidth", "index-bitwidth", "unsigned",
/*default=kDeriveIndexBitwidthFromDataLayout*/ "0",
"Bitwidth of the index type, 0 to use size of machine word">,
@@ -851,9 +862,9 @@ def ConvertMathToROCDL : Pass<"convert-math-to-rocdl", "ModuleOp"> {
let description = [{
This pass converts supported Math ops to ROCDL library calls.
- The chipset option specifies the target AMDGPU architecture. If the chipset
- is empty, none of the chipset-dependent patterns are added, and the pass
- will not attempt to parse the chipset.
+ The triple option specifies the target AMDGPU architecture. If it is empty,
+ none of the target-dependent patterns are added and the pass does not
+ resolve a target; a chip or feature list without a triple is rejected.
}];
let dependentDialects = [
"arith::ArithDialect",
@@ -861,9 +872,14 @@ def ConvertMathToROCDL : Pass<"convert-math-to-rocdl", "ModuleOp"> {
"ROCDL::ROCDLDialect",
"vector::VectorDialect",
];
- let options = [Option<"chipset", "chipset", "std::string",
- /*default=*/"\"\"",
- "Chipset that these operations will run on">];
+ let options = [
+ Option<"arch", "arch", "std::string", /*default=*/"\"\"",
+ "Target architecture, as in Clang, with optional target-ID modifiers "
+ "(e.g. amdgpu9.42-amd-amdhsa, gfx942, gfx942:xnack-). "
+ "If empty, no target-dependent patterns are added">,
+ Option<"chipset", "chipset", "std::string", /*default=*/"\"\"",
+ "Deprecated alias for 'arch'.">,
+ ];
}
//===----------------------------------------------------------------------===//
diff --git a/mlir/include/mlir/Dialect/AMDGPU/Transforms/Passes.h b/mlir/include/mlir/Dialect/AMDGPU/Transforms/Passes.h
index 48e7658568f86..c8ec368231f83 100644
--- a/mlir/include/mlir/Dialect/AMDGPU/Transforms/Passes.h
+++ b/mlir/include/mlir/Dialect/AMDGPU/Transforms/Passes.h
@@ -13,7 +13,7 @@
#ifndef MLIR_DIALECT_AMDGPU_TRANSFORMS_PASSES_H_
#define MLIR_DIALECT_AMDGPU_TRANSFORMS_PASSES_H_
-#include "mlir/Dialect/AMDGPU/Utils/Chipset.h"
+#include "mlir/Dialect/LLVMIR/ROCDLTargetInfo.h"
#include "mlir/IR/PatternMatch.h"
#include "mlir/Pass/Pass.h"
@@ -29,7 +29,7 @@ namespace amdgpu {
void populateAmdgpuEmulateAtomicsPatterns(ConversionTarget &target,
RewritePatternSet &patterns,
- Chipset chipset,
+ const ROCDL::TargetInfo &targetInfo,
PatternBenefit benefit = 1);
void populateAmdgpuResolveStridedMetadataPatterns(RewritePatternSet &patterns,
diff --git a/mlir/include/mlir/Dialect/AMDGPU/Transforms/Passes.td b/mlir/include/mlir/Dialect/AMDGPU/Transforms/Passes.td
index 7dd7ac750a9eb..d9e68a7be4dc7 100644
--- a/mlir/include/mlir/Dialect/AMDGPU/Transforms/Passes.td
+++ b/mlir/include/mlir/Dialect/AMDGPU/Transforms/Passes.td
@@ -16,19 +16,23 @@
include "mlir/Pass/PassBase.td"
def AmdgpuEmulateAtomicsPass : Pass<"amdgpu-emulate-atomics"> {
- let summary = "Emulate atomic operations on chipsets that do not support them";
+ let summary = "Emulate atomic operations the target does not support";
let description = [{
- This pass rewrites any AMDGPU-specific atomic operation that is not supported
- on the given `chipset` into a compare-and-swap loop.
+ This pass rewrites any AMDGPU-specific atomic operation that the target does
+ not support into a compare-and-swap loop.
}];
let dependentDialects = [
"cf::ControlFlowDialect",
"arith::ArithDialect",
"vector::VectorDialect"
];
- let options = [Option<"chipset", "chipset", "std::string",
- /*default=*/"\"gfx000\"",
- "Chipset that these operations will run on">];
+ let options = [
+ Option<"arch", "arch", "std::string",
+ /*default=*/"\"invalid\"",
+ "Target architecture, as in Clang, with optional target-ID modifiers. New-style triples such as amdgpu9.42-amd-amdhsa are preferred, and can be extended to a full target ID like amdgpu9.42-amd-amdhsa--gfx942:xnack+:sramecc-. Bare chip names like gfx1250 or gfx942:xnack- are also supported. Defaults to an invalid target so that one must be given explicitly">,
+ Option<"chipset", "chipset", "std::string", /*default=*/"\"\"",
+ "Deprecated alias for 'arch'.">,
+ ];
}
def AmdgpuResolveStridedMetadataPass : Pass<"amdgpu-resolve-strided-metadata"> {
diff --git a/mlir/include/mlir/Dialect/AMDGPU/Utils/Chipset.h b/mlir/include/mlir/Dialect/AMDGPU/Utils/Chipset.h
index 256065537aeb9..1a964c27a6333 100644
--- a/mlir/include/mlir/Dialect/AMDGPU/Utils/Chipset.h
+++ b/mlir/include/mlir/Dialect/AMDGPU/Utils/Chipset.h
@@ -9,6 +9,7 @@
#define MLIR_DIALECT_AMDGPU_UTILS_CHIPSET_H_
#include "mlir/Support/LLVM.h"
+#include "llvm/Support/Compiler.h"
#include <tuple>
namespace mlir::amdgpu {
@@ -19,6 +20,10 @@ namespace mlir::amdgpu {
/// gfx942 --> major = 9, minor = 0x4, stepping = 0x2
/// gfx90a --> major = 9, minor = 0x0, stepping = 0xa
/// gfx1103 --> major = 11, minor = 0x0, stepping = 0x3
+///
+/// \deprecated Use `mlir::ROCDL::TargetInfo` instead, and rely on target
+/// features rather than about version numbers. Will be removed after one
+/// release.
struct Chipset {
unsigned majorVersion = 0; // The major version (decimal).
unsigned minorVersion = 0; // The minor version (hexadecimal).
@@ -30,6 +35,9 @@ struct Chipset {
/// Parses the chipset version string and returns the chipset on success, and
/// failure otherwise.
+ ///
+ /// \deprecated Use `ROCDL::TargetInfo::get`.
+ LLVM_DEPRECATED("use ROCDL::TargetInfo::get instead", "")
static FailureOr<Chipset> parse(StringRef name);
std::tuple<unsigned, unsigned, unsigned> asTuple() const {
@@ -49,6 +57,10 @@ struct Chipset {
#undef DEFINE_COMP_OPERATOR
};
+/// \deprecated Test `llvm::AMDGPU::FEAT_OCP_FP8_CONVERSION_INSTS` on a
+/// `ROCDL::TargetInfo` instead. This misses gfx11.7, which does have the OCP
+/// fp8 conversions.
+LLVM_DEPRECATED("test FEAT_OCP_FP8_CONVERSION_INSTS on a ROCDL::TargetInfo", "")
inline bool hasOcpFp8(const Chipset &chipset) {
return (chipset.majorVersion == 9 && chipset.minorVersion >= 5) ||
chipset.majorVersion >= 12;
diff --git a/mlir/include/mlir/Dialect/GPU/Pipelines/Passes.h b/mlir/include/mlir/Dialect/GPU/Pipelines/Passes.h
index 6327415d62769..c34cc0ee9fdd6 100644
--- a/mlir/include/mlir/Dialect/GPU/Pipelines/Passes.h
+++ b/mlir/include/mlir/Dialect/GPU/Pipelines/Passes.h
@@ -72,17 +72,16 @@ struct GPUToROCDLPipelineOptions
llvm::cl::desc("Bitwidth of the index type for the host (warning this "
"should be 64 until the GPU layering is fixed)"),
llvm::cl::init(64)};
- PassOptions::Option<std::string> triple{
- *this, "triple",
- llvm::cl::desc("AMDGPU target triple (e.g. amdgcn-amd-amdhsa)."),
- llvm::cl::init("amdgcn-amd-amdhsa")};
- PassOptions::Option<std::string> chip{
- *this, "chip",
+ PassOptions::Option<std::string> arch{
+ *this, "arch",
llvm::cl::desc(
- "AMDGPU target chip (e.g. gfx90a, gfx942, gfx1100). Required: "
- "AMDGCN binaries are not forward-compatible across chip families.")};
- PassOptions::Option<std::string> features{
- *this, "features", llvm::cl::desc("AMDGPU target features."),
+ "AMDGPU target architecture, as in Clang, with optional target-ID "
+ "modifiers (e.g. gfx942, gfx90a:xnack+, "
+ "amdgpu9.0a-amd-amdhsa--gfx90a:xnack-). Required: AMDGCN binaries "
+ "are "
+ "not forward-compatible across chip families.")};
+ PassOptions::Option<std::string> chip{
+ *this, "chip", llvm::cl::desc("Deprecated alias for 'arch'."),
llvm::cl::init("")};
PassOptions::Option<std::string> binaryFormat{
*this, "binary-format",
@@ -93,11 +92,11 @@ struct GPUToROCDLPipelineOptions
*this, "abi",
llvm::cl::desc("AMDHSA ABI version (e.g. \"500\", \"600\")."),
llvm::cl::init("600")};
- PassOptions::Option<bool> wave64{
- *this, "wave64",
- llvm::cl::desc("Use Wave64 mode (default true; wave32 if false, "
- "appropriate for RDNA / gfx10+ where supported)."),
- llvm::cl::init(true)};
+ PassOptions::Option<unsigned> waveSize{
+ *this, "wavesize",
+ llvm::cl::desc("Wavefront size (32 or 64) for targets that run at "
+ "either, or 0 to use the architecture's default."),
+ llvm::cl::init(0)};
PassOptions::Option<int> optLevel{
*this, "opt-level",
llvm::cl::desc("Optimization level for ROCDL/AMDGPU compilation."),
diff --git a/mlir/include/mlir/Dialect/GPU/TransformOps/GPUTransformOps.td b/mlir/include/mlir/Dialect/GPU/TransformOps/GPUTransformOps.td
index 030506bba1076..d56cc1f13afd2 100644
--- a/mlir/include/mlir/Dialect/GPU/TransformOps/GPUTransformOps.td
+++ b/mlir/include/mlir/Dialect/GPU/TransformOps/GPUTransformOps.td
@@ -61,11 +61,18 @@ def ApplyGPUToROCDLConversionPatternsOp : Op<Transform_Dialect,
let description = [{
Collects patterns that convert GPU dialect ops to ROCDL dialect ops. These
patterns require an "LLVMTypeConverter".
+
+ `arch` names the target the way Clang does: a GPU name with optional
+ target-ID modifiers ("gfx942", "gfx942:xnack+", "gfx9-4-generic"), a triple
+ ("amdgpu9.42-amd-amdhsa"), or a full target ID
+ ("amdgpu9.0a-amd-amdhsa--gfx90a:sramecc+:xnack-").
+
+ `wavesize` pins the wavefront size on targets that run at either; omitting
+ it uses the architecture's own default.
}];
- let arguments = (ins StrAttr:$chipset);
- let assemblyFormat = [{
- `chipset` `=` $chipset attr-dict
- }];
+ let arguments = (ins StrAttr:$arch,
+ OptionalAttr<I32Attr>:$wavesize);
+ let assemblyFormat = "prop-dict attr-dict";
}
//===----------------------------------------------------------------------===//
@@ -328,11 +335,15 @...
[truncated]
``````````
</details>
https://github.com/llvm/llvm-project/pull/223563
More information about the llvm-branch-commits
mailing list