MrSidims wrote: But ultimately.... do we have perf numbers with and without the patch? Better not on ML related workloads (where occupancy might be set in stone), but on some HPC applications? https://github.com/llvm/llvm-project/pull/214624