https://github.com/davemgreen approved this pull request. Other than using the vector cost to get the throughput in line with other operations, I think this LGTM. Thanks https://github.com/llvm/llvm-project/pull/202965