Enable host-native architecture tuning for GNU/Clang builds (fixes #209) - #257
manishpaulish wants to merge 3 commits into
Conversation
Only the Intel path set a host-arch flag (-xHost); GNU/Clang got none, so the build targeted baseline x86-64 and never emitted AVX in the gather/scatter kernels. Add SPATTER_ENABLE_NATIVE_ARCH (default ON: -march=native, or -mcpu=native fallback) and SPATTER_ARCH_FLAGS to override; skipped when cross-compiling. Signed-off-by: MANISH PAUL <manishpaul.24@kgpian.iitkgp.ac.in>
Signed-off-by: MANISH PAUL <manishpaul.24@kgpian.iitkgp.ac.in>
Signed-off-by: MANISH PAUL <manishpaul.24@kgpian.iitkgp.ac.in>
|
Hi @jyoung3131, friendly ping on this one. It's a small, self-contained change: it adds the GNU/Clang equivalent of the Intel Since I'm a first-time contributor, the CI workflows are still waiting on maintainer approval. Would you or another maintainer be able to approve them so the checks can run? Happy to adjust anything, or split the doc change out, if that's easier to review. Thanks! |
|
Quick follow-up: the CI workflows on this PR are still showing "awaiting approval", so there are no check results for @jyoung3131 or @plavin to review against. If one of you is able to approve the workflow run, the checks should give you a clearer signal on whether this is safe to merge. No rush on the review itself. Also happy to split the Build.md doc change into a separate PR if that would make this easier to evaluate. |
|
Hi Manish - thank you for your contribution. We will review and get back to you in the next week or so as folks are currently on on travel and vacation. |
Overview
Spatter is a memory microbenchmark, so its kernels need to be compiled for the ISA of the machine being measured. The Intel-compiler path already does this via
-xHost, but the GNU and Clang paths added no architecture flag at all, so the compiler targeted the baselinex86-64(SSE2) ISA and never emitted AVX/AVX2/AVX-512 in the gather/scatter kernels. This is the behavior reported in #209, where-mavxhad to be passed by hand to turnmovsdintovmovsd.Fixes #209.
✨ Change Description/Rationale
cmake/CompilerType.cmakenow adds the GNU/Clang equivalent of-xHost:SPATTER_ENABLE_NATIVE_ARCH(defaultON) adds-march=native, falling back to-mcpu=nativeon toolchains where-march=nativeis unsupported (e.g. Apple arm64). Detected withcheck_cxx_compiler_flag.SPATTER_ARCH_FLAGSlets you pin an explicit target instead, e.g.-DSPATTER_ARCH_FLAGS="-march=sapphirerapids". It takes precedence over native detection.CMAKE_CROSSCOMPILING), and can be turned off with-DSPATTER_ENABLE_NATIVE_ARCH=OFFfor portable/reproducible binaries.Build.mddocuments both options in the CMake Options table.The design mirrors the existing Intel
-xHosthandling and keeps the default "fast on the machine you're benchmarking," while giving CI and packagers an explicit escape hatch for reproducible builds.Testing
Enabling host-native architecture tuning: -march=native, and the flag appears on the kernel compile line (CXX_FLAGS = -march=native -O3 -DNDEBUG ... -fopenmp).-DSPATTER_ARCH_FLAGS=...) is used verbatim and native detection is skipped.-DSPATTER_ENABLE_NATIVE_ARCH=OFF) reproduces the previous baseline flags exactly — no behavior change for anyone who wants it off.cmake --buildsucceeds and the resulting binary runs correctly (./spatter -pUNIFORM:8:1 -l$((2**20))gives expected bandwidth output).To confirm the AVX codegen change on x86_64 (per #209):