Repository navigation
compiler: add early AVX2 SIMD128 multiversioning through Go APIs - #2748
zhouguangyuan0718 wants to merge 8 commits into
Conversation
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
There was a problem hiding this comment.
SIMD128 CPU-guarded function multiversioning
This is a carefully engineered, well-tested change. The dispatch guard is sound: the x86.avx2 query is constant-folded to true only inside the cloned .__llgo_fmv_avx2 variant, while the baseline entry keeps a genuine runtime AVX2() call before the musttail forward — so AVX2 code never runs on an unsupported CPU, and the independent AVX/FMA queries correctly stay dynamic (+avx2 does not imply +fma in LLVM's x86 model). The subtle hazards are handled correctly: deferred erasure of query calls avoids iterator invalidation, UsedIDs is seeded with every existing PC-site ID before hash-remapping so string substitution can't alias, getNumOperands() is cached before addOperand(), and idempotency / symbol-collision / non-amd64 no-op are all explicitly tested. The README additions and code comments accurately describe the behavior.
I found no correctness or security defects. The inline comments below are a compile-time performance concern plus two maintainability/robustness notes — none are blocking.
Reviewed: internal/llvmfmv/*, cl/simd_fmv.go, ssa/simd_fmv.go, cl/{compile,instr}.go, internal/build/build.go, dev/test_native_simd.sh, test/simd/*.
| for (auto &F : M) { | ||
| if (F.hasFnAttribute(Done) || !supported(F)) | ||
| continue; | ||
| Function *Query = F.isDeclaration() ? nullptr : guardQuery(F); |
There was a problem hiding this comment.
[P2] FMV runs a full-module instruction scan on every amd64 build
llvmfmv.Run is invoked for every package module during every build (internal/build/build.go:3288). On amd64, this discovery loop calls guardQuery(F) for every supported() non-declaration function, and guardQuery walks all of F's instructions (fmv.cpp:60) before the cheap hasFnAttribute(Entry) short-circuit on line 207 is consulted. Since supported() is true for essentially any ordinary function (scalar / pointer / ≤128-bit-vector signature), this turns the pass into an O(all instructions in the module) traversal even for packages that contain no SIMD code at all — the overwhelmingly common case.
Consider gating the per-function scan: a function can only be a root if the module declares at least one llgo.cpu.query function, so compute that once up front and skip guardQuery entirely otherwise. Relatedly, UsedIDs is populated by scanning all llgo.pcline metadata (lines 215-220) on every amd64 module even when Originals ends up empty; returning early when there are no candidates avoids that cost too.
| MDNode *Row = Info->getOperand(I); | ||
| if (Row->getNumOperands() < 6) | ||
| continue; | ||
| auto *Name = dyn_cast<MDString>(Row->getOperand(1)); |
There was a problem hiding this comment.
[P2] Metadata operand indices duplicated from ssa/funcinfo.go
cloneSourceInfo hardcodes the operand layout of the llgo.funcinfo (operand 1 = linker symbol, min 6 fields) and llgo.pcline (operand 1 = id, operand 2 = symbol, exactly 6 fields, line 102-103) metadata. That schema is defined independently in ssa/funcinfo.go (EmitFuncInfo/EmitPCLineInfo), with no shared constant or cross-reference. If the field order or count in funcinfo.go ever changes, this C++ silently reads the wrong fields or skips rows via the < 6 / != 6 guards and corrupts clone symbol / PC-site mappings — with no compile-time error. A comment in each file pointing at the other (and documenting the exact layout next to these index accesses) would make the coupling explicit for future maintainers.
| return nullptr; | ||
| char *Result = static_cast<char *>(std::malloc(Error.size() + 1)); | ||
| if (!Result) | ||
| std::abort(); |
There was a problem hiding this comment.
[P3] std::abort() on malloc failure turns a compile error into a crash
llgoRunSIMDFMV hard-aborts the whole compiler process if the small diagnostic-string allocation fails. This path only runs when there is already a recoverable error to report (e.g. a symbol collision), so an OOM here escalates a reportable build error into a process crash that skips cleanup. Returning a static, non-malloc sentinel the Go side can detect would keep the error recoverable. Minor robustness nit, not a correctness or security issue.
LLGo WebAssembly build benchmarks
WebAssembly output sizes
LLGo WebAssembly build measurements
Compared with |
LLGo baseline benchmarks
Program measurements
Core language and compiler benchmarks
Timer runtime benchmarks
Compared with |
An
archsimd.X86.AVX2()guard currently runs the same baseline SIMD128 implementation even when AVX2 is available. Add early function multiversioning implemented in Go through the LLVM Go API, before aggregate ABI lowering and target optimization, including at O0 and without an LTO plugin.The public entry loads an atomic cached implementation and tail-forwards to it. A separate noinline resolver selects the baseline or AVX2 body from the immutable post-
GODEBUGCPU snapshot. Calls before CPU initialization tail-forward to the baseline without caching. GOAMD64=v3/v4 fold the AVX2 query directly and create no redundant versions, resolver, or slot.FMV inlining policy follows GoALLC: public dispatchers and implementations remain eligible for normal, feature-compatible LLVM inlining; the resolver and only its pre-initialization fallback edge are opaque. Explicit source noinline directives and disabled inlining remain honored. Compiler-required physical source frames stay on implementations. The synthetic dispatcher has no debug subprogram, so LLVM preserves the real callsite and outer inline chain without adding a duplicate source frame. LLGo's general runtime inline-frame infrastructure remains separate from GoALLC's GoObj inline tree.
Specialized direct calls use matching entries within and across packages. Function addresses retain the original entry; bodyless assembly/linkname declarations do not promise specialized implementations. Both implementations fold the AVX2 query while independent AVX/FMA query results remain observable under
GODEBUG. Clones retain Go display names, distinct PC-line records, and their own debug linkage names.Windows/amd64 SIMD128 parameters use the native indirect ABI explicitly before FMV. Dispatchers and resolvers forward caller-owned parameter storage through tail jumps instead of passing vector temporaries in a released stack frame; vector returns retain their register ABI. Windows function-entry records are grouped by PE unwind boundaries so inline copies cannot overwrite the physical caller identity. Failed native SIMD CI binaries are retained for diagnosis.
All FMV policy, discovery, cloning decisions, call rewriting, dispatcher construction, and Go metadata remapping are in Go. LLVM's generic cloning and CFG utilities preserve SSA/debug references and remove dead branches even at O0. LLGo has no FMV C++ implementation or private C entry point.
Scope: amd64, one AVX2 profile, and scalar/pointer/SIMD128 signatures. Aggregate signatures, general feature combinations, portable
simd, and 256/512-bit vector ABIs remain follow-up work.Dependency: xgo-dev/llvm#56 supplies generic IR utility bindings with no FMV policy. Its CI matrix covers LLVM 14–22; the PR remains open. This branch pins the exact public fork commit through
go.mod; replace that pin when the APIs become available upstream. There is no local-path dependency.Validation:
cpu.all=offFMV execution; macOS/arm64 O0/O2 SIMD execution; current Linux/amd64 and macOS/arm64 stage-2 self-host builds and SIMD execution.The dependency pin also includes binding review fixes at
f6e797e: checked IR-kind conversions, recoverable invalid-input panics, enum assertions, and explicit borrowed inline-assembly strings. The updated pin passes the full FMV unit suite, Linux/macOS/Windows O0/O2 IR and assembly checks, and macOS/arm64 stage-2 self-host build and SIMD execution.Part of #2568.