Skip to content

compiler: add early AVX2 SIMD128 multiversioning through Go APIs - #2748

Open
zhouguangyuan0718 wants to merge 8 commits into
xgo-dev:mainfrom
zhouguangyuan0718:codex/simd-fmv-20261007
Open

zhouguangyuan0718 wants to merge 8 commits into
xgo-dev:mainfrom
zhouguangyuan0718:codex/simd-fmv-20261007

Conversation

@zhouguangyuan0718

@zhouguangyuan0718 zhouguangyuan0718 commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

An archsimd.X86.AVX2() guard currently runs the same baseline SIMD128 implementation even when AVX2 is available. Add early function multiversioning implemented in Go through the LLVM Go API, before aggregate ABI lowering and target optimization, including at O0 and without an LTO plugin.

The public entry loads an atomic cached implementation and tail-forwards to it. A separate noinline resolver selects the baseline or AVX2 body from the immutable post-GODEBUG CPU snapshot. Calls before CPU initialization tail-forward to the baseline without caching. GOAMD64=v3/v4 fold the AVX2 query directly and create no redundant versions, resolver, or slot.

FMV inlining policy follows GoALLC: public dispatchers and implementations remain eligible for normal, feature-compatible LLVM inlining; the resolver and only its pre-initialization fallback edge are opaque. Explicit source noinline directives and disabled inlining remain honored. Compiler-required physical source frames stay on implementations. The synthetic dispatcher has no debug subprogram, so LLVM preserves the real callsite and outer inline chain without adding a duplicate source frame. LLGo's general runtime inline-frame infrastructure remains separate from GoALLC's GoObj inline tree.

Specialized direct calls use matching entries within and across packages. Function addresses retain the original entry; bodyless assembly/linkname declarations do not promise specialized implementations. Both implementations fold the AVX2 query while independent AVX/FMA query results remain observable under GODEBUG. Clones retain Go display names, distinct PC-line records, and their own debug linkage names.

Windows/amd64 SIMD128 parameters use the native indirect ABI explicitly before FMV. Dispatchers and resolvers forward caller-owned parameter storage through tail jumps instead of passing vector temporaries in a released stack frame; vector returns retain their register ABI. Windows function-entry records are grouped by PE unwind boundaries so inline copies cannot overwrite the physical caller identity. Failed native SIMD CI binaries are retained for diagnosis.

All FMV policy, discovery, cloning decisions, call rewriting, dispatcher construction, and Go metadata remapping are in Go. LLVM's generic cloning and CFG utilities preserve SSA/debug references and remove dead branches even at O0. LLGo has no FMV C++ implementation or private C entry point.

Scope: amd64, one AVX2 profile, and scalar/pointer/SIMD128 signatures. Aggregate signatures, general feature combinations, portable simd, and 256/512-bit vector ABIs remain follow-up work.

Dependency: xgo-dev/llvm#56 supplies generic IR utility bindings with no FMV policy. Its CI matrix covers LLVM 14–22; the PR remains open. This branch pins the exact public fork commit through go.mod; replace that pin when the APIs become available upstream. There is no local-path dependency.

Validation:

  • FMV tests cover dispatch/cache initialization, query folding, recursion/direct calls, indirect addresses, unsupported signatures, metadata, idempotence, and collision rejection before partial mutation.
  • O2, ThinLTO, and Full LTO IR tests verify dispatcher inlining with caller continuations, compatible/incompatible target features, explicit inlining attributes, and nested debug inline chains.
  • Linux/macOS/Windows amd64 LLVM verification and O0/O2 assembly checks; GOAMD64=v1 baseline instruction checks and v3/v4 redundant-version elimination. Windows tests verify caller-owned parameter forwarding and direct/indirect vector ABI conversion.
  • Compiler, SSA, ABI, C ABI, and build regressions pass. macOS LTO initially picked an LLVM 19 linker; the failed group passes with matching LLVM 22 tools.
  • Linux/amd64 and Windows/amd64 (Wine): complete SIMD suites at O0/O2 without LTO, O2 ThinLTO, and O2 Full LTO, with separate AVX2/AVX/FMA/all-feature override processes. Wine uses GCC MinGW libraries and an environment-only setjmp compatibility alias; native Windows validation remains in this PR's CI.
  • Linux GOAMD64=v3 complete SIMD suite and cpu.all=off FMV execution; macOS/arm64 O0/O2 SIMD execution; current Linux/amd64 and macOS/arm64 stage-2 self-host builds and SIMD execution.

The dependency pin also includes binding review fixes at f6e797e: checked IR-kind conversions, recoverable invalid-input panics, enum assertions, and explicit borrowed inline-assembly strings. The updated pin passes the full FMV unit suite, Linux/macOS/Windows O0/O2 IR and assembly checks, and macOS/arm64 stage-2 self-host build and SIMD execution.

Part of #2568.

@codecov

codecov Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.50779% with 8 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
internal/llvmfmv/fmv.go 97.95% 4 Missing ⚠️
internal/llvmfmv/source.go 95.23% 3 Missing ⚠️
ssa/simd_fmv.go 87.50% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SIMD128 CPU-guarded function multiversioning

This is a carefully engineered, well-tested change. The dispatch guard is sound: the x86.avx2 query is constant-folded to true only inside the cloned .__llgo_fmv_avx2 variant, while the baseline entry keeps a genuine runtime AVX2() call before the musttail forward — so AVX2 code never runs on an unsupported CPU, and the independent AVX/FMA queries correctly stay dynamic (+avx2 does not imply +fma in LLVM's x86 model). The subtle hazards are handled correctly: deferred erasure of query calls avoids iterator invalidation, UsedIDs is seeded with every existing PC-site ID before hash-remapping so string substitution can't alias, getNumOperands() is cached before addOperand(), and idempotency / symbol-collision / non-amd64 no-op are all explicitly tested. The README additions and code comments accurately describe the behavior.

I found no correctness or security defects. The inline comments below are a compile-time performance concern plus two maintainability/robustness notes — none are blocking.

Reviewed: internal/llvmfmv/*, cl/simd_fmv.go, ssa/simd_fmv.go, cl/{compile,instr}.go, internal/build/build.go, dev/test_native_simd.sh, test/simd/*.

Comment thread internal/llvmfmv/fmv.cpp Outdated
for (auto &F : M) {
if (F.hasFnAttribute(Done) || !supported(F))
continue;
Function *Query = F.isDeclaration() ? nullptr : guardQuery(F);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] FMV runs a full-module instruction scan on every amd64 build

llvmfmv.Run is invoked for every package module during every build (internal/build/build.go:3288). On amd64, this discovery loop calls guardQuery(F) for every supported() non-declaration function, and guardQuery walks all of F's instructions (fmv.cpp:60) before the cheap hasFnAttribute(Entry) short-circuit on line 207 is consulted. Since supported() is true for essentially any ordinary function (scalar / pointer / ≤128-bit-vector signature), this turns the pass into an O(all instructions in the module) traversal even for packages that contain no SIMD code at all — the overwhelmingly common case.

Consider gating the per-function scan: a function can only be a root if the module declares at least one llgo.cpu.query function, so compute that once up front and skip guardQuery entirely otherwise. Relatedly, UsedIDs is populated by scanning all llgo.pcline metadata (lines 215-220) on every amd64 module even when Originals ends up empty; returning early when there are no candidates avoids that cost too.

Comment thread internal/llvmfmv/fmv.cpp Outdated
MDNode *Row = Info->getOperand(I);
if (Row->getNumOperands() < 6)
continue;
auto *Name = dyn_cast<MDString>(Row->getOperand(1));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Metadata operand indices duplicated from ssa/funcinfo.go

cloneSourceInfo hardcodes the operand layout of the llgo.funcinfo (operand 1 = linker symbol, min 6 fields) and llgo.pcline (operand 1 = id, operand 2 = symbol, exactly 6 fields, line 102-103) metadata. That schema is defined independently in ssa/funcinfo.go (EmitFuncInfo/EmitPCLineInfo), with no shared constant or cross-reference. If the field order or count in funcinfo.go ever changes, this C++ silently reads the wrong fields or skips rows via the < 6 / != 6 guards and corrupts clone symbol / PC-site mappings — with no compile-time error. A comment in each file pointing at the other (and documenting the exact layout next to these index accesses) would make the coupling explicit for future maintainers.

Comment thread internal/llvmfmv/fmv.cpp Outdated
return nullptr;
char *Result = static_cast<char *>(std::malloc(Error.size() + 1));
if (!Result)
std::abort();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] std::abort() on malloc failure turns a compile error into a crash

llgoRunSIMDFMV hard-aborts the whole compiler process if the small diagnostic-string allocation fails. This path only runs when there is already a recoverable error to report (e.g. a symbol collision), so an OOM here escalates a reportable build error into a process crash that skips cleanup. Returning a static, non-malloc sentinel the Go side can detect would keep the error recoverable. Minor robustness nit, not a correctness or security issue.

@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

f74f07d09fc8 | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 149240 B 0 B / +0.0% 88157 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 148511 B 0 B / +0.0% 72985 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 136707 B 0 B / +0.0% 91328 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 154296 B 0 B / +0.0% 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 154979 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3057202 B +203 B / +0.00664% (worse) 130373 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3153157 B +178 B / +0.005645% (worse) 100773 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2819285 B +193 B / +0.006846% (worse) 135853 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2355610 B +162 B / +0.006878% (worse) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2352181 B +155 B / +0.00659% (worse) 0 B 0 B / 0.0%
j32-emscripten/LLGo 148582 B 0 B / +0.0% 88157 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 148002 B 0 B / +0.0% 72985 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 136047 B 0 B / +0.0% 91328 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1481301 B +178 B / +0.01202% (worse) 104917 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1531168 B +189 B / +0.01235% (worse) 89743 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1373428 B +180 B / +0.01311% (worse) 109492 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1284073 B +178 B / +0.01386% (worse) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1281591 B +171 B / +0.01334% (worse) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 153943 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-wasi/LLGo 154626 B 0 B / +0.0% 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 7.539 s +35.21 ms / +0.5% (worse)
j32-goos-js 7.240 s -371.7 ms / -4.9% (better)
j64-emscripten-memory64 6.265 s -56.78 ms / -0.9% (better)
reflectcall/w32-wasi 23.818 s -725.6 ms / -3.0% (better)
w32-goos-wasip1 4.510 s -266.6 ms / -5.6% (better)
w32-wasi 4.281 s -147.1 ms / -3.3% (better)

Compared with d62e6aa9adee measured in the same runner job.

@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

f74f07d09fc8 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 8568 B 0 B / +0.0% 387 B 0 B / +0.0% 1.033 s +62.39 ms / +6.4% (worse) 1.308 ms +20.17 us / +1.6% (worse)
Linux cprintf-lto 8408 B 0 B / +0.0% 368 B 0 B / +0.0% 1.003 s +6.917 ms / +0.7% (worse) 1.333 ms +66.28 us / +5.2% (worse)
Linux fmtprintf 4886704 B +416 B / +0.008514% (worse) 504790 B +29 B / +0.005745% (worse) 7.838 s +539.9 ms / +7.4% (worse) 3.274 ms +16.8 us / +0.5% (worse)
Linux fmtprintf-lto 3616584 B +16 B / +0.0004424% (worse) 443075 B 0 B / +0.0% 16.163 s +282.4 ms / +1.8% (worse) 2.895 ms -4.869 us / -0.2% (better)
Linux println 674920 B -120 B / -0.01778% (better) 16855 B 0 B / +0.0% 1.049 s +66.98 ms / +6.8% (worse) 1.609 ms -71.53 us / -4.3% (better)
Linux println-lto 189920 B -120 B / -0.1% (better) 14273 B 0 B / +0.0% 1.350 s +79.92 ms / +6.3% (worse) 1.595 ms -127.4 us / -7.4% (better)
macOS cprintf 50736 B 0 B / +0.0% 4409 B 0 B / +0.0% 1.443 s +1.232 ms / +0.1% (worse) 7.249 ms +3.877 ms / +115.0% (worse)
macOS cprintf-lto 50496 B 0 B / +0.0% 161 B 0 B / +0.0% 2.083 s +727.1 ms / +53.6% (worse) 10.406 ms +6.045 ms / +138.6% (worse)
macOS fmtprintf 1774544 B +176 B / +0.009919% (worse) 880609 B +80 B / +0.009085% (worse) 8.704 s +967 ms / +12.5% (worse) 16.456 ms +1.741 ms / +11.8% (worse)
macOS fmtprintf-lto 1361232 B 0 B / +0.0% 765033 B +76 B / +0.009935% (worse) 17.298 s +987.8 ms / +6.1% (worse) 8.043 ms +1.919 ms / +31.3% (worse)
macOS println 99344 B 0 B / +0.0% 24216 B 0 B / +0.0% 1.786 s +566.8 ms / +46.5% (worse) 10.275 ms +5.251 ms / +104.5% (worse)
macOS println-lto 83664 B 0 B / +0.0% 21457 B 0 B / +0.0% 1.820 s +377.4 ms / +26.2% (worse) 6.845 ms +2.002 ms / +41.3% (worse)
Windows MinGW cprintf 651264 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.751 s +16.38 ms / +0.9% (worse) 3.664 ms +119.1 us / +3.4% (worse)
Windows MinGW cprintf-lto 44032 B 0 B / +0.0% 4502 B 0 B / +0.0% 1.776 s +28.71 ms / +1.6% (worse) 3.396 ms -116.1 us / -3.3% (better)
Windows MinGW fmtprintf 5441536 B +1024 B / +0.01882% (worse) 604710 B +240 B / +0.0397% (worse) 7.554 s -144.9 ms / -1.9% (better) 8.047 ms -1.056 ms / -11.6% (better)
Windows MinGW fmtprintf-lto 4145152 B +1024 B / +0.02471% (worse) 551830 B +240 B / +0.04351% (worse) 15.168 s -534.3 ms / -3.4% (better) 8.283 ms -339.5 us / -3.9% (better)
Windows MinGW println 705536 B -512 B / -0.1% (better) 25190 B 0 B / +0.0% 1.746 s +21.65 ms / +1.3% (worse) 9.054 ms +2.771 ms / +44.1% (worse)
Windows MinGW println-lto 208896 B 0 B / +0.0% 22070 B 0 B / +0.0% 2.035 s -7.211 ms / -0.4% (better) 6.609 ms -260.7 us / -3.8% (better)
Windows MinGW 386 cprintf 601600 B 0 B / +0.0% 5352 B 0 B / +0.0% 1.403 s +13.92 ms / +1.0% (worse) 3.929 ms -199.2 us / -4.8% (better)
Windows MinGW 386 cprintf-lto 103424 B 0 B / +0.0% 5112 B 0 B / +0.0% 1.398 s +12.72 ms / +0.9% (worse) 4.038 ms +29.5 us / +0.7% (worse)
Windows MinGW 386 fmtprintf 4745216 B -512 B / -0.01079% (better) 472472 B 0 B / +0.0% 6.294 s +337.9 ms / +5.7% (worse) 8.410 ms +170.3 us / +2.1% (worse)
Windows MinGW 386 fmtprintf-lto 4148736 B 0 B / +0.0% 451132 B 0 B / +0.0% 12.217 s +249.3 ms / +2.1% (worse) 8.587 ms +254.4 us / +3.1% (worse)
Windows MinGW 386 println 654336 B 0 B / +0.0% 21500 B 0 B / +0.0% 1.369 s +423.7 us / +0.03095% (worse) 6.856 ms +238.8 us / +3.6% (worse)
Windows MinGW 386 println-lto 258560 B 0 B / +0.0% 19328 B 0 B / +0.0% 1.565 s -32.53 ms / -2.0% (better) 6.927 ms +12.3 us / +0.2% (worse)
Windows MinGW ARM64 cprintf 662016 B 0 B / +0.0% 4436 B 0 B / +0.0% 2.002 s -14.89 ms / -0.7% (better) 7.323 ms -298.7 us / -3.9% (better)
Windows MinGW ARM64 cprintf-lto 43520 B 0 B / +0.0% 4388 B 0 B / +0.0% 2.079 s +51.26 ms / +2.5% (worse) 7.400 ms +100.4 us / +1.4% (worse)
Windows MinGW ARM64 fmtprintf 5352448 B +4096 B / +0.1% (worse) 512868 B +204 B / +0.03979% (worse) 7.470 s -38.09 ms / -0.5% (better) 14.924 ms +127.7 us / +0.9% (worse)
Windows MinGW ARM64 fmtprintf-lto 4313600 B +2048 B / +0.0475% (worse) 479028 B +176 B / +0.03675% (worse) 14.257 s -331.1 ms / -2.3% (better) 14.753 ms +223 us / +1.5% (worse)
Windows MinGW ARM64 println 713216 B -512 B / -0.1% (better) 23912 B 0 B / +0.0% 2.046 s +49.57 ms / +2.5% (worse) 12.951 ms +458.2 us / +3.7% (worse)
Windows MinGW ARM64 println-lto 215552 B 0 B / +0.0% 21272 B 0 B / +0.0% 2.312 s +50.81 ms / +2.2% (worse) 11.795 ms -662 us / -5.3% (better)
Windows MSVC cprintf 895488 B 0 B / +0.0% 66019 B 0 B / +0.0% 1.607 s +37.71 ms / +2.4% (worse) 3.599 ms +210.9 us / +6.2% (worse)
Windows MSVC cprintf-lto 291328 B 0 B / +0.0% 65971 B 0 B / +0.0% 1.597 s -24.27 ms / -1.5% (better) 3.929 ms +417.4 us / +11.9% (worse)
Windows MSVC fmtprintf 5743616 B +1536 B / +0.02675% (worse) 700438 B +224 B / +0.03199% (worse) 7.357 s +96.23 ms / +1.3% (worse) 8.484 ms -1.725 ms / -16.9% (better)
Windows MSVC fmtprintf-lto 4469760 B +512 B / +0.01146% (worse) 651542 B +192 B / +0.02948% (worse) 14.571 s -87.38 ms / -0.6% (better) 10.754 ms +371.6 us / +3.6% (worse)
Windows MSVC println 1018368 B -512 B / -0.1% (better) 121062 B 0 B / +0.0% 1.617 s -235.2 ms / -12.7% (better) 7.694 ms -2.244 ms / -22.6% (better)
Windows MSVC println-lto 529920 B 0 B / +0.0% 118598 B 0 B / +0.0% 1.858 s +34.88 ms / +1.9% (worse) 8.184 ms +866.4 us / +11.8% (worse)
Windows MSVC 386 cprintf 514048 B 0 B / +0.0% 3947 B 0 B / +0.0% 1.527 s -18.27 ms / -1.2% (better) 6.554 ms +700 us / +12.0% (worse)
Windows MSVC 386 cprintf-lto 44032 B 0 B / +0.0% 3901 B 0 B / +0.0% 1.721 s +195.3 ms / +12.8% (worse) 6.132 ms +679.4 us / +12.5% (worse)
Windows MSVC 386 fmtprintf 4480000 B 0 B / +0.0% 455900 B 0 B / +0.0% 7.219 s +214.7 ms / +3.1% (worse) 11.902 ms +950 us / +8.7% (worse)
Windows MSVC 386 fmtprintf-lto 3899904 B 0 B / +0.0% 426667 B 0 B / +0.0% 13.495 s +66.33 ms / +0.5% (worse) 11.366 ms -1.302 ms / -10.3% (better)
Windows MSVC 386 println 567808 B -512 B / -0.1% (better) 20356 B 0 B / +0.0% 1.562 s +26.6 ms / +1.7% (worse) 10.288 ms -196.2 us / -1.9% (better)
Windows MSVC 386 println-lto 199168 B 0 B / +0.0% 18517 B 0 B / +0.0% 1.797 s -31.94 ms / -1.7% (better) 9.671 ms -100.3 us / -1.0% (better)
Windows MSVC ARM64 cprintf 662528 B 0 B / +0.0% 4240 B 0 B / +0.0% 1.811 s +12.27 ms / +0.7% (worse) 8.636 ms -510.5 us / -5.6% (better)
Windows MSVC ARM64 cprintf-lto 48640 B 0 B / +0.0% 4140 B 0 B / +0.0% 1.846 s +21.57 ms / +1.2% (worse) 8.565 ms -339.6 us / -3.8% (better)
Windows MSVC ARM64 fmtprintf 5348864 B +3072 B / +0.1% (worse) 512820 B +208 B / +0.04058% (worse) 7.553 s -95.72 ms / -1.3% (better) 18.205 ms +37.1 us / +0.2% (worse)
Windows MSVC ARM64 fmtprintf-lto 4320256 B +1024 B / +0.02371% (worse) 479712 B +176 B / +0.0367% (worse) 15.011 s -160.6 ms / -1.1% (better) 16.903 ms -2.644 ms / -13.5% (better)
Windows MSVC ARM64 println 715776 B -512 B / -0.1% (better) 23956 B 0 B / +0.0% 1.821 s +21.58 ms / +1.2% (worse) 15.454 ms +409.4 us / +2.7% (worse)
Windows MSVC ARM64 println-lto 220672 B 0 B / +0.0% 21444 B 0 B / +0.0% 2.063 s -6.401 ms / -0.3% (better) 15.431 ms +379.7 us / +2.5% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.500 ns/op +0.09 ns/op / +0.6% (worse)
Linux BenchmarkMergeCompilerFlags 191.600 ns/op -4.3 ns/op / -2.2% (better)
Linux BenchmarkMergeLinkerFlags 127.300 ns/op -0.5 ns/op / -0.4% (better)
Linux BenchmarkChannelBuffered 55.010 ns/op -0.17 ns/op / -0.3% (better)
Linux BenchmarkChannelHandoff 12869 ns/op -1037 ns/op / -7.5% (better)
Linux BenchmarkDefer 49.050 ns/op +0.93 ns/op / +1.9% (worse)
Linux BenchmarkDirectCall 1.582 ns/op +0.065 ns/op / +4.3% (worse)
Linux BenchmarkGlobalRead 1.164 ns/op -0.39 ns/op / -25.1% (better)
Linux BenchmarkGlobalWrite 7.756 ns/op -0.026 ns/op / -0.3% (better)
Linux BenchmarkGoroutine 22982 ns/op -1909 ns/op / -7.7% (better)
Linux BenchmarkInterfaceCall 5.834 ns/op -0.695 ns/op / -10.6% (better)
Linux BenchmarkRuntimeGetG 2.671 ns/op -0.184 ns/op / -6.4% (better)
macOS BenchmarkLookupPCRandom 18.080 ns/op +2.54 ns/op / +16.3% (worse)
macOS BenchmarkMergeCompilerFlags 155 ns/op +7.9 ns/op / +5.4% (worse)
macOS BenchmarkMergeLinkerFlags 112.700 ns/op +24.4 ns/op / +27.6% (worse)
macOS BenchmarkChannelBuffered 28.400 ns/op -11.49 ns/op / -28.8% (better)
macOS BenchmarkChannelHandoff 8055 ns/op -5855 ns/op / -42.1% (better)
macOS BenchmarkDefer 41.240 ns/op -18.03 ns/op / -30.4% (better)
macOS BenchmarkDirectCall 1.164 ns/op -0.197 ns/op / -14.5% (better)
macOS BenchmarkGlobalRead 1.169 ns/op -0.192 ns/op / -14.1% (better)
macOS BenchmarkGlobalWrite 1.239 ns/op -0.682 ns/op / -35.5% (better)
macOS BenchmarkGoroutine 111660 ns/op +32892 ns/op / +41.8% (worse)
macOS BenchmarkInterfaceCall 4.412 ns/op -1.542 ns/op / -25.9% (better)
macOS BenchmarkRuntimeGetG 2.517 ns/op -0.788 ns/op / -23.8% (better)
Windows MinGW BenchmarkLookupPCRandom 12.980 ns/op -0.16 ns/op / -1.2% (better)
Windows MinGW BenchmarkMergeCompilerFlags 630.700 ns/op -27.5 ns/op / -4.2% (better)
Windows MinGW BenchmarkMergeLinkerFlags 548.600 ns/op -44.4 ns/op / -7.5% (better)
Windows MinGW BenchmarkChannelBuffered 30.740 ns/op -0.51 ns/op / -1.6% (better)
Windows MinGW BenchmarkChannelHandoff 946.700 ns/op +69.3 ns/op / +7.9% (worse)
Windows MinGW BenchmarkDefer 57.550 ns/op -1.16 ns/op / -2.0% (better)
Windows MinGW BenchmarkDirectCall 1.547 ns/op 0 ns/op / +0.0%
Windows MinGW BenchmarkGlobalRead 1.549 ns/op +0.001 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGlobalWrite 2.474 ns/op +0.004 ns/op / +0.2% (worse)
Windows MinGW BenchmarkGoroutine 92442 ns/op -288 ns/op / -0.3% (better)
Windows MinGW BenchmarkInterfaceCall 8.071 ns/op +0.286 ns/op / +3.7% (worse)
Windows MinGW BenchmarkRuntimeGetG 2.481 ns/op -0.006 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 21.580 ns/op +0.04 ns/op / +0.2% (worse)
Windows MinGW 386 BenchmarkMergeCompilerFlags 556.100 ns/op -10.9 ns/op / -1.9% (better)
Windows MinGW 386 BenchmarkMergeLinkerFlags 549.600 ns/op +8 ns/op / +1.5% (worse)
Windows MinGW 386 BenchmarkChannelBuffered 33.800 ns/op -0.2 ns/op / -0.6% (better)
Windows MinGW 386 BenchmarkChannelHandoff 761.100 ns/op -99.7 ns/op / -11.6% (better)
Windows MinGW 386 BenchmarkDefer 35.070 ns/op -0.07 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkDirectCall 1.360 ns/op +0.003 ns/op / +0.2% (worse)
Windows MinGW 386 BenchmarkGlobalRead 1.632 ns/op +0.003 ns/op / +0.2% (worse)
Windows MinGW 386 BenchmarkGlobalWrite 6.998 ns/op +0.017 ns/op / +0.2% (worse)
Windows MinGW 386 BenchmarkGoroutine 70944 ns/op -1173 ns/op / -1.6% (better)
Windows MinGW 386 BenchmarkInterfaceCall 7.341 ns/op -0.062 ns/op / -0.8% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 1.629 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.090 ns/op -0.04 ns/op / -0.3% (better)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 576.600 ns/op +11.3 ns/op / +2.0% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 541.700 ns/op -3.1 ns/op / -0.6% (better)
Windows MinGW ARM64 BenchmarkChannelBuffered 38.860 ns/op +0.05 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkChannelHandoff 1943 ns/op +126 ns/op / +6.9% (worse)
Windows MinGW ARM64 BenchmarkDefer 58.290 ns/op -2.06 ns/op / -3.4% (better)
Windows MinGW ARM64 BenchmarkDirectCall 0.663 ns/op +0.0002 ns/op / +0.03015% (worse)
Windows MinGW ARM64 BenchmarkGlobalRead 0.590 ns/op -0.0001 ns/op / -0.01695% (better)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.738 ns/op +0.0004 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkGoroutine 63945 ns/op -2200 ns/op / -3.3% (better)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.230 ns/op +0.079 ns/op / +1.9% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.769 ns/op -0.036 ns/op / -2.0% (better)
Windows MSVC BenchmarkLookupPCRandom 13.140 ns/op +0.08 ns/op / +0.6% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 612.100 ns/op -10.9 ns/op / -1.7% (better)
Windows MSVC BenchmarkMergeLinkerFlags 552.500 ns/op +2.9 ns/op / +0.5% (worse)
Windows MSVC BenchmarkChannelBuffered 28.980 ns/op +0.08 ns/op / +0.3% (worse)
Windows MSVC BenchmarkChannelHandoff 1096 ns/op -5 ns/op / -0.5% (better)
Windows MSVC BenchmarkDefer 57.010 ns/op +1.98 ns/op / +3.6% (worse)
Windows MSVC BenchmarkDirectCall 1.546 ns/op -0.003 ns/op / -0.2% (better)
Windows MSVC BenchmarkGlobalRead 1.550 ns/op +0.002 ns/op / +0.1% (worse)
Windows MSVC BenchmarkGlobalWrite 2.468 ns/op -0.006 ns/op / -0.2% (better)
Windows MSVC BenchmarkGoroutine 89962 ns/op +2535 ns/op / +2.9% (worse)
Windows MSVC BenchmarkInterfaceCall 8.072 ns/op -0.326 ns/op / -3.9% (better)
Windows MSVC BenchmarkRuntimeGetG 1.861 ns/op -0.312 ns/op / -14.4% (better)
Windows MSVC 386 BenchmarkLookupPCRandom 26.590 ns/op +0.14 ns/op / +0.5% (worse)
Windows MSVC 386 BenchmarkMergeCompilerFlags 781.100 ns/op +12.4 ns/op / +1.6% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 691.200 ns/op +15.2 ns/op / +2.2% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 39.150 ns/op -0.19 ns/op / -0.5% (better)
Windows MSVC 386 BenchmarkChannelHandoff 908.400 ns/op +41 ns/op / +4.7% (worse)
Windows MSVC 386 BenchmarkDefer 46.130 ns/op -1.28 ns/op / -2.7% (better)
Windows MSVC 386 BenchmarkDirectCall 1.550 ns/op +0.002 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGlobalRead 1.550 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGlobalWrite 7.781 ns/op +0.003 ns/op / +0.03857% (worse)
Windows MSVC 386 BenchmarkGoroutine 110781 ns/op -914 ns/op / -0.8% (better)
Windows MSVC 386 BenchmarkInterfaceCall 8.134 ns/op +0.004 ns/op / +0.0492% (worse)
Windows MSVC 386 BenchmarkRuntimeGetG 1.928 ns/op -0.01 ns/op / -0.5% (better)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.120 ns/op +0.05 ns/op / +0.4% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 573.300 ns/op -11.8 ns/op / -2.0% (better)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 556.200 ns/op +4 ns/op / +0.7% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 38.580 ns/op +0.01 ns/op / +0.02593% (worse)
Windows MSVC ARM64 BenchmarkChannelHandoff 3951 ns/op +1044 ns/op / +35.9% (worse)
Windows MSVC ARM64 BenchmarkDefer 73.970 ns/op +7.11 ns/op / +10.6% (worse)
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op +0.0005 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkGlobalRead 0.590 ns/op -0.0003 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.834 ns/op +0.002 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkGoroutine 66295 ns/op +2820 ns/op / +4.4% (worse)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.141 ns/op +0.001 ns/op / +0.02415% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.771 ns/op +0.001 ns/op / +0.1% (worse)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 904.100 ns/op +4.6 ns/op / +0.5% (worse)
Linux AfterFuncZeroDelivery/LLGo 40974 ns/op +2454 ns/op / +6.4% (worse)
Linux CreateStop/Go 289.400 ns/op +0.3 ns/op / +0.1% (worse)
Linux CreateStop/LLGo 1679 ns/op -232 ns/op / -12.1% (better)
Linux RearmStopped/Go 114.500 ns/op 0 ns/op / +0.0%
Linux RearmStopped/LLGo 1415 ns/op +144 ns/op / +11.3% (worse)
Linux ResetActive/Go 67.450 ns/op +0.06 ns/op / +0.1% (worse)
Linux ResetActive/LLGo 753 ns/op +33 ns/op / +4.6% (worse)
Linux ResetHeap1024/Go 67.050 ns/op -0.03 ns/op / -0.04472% (better)
Linux ResetHeap1024/LLGo 178 ns/op -0.2 ns/op / -0.1% (better)
macOS AfterFuncZeroDelivery/Go 734.800 ns/op +155.2 ns/op / +26.8% (worse)
macOS AfterFuncZeroDelivery/LLGo 136881 ns/op +30408 ns/op / +28.6% (worse)
macOS CreateStop/Go 212.300 ns/op +24.8 ns/op / +13.2% (worse)
macOS CreateStop/LLGo 934.700 ns/op +288.7 ns/op / +44.7% (worse)
macOS RearmStopped/Go 90.600 ns/op +19.98 ns/op / +28.3% (worse)
macOS RearmStopped/LLGo 682.600 ns/op +12.2 ns/op / +1.8% (worse)
macOS ResetActive/Go 65.720 ns/op -2.53 ns/op / -3.7% (better)
macOS ResetActive/LLGo 264.500 ns/op +52.3 ns/op / +24.6% (worse)
macOS ResetHeap1024/Go 63.210 ns/op +3.64 ns/op / +6.1% (worse)
macOS ResetHeap1024/LLGo 94.260 ns/op -54.94 ns/op / -36.8% (better)
Windows MinGW AfterFuncZeroDelivery/Go 561.600 ns/op +11.8 ns/op / +2.1% (worse)
Windows MinGW AfterFuncZeroDelivery/LLGo 182519 ns/op -1391 ns/op / -0.8% (better)
Windows MinGW CreateStop/Go 122.600 ns/op +7.1 ns/op / +6.1% (worse)
Windows MinGW CreateStop/LLGo 438.100 ns/op -16.5 ns/op / -3.6% (better)
Windows MinGW RearmStopped/Go 31.580 ns/op -0.09 ns/op / -0.3% (better)
Windows MinGW RearmStopped/LLGo 275.200 ns/op -108.1 ns/op / -28.2% (better)
Windows MinGW ResetActive/Go 20.060 ns/op +0.01 ns/op / +0.04988% (worse)
Windows MinGW ResetActive/LLGo 158.500 ns/op -95.7 ns/op / -37.6% (better)
Windows MinGW ResetHeap1024/Go 20.330 ns/op -0.23 ns/op / -1.1% (better)
Windows MinGW ResetHeap1024/LLGo 125.600 ns/op -1.8 ns/op / -1.4% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 779.500 ns/op +15.6 ns/op / +2.0% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 129803 ns/op -438 ns/op / -0.3% (better)
Windows MinGW 386 CreateStop/Go 167.500 ns/op -1.3 ns/op / -0.8% (better)
Windows MinGW 386 CreateStop/LLGo 392.700 ns/op -10.8 ns/op / -2.7% (better)
Windows MinGW 386 RearmStopped/Go 56.600 ns/op -0.2 ns/op / -0.4% (better)
Windows MinGW 386 RearmStopped/LLGo 286.700 ns/op +15.7 ns/op / +5.8% (worse)
Windows MinGW 386 ResetActive/Go 32.570 ns/op +0.04 ns/op / +0.1% (worse)
Windows MinGW 386 ResetActive/LLGo 845.200 ns/op -40.8 ns/op / -4.6% (better)
Windows MinGW 386 ResetHeap1024/Go 32.800 ns/op -0.09 ns/op / -0.3% (better)
Windows MinGW 386 ResetHeap1024/LLGo 152.100 ns/op +1.9 ns/op / +1.3% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 665.200 ns/op -2.6 ns/op / -0.4% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 151916 ns/op -4724 ns/op / -3.0% (better)
Windows MinGW ARM64 CreateStop/Go 199.800 ns/op +3.5 ns/op / +1.8% (worse)
Windows MinGW ARM64 CreateStop/LLGo 439.200 ns/op +26.8 ns/op / +6.5% (worse)
Windows MinGW ARM64 RearmStopped/Go 70.570 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 RearmStopped/LLGo 259.700 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 ResetActive/Go 31.040 ns/op -0.09 ns/op / -0.3% (better)
Windows MinGW ARM64 ResetActive/LLGo 138.400 ns/op +3.5 ns/op / +2.6% (worse)
Windows MinGW ARM64 ResetHeap1024/Go 31.150 ns/op +0.02 ns/op / +0.1% (worse)
Windows MinGW ARM64 ResetHeap1024/LLGo 124.800 ns/op +0.1 ns/op / +0.1% (worse)
Windows MSVC AfterFuncZeroDelivery/Go 557.600 ns/op +2.4 ns/op / +0.4% (worse)
Windows MSVC AfterFuncZeroDelivery/LLGo 177207 ns/op +1682 ns/op / +1.0% (worse)
Windows MSVC CreateStop/Go 115.800 ns/op +0.2 ns/op / +0.2% (worse)
Windows MSVC CreateStop/LLGo 412.200 ns/op +8.6 ns/op / +2.1% (worse)
Windows MSVC RearmStopped/Go 31.560 ns/op +0.13 ns/op / +0.4% (worse)
Windows MSVC RearmStopped/LLGo 259.900 ns/op +11 ns/op / +4.4% (worse)
Windows MSVC ResetActive/Go 20.070 ns/op -0.01 ns/op / -0.0498% (better)
Windows MSVC ResetActive/LLGo 159.300 ns/op +3.3 ns/op / +2.1% (worse)
Windows MSVC ResetHeap1024/Go 20.480 ns/op +0.11 ns/op / +0.5% (worse)
Windows MSVC ResetHeap1024/LLGo 124.200 ns/op +0.3 ns/op / +0.2% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/Go 948.200 ns/op +5.3 ns/op / +0.6% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 209034 ns/op +3142 ns/op / +1.5% (worse)
Windows MSVC 386 CreateStop/Go 195.300 ns/op +2.1 ns/op / +1.1% (worse)
Windows MSVC 386 CreateStop/LLGo 467.100 ns/op +3.5 ns/op / +0.8% (worse)
Windows MSVC 386 RearmStopped/Go 63.510 ns/op +0.16 ns/op / +0.3% (worse)
Windows MSVC 386 RearmStopped/LLGo 317.600 ns/op -6.8 ns/op / -2.1% (better)
Windows MSVC 386 ResetActive/Go 39.150 ns/op +0.18 ns/op / +0.5% (worse)
Windows MSVC 386 ResetActive/LLGo 964.900 ns/op +5.8 ns/op / +0.6% (worse)
Windows MSVC 386 ResetHeap1024/Go 39.510 ns/op +0.2 ns/op / +0.5% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 173.400 ns/op +4.7 ns/op / +2.8% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 679.200 ns/op +10.3 ns/op / +1.5% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 161238 ns/op -5361 ns/op / -3.2% (better)
Windows MSVC ARM64 CreateStop/Go 199 ns/op -1.5 ns/op / -0.7% (better)
Windows MSVC ARM64 CreateStop/LLGo 431.400 ns/op +47.8 ns/op / +12.5% (worse)
Windows MSVC ARM64 RearmStopped/Go 70.560 ns/op -0.08 ns/op / -0.1% (better)
Windows MSVC ARM64 RearmStopped/LLGo 276.600 ns/op +0.2 ns/op / +0.1% (worse)
Windows MSVC ARM64 ResetActive/Go 30.960 ns/op -0.2 ns/op / -0.6% (better)
Windows MSVC ARM64 ResetActive/LLGo 141 ns/op +7.5 ns/op / +5.6% (worse)
Windows MSVC ARM64 ResetHeap1024/Go 31.070 ns/op -0.13 ns/op / -0.4% (better)
Windows MSVC ARM64 ResetHeap1024/LLGo 137.800 ns/op -0.2 ns/op / -0.1% (better)

Compared with d62e6aa9adee measured in the same runner job.

@zhouguangyuan0718 zhouguangyuan0718 changed the title compiler: add early AVX2 multiversioning for SIMD128 compiler: add early AVX2 SIMD128 multiversioning through Go APIs Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant