You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Compile time: fewer kernel specialisations per (T, INT) #91
Context: PR #90 (CI time). After #87–#89 the test suite is dominated by compilation, not by tests. CI went from ~21 min to ~52 min. With the parallel runner and the CPU job on kkt it is ~34 min, and the floor is the cold compile of one worker.
Measurements (owner's machine, CPU backend unless noted)
Splitting a test file by element type only partly helps, because part of the compile doesn't depend on T. test_numeric_cholesky_c: one element type alone takes 157 s, all four 168 s. test_numeric_ldlt: 814 s → 226 s per element type.
Where the specialisations come from
Kernels are specialised on T (4) × INT (2) and, in addition, on several compile-time classes chosen with explicit Val branches:
_with_width_class (src/numeric/front.jl): regime-B width classes W ∈ (8, 16, 32, 64), used by the Cholesky (front.jl) and LDLᵀ (ldlt.jl) fused kernels. The branches are static, so the launcher compiles all four variants with its caller, even if the matrix only has one width class.
_with_local_bytes (src/numeric/subtree.jl): 4 regime-A budget classes (SUBTREE_LOCAL_SIZES), Cholesky and LDLᵀ.
Further Vals in the LDLᵀ kernels: HERM, the packed/unpacked panel (_lt_pk(Val(W))), Val(ASM), Val(NL), Val(W*(W+1)÷2), and the three pivot passes Val(1..3).
Solve sweeps (src/solve/sweeps.jl): 3 solve kinds × sub/det/forward variants, also through explicit branches.
Dense implementations: :generic, :vendor, :ka, plus strided-batched variants.
On the CPU backend a KA kernel is an ordinary Julia function, and every statically reachable variant compiles with its caller. On GPUs a kernel compiles at its first launch, so only the variants actually launched are compiled. But each of them is a separate GPUCompiler/LLVM run per T, INT.
Possible levers (not sure any is acceptable)
Keep @localmem sizes compile-time (required by KA), but don't compile unused classes on the CPU backend. For example, put a dynamic barrier at the class branch, like the structure barrier in PR CI: parallel test runner, CPU job on kkt, LDLᵀ compiled only when used #90. It must not allocate in the numeric phase: the argument tuples hold immutable Numeric/Symbolic, which get boxed. A mutable carrier or a precomputed launcher table could avoid that.
Fewer classes. On the CPU backend, a single width class (64) and a single budget class would do. Cost: the CPU tests would no longer exercise the smaller-class code paths that the GPU uses.
Merge the Val flags that only select small code paths (HERM, pivot pass, packed panel) into run-time branches where the GPU cost is negligible. That needs benchmarking per flag (bench/compare.jl).
Test side: not every testset needs all four element types and both index types. That would change the test conventions in AGENTS.md, so it's the owner's decision.
Open question: whether 1–3 can be done without losing GPU performance. The width-class and local-memory specialisations are the core of the regime-A/B design (PLAN §2.4), so this may turn out not to be worth it.
Context: PR #90 (CI time). After #87–#89 the test suite is dominated by compilation, not by tests. CI went from ~21 min to ~52 min. With the parallel runner and the CPU job on kkt it is ~34 min, and the floor is the cold compile of one worker.
Measurements (owner's machine, CPU backend unless noted)
test_numeric_ldlttakes 1155–1787 s on CI (351 s before perf: cooperative pivot search in the device LDLᵀ kernels (#75) #87–perf: regime-C LDLᵀ through blocked pivot steps and GEMMs (#75) #89). This is GPU kernel compilation of the new LDLᵀ variants.T.test_numeric_cholesky_c: one element type alone takes 157 s, all four 168 s.test_numeric_ldlt: 814 s → 226 s per element type.Where the specialisations come from
Kernels are specialised on
T(4) ×INT(2) and, in addition, on several compile-time classes chosen with explicitValbranches:_with_width_class(src/numeric/front.jl): regime-B width classesW ∈ (8, 16, 32, 64), used by the Cholesky (front.jl) and LDLᵀ (ldlt.jl) fused kernels. The branches are static, so the launcher compiles all four variants with its caller, even if the matrix only has one width class._with_local_bytes(src/numeric/subtree.jl): 4 regime-A budget classes (SUBTREE_LOCAL_SIZES), Cholesky and LDLᵀ.Vals in the LDLᵀ kernels:HERM, the packed/unpacked panel (_lt_pk(Val(W))),Val(ASM),Val(NL),Val(W*(W+1)÷2), and the three pivot passesVal(1..3).src/solve/sweeps.jl): 3 solve kinds × sub/det/forward variants, also through explicit branches.:generic,:vendor,:ka, plus strided-batched variants.On the CPU backend a KA kernel is an ordinary Julia function, and every statically reachable variant compiles with its caller. On GPUs a kernel compiles at its first launch, so only the variants actually launched are compiled. But each of them is a separate GPUCompiler/LLVM run per
T,INT.Possible levers (not sure any is acceptable)
@localmemsizes compile-time (required by KA), but don't compile unused classes on the CPU backend. For example, put a dynamic barrier at the class branch, like the structure barrier in PR CI: parallel test runner, CPU job on kkt, LDLᵀ compiled only when used #90. It must not allocate in the numeric phase: the argument tuples hold immutableNumeric/Symbolic, which get boxed. A mutable carrier or a precomputed launcher table could avoid that.Valflags that only select small code paths (HERM, pivot pass, packed panel) into run-time branches where the GPU cost is negligible. That needs benchmarking per flag (bench/compare.jl).Open question: whether 1–3 can be done without losing GPU performance. The width-class and local-memory specialisations are the core of the regime-A/B design (PLAN §2.4), so this may turn out not to be worth it.