Skip to content

Compile time: fewer kernel specialisations per (T, INT) #91

Description

@michel2323

Context: PR #90 (CI time). After #87–#89 the test suite is dominated by compilation, not by tests. CI went from ~21 min to ~52 min. With the parallel runner and the CPU job on kkt it is ~34 min, and the floor is the cold compile of one worker.

Measurements (owner's machine, CPU backend unless noted)

Where the specialisations come from

Kernels are specialised on T (4) × INT (2) and, in addition, on several compile-time classes chosen with explicit Val branches:

  • _with_width_class (src/numeric/front.jl): regime-B width classes W ∈ (8, 16, 32, 64), used by the Cholesky (front.jl) and LDLᵀ (ldlt.jl) fused kernels. The branches are static, so the launcher compiles all four variants with its caller, even if the matrix only has one width class.
  • _with_local_bytes (src/numeric/subtree.jl): 4 regime-A budget classes (SUBTREE_LOCAL_SIZES), Cholesky and LDLᵀ.
  • Further Vals in the LDLᵀ kernels: HERM, the packed/unpacked panel (_lt_pk(Val(W))), Val(ASM), Val(NL), Val(W*(W+1)÷2), and the three pivot passes Val(1..3).
  • Solve sweeps (src/solve/sweeps.jl): 3 solve kinds × sub/det/forward variants, also through explicit branches.
  • Dense implementations: :generic, :vendor, :ka, plus strided-batched variants.

On the CPU backend a KA kernel is an ordinary Julia function, and every statically reachable variant compiles with its caller. On GPUs a kernel compiles at its first launch, so only the variants actually launched are compiled. But each of them is a separate GPUCompiler/LLVM run per T, INT.

Possible levers (not sure any is acceptable)

  1. Keep @localmem sizes compile-time (required by KA), but don't compile unused classes on the CPU backend. For example, put a dynamic barrier at the class branch, like the structure barrier in PR CI: parallel test runner, CPU job on kkt, LDLᵀ compiled only when used #90. It must not allocate in the numeric phase: the argument tuples hold immutable Numeric/Symbolic, which get boxed. A mutable carrier or a precomputed launcher table could avoid that.
  2. Fewer classes. On the CPU backend, a single width class (64) and a single budget class would do. Cost: the CPU tests would no longer exercise the smaller-class code paths that the GPU uses.
  3. Merge the Val flags that only select small code paths (HERM, pivot pass, packed panel) into run-time branches where the GPU cost is negligible. That needs benchmarking per flag (bench/compare.jl).
  4. Test side: not every testset needs all four element types and both index types. That would change the test conventions in AGENTS.md, so it's the owner's decision.

Open question: whether 1–3 can be done without losing GPU performance. The width-class and local-memory specialisations are the core of the regime-A/B design (PLAN §2.4), so this may turn out not to be worth it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    performanceSpeed or memory of the numeric and solve phases; tracked in PERFORMANCE.mdtriagedOwner has looked at this found-by-agent issue; it no longer holds the task chain

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions