Skip to content

CI: parallel test runner, CPU job on kkt, LDLᵀ compiled only when used - #90

Merged
michel2323 merged 2 commits into
mainfrom
ci/faster-tests
Oct 6, 2026
Merged

michel2323 merged 2 commits into
mainfrom
ci/faster-tests

Conversation

@michel2323

@michel2323 michel2323 commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Not a task PR: CI time went from ~21 min (Oct 5) to ~52 min (Oct 6) after the regime-B/C LDLᵀ PRs (#87, #88, #89).

Why it got slower. Almost all of it is compilation; coverage instrumentation makes no difference (measured). Per-file times on CI, before → after the three PRs:

file job before after
test_api CPU 658 s 2416 s
test_numeric_ldlt CUDA 351 s 1155 s
test_numeric_cholesky_a / _c CUDA 98 / 60 s 297 / 257 s

On the CPU backend, KA kernels are ordinary Julia functions and compile together with the code that calls them. The solver's numeric phase called factorize!, which picks Cholesky or LDLᵀ from the structure at run time with a static call. So every SPD/HPD solver also compiled the whole LDLᵀ path. test_api runs SPD/HPD only, over 4 element types × 2 index types, and paid for LDLᵀ eight times.

CI result (kkt, head ff5bf68): CPU job 16.5 min, CUDA job 17 min. Both green with unchanged counts (62938 / 69499 pass, 1 broken). On main the two jobs took 52 and 45 min. The slowest parts on kkt are the complex-element-type parts at ~12.5 min each.

Changes

  • src/solver.jl: the solver's numeric phase calls the Cholesky or the LDLᵀ path through an inference barrier, so only the path that runs gets compiled. The barrier takes only the mutable solver handle. A barrier at the factorize!(N, S, …) level would box the immutable Numeric/Symbolic (~1 KB per call) and break the numeric-phase allocation budgets, so I left that level alone. factorize_cholesky! is the Cholesky path of factorize!; factorize! itself behaves as before.
    • First factorization per (T, INT) through the API on the CPU backend: 42–84 s → 20–36 s.
  • test/runtests.jl: switched to ParallelTestRunner.jl, as CUDA.jl and oneAPI.jl use.
    • Each test_*.jl runs in its own module on a pool of worker processes; the shared helpers are loaded in init_code.
    • The seed 666 is set per test file, because ParallelTestRunner reseeds with 1 after init_code. Without this, test_refinement fails.
    • SDS_TEST_ONLY, SDS_TEST_SKIP, SDS_TEST_CPU and SDS_TEST_GPU still work. Test arguments (prefix match, !name to exclude) and PTR_NUM_JOBS / --jobs=N are new.
    • The backend banner prints once, from the main process.
  • Split by element type (second commit): the long files (SPLIT_FILES in test/runtests.jl) run as four parts each, test_api[Float32] and so on, each part in a worker of its own.
    • In a part, ELTYPES, REAL_ELTYPES and COMPLEX_ELTYPES are that part's subset.
    • Testsets that don't loop over the element types run in the Float64 part only, behind RUN_SHARED && @testset …. Test counts are unchanged.
    • test_numeric_ldlt cold: 814 s unsplit → 226 s for its slowest part.
  • ci.yml:
    • The CPU-only job runs on the self-hosted kkt machine. The job id and matrix are unchanged, so the required check name test-github-cpuonly (ubuntu-latest, 1, x64) still matches the ruleset.
    • PTR_NUM_JOBS is set to 32 for the CPU job and 12 for the GPU job (kkt: 192 threads, 2 TB, 3× H200, shared by its four runners).
    • A "Machine" step prints nproc, memory and GPUs, so the caps can be tuned.
  • AGENTS.md: documents the new runner and where the CPU suite now runs.

Tests (owner's machine, 24 threads, RTX 4080)

  • CPU (SDS_TEST_GPU=0 PTR_NUM_JOBS=8): 62938 pass, 0 fail, 1 broken (unchanged); 13m52s wall.
  • CUDA as CI runs it (SDS_TEST_CPU=0 SDS_TEST_SKIP=test_aqua PTR_NUM_JOBS=4): 69499 pass, 0 fail, 1 broken; 16m23s wall. This first run had no duration history; later runs start the longest file (test_numeric_ldlt, ~800 s) first.

Local CPU wall time, same machine (SDS_TEST_GPU=0):

run wall
main, old serial runner 18m34s
this branch, 1 worker (PTR_NUM_JOBS=1) 18m45s
this branch, 8 workers 13m52s
this branch + element-type split, 16 workers 9m30s

CUDA as CI runs it: 16m23s (4 workers, unsplit) → 10m31s (8 workers, split), 69499 pass, 1 broken.

In a single process the barrier saves nothing overall, because the suite compiles both paths anyway. It only helps workers that never need LDLᵀ. The ~12 min cold compile of everything is the floor. The 52 min on CI was this compile on a GitHub runner, roughly 2.8× slower than this machine, so moving to kkt is probably the bigger part of the gain. The real lever beyond this PR is the number of kernel specialisations compiled per (T, INT).

What limits it now (follow-up: #91). Each worker compiles the solver for every (T, INT) it tests, so the wall time can't go below the slowest file's cold compile, roughly 12–13 min locally. test_numeric_ldlt on CUDA is mostly GPU kernel compilation of the new LDLᵀ kernel variants. Splitting files doesn't help: I tried running each test/ported/*.jl as its own test, and every part paid the full compile.

No label added, so the agent pipeline (review and auto-merge) doesn't pick this PR up.

🤖 Generated with Claude Code

The CI jobs went from ~21 to ~52 min with the regime-B/C LDLᵀ kernels (#87-#89).
Almost all of it is compilation: on the CPU backend the KA kernels compile
with their caller, and the solver's numeric phase called factorize!, which
branches on the structure at run time and so compiled the LDLᵀ path (all its
kernels) for every Cholesky solver as well. test_api (SPD/HPD only, 4 T x 2
INT) went from 658 s to 2416 s.

* src/solver.jl: the numeric phase dispatches on the structure behind an
  inference barrier that takes the mutable solver handle only (no boxing, so
  the numeric-phase allocation budgets hold); factorize_cholesky! is the
  Cholesky path of factorize!, which keeps its static branch. First solver
  factorization per (T, INT) on the CPU backend: 42-84 s -> 20-36 s.
* test/runtests.jl: ParallelTestRunner (as CUDA.jl and oneAPI.jl); every
  test_*.jl runs in its own module on a worker pool with the shared helpers
  as init_code. The seed 666 goes with each test file since the runner
  reseeds with 1 after init_code. SDS_TEST_ONLY/SDS_TEST_SKIP still work.
* ci.yml: the CPU-only job runs on the self-hosted kkt machine (same check
  name, it is required by the ruleset); PTR_NUM_JOBS caps the workers (8 CPU,
  4 GPU) since the four kkt runners share the machine.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSF5MG6xxpF2D1TCFkb3UZ
A worker's cold compilation sets the wall time, and it grows with the element
types a test file covers. The long files (SPLIT_FILES in test/runtests.jl) now
run as four parts, test_api[Float32] etc., each in a worker of its own:
ELTYPES/REAL_ELTYPES/COMPLEX_ELTYPES are the part's subset (test/utils.jl, set
through Main.SDS_TEST_PART before the helpers are included), and the testsets
that do not loop over the element types run in the Float64 part only
(RUN_SHARED). Test counts are unchanged (CPU 62938, CUDA 69499, 1 broken).

Owner's machine: CPU suite 13m52s -> 9m30s (16 workers), CUDA 16m23s ->
10m31s (8 workers); test_numeric_ldlt's slowest part 226 s against 814 s
unsplit. ci.yml: 32 CPU workers and 12 CUDA workers on kkt.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSF5MG6xxpF2D1TCFkb3UZ
@michel2323 michel2323 added the claude:pr PR opened by the Claude implementer; handled by the Claude pipeline label Oct 6, 2026

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

VERDICT: APPROVE (head ff5bf68)

Not a task PR (owner infrastructure change, no TNN / Report block), so reviewed against AGENTS.md, PLAN.md and the diff rather than a TASKS.md section.

Checked

  • Scope: PLAN.md and TASKS.md untouched, no Manifest.toml; public names, parameter and phase strings unchanged. factorize! keeps its behaviour; factorize_cholesky! is the extracted Cholesky path with a docstring. AGENTS.md edits only document the new runner and where the CPU suite runs.
  • src/solver.jl: the _numeric_phase! barrier takes only the mutable DirectSolver, calls _numeric_phase_cholesky! / _numeric_phase_ldlt! through Base.inferencebarrier, returns nothing; no boxing, and _factorize! never used the old return value. _flat reproduces the previous vector/matrix nzval handling. The CPU allocation budget test through the API (test_api, MadNLP-like loop) still passes on CI, so the dynamic call adds no allocation to the numeric phase.
  • test/runtests.jl against ParallelTestRunner v2.9.1 source: init_code is evaluated per sandbox module, init_worker_code once per worker, filter_tests! returns true only when no positional filter / --list was given (so the env filters compose correctly), the worker reseeds with 1 before the test expression, so Random.seed!(666) inside test_expr after the helper includes is the right place. history_key is a documented kwarg.
  • Element-type split: audited every top-level testset in the 11 SPLIT_FILES and in test/ported/*.jl; each either loops over ELTYPES / REAL_ELTYPES / COMPLEX_ELTYPES (now a per-part subset) or sits behind RUN_SHARED, so no testset runs four times and none is dropped. The capability audit in test_dense guards only the all-types table. @testset for restores the RNG state per iteration, so a part sees the same random draws for its T as the unsplit file did.
  • Tests not weakened: the FGMRES uniform-batch set now filters (Float64, ComplexF32) through ELTYPES, same coverage; no @test_broken added.
  • CI: both jobs green on this head. Log totals match the PR body exactly: CPU on kkt 62938 pass / 1 broken (15m40s), CUDA (H200) 69499 pass / 1 broken (15m51s). Follow-up #91 exists and is open.
  • test/Project.toml: ParallelTestRunner = "2" is a test-only dependency, justified in the PR body; latest release is v2.9.1.

Non-blocking

  • history_key is "gpu" only when SDS_TEST_CPU=0; the owner's default run (both backends) and SDS_TEST_GPU=0 runs share one duration history, which only affects scheduling order.

No blocking findings.

@michel2323
michel2323 merged commit 0d069dd into main Oct 6, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

claude:pr PR opened by the Claude implementer; handled by the Claude pipeline

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant