Repository navigation
CI: parallel test runner, CPU job on kkt, LDLᵀ compiled only when used - #90
Merged
Merged
Conversation
The CI jobs went from ~21 to ~52 min with the regime-B/C LDLᵀ kernels (#87-#89). Almost all of it is compilation: on the CPU backend the KA kernels compile with their caller, and the solver's numeric phase called factorize!, which branches on the structure at run time and so compiled the LDLᵀ path (all its kernels) for every Cholesky solver as well. test_api (SPD/HPD only, 4 T x 2 INT) went from 658 s to 2416 s. * src/solver.jl: the numeric phase dispatches on the structure behind an inference barrier that takes the mutable solver handle only (no boxing, so the numeric-phase allocation budgets hold); factorize_cholesky! is the Cholesky path of factorize!, which keeps its static branch. First solver factorization per (T, INT) on the CPU backend: 42-84 s -> 20-36 s. * test/runtests.jl: ParallelTestRunner (as CUDA.jl and oneAPI.jl); every test_*.jl runs in its own module on a worker pool with the shared helpers as init_code. The seed 666 goes with each test file since the runner reseeds with 1 after init_code. SDS_TEST_ONLY/SDS_TEST_SKIP still work. * ci.yml: the CPU-only job runs on the self-hosted kkt machine (same check name, it is required by the ruleset); PTR_NUM_JOBS caps the workers (8 CPU, 4 GPU) since the four kkt runners share the machine. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QSF5MG6xxpF2D1TCFkb3UZ
A worker's cold compilation sets the wall time, and it grows with the element types a test file covers. The long files (SPLIT_FILES in test/runtests.jl) now run as four parts, test_api[Float32] etc., each in a worker of its own: ELTYPES/REAL_ELTYPES/COMPLEX_ELTYPES are the part's subset (test/utils.jl, set through Main.SDS_TEST_PART before the helpers are included), and the testsets that do not loop over the element types run in the Float64 part only (RUN_SHARED). Test counts are unchanged (CPU 62938, CUDA 69499, 1 broken). Owner's machine: CPU suite 13m52s -> 9m30s (16 workers), CUDA 16m23s -> 10m31s (8 workers); test_numeric_ldlt's slowest part 226 s against 814 s unsplit. ci.yml: 32 CPU workers and 12 CUDA workers on kkt. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QSF5MG6xxpF2D1TCFkb3UZ
Contributor
There was a problem hiding this comment.
VERDICT: APPROVE (head ff5bf68)
Not a task PR (owner infrastructure change, no TNN / Report block), so reviewed against AGENTS.md, PLAN.md and the diff rather than a TASKS.md section.
Checked
- Scope:
PLAN.mdandTASKS.mduntouched, noManifest.toml; public names, parameter and phase strings unchanged.factorize!keeps its behaviour;factorize_cholesky!is the extracted Cholesky path with a docstring.AGENTS.mdedits only document the new runner and where the CPU suite runs. src/solver.jl: the_numeric_phase!barrier takes only the mutableDirectSolver, calls_numeric_phase_cholesky!/_numeric_phase_ldlt!throughBase.inferencebarrier, returnsnothing; no boxing, and_factorize!never used the old return value._flatreproduces the previous vector/matrixnzvalhandling. The CPU allocation budget test through the API (test_api, MadNLP-like loop) still passes on CI, so the dynamic call adds no allocation to the numeric phase.test/runtests.jlagainst ParallelTestRunner v2.9.1 source:init_codeis evaluated per sandbox module,init_worker_codeonce per worker,filter_tests!returnstrueonly when no positional filter /--listwas given (so the env filters compose correctly), the worker reseeds with 1 before the test expression, soRandom.seed!(666)insidetest_exprafter the helper includes is the right place.history_keyis a documented kwarg.- Element-type split: audited every top-level testset in the 11
SPLIT_FILESand intest/ported/*.jl; each either loops overELTYPES/REAL_ELTYPES/COMPLEX_ELTYPES(now a per-part subset) or sits behindRUN_SHARED, so no testset runs four times and none is dropped. The capability audit intest_denseguards only the all-types table.@testset forrestores the RNG state per iteration, so a part sees the same random draws for itsTas the unsplit file did. - Tests not weakened: the FGMRES uniform-batch set now filters
(Float64, ComplexF32)throughELTYPES, same coverage; no@test_brokenadded. - CI: both jobs green on this head. Log totals match the PR body exactly: CPU on kkt 62938 pass / 1 broken (15m40s), CUDA (H200) 69499 pass / 1 broken (15m51s). Follow-up #91 exists and is open.
test/Project.toml:ParallelTestRunner = "2"is a test-only dependency, justified in the PR body; latest release is v2.9.1.
Non-blocking
history_keyis"gpu"only whenSDS_TEST_CPU=0; the owner's default run (both backends) andSDS_TEST_GPU=0runs share one duration history, which only affects scheduling order.
No blocking findings.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Not a task PR: CI time went from ~21 min (Oct 5) to ~52 min (Oct 6) after the regime-B/C LDLᵀ PRs (#87, #88, #89).
Why it got slower. Almost all of it is compilation; coverage instrumentation makes no difference (measured). Per-file times on CI, before → after the three PRs:
test_apitest_numeric_ldlttest_numeric_cholesky_a/_cOn the CPU backend, KA kernels are ordinary Julia functions and compile together with the code that calls them. The solver's numeric phase called
factorize!, which picks Cholesky or LDLᵀ from the structure at run time with a static call. So every SPD/HPD solver also compiled the whole LDLᵀ path.test_apiruns SPD/HPD only, over 4 element types × 2 index types, and paid for LDLᵀ eight times.CI result (kkt, head ff5bf68): CPU job 16.5 min, CUDA job 17 min. Both green with unchanged counts (62938 / 69499 pass, 1 broken). On
mainthe two jobs took 52 and 45 min. The slowest parts on kkt are the complex-element-type parts at ~12.5 min each.Changes
src/solver.jl: the solver's numeric phase calls the Cholesky or the LDLᵀ path through an inference barrier, so only the path that runs gets compiled. The barrier takes only the mutable solver handle. A barrier at thefactorize!(N, S, …)level would box the immutableNumeric/Symbolic(~1 KB per call) and break the numeric-phase allocation budgets, so I left that level alone.factorize_cholesky!is the Cholesky path offactorize!;factorize!itself behaves as before.test/runtests.jl: switched to ParallelTestRunner.jl, as CUDA.jl and oneAPI.jl use.test_*.jlruns in its own module on a pool of worker processes; the shared helpers are loaded ininit_code.init_code. Without this,test_refinementfails.SDS_TEST_ONLY,SDS_TEST_SKIP,SDS_TEST_CPUandSDS_TEST_GPUstill work. Test arguments (prefix match,!nameto exclude) andPTR_NUM_JOBS/--jobs=Nare new.SPLIT_FILESintest/runtests.jl) run as four parts each,test_api[Float32]and so on, each part in a worker of its own.ELTYPES,REAL_ELTYPESandCOMPLEX_ELTYPESare that part's subset.Float64part only, behindRUN_SHARED && @testset …. Test counts are unchanged.test_numeric_ldltcold: 814 s unsplit → 226 s for its slowest part.ci.yml:kktmachine. The job id and matrix are unchanged, so the required check nametest-github-cpuonly (ubuntu-latest, 1, x64)still matches the ruleset.PTR_NUM_JOBSis set to 32 for the CPU job and 12 for the GPU job (kkt: 192 threads, 2 TB, 3× H200, shared by its four runners).nproc, memory and GPUs, so the caps can be tuned.AGENTS.md: documents the new runner and where the CPU suite now runs.Tests (owner's machine, 24 threads, RTX 4080)
SDS_TEST_GPU=0 PTR_NUM_JOBS=8): 62938 pass, 0 fail, 1 broken (unchanged); 13m52s wall.SDS_TEST_CPU=0 SDS_TEST_SKIP=test_aqua PTR_NUM_JOBS=4): 69499 pass, 0 fail, 1 broken; 16m23s wall. This first run had no duration history; later runs start the longest file (test_numeric_ldlt, ~800 s) first.Local CPU wall time, same machine (
SDS_TEST_GPU=0):main, old serial runnerPTR_NUM_JOBS=1)CUDA as CI runs it: 16m23s (4 workers, unsplit) → 10m31s (8 workers, split), 69499 pass, 1 broken.
In a single process the barrier saves nothing overall, because the suite compiles both paths anyway. It only helps workers that never need LDLᵀ. The ~12 min cold compile of everything is the floor. The 52 min on CI was this compile on a GitHub runner, roughly 2.8× slower than this machine, so moving to kkt is probably the bigger part of the gain. The real lever beyond this PR is the number of kernel specialisations compiled per (T, INT).
What limits it now (follow-up: #91). Each worker compiles the solver for every (T, INT) it tests, so the wall time can't go below the slowest file's cold compile, roughly 12–13 min locally.
test_numeric_ldlton CUDA is mostly GPU kernel compilation of the new LDLᵀ kernel variants. Splitting files doesn't help: I tried running eachtest/ported/*.jlas its own test, and every part paid the full compile.No label added, so the agent pipeline (review and auto-merge) doesn't pick this PR up.
🤖 Generated with Claude Code