Skip to content

bench: cuDSS vs SparseDirectSolver.jl comparison per planned feature - #78

Merged
michel2323 merged 3 commits into
mainfrom
bench/cudss-comparison
Oct 2, 2026
Merged

michel2323 merged 3 commits into
mainfrom
bench/cudss-comparison

Conversation

@michel2323

@michel2323 michel2323 commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

No issue: owner-requested benchmark tooling, not a TASKS.md task. No claude:pr label, so the agent review does not run.

Task: none.

What was built:

  • bench/features.jl: one comparison feature per capability planned in TASKS.md (Cholesky, 16 right-hand sides, LDLᵀ, refinement, uniform batch, LU, Schur, matching, non-uniform batch, partitioned-inverse solve, hybrid memory, mixed precision). A feature counts as done when its task is marked done in TASKS.md.
  • bench/compare.jl: times analysis, factorization, refactorization and solve per feature and matrix with BenchmarkTools. It runs one solver per process (--solver=cudss or --solver=sds) and merges the rows into bench/comparison/<solver>.csv. The SDS run skips pending features, so only those need rerunning when a task lands.
  • bench/compare_report.jl: writes bench/comparison/comparison.md (an overview plus a table per feature, with pending features left blank) and comparison.png (SDS/cuDSS ratio per feature and phase). It uses the new bench/report/ environment with CairoMakie.
  • The README shows the plot. bench/README.md and AGENTS.md document the commands.

Current results (RTX 4080, cuDSS 0.8.0, rebased on T16; geometric mean of SDS time / cuDSS time):

feature matrices analysis factorization refactorization solve
Cholesky, Float64 14 0.60× 5.64× 9.06× 8.93×
Cholesky, 16 right-hand sides 14 0.60× 6.06× 9.52× 7.28×
LDLᵀ, static pivoting (T15) 23 1.12× 20.44× 29.78× 13.42×
LDLᵀ + 2 refinement steps, K2 dumps (T16) 9 2.67× 23.40× 35.99× 33.25×

Cholesky residuals match cuDSS. For LDLᵀ, residuals on the K2 dumps are lower than cuDSS's with default settings. With refinement they reach 1e-9 or better on case14 and case118, and stay large on case1354 for both solvers. The T15 LDLᵀ factorization is slow on large fronts, taking 14.6 s on lap3d_40 and 91 s on apache2, as the T25 owner note expects. Those two rows are timed from a single run (--single-run-above, default 5 s) rather than a BenchmarkTools trial.

Tests: SDS_TEST_ONLY="test_bench_smoke,test_bench_compare" SDS_TEST_GPU=0 julia --project=. -e 'using Pkg; Pkg.test()' gives 123 pass / 0 fail / 0 broken on the KA CPU backend. The new test/test_bench_compare.jl covers the feature table, the TASKS.md gate and the Markdown rendering on synthetic CSVs. It runs on the CPU only and does no plotting. CUDA: pending CI.

Deviations from PLAN.md / TASKS.md:

  • Each timed sample spins the GPU for 0.2 s before the phase. Without the spin, the GPU drops to a low clock during the host-side setup, and lap2d_300 factorization varied between 6 and 19 ms. With it the time is a stable 5.9 ms, against 6.05 ms in the T04 baseline.
  • Plain Float32 is not compared, because cuDSS fails on every condensed KKT dump in Float32.
  • The SDS calls for uniform and non-uniform batches assume APIs that do not exist yet. They need a check when T17 and T22 land.
  • New dependencies, all on the bench and test side:
    • bench env: BenchmarkTools for phase timing, Metis for the default ordering, and SparseDirectSolver through [sources].
    • bench/report: CairoMakie for the plot, DelimitedFiles for the CSV, Printf for formatting.
    • test env: DelimitedFiles and Printf for the report test.

Follow-up issues opened: none.

🤖 Generated with Claude Code

https://claude.ai/code/session_01TJrmq7e9jiW4VxBMEVup2E

michel2323 and others added 3 commits October 2, 2026 15:58
bench/features.jl lists one comparison feature per capability planned in
TASKS.md (Cholesky, nrhs = 16, LDLT, refinement, uniform batch, LU, Schur,
matching, non-uniform batch, partitioned-inverse solve, hybrid memory,
mixed precision). bench/compare.jl times the four phases per feature and
matrix with BenchmarkTools, one solver per process, and merges the rows
into bench/comparison/<solver>.csv; the SparseDirectSolver run skips
features whose task is not marked done. bench/compare_report.jl renders
bench/comparison/comparison.md (overview + per-feature tables, pending
features blank) and comparison.png (CairoMakie, bench/report env).

Each timed sample starts on a fresh solver and spins the GPU for 0.2 s so
it leaves its low-clock power state (lap2d_300 factorization: 6-19 ms
without, a stable 5.9 ms with; T04 baseline 6.05 ms). Plain Float32 is not
compared (cuDSS fails on every condensed KKT dump in Float32).

Dependencies: BenchmarkTools and Metis plus SparseDirectSolver via
[sources] in the bench env; CairoMakie, DelimitedFiles, Printf in the new
bench/report env; DelimitedFiles and Printf in the test env.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TJrmq7e9jiW4VxBMEVup2E
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TJrmq7e9jiW4VxBMEVup2E
… rows

The SDS side now covers LDLT (T15) and LDLT with two refinement steps
(T16). The T15 LDLT factorization takes 14.6 s on lap3d_40 and 91 s on
apache2, and every BenchmarkTools sample refactorizes in its setup, so a
row would cost twenty factorizations. compare.jl now runs all phases once
(compilation, info, statistics, residual), times a second run, and records
that run's times when its factorization exceeds --single-run-above
(default 5 s) instead of a trial. A new `samples` column records which;
the report lists single-run rows under the feature's table. The CSV is
rewritten after every row so an interrupted run keeps its rows. Numeric
CSV fields keep their types (nnz(L), samples were promoted to Float64).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TJrmq7e9jiW4VxBMEVup2E
@michel2323
michel2323 force-pushed the bench/cudss-comparison branch from b9909c5 to 3765878 Compare October 2, 2026 22:03
@michel2323
michel2323 merged commit 404b1f4 into main Oct 2, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant