bench: cuDSS vs SparseDirectSolver.jl comparison per planned feature - #78
Merged
Merged
Conversation
bench/features.jl lists one comparison feature per capability planned in TASKS.md (Cholesky, nrhs = 16, LDLT, refinement, uniform batch, LU, Schur, matching, non-uniform batch, partitioned-inverse solve, hybrid memory, mixed precision). bench/compare.jl times the four phases per feature and matrix with BenchmarkTools, one solver per process, and merges the rows into bench/comparison/<solver>.csv; the SparseDirectSolver run skips features whose task is not marked done. bench/compare_report.jl renders bench/comparison/comparison.md (overview + per-feature tables, pending features blank) and comparison.png (CairoMakie, bench/report env). Each timed sample starts on a fresh solver and spins the GPU for 0.2 s so it leaves its low-clock power state (lap2d_300 factorization: 6-19 ms without, a stable 5.9 ms with; T04 baseline 6.05 ms). Plain Float32 is not compared (cuDSS fails on every condensed KKT dump in Float32). Dependencies: BenchmarkTools and Metis plus SparseDirectSolver via [sources] in the bench env; CairoMakie, DelimitedFiles, Printf in the new bench/report env; DelimitedFiles and Printf in the test env. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TJrmq7e9jiW4VxBMEVup2E
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TJrmq7e9jiW4VxBMEVup2E
… rows The SDS side now covers LDLT (T15) and LDLT with two refinement steps (T16). The T15 LDLT factorization takes 14.6 s on lap3d_40 and 91 s on apache2, and every BenchmarkTools sample refactorizes in its setup, so a row would cost twenty factorizations. compare.jl now runs all phases once (compilation, info, statistics, residual), times a second run, and records that run's times when its factorization exceeds --single-run-above (default 5 s) instead of a trial. A new `samples` column records which; the report lists single-run rows under the feature's table. The CSV is rewritten after every row so an interrupted run keeps its rows. Numeric CSV fields keep their types (nnz(L), samples were promoted to Float64). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TJrmq7e9jiW4VxBMEVup2E
michel2323
force-pushed
the
bench/cudss-comparison
branch
from
October 2, 2026 22:03
b9909c5 to
3765878
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No issue: owner-requested benchmark tooling, not a TASKS.md task. No
claude:prlabel, so the agent review does not run.Task: none.
What was built:
bench/features.jl: one comparison feature per capability planned in TASKS.md (Cholesky, 16 right-hand sides, LDLᵀ, refinement, uniform batch, LU, Schur, matching, non-uniform batch, partitioned-inverse solve, hybrid memory, mixed precision). A feature counts as done when its task is marked done in TASKS.md.bench/compare.jl: times analysis, factorization, refactorization and solve per feature and matrix with BenchmarkTools. It runs one solver per process (--solver=cudssor--solver=sds) and merges the rows intobench/comparison/<solver>.csv. The SDS run skips pending features, so only those need rerunning when a task lands.bench/compare_report.jl: writesbench/comparison/comparison.md(an overview plus a table per feature, with pending features left blank) andcomparison.png(SDS/cuDSS ratio per feature and phase). It uses the newbench/report/environment with CairoMakie.bench/README.mdandAGENTS.mddocument the commands.Current results (RTX 4080, cuDSS 0.8.0, rebased on T16; geometric mean of SDS time / cuDSS time):
Cholesky residuals match cuDSS. For LDLᵀ, residuals on the K2 dumps are lower than cuDSS's with default settings. With refinement they reach 1e-9 or better on case14 and case118, and stay large on case1354 for both solvers. The T15 LDLᵀ factorization is slow on large fronts, taking 14.6 s on lap3d_40 and 91 s on apache2, as the T25 owner note expects. Those two rows are timed from a single run (
--single-run-above, default 5 s) rather than a BenchmarkTools trial.Tests:
SDS_TEST_ONLY="test_bench_smoke,test_bench_compare" SDS_TEST_GPU=0 julia --project=. -e 'using Pkg; Pkg.test()'gives 123 pass / 0 fail / 0 broken on the KA CPU backend. The newtest/test_bench_compare.jlcovers the feature table, the TASKS.md gate and the Markdown rendering on synthetic CSVs. It runs on the CPU only and does no plotting. CUDA: pending CI.Deviations from PLAN.md / TASKS.md:
[sources].bench/report: CairoMakie for the plot, DelimitedFiles for the CSV, Printf for formatting.Follow-up issues opened: none.
🤖 Generated with Claude Code
https://claude.ai/code/session_01TJrmq7e9jiW4VxBMEVup2E