Milestones
List view
Definition of done (PLAN.md §5): A posteriori threshold pivoting with optional delayed pivots and host re-analysis; `factor_precision = Float32` + FGMRES-IR mixed precision (also the Metal Float64 story); CPU MA27/MA57 fallback guidance documented for hard nonconvex NLPs. --- _Generated by [Claude Code](https://claude.ai/code)_
No due date•0/1 issues closedDefinition of done (PLAN.md §5): Host-resident panels (`hybrid_memory_mode`, limits, pinned memory), CPU-backend execution of small levels (`hybrid_execute_mode`, `host_nthreads`). --- _Generated by [Claude Code](https://claude.ai/code)_
No due date•0/1 issues closedDefinition of done (PLAN.md §5): Partitioned-inverse solve option, sync-free sweeps and graph capture in the CUDA/ROCm extensions, level merging, amalgamation and bin tuning per backend, warp-level fast paths once KernelInterface stabilizes, benchmark tracking vs cuDSS, CHOLMOD and MA57. **Target**: match or beat cuDSS refactorization+solve on ACOPF KKT matrices on CUDA; parity ratios reported for AMD/Intel. --- _Generated by [Claude Code](https://claude.ai/code)_
No due date•4/5 issues closedDefinition of done (PLAN.md §5): `BatchedDirectSolver`, block-diagonal packing, `test_nonuniform_batch_cudss.jl` ported. --- _Generated by [Claude Code](https://claude.ai/code)_
No due date•0/1 issues closedDefinition of done (PLAN.md §5): MC64-style jobs (max product first), `perm_matching`, `scale_row/col`, inertia correct with matching. --- _Generated by [Claude Code](https://claude.ai/code)_
No due date•3/3 issues closedDefinition of done (PLAN.md §5): `schur_mode`, `user_schur_indices`, `schur_shape`, dense and CSR `schur_matrix`, `solve_*_schur`; `test_schur_cudss.jl` ported and enabled. --- _Generated by [Claude Code](https://claude.ai/code)_
No due date•2/2 issues closedDefinition of done (PLAN.md §5): Symmetric-pattern multifrontal LU with in-block pivoting and GESP-style perturbation, `lu`/`lu!`, `perm_row/col`, transpose/adjoint `solve_mode`; optional up-looking row-per-workgroup kernel for extremely sparse factors (Ginkgo style) if the harness shows a gap. --- _Generated by [Claude Code](https://claude.ai/code)_
No due date•2/2 issues closedDefinition of done (PLAN.md §5): Strided/3-D arrays, `ubatch_size/index/mask`, per-member `info`/stats, `nrhs > 1` at any batch size, generic API auto-detect, strided-batched vendor calls on root fronts. MadNLP two-stage Schur KKT and ExaModels multi-scenario models run on it. --- _Generated by [Claude Code](https://claude.ai/code)_
No due date•2/2 issues closedDefinition of done (PLAN.md §5): IR (`ir_n_steps`, `ir_tol`, actual steps), FGMRES-IR extension, all solve sub-phases, `solve_mode`, `perm_*` getters, `user_host_interrupt`, logging. --- _Generated by [Claude Code](https://claude.ai/code)_
No due date•4/4 issues closedDefinition of done (PLAN.md §5): In-front Bunch–Kaufman, `pivot_threshold`, `pivot_epsilon(_alg)`, `pivot_sign`, `inertia`, `pivot_stats`, `npivots`, `diag`, `solve_diag`, `ldlt`/`ldlt!`; vendor `sytrf` with post-check on root fronts. MadNLPGPU gets a `SparseDirectSolver` option; validated on OPF/ExaModels KKTs (MadNLP K2/K2r, MadIPM, MadNCL settings) against the cuDSS path, including refinement counts against MA27/MA57. --- _Generated by [Claude Code](https://claude.ai/code)_
No due date•8/9 issues closedDefinition of done (PLAN.md §5): All phases, single and multi RHS, `info`, `cholesky`/`cholesky!`/`ldiv!`/`\`, `Hermitian` wrapper, async flag; subtree/level solve sweeps. Ported `test_cudss.jl` subsets pass on all backends. **Target**: within 1.5× of cuDSS Cholesky refactorization+solve on condensed pglib-opf systems on CUDA, running on AMD and Intel. --- _Generated by [Claude Code](https://claude.ai/code)_
No due date•4/4 issues closedDefinition of done (PLAN.md §5): Regime A fused subtree kernels, regime B fused per-front kernels (Cholesky first, LDLᵀ/LU hooks), regime C vendor bindings in all four extensions with KA tiled fallbacks, assembly/extend-add kernels. Unit tests against the CPU reference on every backend; per-bin micro-benchmarks deciding where vendor batched calls replace regime B. --- _Generated by [Claude Code](https://claude.ai/code)_
No due date•6/7 issues closedDefinition of done (PLAN.md §5): Pattern, orderings with the AMD/ND cost model, etree, column counts, GPU-tuned amalgamation, subtree partition, size bins, level lists, static layout, device maps, `memory_estimates`, `flops`, `lu_nnz`, `nsuperpanels`, ND-tree I/O, `max_lu_nnz`. Validated against CHOLMOD's nnz(L) and etree; supernode/level statistics compared with the cuDSS baseline on the harness matrices. A CPU reference numeric factorization (plain Julia, same layout) for testing. --- _Generated by [Claude Code](https://claude.ai/code)_
No due date•8/9 issues closedDefinition of done (PLAN.md §5): Package, four backend extensions with CSR adapters, CI matrix (GitHub Actions CPU; Buildkite juliagpu queue for CUDA/AMDGPU/oneAPI/Metal; self-hosted `kkt`), `Options` with the full name tables, error types, Aqua. **Capability audit** script filling the §2.6 table per backend and eltype (kept as a test). **Benchmark harness**: NREL opf_matrices, pglib-opf KKT and condensed matrices dumped from MadNLP, a CUTEst subset; cuDSS analysis/factorization/solve times, `flops`, supernode statistics and MA57 refinement counts recorded as the baseline every later milestone is measured against. --- _Generated by [Claude Code](https://claude.ai/code)_
No due date•7/7 issues closed