Skip to content

Settle MATH SFT/CISPO through sidecar-owned mlx-rl - #89

Closed
JoshuaPurtell wants to merge 1 commit into
mainfrom
fix/math-sft-cispo-settle
Closed

JoshuaPurtell wants to merge 1 commit into
mainfrom
fix/math-sft-cispo-settle

Conversation

@JoshuaPurtell

Copy link
Copy Markdown
Contributor

Summary

  • Admit local SFT/CISPO from projected optimization_algorithms (and the older algorithms handshake).
  • Persist the SFT train JSONL digest as sft.dataset.validated, retry mlx-rl event polls across GPU-blocked HTTP resets, and settle CISPO checkpoint identity from sidecar item id.

This is the MATH-only settlement from the acceptance-run work, ported onto current main. It is not the Craftax acceptance branch.

Test plan

  • optimizers::sidecar_training::tests::training_ready_reads_projected_optimization_algorithms
  • optimizers::kernel::algorithms::cispo::tests::sidecar_checkpoint_item_id_settles
  • Live MATH SFT → CISPO on a build from this branch (Apple Silicon + resident 2B)

Made with Cursor

Admit training from projected optimization_algorithms, persist the SFT
dataset digest, retry GPU-blocked mlx-rl event polls, and read CISPO
checkpoint identity from sidecar item payloads.

Co-authored-by: Cursor <cursoragent@cursor.com>
@JoshuaPurtell

Copy link
Copy Markdown
Contributor Author

Superseded by the current internal Workshop release line. The projected optimizer algorithm admission, SFT dataset validation, and bounded sidecar polling are already present there; the remaining mlx-rl item-id checkpoint behavior was ported with a focused regression in merged PR #117. Closing this conflicting public-main PR so it is not mistaken for release authority.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant