Skip to content

feat: add reproducible reasoning SFT on LUMI - #61

Open
BirgerMoell wants to merge 1 commit into
OpenEuroLLM:mainfrom
BirgerMoell:feat/reasoning-sft-reproducibility
Open

BirgerMoell wants to merge 1 commit into
OpenEuroLLM:mainfrom
BirgerMoell:feat/reasoning-sft-reproducibility

Conversation

@BirgerMoell

Copy link
Copy Markdown

Summary

Adds the missing shared infrastructure needed to express a reproducible reasoning-SFT continuation on LUMI:

  • supports a full accelerate.config_file launch profile and freezes it beside the generated SLURM job, enabling the tested multi-node FSDP shape without a bespoke launcher;
  • adds a LUMI FSDP profile and a production-oriented 9B/16K reasoning-SFT config for the pinned Anneal-300B instruction checkpoint;
  • adds dataset revision pinning through loading, offline prefetching, and data inspection;
  • includes revision, subset, split, transform, and other row-defining fields in automatic run identity, including single-dataset runs;
  • documents the assistant-only masking, tokenize-only gate, checkpoint/evaluation workflow, and the contract for an immutable token-balanced reasoning artifact.

The guide deliberately does not present row-weighted source mixing as equivalent to the current production recipe. Token allocation, cross-source deduplication, loop rejection, decontamination, and manifest creation remain explicit upstream data-curation responsibilities; this trainer consumes their materialized Parquet artifact.

Type of change

  • Bug fix
  • New feature
  • Refactor
  • Performance
  • Documentation
  • Maintenance

Validation

  • Full test suite: 160 passed
  • Ruff 0.9.10 lint and format checks passed
  • Black check passed
  • Python compilation and git diff --check passed
  • Rendered the reasoning profile end to end and verified that the FSDP config is frozen into the run directory, forwarded with --config_file, and no DeepSpeed flags are emitted
  • Verified that the reference token budget resolves to exactly 500 optimizer steps (524,288,000 packed tokens at global batch 64 and sequence length 16,384)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant