Repository navigation
chore: default to stratified_group kfold everywhere - #21
Open
vojtech-cifka wants to merge 1 commit into
Open
vojtech-cifka wants to merge 1 commit into
vojtech-cifka wants to merge 1 commit into
Conversation
Point both kfold run ids at the splits recomputed over the corrected filter_tiles output (#20). The kfold convention was inconsistent: the base tasks defaulted to plain stratified, while train_*_group_kfold and final_linear_provgigapath_* overrode to stratified_group. final_linear_virchow2_* overrode nothing, so it read stratified_kfold_run_id while its probe used the group split. That mismatch did not change the final training data — final_embedding_tiles sets neither include_folds nor exclude_folds, so the final fit uses every row of kfold_tiles.parquet and ignores the fold column, and both artifacts hold the same tile set. It mattered because the two run ids have to be kept current independently: the stratified one still pointed at a split built over the pre-fix filter_tiles, so the final virchow2 fit would have trained on the old tile set, missing the 30% of annotated tiles the scale bug dropped. The run also recorded a kfold run id its probe never used. Make stratified_group the default in both base tasks and drop the now redundant per-experiment overrides, so one run id drives every experiment. All nine ml experiments resolve to the same kfold run id, verified by composing each config. Plain stratified folds remain available by overriding kfold_strategy and kfold_run_id together. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configuration
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
matejpekar
approved these changes
Oct 9, 2026
vejtek
approved these changes
Oct 9, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Point both kfold run ids at the splits recomputed over the corrected filter_tiles output (#20).
Make stratified_group the default in both base tasks and drop the now redundant per-experiment overrides, so one run id drives every experiment. All nine ml experiments resolve to the same kfold run id, verified by composing each config. Plain stratified folds remain available by overriding kfold_strategy and kfold_run_id together.