Skip to content

Thesis - #13

Draft
Adames4 wants to merge 54 commits into
masterfrom
thesis
Draft

Adames4 wants to merge 54 commits into
masterfrom
thesis

Conversation

@Adames4

@Adames4 Adames4 commented Jul 21, 2026

Copy link
Copy Markdown
Member

No description provided.

@coderabbitai

coderabbitai Bot commented Jul 21, 2026 •

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch thesis

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request establishes the core configuration infrastructure for a machine learning pipeline designed to automate the histological grading of ulcerative colitis. It introduces a structured Hydra configuration system to manage diverse experimental setups, foundation model integrations, and dataset processing pipelines. Additionally, it provides detailed documentation to facilitate reproducibility and understanding of the project's methodology.

Highlights

  • Documentation: Added a comprehensive README.md detailing the machine learning pipeline for ulcerative colitis grading, including setup, pipeline stages, and experiment descriptions.
  • Configuration Infrastructure: Implemented an extensive Hydra configuration tree for managing datasets, checkpoints, and various ML experiments.
  • Linting: Updated the .ruff.toml configuration to use the correct rule identifier for type-checking blocks.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request sets up the machine learning pipeline and Hydra configurations for automated grading of ulcerative colitis from H&E whole slide images, including a detailed README. Feedback on these changes highlights an invalid Ruff rule code (TC002 instead of TCH002) in .ruff.toml that would cause configuration parsing errors. Additionally, typos were found in several knl_patos embedding configuration files referencing a non-existent default path, and the README table is missing the configs/ prefix for the listed configuration paths.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread .ruff.toml
"D107", # missing docstring in __init__
"N812", # lowercase imported as non lowercase
"TCH002", # move third-party into a type-checking block
"TC002", # move third-party into a type-checking block

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The rule code TC002 is invalid in Ruff. The correct rule code for ignoring third-party imports in type-checking blocks is TCH002 (as part of the TCH / flake8-type-checking ruleset). Changing this to TC002 will cause Ruff to fail with a configuration parsing error.

Suggested change
"TC002", # move third-party into a type-checking block
"TCH002", # move third-party into a type-checking block

@@ -0,0 +1,8 @@
defaults:
- /dataset/tiled/knl_patos/0_224@_here_

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The default config path /dataset/tiled/knl_patos/0_224 does not exist. The repository contains 1_224.yaml under configs/dataset/tiled/knl_patos/. This typo will cause Hydra to fail when composing the configuration.

  - /dataset/tiled/knl_patos/1_224@_here_
References
  1. Hydra Configs (configs/): Reproducibility is paramount. Ensure configurations are correct and complete. (link)

@@ -0,0 +1,8 @@
defaults:
- /dataset/tiled/knl_patos/0_224@_here_

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The default config path /dataset/tiled/knl_patos/0_224 does not exist. The repository contains 1_224.yaml under configs/dataset/tiled/knl_patos/. This typo will cause Hydra to fail when composing the configuration.

  - /dataset/tiled/knl_patos/1_224@_here_
References
  1. Hydra Configs (configs/): Reproducibility is paramount. Ensure configurations are correct and complete. (link)

@@ -0,0 +1,8 @@
defaults:
- /dataset/tiled/knl_patos/0_224@_here_

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The default config path /dataset/tiled/knl_patos/0_224 does not exist. The repository contains 1_224.yaml under configs/dataset/tiled/knl_patos/. This typo will cause Hydra to fail when composing the configuration.

  - /dataset/tiled/knl_patos/1_224@_here_
References
  1. Hydra Configs (configs/): Reproducibility is paramount. Ensure configurations are correct and complete. (link)

Comment thread README.md
Comment on lines +127 to +135
| I | Baseline Classification | `experiment/ml/experiment_i_baseline_classification` | VGG16 tile-level classifier trained on raw pixel tiles |
| II | Ordered Regression | `experiment/ml/experiment_ii_ordered_regression` | Cumulative link loss for NHI ordinal structure |
| III | Pathology Foundation Models | `experiment/ml/experiment_iii_pathology_foundation_model` | Linear probe on frozen prov-gigapath embeddings |
| IV | MIL & Hierarchical Modeling | `experiment/ml/experiment_iv_mil_and_hierarchical_modeling` | Attention MIL; three-task hierarchical decomposition |
| V | More Foundation Models | `experiment/ml/experiment_v_more_foundation_models` | MIL sweep over all foundation models |
| VI | Neutrophil Detection | `experiment/ml/experiment_vi_neutrophil_detection` | Binary tile-level neutrophil classifier |
| VII | Ensembling | `experiment/postprocessing/experiment_vii_ensembling` | Soft majority vote vs. hierarchical routing |
| VIII | Markov Chain Confidence | `experiment/postprocessing/experiment_viii_markov_chain_model_aggregation` | Absorption distribution confidence (entropy / herfindahl / std) |
| — | Final Model | `experiment/ml/final` | Attention MIL (virchow2) trained on all three institutions |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The configuration paths listed in the table are missing the configs/ prefix. According to the repository structure, the experiment configurations are located under configs/experiment/.... Updating these paths will make it easier for users to locate the files in the repository.

Suggested change
| I | Baseline Classification | `experiment/ml/experiment_i_baseline_classification` | VGG16 tile-level classifier trained on raw pixel tiles |
| II | Ordered Regression | `experiment/ml/experiment_ii_ordered_regression` | Cumulative link loss for NHI ordinal structure |
| III | Pathology Foundation Models | `experiment/ml/experiment_iii_pathology_foundation_model` | Linear probe on frozen prov-gigapath embeddings |
| IV | MIL & Hierarchical Modeling | `experiment/ml/experiment_iv_mil_and_hierarchical_modeling` | Attention MIL; three-task hierarchical decomposition |
| V | More Foundation Models | `experiment/ml/experiment_v_more_foundation_models` | MIL sweep over all foundation models |
| VI | Neutrophil Detection | `experiment/ml/experiment_vi_neutrophil_detection` | Binary tile-level neutrophil classifier |
| VII | Ensembling | `experiment/postprocessing/experiment_vii_ensembling` | Soft majority vote vs. hierarchical routing |
| VIII | Markov Chain Confidence | `experiment/postprocessing/experiment_viii_markov_chain_model_aggregation` | Absorption distribution confidence (entropy / herfindahl / std) |
| — | Final Model | `experiment/ml/final` | Attention MIL (virchow2) trained on all three institutions |
| I | Baseline Classification | `configs/experiment/ml/experiment_i_baseline_classification` | VGG16 tile-level classifier trained on raw pixel tiles |
| II | Ordered Regression | `configs/experiment/ml/experiment_ii_ordered_regression` | Cumulative link loss for NHI ordinal structure |
| III | Pathology Foundation Models | `configs/experiment/ml/experiment_iii_pathology_foundation_model` | Linear probe on frozen prov-gigapath embeddings |
| IV | MIL & Hierarchical Modeling | `configs/experiment/ml/experiment_iv_mil_and_hierarchical_modeling` | Attention MIL; three-task hierarchical decomposition |
| V | More Foundation Models | `configs/experiment/ml/experiment_v_more_foundation_models` | MIL sweep over all foundation models |
| VI | Neutrophil Detection | `configs/experiment/ml/experiment_vi_neutrophil_detection` | Binary tile-level neutrophil classifier |
| VII | Ensembling | `configs/experiment/postprocessing/experiment_vii_ensembling` | Soft majority vote vs. hierarchical routing |
| VIII | Markov Chain Confidence | `configs/experiment/postprocessing/experiment_viii_markov_chain_model_aggregation` | Absorption distribution confidence (entropy / herfindahl / std) |
| — | Final Model | `configs/experiment/ml/final` | Attention MIL (virchow2) trained on all three institutions |

Adames4 and others added 9 commits August 29, 2026 17:54
* feat: add ikem validation configuration

* fix: update institution name in ikem validation configuration

* feat: add drop_missing parameter to create_dataset function

* fix: correct typo in dataset configuration parameter

* fix: update condition for institution check in create_dataset function

* feat: add ikem validation configuration file

* feat: add ikem validation configuration and update storage parameters in quality control script

* feat: implement filtering of dataset by QC errors using qc_errors_uri

* fix: improve line handling in filter_dataset_by_qc_errors function

* fix: update qc_main to handle qc_errors.log more efficiently

* feat: enhance QC handling in tiling process with NaN filling and mask availability checks

* fix: update qc_mask URI in ikem validation configuration

* feat: add ikem validation configuration for dataset processing
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant