Revert "Always Pre-Split Microbatches for PP" - #4042
Open
wconstab wants to merge 1 commit into
Open
Conversation
This reverts commit 9228564.
wconstab
requested review from
IvanKobzarev,
SherlockNoMad,
aditvenk,
fegin,
felipemello1,
sanketpurandare,
tianyu-l,
wwwjn and
xmfan
as code owners
July 31, 2026 23:12
Contributor
|
@wconstab the upstream change had already been landed: pytorch/pytorch#188500 |
wconstab
added a commit
that referenced
this pull request
Aug 1, 2026
Fixes the shared GraphTrainer failure seen on main and the PP revert PR: - Main GraphTrainer 8 GPU Integration Tests: https://github.com/pytorch/torchtitan/actions/runs/30675722021/job/91302425570 - Revert PR #4042 GraphTrainer 8 GPU Integration Tests: https://github.com/pytorch/torchtitan/actions/runs/30672227677/job/91292165456 Both fail in `TestMetadataPropagation::test_backward_nodes_have_stack_trace` with `AssertionError: 23 != 24`, while `bwd_nodes_missing_stack_trace = []`. The exact number of eligible FX/autograd nodes is not the behavior this test needs to lock down; it can change as PyTorch tracing/decomposition/autograd internals change. The invariant is that backward nodes corresponding to forward nodes with stack traces also have stack traces. This PR keeps that assertion and only replaces the brittle exact graph-shape count with a sanity check that the test actually examined at least one backward node. Validation: - `python3 -m py_compile torchtitan/experiments/graph_trainer/tests/test_trace_module.py` Note: local `pytest` was unavailable in the default Python environment on my host, so CI should be used for the targeted runtime check. This PR does not address the separate H100 GraphTrainer integration failures.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Revert #3856 because it broke PP model tests by passing pre-split microbatches through
pp_schedule.step()asarg_mbs/kwarg_mbs/target_mbs.step()is the public whole-batch API and re-splits its inputs internally. The pre-split arguments are only accepted by the private_step_microbatches()path, so the scheduler sees onemicrobatch and fails with errors like
ValueError: Expecting 8 arg_mbs but got 1.This restores main while we work out a proper pre-split PP API/path for varlen metadata.
@sanketpurandare can you confirm? Possibly this was coordinated with an upstream pytorch change that we're not using yet?