Large-model training increasingly relies on multi-dimensional parallelism, including combinations of DP, TP, EP, PP, CP, SP, ZeRO-related sharding, and different pipeline scheduling strategies.
In practice, selecting a good parallel strategy is difficult because:
the search space grows quickly with the number of parallel dimensions
different dimensions interact with each other and cannot be tuned independently
memory feasibility and performance efficiency must be considered together
manual exploration is slow, costly, and often hardware-dependent
For many realistic workloads, users need an automatic way to evaluate candidate parallel configurations before launching large-scale training jobs.
ND is intended to address this problem by providing symbolic estimation for N-dimensional parallelism. Its analytical nature enables:
exhaustive exploration over candidate configurations
fast exploration in seconds
no dependency on the execution cluster during search time
This issue proposes to introduce ND into hyper_parallel/auto_parallel as a symbolic planning module for multi-dimensional parallel strategy exploration.
Motivation
The goal of ND is to provide a systematic way to search and evaluate candidate parallel strategies for large-model training.
More specifically, ND is intended to help with the following scenarios:
users need to explore multiple parallel dimensions jointly instead of tuning one dimension at a time
users want a fast and explainable way to estimate whether a strategy is feasible
users need both memory estimation and performance estimation to guide strategy selection
users want to reduce manual trial-and-error on real clusters
Compared with manual tuning, ND aims to provide:
lower strategy search cost
faster exploration of candidate configurations
better visibility into memory and performance trade-offs
a reusable planning workflow for different models and hardware settings
Proposal
We propose to introduce ND as a symbolic estimation module for multi-dimensional parallel strategy exploration.
ND will focus on two core capabilities:
Memory modeling
estimate memory usage under different parallel strategies
reason about peak memory, static memory, dynamic memory, and stage-level memory behavior
support detailed memory insight and optional artifacts useful for downstream modules such as PPB
Performance modeling
estimate performance-related cost under different parallel strategies
support fast ranking or filtering of candidate strategies
enable symbolic exploration of feasible and efficient configurations
The overall goal is to use symbolic estimation to search the strategy space efficiently and return a set of promising parallel configurations under user-provided constraints.
Design Overview
1. Core idea
ND takes model configuration and strategy constraints as input, then symbolically evaluates candidate parallel strategies across multiple dimensions.
The exploration process should be:
fast, so that large candidate spaces can be explored in seconds or minutes rather than by repeated real execution
explainable, so that estimated memory/performance signals can be inspected
modular, so that memory modeling and performance modeling can evolve independently while still contributing to a unified ND workflow
2. Memory modeling
The memory modeling part is responsible for estimating memory usage under different parallel strategies.
Based on the current memory estimation design, the required capabilities include:
estimating peak memory usage under parallelism
distinguishing static and dynamic memory contributions
providing stage-level memory insight
supporting different pipeline schedulings and recomputation-related settings
optionally exporting layer descriptions for downstream pipeline balancing
This part should serve both:
ND’s internal strategy search
downstream modules such as PPB
3. Performance modeling
The performance modeling part is responsible for estimating the execution-related quality of candidate strategies.
The required capabilities include:
evaluating candidate strategies symbolically without relying on the execution cluster at search time
comparing configurations across multiple dimensions
ranking or filtering configurations based on estimated performance-related cost
supporting ND’s top-k or best-strategy style outputs
Background
Large-model training increasingly relies on multi-dimensional parallelism, including combinations of DP, TP, EP, PP, CP, SP, ZeRO-related sharding, and different pipeline scheduling strategies.
In practice, selecting a good parallel strategy is difficult because:
For many realistic workloads, users need an automatic way to evaluate candidate parallel configurations before launching large-scale training jobs.
ND is intended to address this problem by providing symbolic estimation for N-dimensional parallelism. Its analytical nature enables:
This issue proposes to introduce ND into
hyper_parallel/auto_parallelas a symbolic planning module for multi-dimensional parallel strategy exploration.Motivation
The goal of ND is to provide a systematic way to search and evaluate candidate parallel strategies for large-model training.
More specifically, ND is intended to help with the following scenarios:
Compared with manual tuning, ND aims to provide:
Proposal
We propose to introduce ND as a symbolic estimation module for multi-dimensional parallel strategy exploration.
ND will focus on two core capabilities:
Memory modeling
Performance modeling
The overall goal is to use symbolic estimation to search the strategy space efficiently and return a set of promising parallel configurations under user-provided constraints.
Design Overview
1. Core idea
ND takes model configuration and strategy constraints as input, then symbolically evaluates candidate parallel strategies across multiple dimensions.
The exploration process should be:
2. Memory modeling
The memory modeling part is responsible for estimating memory usage under different parallel strategies.
Based on the current memory estimation design, the required capabilities include:
This part should serve both:
3. Performance modeling
The performance modeling part is responsible for estimating the execution-related quality of candidate strategies.
The required capabilities include:
4. ND workflow
At a high level, the expected workflow of ND is:
Inputs / Outputs
Inputs
ND is expected to take inputs including:
model configuration
search space / strategy constraints
estimation options
Outputs
ND is expected to provide outputs including:
candidate strategy results
memory-related analysis
performance-related analysis
optional artifacts
Expected Benefits
After introducing ND, we expect the following benefits:
Development Plan
This work is expected to be delivered through multiple PRs, rather than a single large PR.
The initial plan is to split the implementation into at least the following stages:
PR 1: ND memory modeling
This PR will focus on the memory estimation side of ND, including:
PR 2: ND performance modeling
This PR will focus on the performance estimation side of ND, including: