| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
fix: align CP metrics and TP grad norm metadata (#1497) * fix: align CP metrics and TP grad norm metadata Two independent correctness fixes for the Megatron engine. 1. CP metrics alignment (stats_tracker): Per-key reduce_group overrides may be wider than the default DP export group (e.g. DP+CP for token-level metrics). Add a `key_sync_group` argument so override keys sync their key set / metadata and reduce over DP+CP, while default keys keep the DP alignment. Aggregation helpers now emit identity placeholder tensors (0 / +inf / -inf) for keys absent on a rank, so every rank still participates in the collective and avoids hangs / mismatched reductions when CP > 1. 2. TP grad norm metadata (megatron_engine): `_mark_duplicated_params` now also clears `param.tensor_model_parallel` on replicated (tp_size == 1, non-expert) params. Megatron's optimizer uses that attribute to decide which TP ranks contribute to grad norm/clipping, so leaving it True double-counts duplicated params when TP > 1. Clearing it is also consistent with `all_gather_param` and `hf_save`, which already treat non-TP params as replicated. Adds unit tests for both paths (tests/test_stats_tracker.py and a new case in tests/test_megatron_engine.py). * test: cover all_gather_param routing and duplicated-param grad norm Add two tests backing the TP grad-norm metadata fix: - tests/test_all_gather_param.py: assert `all_gather_param` returns the param as-is (no TP all-gather) when `tensor_model_parallel` is False or the name is in `duplicated_param_names`, and only all-gathers genuine TP-sharded params. Uses `pytest.importorskip` so it runs in the megatron CI env and skips gracefully elsewhere. - tests/test_grad_norm_duplicated.py: a CPU/gloo 2- and 4-rank test showing a replicated param is counted once (correct) when tensor_model_parallel is False, and inflated by sqrt(tp) when it is (incorrectly) left True. Mirrors Megatron's param_is_not_tensor_parallel_duplicate selection. * test: add 2-GPU grad-norm TP-invariance integration test End-to-end guard for the TP grad-norm metadata fix: run the real MegatronEngine at TP=1 and TP=2 on identical deterministic input via torchrun and assert the reported grad norm is TP-invariant. Double-counting replicated params (tensor_model_parallel left True) would inflate it as TP grows. Also asserts the fix demoted at least one real duplicated param at TP=2. Marked multi_gpu/slow and skipped without >= 2 GPUs. * refactor(stats): dedup identity tensors, avoid min/max GPU sync Address review feedback on the CP metrics fix: - Generalize the placeholder-tensor helper to `_placeholder_scalar(fill=...)` so the SUM/AVG (0.0) and MIN/MAX (+/-inf) empty-value branches share one device-aware constructor instead of four inline copies. - In `_min_of`/`_max_of`, reduce the per-shard list with `torch.stack(xs).min()` / `.max()` instead of Python `min()`/`max()`, which forced a CPU-GPU sync. - In the SCALAR branch, read values via `self.stats.get(key, [])` (a rank may learn a key only through metadata sync) and guard the `value / cnt` division against `cnt == 0` so a key absent everywhere yields 0.0 instead of NaN. - Add a regression test for the missing-on-this-rank SCALAR path. | 2 个月前 | |
fix: align CP metrics and TP grad norm metadata (#1497) * fix: align CP metrics and TP grad norm metadata Two independent correctness fixes for the Megatron engine. 1. CP metrics alignment (stats_tracker): Per-key reduce_group overrides may be wider than the default DP export group (e.g. DP+CP for token-level metrics). Add a `key_sync_group` argument so override keys sync their key set / metadata and reduce over DP+CP, while default keys keep the DP alignment. Aggregation helpers now emit identity placeholder tensors (0 / +inf / -inf) for keys absent on a rank, so every rank still participates in the collective and avoids hangs / mismatched reductions when CP > 1. 2. TP grad norm metadata (megatron_engine): `_mark_duplicated_params` now also clears `param.tensor_model_parallel` on replicated (tp_size == 1, non-expert) params. Megatron's optimizer uses that attribute to decide which TP ranks contribute to grad norm/clipping, so leaving it True double-counts duplicated params when TP > 1. Clearing it is also consistent with `all_gather_param` and `hf_save`, which already treat non-TP params as replicated. Adds unit tests for both paths (tests/test_stats_tracker.py and a new case in tests/test_megatron_engine.py). * test: cover all_gather_param routing and duplicated-param grad norm Add two tests backing the TP grad-norm metadata fix: - tests/test_all_gather_param.py: assert `all_gather_param` returns the param as-is (no TP all-gather) when `tensor_model_parallel` is False or the name is in `duplicated_param_names`, and only all-gathers genuine TP-sharded params. Uses `pytest.importorskip` so it runs in the megatron CI env and skips gracefully elsewhere. - tests/test_grad_norm_duplicated.py: a CPU/gloo 2- and 4-rank test showing a replicated param is counted once (correct) when tensor_model_parallel is False, and inflated by sqrt(tp) when it is (incorrectly) left True. Mirrors Megatron's param_is_not_tensor_parallel_duplicate selection. * test: add 2-GPU grad-norm TP-invariance integration test End-to-end guard for the TP grad-norm metadata fix: run the real MegatronEngine at TP=1 and TP=2 on identical deterministic input via torchrun and assert the reported grad norm is TP-invariant. Double-counting replicated params (tensor_model_parallel left True) would inflate it as TP grows. Also asserts the fix demoted at least one real duplicated param at TP=2. Marked multi_gpu/slow and skipped without >= 2 GPUs. * refactor(stats): dedup identity tensors, avoid min/max GPU sync Address review feedback on the CP metrics fix: - Generalize the placeholder-tensor helper to `_placeholder_scalar(fill=...)` so the SUM/AVG (0.0) and MIN/MAX (+/-inf) empty-value branches share one device-aware constructor instead of four inline copies. - In `_min_of`/`_max_of`, reduce the per-shard list with `torch.stack(xs).min()` / `.max()` instead of Python `min()`/`max()`, which forced a CPU-GPU sync. - In the SCALAR branch, read values via `self.stats.get(key, [])` (a rank may learn a key only through metadata sync) and guard the `value / cnt` division against `cnt == 0` so a key absent everywhere yields 0.0 instead of NaN. - Add a regression test for the missing-on-this-rank SCALAR path. | 2 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 2 个月前 | ||
| 2 个月前 |