| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
refactor!: reorganize auto_models into top-level models/components/distributed/trainer/data Squashed 58 commits of the AutoModels directory restructure (M0-M5, S4a-S8d, top-level split): - characterization tests first (M0): UT collection baseline, public API contract snapshots, build-order golden cases, shared fixtures, training YAML snapshots - distributed split (S4a-S4g): sharding_config/planner/applier/injection split into distributed/{plan,recipe_spec,apply,compile, activation_checkpoint} + private _builder/; CP/EP into context_parallel/ and expert_parallel/; fsdp2 into fsdp_adapter + source_shard - models adapter layer (M1-M5): adapter_spec/registry skeleton, Qwen3-MoE replacements and CP/EP contracts migrated into models/qwen3_moe/adapter, single recipe converged onto models/qwen3_moe/recipes/train.yaml, legacy fusion files deleted; per-family sharding_rules declared via ModelAdapterSpec with directory-convention auto-discovery registration - build pipeline (S5a-S5g): checkpoint finalization into checkpoint_loader, plan DTO boundary normalized, god module split into model_builder, registry split into config_resolver + models/registry - data stack (S6a-S6f): indexed/text/online/batching/parallel/vlm/tools migrated to top-level hyper_parallel/data - trainer (S7a-S7e): config tree/callbacks/runtime/base+text+vlm migrated to top-level hyper_parallel/trainer, utils split eliminated - closeout (S8a-S8d): legacy core/shard API, legacy hyper_parallel/models and components compatibility package deleted; import boundary gates - top-level split: auto_models/ dissolved into models/ (adapter layer + api/build_options/replacement/_transformers, external interface), components/ (modules/functional/quantization/losses/optim/checkpoint) and distributed/; references rewritten across py/yaml/sh and setup.py; boundary gates updated (distributed -> models limited to the build_options leaf DTO plus the planner's function-level registry lookup) Verified: full UT 4562 passed / 39 failed (identical to the pre-refactor environment-failure baseline), characterization suites green. | 1 个月前 | |
refactor: split trainer config parsing and align resolver naming - move the YAML/CLI entry point into trainer/config/parser.py; resolver.py keeps the dotted-override engine and is grouped into normalization / dotted-override / node-resolution sections - rename helpers to match the sections (coerce_value -> normalize_value, resolve_root -> resolve_config, _replace_path -> replace_override_path, _coerce_* -> _normalize_*, _normalize_target_args -> _resolve_target_args) - add the read-only Target.callable property and read it from the override path instead of _target_ - let ConfigResolutionError own the path prefix; drop the module-level _fail helper - migrate importers from config.manager to config.parser - add tests/ut/trainer/test_config_overrides.py (12 cases) - finish the AutoModels -> HyperParallel naming sweep in docs and READMEs No behavior change: pylint/lizard clean, differential probes and the new UT pass. | 19 天前 | |
feat: optimize checkpoint save boundaries (#307) | 1 个月前 | |
docs(activation): align checkpoint/swap docs with implementation Update API reference, FAQ, and agent skills per review. Exclude activation_checkpoint.md, pipeline_parallel PP swap section, group_swap, and LlamaFactory activation docs from user-facing materials. | 3 个月前 | |
feat: add chunked causal language model loss | 30 天前 | |
refactor: remove legacy custom ops feature | 28 天前 | |
refactor: 去除 platform 抽象,改用 torch 原生接口 删除 hyper_parallel/platform/ 整个双后端抽象层,相关实现直接使用 torch 原生接口;仅保留的 Torch-only 组件迁至 components/。 主要变更: - 删除 hyper_parallel/platform/**(mindspore 后端、platform.py、 swap_optimizer、activation_checkpoint、loss_parallel_ops、 function_override、init_weights 等) - hyper_parallel/platform/torch/common/moe.py 重命名至 hyper_parallel/components/modules/moe.py(内容不变),并同步更新 examples/torch、tests/torch、核心模块中的 20 余处 import - hyper_parallel/__init__.py 移除 get_platform 导出 - destroy_process_group 中 _P2P_MULTI_STREAM_GROUPS 改从 hyper_parallel.core.pipeline_parallel._p2p 导入 - 删除 examples/mindspore/**、tests/ut/platform/** - scripts/pylint_hyperparallel.py:PLATFORM_ALLOWED_PARTS 移除 hyper_parallel/platform/;删除 instance-platform-assignment (C9003) 规则 - 更新 .agent 规则、docs、docker、CODEOWNERS 等配套内容 327 files changed, 648 insertions(+), 20321 deletions(-) | 21 天前 | |
docs: 整理多核并行、专家并行与 DFunction 文档 - 多核并行:移除无关的 sinkhorn 算子描述,修正文档链接,补充融合 kernel/RATR 说明 - DFunction:统一以 preprocess + infer_layout 描述接口 - 专家并行:移除 zero-overhead activation storage 描述 - Release Notes:调整特性归类 | 3 个月前 | |
docs(fsdp): update FSDP/HSDP guide | 3 个月前 | |
fix: rename LlamaFactory SFT guide | 3 个月前 | |
fix_swap_optimizer | 1 个月前 | |
[doc] sapp & mpipe guide | 3 个月前 | |
refactor: make pipeline parallel torch native | 26 天前 | |
fix: resolve markdownlint MD022/MD032/MD040 violations | 3 个月前 | |
docs: add v1.0.0 comprehensive documentation Add complete documentation for HyperParallel v1.0.0 first official release, covering all implemented features across four document categories: - Release Notes: v1.0.0 full feature list, v0.2.0→v1.0.0 incremental changes, contributor list, merged PRs, known limitations, upgrade guide - README: Updated Chinese and English versions with v1.0.0 full feature checkboxes, new Quick Start examples (Optimizer, AC/Swap, TP Styles), new Documentation section, Optimizer/Trainer/Integration feature sections - User Manual: structured docs layout (getting_started/guide/api/contributing), installation guide, 10 feature usage guides, FAQ, dev/test/release norms - API Reference: per-module interface documentation covering HSDP/FSDP, DTensor, Shard/TP, TP Styles, PP, CP, EP, Process Group, Optimizer, Activation Checkpoint/Swap with full parameter signatures | 3 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 1 个月前 | ||
| 19 天前 | ||
| 1 个月前 | ||
| 3 个月前 | ||
| 30 天前 | ||
| 28 天前 | ||
| 21 天前 | ||
| 3 个月前 | ||
| 3 个月前 | ||
| 3 个月前 | ||
| 1 个月前 | ||
| 3 个月前 | ||
| 26 天前 | ||
| 3 个月前 | ||
| 3 个月前 |