| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
feat(api): add unified RejectionSamplingConfig for async training (#1088) Replace behave_imp_weight_cap/behave_imp_weight_mode with unified RejectionSamplingConfig supporting multiple metrics (ratio, kl_k1, kl_k2, kl_k3), levels (token/sequence), and actions (mask/clamp). Key changes: - Add RejectionSamplingConfig dataclass with comprehensive validation - Implement apply_rejection_sampling for 1D packed and 2D padded formats - Fix loss denominator scaling bug in mask mode (save count before filtering) - Use geometric mean for sequence-level ratio aggregation (matching GSPO) - Broadcast sequence-level geometric mean as uniform behave_imp_weight - Warn when use_decoupled_loss=True but rejection_sampling is None - Update ppo_actor_loss_fn and grpo_loss_fn to use new config - Migrate 40 example configs to new rejection_sampling field - Add 43 unit tests covering all modes, metrics, and edge cases Refs: #1052 | 5 个月前 | |
refactor: flatten sub-module imports to use parent package re-exports (#996) Add __init__.py with lazy re-exports (__getattr__ + __all__) for areal/api, areal/engine, areal/reward, areal/workflow, and __all__ for areal/dataset, then rewrite all external imports across the codebase to use the shorter parent-package form (e.g. `from areal.api import TrainEngine`). Key changes: - All re-exports are fully lazy via __getattr__ (no eager imports) - Flatten ~100 files across areal/, tests/, examples/ - Preserve cli_args deep imports (dozens of config classes, including SchedulingSpec which stays in cli_args) - Preserve intra-package relative imports to avoid circular deps - Reward submodules use sibling-relative imports (from . import ...) - Preserve non-exported symbols (DeviceRuntimeInfo, HttpRequest, etc.) Co-authored-by: Wentai Zhang <zhangwentai.zwt@antgroup.com> | 6 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 5 个月前 | ||
| 6 个月前 |