| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
perf: reduce Megatron training memory peaks (#1555) * perf: reduce Megatron training memory peaks Add an SFT profiling workflow and use its memory snapshots to remove full-sequence vocabulary and optimizer gradient peaks from Megatron training. Key changes: - Add rank-aware kernel and memory profiling for packed SFT workloads - Fuse FP32 vocab-parallel logprob storage with LM head backward - Add optional true chunked LM head loss with recomputed backward - Configure precision-aware optimizer fields before Megatron validation - Cover BF16/FP32 numerical parity and distributed TP/SP behavior * fix(models): avoid private storage identity checks Track the LM head output tensor weakly and compare storage through the public data_ptr API. This preserves allocator-address reuse protection without depending on PyTorch's private storage _cdata field. * test: compare parameter storage without object identity Parameter.data may return a fresh Tensor wrapper on each access. Verify that replicated parameters retain their data pointer and storage offset instead of comparing transient Python objects. * test: make recycled CUDA storage check deterministic Construct the replacement tensor from the original storage instead of relying on the caching allocator to immediately reuse a freed address after the full CI suite. * fix(engine): guard AReaL LM Head storage reuse Keep entropy differentiable unless its gradients are disabled, and make the optimized LM Head path opt-in. Warn when destructive storage reuse makes entropy non-differentiable, while rejecting unsupported NPU and tree-training combinations. * fix(engine): export standard FSDP LoRA adapter keys PEFT keeps the adapter name in live parameter FQNs, while serving engines expect serialized LoRA keys without it. Normalize per-parameter FSDP exports and align the SGLang best-effort assertion with its load behavior. * feat(engine): support chunked logits for padded models Enable chunked LM Head loss for text-only padded BSHD models such as Qwen3.5 and rename the public toggle to enable_chunked_logits so the configuration reflects its behavior. Key changes: - Add padded label construction and output repacking - Add Qwen3.5 and updated Qwen3 MoE profile recipes - Update CLI docs, validation, and regression coverage * fix(engine): configure logprob chunking explicitly Replace the profile-only environment override with a validated train-engine option so FSDP, Megatron, Archon, and tree paths use the same explicit value. Key changes: - add and document TrainEngineConfig.logprobs_chunk_size - pass the setting through every engine logprob path - translate the profile guide and remove out-of-scope FSDP LoRA changes - add config, launcher, and explicit chunk-size tests Refs: #1555 | 1 个月前 | |
chore: enforce license (#1171) | 5 个月前 | |
chore: enforce license (#1171) | 5 个月前 | |
fix(recover): publish immutable checkpoints via latest pointer (#1616) * fix(recover): publish immutable checkpoints via latest pointer Write recovery state into immutable per-step generations and publish LATEST only after payload completion. Defer Megatron async publication to MCore finalization, retain legacy checkpoint reads, and preserve existing SPMD writers on the legacy layout. Signed-off-by: jiawei <jiaweibit@gmail.com> * fix(recover): coordinate multi-engine async saves Drain non-publishing async engine saves before scheduling the final publisher so LATEST advances only after every payload is durable. * fix(recover): load PPO critic checkpoint Use one stable actor/critic engine mapping for both recovery save and load so transactional multi-engine checkpoints can be restored. * test(megatron): isolate async finalize callback * refactor(recover): clarify checkpoint publication naming Name the finalize helper after its checkpoint publication responsibility and use the full checkpoint_pointer module name instead of the ambiguous cp alias. --------- Signed-off-by: jiawei <jiaweibit@gmail.com> | 17 天前 | |
fix(infra): prevent leaks after RTensor delete failures (#1633) * fix(infra): prevent leaks after RTensor delete failures Preserve failed shard IDs across one clear_batches call so transient storage-node outages can recover. Stop training after a second failure instead of silently accumulating CPU tensors. Key changes: - send DELETE payloads through the existing bounded retry helper - surface per-node failures and storage cleanup statistics - retry failed shards across one step, then fail after worker cleanup - preserve the whole pending batch atomically on cancellation Refs: #1581 * fix(infra): make batch cleanup failure-safe Keep exhausted RTensor shards pending until worker buffers drain, so a worker RPC error cannot erase storage leak state. Run every trainer role cleanup before propagating the first failure. Refs: #1581 | 1 个月前 | |
feat: add Arena single- and multi-stream rollouts (#1702) * feat: port Arena single- and multi-stream rollouts to main Run Arena Harness tasks through the existing rollout proxy, with weighted prompt mixtures, per-stream rewards, session-bound gateway routing, failure classification, and worker-scoped registration cleanup. Preserve main's processor cache and fail-fast grouped cancellation. Include portable two-node FSDP/SGLang examples and integration guidance. Keep PRM, RewardSystem, AWEX, and mean-only normalization out of scope. Port selected changes from swe-dev: - da1da65c7: initial Arena Stream integration - 1d0fc7ba0: configurable reward transforms - 714f733a0: episode and grouped metrics - 586d7cc06: separate task and proxy timeouts - 84f7bb341: recover typed context overflow - ddb7d7f3c: example-level continuous reward transform - 2b994c10c: preserve original rewards for debugging - 9ecbde97c: multi-stream routing and lifecycle support Validation: 227 targeted tests passed; full pre-commit stack run. * refactor: consolidate Arena config loading and test scenarios * feat(examples): add environment-configured Arena submission script * style(examples): align Terminal Bench imports with package layout * fix(examples): leave token headroom in Arena smoke profiles * fix(utils): decouple port allocation from training seed * fix(engine): synchronize uneven FSDP microbatch execution * test: remove Arena launcher and profile smoke tests * feat(examples): add Flash V3 Arena launch profiles * fix(examples): use default Arena score handling * fix(experimental): preprocess proxy messages only once --------- Co-authored-by: yulangz <yulangz@users.noreply.github.com> | 14 天前 | |
feat(ppo): add reuse_train_logp proximal logp method (#1453) Add a reuse_train_logp option to prox_logp_method for decoupled PPO. It reuses the training forward-pass logprobs (detached) as the proximal logp, skipping the extra proximal forward pass and its memory/compute cost. This is only valid with ppo_n_minibatches=1: with a single minibatch the training forward still reflects the policy that generated the rollout, so its logprobs equal the proximal policy. With multiple minibatches the weights change between steps, so PPOActorConfig.__post_init__ rejects that combination. Add tests for the enum/constant wiring, skips_forward_pass, and the ppo_n_minibatches validation. | 2 个月前 | |
perf: reduce VLM CPU broadcast and microbatch memory overhead (#1697) * perf: reduce VLM CPU broadcast and microbatch memory overhead * perf(v2): preserve aliases in CPU-staged vision broadcasts * docs: cover v2 multimodal broadcast validation * docs: remove multimodal memory validation guide * test: remove multimodal memory validation scripts * test: complete Megatron staging fixture configuration | 11 天前 | |
feat(trainer): add multi-teacher on-policy distillation (#1592) Add phase-scoped teacher scoring and weighted reverse-KL targets for heterogeneous dataset mixtures while sharing actor GPUs with SGLang and Megatron teachers. Key changes: - Route dataset teacher groups to weighted teacher checkpoints - Add persistent teacher residency and local-memory checkpoint staging - Coordinate AWEX weight transfer and phase-scoped GPU offload - Include a local Qwen3 14B-to-0.6B GSM8K example and tests | 22 天前 | |
feat(awex): add separation AdamW delta weight transfer (#1623) * feat(config): add separation DTE topology gates * feat(awex): add separation AdamW delta weight transfer * docs(examples): add DTE separation GSM8K example * docs(examples): expand DTE GSM8K configuration * fix(awex): address separation DTE review feedback * fix(awex): harden separation delta transfer protocol * refactor(config): integrate delta transfer into train engine * refactor(config): simplify delta weight update switch * test(awex): isolate optional DTE dependency | 1 个月前 | |
chore: enforce license (#1171) | 5 个月前 | |
feat(trainer): add multi-teacher on-policy distillation (#1592) Add phase-scoped teacher scoring and weighted reverse-KL targets for heterogeneous dataset mixtures while sharing actor GPUs with SGLang and Megatron teachers. Key changes: - Route dataset teacher groups to weighted teacher checkpoints - Add persistent teacher residency and local-memory checkpoint staging - Coordinate AWEX weight transfer and phase-scoped GPU offload - Include a local Qwen3 14B-to-0.6B GSM8K example and tests | 22 天前 | |
chore: enforce license (#1171) | 5 个月前 | |
fix(trainer): run initial evaluation before first update (#1636) `eval_before_train` currently relies on the first scheduled evaluator check, which trainers perform only after one optimization step. The reported baseline can therefore contain updated weights. Run the one-shot evaluation separately so its metrics are logged before the first update without advancing periodic evaluation cadence. Skip the startup evaluation on recovery and clear legacy deferred triggers when loading evaluator state. Key changes: - Invoke the startup baseline from PPO, SFT, DPO, and reward trainers - Preserve periodic evaluation cadence and recovery behavior - Cover call ordering, logging steps, legacy state, and PPO offload Refs: #1232 Signed-off-by: Bo Yang <yb550079@antgroup.com> | 1 个月前 | |
chore: enforce license (#1171) | 5 个月前 | |
feat: add Bailing V3 SWE SFT support (#1598) * feat: add Bailing V3 SWE SFT support Port the reviewed internal implementation to the public main branch while preserving newer upstream engine behavior. Key changes: - Add Bailing V3 KDA, gated MLA, and MoE model support - Add SWE SFT dataset loading and cache handling - Add focused model, loader, and dataset tests Refs: inclusionAI/AReaL#2188 Signed-off-by: chucai.dzq <chucai.dzq@alibaba-inc.com> * refactor(dataset): split SWE SFT loader into modules Separate message processing, tokenization, pipeline orchestration, and CLI code so each concern can evolve without growing a single dataset module. Signed-off-by: chucai.dzq <chucai.dzq@alibaba-inc.com> * fix(dataset): align SWE SFT with adaptive templates Port the follow-up SWE SFT fixes from swe-dev so Bailing V3 adaptive chat templates use structural assistant masks and consistent thinking modes. Signed-off-by: chucai.dzq <chucai.dzq@alibaba-inc.com> * test(dataset): restore modules after SWE loader tests Prevent collection-time stubs from replacing the real areal.dataset package for subsequent data-service tests. * refactor(engine): remove precision dump hooks Keep Bailing V3 support focused on production training behavior by removing out-of-scope routing and log-probability dump paths. * test(models): run zigzag coverage in CI Place the CP zigzag unit tests under the root test pattern used by the GCP unit-test workflow. * fix: preserve Bailing V3 HF export metadata Keep runtime model metadata valid across fast, native, direct, and in-place mbridge exports while retaining the production local-source fallback. Key changes: - snapshot and validate HF config metadata before exporters overwrite it - preserve source assets and support native mbridge finalization - forward SWE preprocessing kwargs through the remote dataset path - add config round-trip, Saver, and controller regression coverage Refs: #1598 Constraint: Preserve swe-dev Bailing V3 architecture-based bridge dispatch Confidence: high Scope-risk: moderate Not-tested: Real multi-rank Bailing V3 HF save/load canary * test(engine): initialize native save fixture correctly Use the MegatronEngine backing process-group fields so the native mbridge finalization test can exercise the real cpu_group property. Refs: #1598 Constraint: Keep the production save path unchanged Confidence: high Scope-risk: narrow Not-tested: Full suite rerun pending on GCP * refactor(dataset): remove unused SWE augmentation options Keep the public SWE SFT loader focused on the pair and trajectory behavior exercised by production recipes, without shipping disabled sampling and truncation branches. Key changes: - remove random thinking variants and ratio balancing - remove task-notification truncation and the dead trajectory-copy CLI - preserve canonical pair, pre-split, and trajectory outputs with tests Refs: #1598 Constraint: Preserve tracked swe-dev production defaults and trajectory mode Rejected: Broad SWE substring dispatch | reintroduces c84db0bbf false matches Confidence: high Scope-risk: moderate Not-tested: Real production JSONL end-to-end run * fix(dataset): harden SWE SFT preprocessing Use structural assistant masks so literal template delimiters cannot silently drop supervised tokens. Reject unsupervised rows before training and pass data-worker cache topology explicitly. Key changes: - Make Qwen and Bailing template patches idempotent and fail closed - Filter malformed and all-zero masks, including pre-tokenized data - Coordinate shared caches with explicit data-worker rank metadata - Restore the split SWE preprocessing CLI and add regression coverage Refs: #1598 * fix: harden Bailing V3 and SWE preprocessing Keep the pull request focused on Bailing V3 SWE SFT correctness while retaining cache invalidation and export metadata fixes requested in review. Generic Hugging Face checkpoint publication hardening belongs in a separate change. --------- Signed-off-by: chucai.dzq <chucai.dzq@alibaba-inc.com> Co-authored-by: 楚财 <chucai.dzq@alibaba-inc.com> | 26 天前 | |
chore: enforce license (#1171) | 5 个月前 | |
chore: enforce license (#1171) | 5 个月前 | |
feat(engine): support native MTP-only and packed Qwen SFT (#1719) * feat(engine): support MTP-only training with a frozen backbone Signed-off-by: chucai.dzq <chucai.dzq@antgroup.com> * fix(engine): preserve MTP-only gradients during recomputation Signed-off-by: chucai.dzq <chucai.dzq@antgroup.com> * feat(engine): support packed Qwen MTP SFT with pipeline parallelism * fix(api): validate MTP-only training runtime versions * fix(engine): scope Qwen packed sequence support to active bridge --------- Signed-off-by: chucai.dzq <chucai.dzq@antgroup.com> | 7 天前 | |
feat(engine): support fixed warmup steps (#1597) Allow optimizer configs to select an absolute scheduler warmup while preserving the proportional fallback across FSDP, Megatron, and Archon. Validate shared scheduler boundaries and keep Megatron resume initialization independent of the active warmup config. Key changes: - Add fixed warmup resolution with cross-engine validation - Align Megatron decay and resume scheduler parameters - Add boundary and Megatron integration regression tests Co-authored-by: 峯回 <dh183333@antgroup.com> | 1 个月前 | |
chore: enforce license (#1171) | 5 个月前 | |
feat(infra): add HTTP-based Ray Scheduler (#1441) * fix(infra): preserve dataclass state over RPC * feat(engine): support headless vLLM server mode * fix(utils): use fixed Ray name resolve namespace * refactor(engine): expose backend server env builders * refactor(infra): remove Ray-native scheduler Drop the single-controller Ray scheduler path. Key changes: - Remove RayScheduler, Ray RPC actors, and Ray vLLM remote launcher - Make single-controller trainers dispatch only local or slurm schedulers - Simplify RTensor and RPC serialization to the HTTP backend * feat(infra): add HTTP-based Ray scheduler Add a Ray-backed scheduler that allocates accelerator placement groups while keeping worker and inference engine traffic on the existing HTTP RPC path, with batched launcher operations and multi-node rollout support. Key changes: - Add RayScheduler and RayWorkerProcessLauncher for Ray-managed HTTP workers - Wire scheduler.type=ray into infra exports, trainer initialization, logging, docs, and examples - Batch worker startup and status checks by Ray launcher to reduce per-worker actor calls - Split multi-node rollout backend launch and cleanup into a dedicated coordinator - Tighten Ray launcher lifecycle handling for worker shutdown, placement groups, and backend process cleanup * test(infra): add Ray scheduler tests * fix(trainer): allow proxy workers with Ray scheduler * fix(infra): reuse existing Ray cluster on init * fix(api): restore deprecated Ray placement config --------- Co-authored-by: Ge Shi <utashih@gmail.com> | 2 个月前 | |
chore: enforce license (#1171) | 5 个月前 | |
fix(utils): honor custom IPv6 probe addresses (#1740) Resolve the configured probe host before route detection so IPv6-only environments do not silently fall back to an unrelated hardcoded endpoint. Signed-off-by: Andrew9603 <Andrew9603@users.noreply.github.com> Co-authored-by: Andrew9603 <Andrew9603@users.noreply.github.com> | 7 天前 | |
chore: enforce license (#1171) | 5 个月前 | |
perf: reduce Megatron training memory peaks (#1555) * perf: reduce Megatron training memory peaks Add an SFT profiling workflow and use its memory snapshots to remove full-sequence vocabulary and optimizer gradient peaks from Megatron training. Key changes: - Add rank-aware kernel and memory profiling for packed SFT workloads - Fuse FP32 vocab-parallel logprob storage with LM head backward - Add optional true chunked LM head loss with recomputed backward - Configure precision-aware optimizer fields before Megatron validation - Cover BF16/FP32 numerical parity and distributed TP/SP behavior * fix(models): avoid private storage identity checks Track the LM head output tensor weakly and compare storage through the public data_ptr API. This preserves allocator-address reuse protection without depending on PyTorch's private storage _cdata field. * test: compare parameter storage without object identity Parameter.data may return a fresh Tensor wrapper on each access. Verify that replicated parameters retain their data pointer and storage offset instead of comparing transient Python objects. * test: make recycled CUDA storage check deterministic Construct the replacement tensor from the original storage instead of relying on the caching allocator to immediately reuse a freed address after the full CI suite. * fix(engine): guard AReaL LM Head storage reuse Keep entropy differentiable unless its gradients are disabled, and make the optimized LM Head path opt-in. Warn when destructive storage reuse makes entropy non-differentiable, while rejecting unsupported NPU and tree-training combinations. * fix(engine): export standard FSDP LoRA adapter keys PEFT keeps the adapter name in live parameter FQNs, while serving engines expect serialized LoRA keys without it. Normalize per-parameter FSDP exports and align the SGLang best-effort assertion with its load behavior. * feat(engine): support chunked logits for padded models Enable chunked LM Head loss for text-only padded BSHD models such as Qwen3.5 and rename the public toggle to enable_chunked_logits so the configuration reflects its behavior. Key changes: - Add padded label construction and output repacking - Add Qwen3.5 and updated Qwen3 MoE profile recipes - Update CLI docs, validation, and regression coverage * fix(engine): configure logprob chunking explicitly Replace the profile-only environment override with a validated train-engine option so FSDP, Megatron, Archon, and tree paths use the same explicit value. Key changes: - add and document TrainEngineConfig.logprobs_chunk_size - pass the setting through every engine logprob path - translate the profile guide and remove out-of-scope FSDP LoRA changes - add config, launcher, and explicit chunk-size tests Refs: #1555 | 1 个月前 | |
fix(utils): treat missing packages as not matching version checks (#1689) * fix(utils): lazy test model dict and safe package version checking - Use lazy evaluation in testing model path dictionaries to avoid eager downloads on test collection - Safely handle PackageNotFoundError in version check helper functions * fix(utils): drop lazy model dict already landed in #1568 Keep PackageNotFoundError handling in package version checks and add unit coverage for missing and installed packages. | 7 天前 | |
chore: enforce license (#1171) | 5 个月前 | |
fix(engine): preserve colocate state across offload and recovery (#1749) * fix(engine): preserve colocate state across offload and recovery * test: align regression fixtures with current engine contracts | 5 天前 | |
chore: enforce license (#1171) | 5 个月前 | |
feat: add Bailing V3 SWE SFT support (#1598) * feat: add Bailing V3 SWE SFT support Port the reviewed internal implementation to the public main branch while preserving newer upstream engine behavior. Key changes: - Add Bailing V3 KDA, gated MLA, and MoE model support - Add SWE SFT dataset loading and cache handling - Add focused model, loader, and dataset tests Refs: inclusionAI/AReaL#2188 Signed-off-by: chucai.dzq <chucai.dzq@alibaba-inc.com> * refactor(dataset): split SWE SFT loader into modules Separate message processing, tokenization, pipeline orchestration, and CLI code so each concern can evolve without growing a single dataset module. Signed-off-by: chucai.dzq <chucai.dzq@alibaba-inc.com> * fix(dataset): align SWE SFT with adaptive templates Port the follow-up SWE SFT fixes from swe-dev so Bailing V3 adaptive chat templates use structural assistant masks and consistent thinking modes. Signed-off-by: chucai.dzq <chucai.dzq@alibaba-inc.com> * test(dataset): restore modules after SWE loader tests Prevent collection-time stubs from replacing the real areal.dataset package for subsequent data-service tests. * refactor(engine): remove precision dump hooks Keep Bailing V3 support focused on production training behavior by removing out-of-scope routing and log-probability dump paths. * test(models): run zigzag coverage in CI Place the CP zigzag unit tests under the root test pattern used by the GCP unit-test workflow. * fix: preserve Bailing V3 HF export metadata Keep runtime model metadata valid across fast, native, direct, and in-place mbridge exports while retaining the production local-source fallback. Key changes: - snapshot and validate HF config metadata before exporters overwrite it - preserve source assets and support native mbridge finalization - forward SWE preprocessing kwargs through the remote dataset path - add config round-trip, Saver, and controller regression coverage Refs: #1598 Constraint: Preserve swe-dev Bailing V3 architecture-based bridge dispatch Confidence: high Scope-risk: moderate Not-tested: Real multi-rank Bailing V3 HF save/load canary * test(engine): initialize native save fixture correctly Use the MegatronEngine backing process-group fields so the native mbridge finalization test can exercise the real cpu_group property. Refs: #1598 Constraint: Keep the production save path unchanged Confidence: high Scope-risk: narrow Not-tested: Full suite rerun pending on GCP * refactor(dataset): remove unused SWE augmentation options Keep the public SWE SFT loader focused on the pair and trajectory behavior exercised by production recipes, without shipping disabled sampling and truncation branches. Key changes: - remove random thinking variants and ratio balancing - remove task-notification truncation and the dead trajectory-copy CLI - preserve canonical pair, pre-split, and trajectory outputs with tests Refs: #1598 Constraint: Preserve tracked swe-dev production defaults and trajectory mode Rejected: Broad SWE substring dispatch | reintroduces c84db0bbf false matches Confidence: high Scope-risk: moderate Not-tested: Real production JSONL end-to-end run * fix(dataset): harden SWE SFT preprocessing Use structural assistant masks so literal template delimiters cannot silently drop supervised tokens. Reject unsupervised rows before training and pass data-worker cache topology explicitly. Key changes: - Make Qwen and Bailing template patches idempotent and fail closed - Filter malformed and all-zero masks, including pre-tokenized data - Coordinate shared caches with explicit data-worker rank metadata - Restore the split SWE preprocessing CLI and add regression coverage Refs: #1598 * fix: harden Bailing V3 and SWE preprocessing Keep the pull request focused on Bailing V3 SWE SFT correctness while retaining cache invalidation and export metadata fixes requested in review. Generic Hugging Face checkpoint publication hardening belongs in a separate change. --------- Signed-off-by: chucai.dzq <chucai.dzq@alibaba-inc.com> Co-authored-by: 楚财 <chucai.dzq@alibaba-inc.com> | 26 天前 | |
fix: make rollout sampling deterministic (#1625) | 1 个月前 | |
fix(rollout): train safely on incomplete groups (#1563) * fix(engine): make ragged transport padding objective-safe * fix(rollout): train safely on incomplete groups * fix(infra): batch ragged gathers and bound rollout stalls Address review on the incomplete-group transport and collection paths: - Replace the serial per-item broadcast fallback for ragged trajectory lists with one metadata all-gather plus one padded all-gather per dtype/device bucket, so unequal group counts cost a handful of collectives instead of sum(lengths) container broadcasts. Ranks with zero trajectories join the same collectives via the gathered metadata, and gathered tensors are cloned out of the padded buffers so a skewed distribution does not pin world_size x max-payload memory. - Split redistribute_trajectories into its gather phase and a pure-local packing phase, and recover only the latter: every head holds the same gathered data there, so failures are symmetric and always reach the error sync; a failure inside a collective now propagates instead of posting mismatched operations at peers. - Abort dynamic preparation only on a sustained stall: at least eight consecutive rounds that add no trainable group AND thirty minutes without progress. Both signals ride the existing per-round all-gather, so every head raises on the same iteration and the coordinated error path turns a silent SPMD spin into a terminal error on all ranks, while fast legitimate all-reject bursts (e.g. a staleness flush after a weight update) stay alive. Refs: #1563 Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com> * fix(api): keep singleton groups trainable and state the slot contract Address review on estimator-owned group minimums and the tightened workflow contract: - NormConfig gains uses_group_statistics and PPOActorConfig owns minimum_usable_group_size, replacing the trainer-side free function and its ad-hoc shape duck-typing. - A singleton target group is complete by definition, so group statistics no longer hard-fail n_samples=1 configs; the previous ValueError broke examples/openclaw on startup while the v2 rollout path skipped the check entirely. Config-level degeneracy stays a UserWarning, as on main. - The trainer collection loop shares the empty-round streak plus wall-clock stall bound and rate-limits its progress warnings. - WorkflowContractError now names the two remedies (one sample per episode or batch-level normalization), and the one-sample-per-slot contract is documented in the RolloutWorkflow docstring and the grouped-rollout reference (EN/ZH), scoped honestly to grouped rollouts: n_samples=1 installs no group wrapper and is not checked. - Correct docstrings that still described divisibility-based dispatch: balanced_greedy_partition, _pad_eval_batch, and the v2 data-proxy dispatcher module. Refs: #1563 Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com> * feat(api): make min_usable_group_size configurable Reviewer request on #1563: expose the partial-group trainability threshold instead of only deriving it. PPOActorConfig gains min_usable_group_size (default None); None keeps the derived value (2 when reward_norm/adv_norm uses group statistics, 1 for singleton target groups, else 1), an explicit value replaces it and is still validated against group_size at workflow setup. Signed-off-by: Max Wang <maxwill@vmax.ai> * fix(examples): center openclaw rewards at batch level With n_samples=1 every prompt group is a singleton, so group mean centering erases the task reward. Batch centering keeps a live signal; singleton group std is already pinned to 1 and stays a no-op. Signed-off-by: Max Wang <maxwill@vmax.ai> * fix(api): reject explicit singleton minimum under group statistics Reviewer request on #1563: an explicit min_usable_group_size below 2 combined with group-relative normalization would train lone survivors that have no group peers to normalize against, so PPOActorConfig now rejects it at construction. The derived default is unaffected. Signed-off-by: Max Wang <maxwill@vmax.ai> * fix(examples): use batch reward std in openclaw demo std_level: group is not the no-op I claimed on the review thread: the OpenAI-proxy agent exports one row per interaction, so compute_advantages passes per-episode row counts as group sizes and a k-turn episode's identical rewards collapse to sign(r - mean) * sqrt((k-1)/k), discarding reward magnitude. Batch std matches adv_norm and keeps the signal; the group_size line had no remaining consumer. Signed-off-by: Max Wang <maxwill@vmax.ai> * fix(api): state min_usable_group_size scope and trigger exactly Two exactness gaps around the new config field. The v2 rollout path never consumes it, so an explicit setting now fails fast there (matching RolloutControllerV2's rejection of reward_normalization and drop_incomplete_group) and the help text names the v1 scope. The one-sample-per-slot contract is armed by the resolved minimum, not by group normalization itself, so the error message, arun_episode docstring, and reference docs now name the real trigger and include unsetting actor.min_usable_group_size among the remedies; the agent tutorial and export-style reference gain the same caveat where they teach export_style: individual. Signed-off-by: Max Wang <maxwill@vmax.ai> * fix(ppo): drop implicit partial-group guard duplicated by #1415 Keep the incomplete-group path's variable group_sizes handling here. The fixed-size divisibility raise belongs in #1415 so the two PRs can land independently. Co-authored-by: Cursor <cursoragent@cursor.com> * test(engine): trim redundant transport test assertions Fail directly when a transport row reaches an objective callback. Keep the collective-ordering regression without duplicate mock assertions. Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com> * fix(rollout): normalize rewards over usable group members * fix: stabilize logical rollout reward normalization Keep centered row rewards when reference scale is within epsilon, and supply the built-in individual agent export's terminal reward as its normalization reference. Preserve discounted rows and explicit references. --------- Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com> Signed-off-by: Max Wang <maxwill@vmax.ai> Co-authored-by: EazyReal <8047065+EazyReal@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: sitabulaixizawaluduo <ljl2020110773@gmail.com> | 13 天前 | |
feat: add Arena single- and multi-stream rollouts (#1702) * feat: port Arena single- and multi-stream rollouts to main Run Arena Harness tasks through the existing rollout proxy, with weighted prompt mixtures, per-stream rewards, session-bound gateway routing, failure classification, and worker-scoped registration cleanup. Preserve main's processor cache and fail-fast grouped cancellation. Include portable two-node FSDP/SGLang examples and integration guidance. Keep PRM, RewardSystem, AWEX, and mean-only normalization out of scope. Port selected changes from swe-dev: - da1da65c7: initial Arena Stream integration - 1d0fc7ba0: configurable reward transforms - 714f733a0: episode and grouped metrics - 586d7cc06: separate task and proxy timeouts - 84f7bb341: recover typed context overflow - ddb7d7f3c: example-level continuous reward transform - 2b994c10c: preserve original rewards for debugging - 9ecbde97c: multi-stream routing and lifecycle support Validation: 227 targeted tests passed; full pre-commit stack run. * refactor: consolidate Arena config loading and test scenarios * feat(examples): add environment-configured Arena submission script * style(examples): align Terminal Bench imports with package layout * fix(examples): leave token headroom in Arena smoke profiles * fix(utils): decouple port allocation from training seed * fix(engine): synchronize uneven FSDP microbatch execution * test: remove Arena launcher and profile smoke tests * feat(examples): add Flash V3 Arena launch profiles * fix(examples): use default Arena score handling * fix(experimental): preprocess proxy messages only once --------- Co-authored-by: yulangz <yulangz@users.noreply.github.com> | 14 天前 | |
feat(trainer): add MiMo-inspired RL training metrics (#1736) * feat(trainer): add MiMo-inspired RL training metrics Align AReaL with public MiMo RL observability dimensions for training-inference consistency, policy staleness, and learning-signal composition without claiming parity with unpublished collectors. Compare recomputed and rollout logprobs before PPO overwrites them, including F(tau)-equivalent ratio tails, KL estimators, NLL values, and nonfinite input rates. Correct staleness alignment and aggregation, and report advantage sign occupancy before behavioral filtering. Key changes: - Add train-inference logprob diagnostics and ratio tails - Align rollout versions to prediction positions for staleness metrics - Use token-weighted means and global min/max reductions - Cover unequal and missing DP populations with CPU Gloo tests - Document definitions, sampling assumptions, and memory overhead Validation: 52 targeted tests, including two-process CPU Gloo Validation: pre-commit run --all-files Not-tested: GPU multi-node CP/TP/PP integration and memory overhead Scope-risk: moderate * fix(trainer): bound RL diagnostic storage and clarify version age Aggregate new token diagnostics into masked scalar summaries instead of retaining full sequences until stats export. This preserves weighted DP averages and global extrema while avoiding long-context memory growth. Compute k3 in float64 and expose a separate overflow fraction for finite inputs. Keep staleness version alignment local to diagnostics, including pure-MOPD batches; leave the existing loglinear approximation unchanged until its optimizer-step semantics can be validated independently. Validation: 134 targeted tests, including CPU Gloo rank-imbalance cases Validation: pre-commit run --all-files Not-tested: GPU multi-node CP/TP/PP training or peak transient memory --------- Co-authored-by: chucai.dzq <chucai.dzq@alibaba-inc.com> | 8 天前 | |
fix(utils): resolve test model paths lazily (#1568) * fix(utils): resolve test model paths lazily Importing shared testing utilities should not download every configured model before pytest can apply GPU and slow-test skips. Preserve the registry mapping interface while resolving and caching only requested keys. Key changes: - replace eager model path registries with lazy mappings - defer VLM model lookup to fixture setup after skip evaluation - cover offline import, collection, and per-key caching behavior * test(utils): focus lazy path membership coverage Exercise membership on a fresh mapping once instead of repeating the same behavior across registries and every configured key. Drop the fixture type assertion; resolving its requested key is sufficient. Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com> --------- Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com> Co-authored-by: EazyReal <8047065+EazyReal@users.noreply.github.com> | 14 天前 | |
fix(utils): handle time checks without distributed init (#1715) Signed-off-by: Andrew9603 <Andrew9603@users.noreply.github.com> Co-authored-by: Andrew9603 <Andrew9603@users.noreply.github.com> | 12 天前 | |
fix(inference): reject incomplete sampling evidence (#1554) Validate normalized token/logprob evidence before rollout accumulation, reject vLLM ambiguous sentinel values, and preserve the explicit SGLang abort-before-prefill result. Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com> Co-authored-by: EazyReal <8047065+EazyReal@users.noreply.github.com> | 2 个月前 | |
chore: enforce license (#1171) | 5 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 1 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 17 天前 | ||
| 1 个月前 | ||
| 14 天前 | ||
| 2 个月前 | ||
| 11 天前 | ||
| 22 天前 | ||
| 1 个月前 | ||
| 5 个月前 | ||
| 22 天前 | ||
| 5 个月前 | ||
| 1 个月前 | ||
| 5 个月前 | ||
| 26 天前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 7 天前 | ||
| 1 个月前 | ||
| 5 个月前 | ||
| 2 个月前 | ||
| 5 个月前 | ||
| 7 天前 | ||
| 5 个月前 | ||
| 1 个月前 | ||
| 7 天前 | ||
| 5 个月前 | ||
| 5 天前 | ||
| 5 个月前 | ||
| 26 天前 | ||
| 1 个月前 | ||
| 13 天前 | ||
| 14 天前 | ||
| 8 天前 | ||
| 14 天前 | ||
| 12 天前 | ||
| 2 个月前 | ||
| 5 个月前 |