| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
feat(inference): support PRM scorers in v2 trajectory export (#1735) * feat: support PRM scorers in v2 trajectory export Reuse the existing v1 scoring contracts when exporting complete v2 trajectories, so configured process rewards and metrics reach the training workflow without changing reward or advantage formulas. Key changes: - Forward PRM configuration to inference data proxies. - Share branch-isolated scoring and existing metric recorders with v1. - Handle concurrent group scoring, failures and session cleanup. - Cover v1 parity, export retries and multimodal tensor transport. Signed-off-by: 杨博 <yb550079@antgroup.com> * fix(inference): preserve online sessions during PRM export Retain the persistent HITL session while scoring consumes a ready trajectory, so concurrent interactions survive success, failure, and cancellation. Ordinary session keys still clean up and remain refreshable. Restore export-schema regression coverage and complete the Chinese PRM documentation. Add real workflow/export interleaving tests for the review finding, alongside PRM-enabled and disabled lifecycle coverage. Signed-off-by: 杨博 <yb550079@antgroup.com> * fix(inference): close owned PRM scorers during shutdown Flush scorer-owned clients and audit writers before the v2 data proxy releases its service resources. A failed scorer cleanup must not prevent other resources from being closed. Key changes: - Add a default async cleanup hook and runner-owned scorer cleanup - Pair runner construction and cleanup with the v2 service lifespan - Cover awaited flushes, ownership, failure paths, and backend replacement - Document the lifecycle contract in English and Chinese Refs: #1735 Signed-off-by: 杨博 <yb550079@antgroup.com> * fix(inference): validate PRM mode and roll back failed startup Reject tokenless v2 external-model scoring before allocating resources. Use an async PRM factory to await owned-scorer rollback while preserving the original startup failure and existing synchronous v1 construction. Signed-off-by: 杨博 <yb550079@antgroup.com> --------- Signed-off-by: 杨博 <yb550079@antgroup.com> | 9 天前 | |
feat: add process rewards for agent training (#1701) * feat: add process rewards for agent training * fix(examples): use local truncation check for PRM filter * fix: avoid process reward overhead when scoring is disabled * fix: preserve PRM termination semantics and non-agent training Honor explicit truncation metadata and allow PPO rollouts without agents. Repair baseline test fixtures for model residency, MOPD, and RPC storage. Not-tested: Megatron tests require mbridge unavailable locally --------- Co-authored-by: chucai.dzq <chucai.dzq@alibaba-inc.com> | 16 天前 | |
fix(reward): guard clevr_count_70k_reward_fn against scoring failures (#1430) * fix(reward): score non-string answers in clevr_count_70k_reward_fn clevr_count_70k_reward_fn did not str()-coerce its inputs or guard against errors, unlike the sibling reward fns (gsm8k, geometry3k). A non-string answer (e.g. an int) made ans.strip() raise AttributeError, which WorkflowExecutor catches and uses to reject the whole trajectory instead of scoring the sample — and even a matching completion was lost rather than scored 1.0. Coerce completions and answer to str and wrap the body in try/except, matching the sibling reward fns. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(example): apply the same guard to the clevr GRPO example reward fn The example defines its own clevr_count_70k_reward_fn (referenced via workflow_kwargs) that still had the old logic. Mirror the built-in fix: str-coerce completions/answer and guard, so a non-string answer is scored rather than raising and dropping the trajectory. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> | 3 个月前 | |
chore: enforce license (#1171) | 5 个月前 | |
chore: enforce license (#1171) | 5 个月前 | |
feat(trainer): add multi-teacher on-policy distillation (#1592) Add phase-scoped teacher scoring and weighted reverse-KL targets for heterogeneous dataset mixtures while sharing actor GPUs with SGLang and Megatron teachers. Key changes: - Route dataset teacher groups to weighted teacher checkpoints - Add persistent teacher residency and local-memory checkpoint staging - Coordinate AWEX weight transfer and phase-scoped GPU offload - Include a local Qwen3 14B-to-0.6B GSM8K example and tests | 24 天前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 9 天前 | ||
| 16 天前 | ||
| 3 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 24 天前 |