| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
fix(workflow): account for thinking prefill in Geometry3K rewards (#1708) * fix(workflow): account for thinking prefill in Geometry3K rewards * fix(workflow): use explicit thinking prefill for Geometry3K rewards * fix(dataset): derive thinking prefill from chat template output * test: defer Geometry3K dataset prefill coverage * test: restore baseline Geometry3K agent coverage | 16 天前 | |
feat(vlm): add Qwen3.6 LoRA GRPO training support for 27B and 35B-A3B (#1444) - Add VLM geometry3k GRPO configs for Qwen3.6-27B (dense) and 35B-A3B (MoE) - Add sft_train_batch to FSDPPPOActor for areal-mint SFT - Fix GPU group separation for actor and rollout (avoid colocation) - Fix max_tokens handling in sglang_remote - Add vlm_math_agent and train.py examples | 2 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 16 天前 | ||
| 2 个月前 |