| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
fix: prevent empty layer range and too many neighbors in ARA (#478) * Cap ARA's neighbor count at the number of prompts The nearest neighbors are selected from the outputs for the good and bad prompts, which have one row per prompt. With fewer than 15 prompts in either dataset, the optimizer could pick a neighbor count larger than that, and torch.topk would crash with "selected index k out of range". Bounding the search range instead of clamping the value keeps the optimizer from wasting trials on counts that would all behave the same. * Render chat templates with a fixed date Some chat templates, including gpt-oss's, put the current date in the system prompt. That makes the residuals, and with them the modified model, depend on the day Heretic runs, so a model could not be reproduced on a different day. It also made the gpt-oss test's hash change daily once ARA started modifying the model. Prompts used for computing residuals and for evaluation are now rendered with a fixed date. Interactive chat keeps the real one. * Prevent ARA from sampling an empty layer range The start index was sampled from [0, L//2] and the exclusive end index from [L//2, L], so both could land on L//2 and leave no layers to modify, wasting the trial. Starting the end range at L//2 + 1 rules this out while keeping the bounds fixed, as multivariate TPE requires. On the 2-layer test models, every ARA trial had picked this empty range, so the seed-oss and gpt-oss tests never ran the optimization. They now modify layer 1, which changes their expected hashes. The seed-oss output also differs between the CPUs of GitHub's runners, because L-BFGS amplifies small floating-point differences, so it gets CI hashes as well. | 1 天前 |