| fix(fleet): bound per-tick heartbeat and memory sampling load Four Fleet manager lifecycle tests failed on macOS CI under full-suite load ("first Fleet attempt never started", "fake worker never started", managers never converging) and passed alone from the same CI binary in under 1.2 s. In the concurrent-managers failure the standby executor held the worker, yet the child's first shell builtin never ran within 15 s. Timeouts were already raised from 5 s to 15 s without effect, so this does not raise them again. Product side: every scheduler tick appended a durable heartbeat per leased task (sync_data, which is F_FULLFSYNC on macOS) and spawned `ps` to sample worker memory. At the tests' 10 ms tick that is ~100 full-drive flushes and ~100 process spawns a second per test, against a 300 s stale window. Heartbeats are now skipped when the worker already has one stamped this second (timestamps are whole-second), and memory is sampled at most once a second per worker, keeping the last value. Test side: the fleet::manager module runs in a max-threads = 1 nextest group, the same remedy as the exec-persistent-service tests (#5355). Not established: that this is the whole cause. Late spawn vs. late child start was not distinguished; acceptance is a green loaded hosted macOS run. Evidence: 431 passed, 0 failed (12,828 skipped) on the fleet:: selection; `nextest show-config test-groups` lists the manager tests in the new group. TUI all-target/all-feature Clippy with CI flags and fmt passed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> | 5 天前 |