| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
Fix multi-image terminal drops (#3769) * Add multi-image drop regressions * Fix multi-image terminal drops * Add CLI stderr crash regression * Fix CLI stderr crash on closed pipes * Add Claude image drop paste regression * Use one paste transaction for local image drops * Remove unused local image type helpers * Copy pasteboard item before fallback rendering * Add Claude multi-image paste regression * Fix Claude multi-image drops --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 4 个月前 | |
CmuxFoundation: dissolve namespace-enums into value types / owning-type members (#6128) Remove every caseless namespace-enum in CmuxFoundation (the lint:allow namespace-type suppressions), giving each type genuine instance surface per the owner ruling. Logic, Defaults keys, file paths, and notification names are byte-identical; this is a shape change only. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> | 3 个月前 | |
Fix ssh stack review regressions | 6 个月前 | |
Fix background workspace PTY startup for socket-created surfaces (#3876) * Expose focus-coupled PTY startup for background split sends Socket clients can create terminal surfaces in workspaces that are not selected. The existing workspace-create coverage only proves the first hidden terminal starts; it does not cover a later split surface that receives input while its view is still off-window. This regression test drives the V2 socket API end to end: create a workspace without selecting it, split a terminal in that workspace, send text to the new surface, and assert the command executed by observing a marker file. On the 0.64.x regression path the input is accepted but remains queued until the user selects the workspace. Constraint: Repository policy runs socket/UI tests in CI or VM, not locally. Confidence: high Scope-risk: narrow Tested: Not run locally per repository policy; test is intended to fail before the fix in CI/VM. Not-tested: Local socket execution. * Start background terminal PTYs from runtime demand, not window focus Socket-driven terminal operations need to work as a control plane for background workspaces, but the previous eager off-window attach path made every model-only TerminalPanel start a PTY in unit tests. Keep the visual attach lifecycle lazy until an NSWindow exists, and route socket/read/send readiness through requestBackgroundSurfaceStartIfNeeded so explicit runtime demand can still create from the attached Ghostty view without selecting the workspace. Constraint: Do not run local tests or xcodebuild; validation must come from CI and lightweight local diff checks. Constraint: Background socket API must not mutate workspace focus. Rejected: Auto-select the target workspace before send_text | violates the focus allowlist and steals user focus. Rejected: Eagerly create every off-window attach | spawns PTYs for model-only/unit-test panels and timed out CircleCI unit tests. Confidence: high Scope-risk: moderate Directive: Do not reintroduce NSWindow membership as a prerequisite in requestBackgroundSurfaceStartIfNeeded; visual attach can be window-lazy, but explicit socket/API runtime demand must be able to start off-window. Tested: git diff --check Not-tested: Local socket/UI/unit tests per repository policy; CI will rerun the regression and unit suites. * Fix background split test cleanup logging * Tighten background terminal port ownership * Tighten background send text regression assertions * Remove redundant background send assertion | 4 个月前 | |
Avoid idle background terminal surface priming (#4184) * Add regression for idle background surface priming * Avoid idle background terminal surface priming * Simplify background prime guards * test: surface background workspace cleanup failures * Fix background template startup detection * Fix deferred startup predicate return | 4 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Flaky-test follow-up: review fixes + second-pass team sweep (#6452) * Address review: fail-loud guards + scoped soft-skips in de-flaked tests - test_browser_goto_split.py: _wait_url_loaded now raises on timeout instead of silently returning (false pass), and the local HTTP server is managed via a "with" block so the thread/socket can't leak on setup exceptions. - test_surface_move_reorder_api.py: assert the cleanup re-select of ws0 actually converges instead of ignoring the _wait() result (state leak across runs). - test_homebrew_sha.sh: only soft-skip transient transport failures (000/408/429/5xx); fail hard on deterministic client errors like 404 (missing release asset). - test_visual_typing_char_by_char.py: validate the typed glyph against the post-prompt region of the last line, so a "cmux" already in the prompt/path can't satisfy the check; robust to trailing zsh-autosuggestion glyphs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 5d4c282535a497ad263128b21f8805a34556f55d) * Address review round 2: critical walrus bug + 3 fail-loud/correctness guards - test_pane_break_swap_preserve_focus.py: fix NameError -- the lambda referenced `p` in the ternary condition before the walrus assigned it. Bind the walrus inside the condition so the predicate reads panes once and returns them. - perf-activation-session.py: a missing measurement (actual is None) is now always a blocking failure, not advisory -- absent data is a benchmark-contract violation, distinct from load-sensitive over-budget timing. - perf-activation-session.py: copy the real-scrollback measurement before storing it under snapshot_with_scrollback, so best_of_snapshot_timing's in-place mutation no longer clobbers the raw snapshot_with_real_scrollback capture. - NotificationAndMenuBarTests.swift: on the stall-timeout path, fail fast via a lock-guarded result box instead of awaiting evaluationTask.value, which could hang until the suite timeout when the hook ignores cancellation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7e75ad467bc95cef0436236d94183f43b9aaffb2) * Add reports.md: consolidated flaky-test discovery (in-session workflow) Output of a parallel dynamic workflow (20 agents over the 419 signal-bearing test files of 2080) verifying residual flakiness on this already-de-flaked branch. 101 candidate findings across 86 files, dominated by wall-clock timing asserts, sleep-as-sync, and async races -- concentrated in Packages/ unit tests the first 44-file sweep did not cover. Each entry is a candidate; the fix phase adversarially verifies every one before changing code (repo rules allow deterministic test sleeps, so those are left alone). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit e1952d01faee805298bdff215af1d0897f63abcf) * Fix 37 verified residual flaky tests (coordinated area-team workflow) Second-pass de-flake driven by the in-session discovery + team-fix workflows (see reports.md). Adversarial per-area verification confirmed genuine flakiness before any edit; 9 candidates were refuted (deterministic-allowed sleeps, factually-wrong premises, fixes needing risky production changes) and left. Highlights: - Packages/ unit tests (missed by the first 44-file sweep): AuthCoordinator and HostBrowserSignIn now drive timeouts off an injected ManualTestClock instead of real ContinuousClock; ChatConversationStore poller results are asserted with #expect instead of discarded (fail-loud); RemoteProxyBroker / RemoteCLIRelay / CommandPaletteSearchEngine timing/race tightened. - Integration tests: per-PID/UUID socket+port paths (terminal_focus_routing, ctrl_socket, ssh_remote_* disjoint port ranges, CMUXCLIErrorOutput socket) to kill cross-run collisions; order-dependent session-restore checks now scan all workspaces instead of trusting seeding order; latent NameErrors fixed (pane_resize missing `import time` / `must`); fail-open waits made loud (new_tab tmp-write now raises); surface-targeted send to avoid focus races (tab_dragging); fd/temp-file leaks closed. - App-target Swift: removed/loosened flaky wall-clock asserts while keeping the behavior assert (TerminalAndGhostty paste, ShellStartupMatrix budget), widened burst spacing (GhosttyNotificationDispatcher), and routed UITest socket/condition waits through the shared waitForControlSocketReady helper. All Python pass py_compile; all 4 touched packages pass swift build --build-tests; app-target Swift is CI-compiled. No test run against a socket. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7a374ac84c2c96d175df7656b9fd6657a788ea68) * Fix MobileCoreRPC cancelled-while-queued race deterministically (DEBUG-only seam) The remaining discovery finding: MobileCoreRPCClientTests spun a fixed 100 `Task.yield()`s to "wait" for a queued request to reach the session writer gate before cancelling it. Under scheduler load the queued task may not have reached the gate yet, so cancellation fires before `session.send` registers it and the cancelled-while-queued invariant is never exercised (false pass). That gate state (`queuedRequestIDs`) is private to the production `MobileCoreRPCSession` actor, so there is no deterministic test-only signal. Fix adds a `#if DEBUG`, read-only `debugQueuedRequestCount()` to the session and a thin client wrapper, placed in the file's existing `#if DEBUG` test-support extension (alongside `debugWithRequestTimeout`). The test now polls that real signal until the request is registered at the gate, then cancels. The accessor is `#if DEBUG` and read-only: it is compiled out of Release builds and changes no shipping-build behavior, so it needs no dogfood. `swift build --build-tests` (debug) compiles and links the package + test target cleanly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit c56f53ece49ae1c98072bb24996c55e6d15b4fa6) * Refresh Swift file-length budget for the de-flaked test files The fail-loud guards and ManualTestClock/RunLoop-poll rewrites grew five test files past their budgets (NotificationAndMenuBar +31, CommandPaletteSearchEngine +21, MultiWindowNotifications +9, BrowserPaneNavigationKeybind +3, AgentHibernation +1) and shrank two (TerminalAndGhostty -7 after dropping a flaky wall-clock assert, BrowserFixtureInteraction -6 after routing through the shared readiness helper). Budget updated for exactly those seven files. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: drop reintroduced wall-clock asserts / fixed sleeps CodeRabbit/Greptile caught spots where the second-pass fixes leaned on the very patterns this PR is removing. All test-only: - ShellStartupMatrixTests: drop the redundant `duration < 5.0` ceiling; the bootstrap already runs under a 5s `runProcess` timeout, so status==0 + !timedOut is the causal completion signal (no wall-clock assert). - MultiWindowNotificationsUITests: assert no-foreground at the causal point (right after `waitForCommandCompletionWhileBackgrounded`) instead of polling a fixed 2s window. - test_tab_dragging.py: remove two `time.sleep(1.5)` waits; the file-content poll loop right below is the readiness signal. - test_terminal_focus_routing.py: `tempfile.mktemp` -> `mkstemp` (atomic, no TOCTOU; clears Ruff S306). - test_session_restore_stress_kill_cycles.py: replace the fixed 0.1s settle with a deadline-bounded poll on the real is-selected signal. - RemoteCLIRelayServerTests: capture errno before close() so thrown bind/listen diagnostics report the real failure code. py_compile + swift build --build-tests (CmuxRemoteWorkspace) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Make CommandPalette search benchmarks advisory instead of wall-clock-gated Greptile/CodeRabbit correctly flagged that loosening these optimized-vs-reference ratio asserts (e.g. 0.80 -> 1.25) no longer validates anything: a band tight enough to prove the optimized path is faster flakes on shared CI, and a band wide enough not to flake passes even when the optimized path regresses. The engine exposes no preparation/work counter, so there is no causal (non-wall- clock) signal to assert on here. Per the repo test-time policy (no wall-clock latency asserts on shared CI) and matching the activation-session perf gate's advisory-timing approach, drop the flaky `#expect` ratio/dropped-frame assertions across all four benchmarks and keep the `BENCH ...` diagnostic prints for trend tracking. Each test still exercises both code paths end to end; real activation latency/frame-budget regressions are gated by the dedicated activation-session job. swift build --build-tests (CmuxCommandPalette) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * MobileCoreRPC: observe queue gate via @testable, not a production debug hook Per review (Aziz): test/debug extensions should not live in production source. Remove the `#if DEBUG debugQueuedRequestCount()` accessor from MobileCoreRPCClient/Session and instead widen `session` and `queuedRequestIDs` from `private` to `internal` so the cancellation test reads the writer-gate state directly through its existing `@testable import CmuxMobileRPC`. All test scaffolding now lives in the test target; production source carries only the two access-level changes (still module-internal, no shipping behavior change). swift build --build-tests (CmuxMobileRPC) clean; no debug funcs remain in Sources/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
feat: support resizing WKWebView browser viewports (#8072) * test(browser): cover WKWebView viewport emulation * feat(browser): emulate WKWebView viewports * docs(browser): explain WKWebView viewport sizing * fix(browser): preserve viewport across zoom and capture * fix(browser): close viewport review findings * test(browser): preserve reconciled host after capture * fix(browser): preserve current host after capture * test(browser): clean up viewport screenshots * fix(browser): import viewport geometry types * test(browser): cover viewport scale boundaries * fix(browser): bound and normalize viewport rendering * test(browser): cover viewport safety policies * fix(browser): bound viewport capture and inspector transitions * test(browser): cover inspector and zoomed full-page transitions * fix(browser): reconcile inspector and snapshot coordinates * test(browser): cover fullscreen and native zoom capture * fix(browser): reconcile fullscreen and native zoom capture * test(browser): assert runtime viewport dimensions * test(browser): gate runtime viewport regression * test(browser): await runtime viewport metrics * test: cover native viewport rounding at page zoom * fix: match native viewport rounding to WebKit * fix: preserve browser viewport presentation geometry * test: cover emulated viewport autoresizing mask * fix: disable raw viewport autoresizing while emulated * test: cover emulated viewport precision at page zoom * fix: preserve emulated viewport precision at page zoom * test: cover emulated viewport precision at default zoom * fix: preserve emulated viewport precision at default zoom * test: reject stale detached browser host restoration * fix: leave detached browser hosts unowned after capture * refactor: split browser viewport support types * test: report browser zoom render limits * fix: report browser zoom render limits * docs: keep viewport mode docs on public symbol * test: keep native browsers outside viewport host * fix: activate browser viewport host lazily * test: keep native portal layout passive * fix: keep native browser portal layout passive * test: preserve fractional zoom viewport dimensions * fix: preserve exact viewport dimensions at page zoom * test: cover near-integer zoom viewport rounding * ci: run CmuxBrowser package tests * fix: preserve near-integer zoom viewport dimensions * fix: keep viewport limit errors browser-generic | 2 个月前 | |
feat: support resizing WKWebView browser viewports (#8072) * test(browser): cover WKWebView viewport emulation * feat(browser): emulate WKWebView viewports * docs(browser): explain WKWebView viewport sizing * fix(browser): preserve viewport across zoom and capture * fix(browser): close viewport review findings * test(browser): preserve reconciled host after capture * fix(browser): preserve current host after capture * test(browser): clean up viewport screenshots * fix(browser): import viewport geometry types * test(browser): cover viewport scale boundaries * fix(browser): bound and normalize viewport rendering * test(browser): cover viewport safety policies * fix(browser): bound viewport capture and inspector transitions * test(browser): cover inspector and zoomed full-page transitions * fix(browser): reconcile inspector and snapshot coordinates * test(browser): cover fullscreen and native zoom capture * fix(browser): reconcile fullscreen and native zoom capture * test(browser): assert runtime viewport dimensions * test(browser): gate runtime viewport regression * test(browser): await runtime viewport metrics * test: cover native viewport rounding at page zoom * fix: match native viewport rounding to WebKit * fix: preserve browser viewport presentation geometry * test: cover emulated viewport autoresizing mask * fix: disable raw viewport autoresizing while emulated * test: cover emulated viewport precision at page zoom * fix: preserve emulated viewport precision at page zoom * test: cover emulated viewport precision at default zoom * fix: preserve emulated viewport precision at default zoom * test: reject stale detached browser host restoration * fix: leave detached browser hosts unowned after capture * refactor: split browser viewport support types * test: report browser zoom render limits * fix: report browser zoom render limits * docs: keep viewport mode docs on public symbol * test: keep native browsers outside viewport host * fix: activate browser viewport host lazily * test: keep native portal layout passive * fix: keep native browser portal layout passive * test: preserve fractional zoom viewport dimensions * fix: preserve exact viewport dimensions at page zoom * test: cover near-integer zoom viewport rounding * ci: run CmuxBrowser package tests * fix: preserve near-integer zoom viewport dimensions * fix: keep viewport limit errors browser-generic | 2 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Eliminate flaky tests: deterministic rewrites across Swift, Python, shell (#6392) * Fix flaky tests across Swift, Python, and shell suites Replace nondeterministic test patterns with waits on real signals so the same correct code stops failing under CI/VM load. Found by a parallel static audit of ~950 test files; 61 fixes across 44 files. - sleep-as-sync: fixed sleep + single assert replaced with a bounded poll on the actual readiness signal (existing in-file wait helpers preferred). - tight-timeout: expectation/wait timeouts widened to generous bounds; only the failure path waits longer, passing runs return as soon as the signal fires. - wall-clock-assert: single-shot timing ratios replaced with best-of-N minima plus deterministic work-count assertions; perf benchmarks no longer flip on a single scheduler spike. - order-dependence: shared static recorder suite serialized; shared state reset per test. - network/port races: external endpoints swapped for local/ephemeral servers; real-network download hardened with retries, timeouts, and a transient-failure soft-skip. Deterministic test sleeps (allowed by the repo review rules) are left intact. All changed Python/shell files pass py_compile/bash -n; all five changed SPM package test targets compile via swift build --build-tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for de-flaked test files The poll helpers and best-of-N scaffolding that replaced fixed sleeps added a few lines to 6 Swift test files, crossing their per-file budgets. Update only those rows (and track HostBrowserSignInFlowTests, newly over 500). No other files touched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Deflake activation-session perf gate: best-of-N timing, advisory budgets The activation benchmark failed nondeterministically on snapshot_with_scrollback .elapsed_ms (e.g. 1827ms > 1500ms): a single wall-clock measurement of a load-sensitive ~1.3M-char snapshot, compared to a hard absolute ceiling, run on shared/Depot GUI runners. A loaded runner trivially adds 20%+, so the same correct code flipped red across unrelated PRs. Two changes, matching the de-flake philosophy used elsewhere in this PR (performance is measured, not gated on a fixed wall-clock ceiling): - best_of_snapshot_timing: re-measure each snapshot timing N times (--budget-snapshot-samples, default 3; persist=False so reruns have no session side effects) and keep the MINIMUM elapsed_ms. A real regression slows every sample so the min still regresses; transient contention only inflates some samples, so the min is stable. Shape/char counts are preserved. - Wall-clock timing budgets (launch/restore socket-ready, snapshot elapsed_ms) are now advisory: reported as budget_warnings, non-blocking. Deterministic structural budgets (min scrollback chars, min terminal surfaces, restored workspace/terminal counts) stay blocking. Pass --fail-on-timing-budget to hard-gate timing if a dedicated, controlled perf runner is used. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Rewrite remaining flaky tests to logic/signal/virtual-clock based Eliminates the 37 residual time-dependent test primitives the determinism checker flags (27 sleep-then-assert, 10 assert-on-duration) across 17 files, applying the two principles: - Assert on causality, not latency: duration assertions (elapsed < N, Date().timeIntervalSince < N) replaced with assertions on the logical outcome the timing proxied (operation completed / right value / right state / right order), or a generous deadline-bounded wait on the real completion signal where a "did not hang" guard is genuinely needed. - Invert the time dependency: sleep-then-assert replaced with a wait on the real signal: an injected virtual clock advanced by hand (e.g. ManualTestClock via existing makeHarness(clock:) seam), a completion signal awaited directly (semaphore/continuation/main-queue drain), or a deadline-bounded poll of the real state predicate where the system exposes no completion event. No sleeps or timeouts were merely widened; no assertions were weakened. The determinism checker now reports 0 active findings. All four changed SPM package test targets compile via swift build --build-tests; Python files pass py_compile. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for rewritten test files The deterministic rewrites (virtual clocks, poll helpers, signal waits) added lines to 7 Swift test files. Bump only those rows. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Flaky-test follow-up: review fixes + second-pass team sweep (#6452) * Address review: fail-loud guards + scoped soft-skips in de-flaked tests - test_browser_goto_split.py: _wait_url_loaded now raises on timeout instead of silently returning (false pass), and the local HTTP server is managed via a "with" block so the thread/socket can't leak on setup exceptions. - test_surface_move_reorder_api.py: assert the cleanup re-select of ws0 actually converges instead of ignoring the _wait() result (state leak across runs). - test_homebrew_sha.sh: only soft-skip transient transport failures (000/408/429/5xx); fail hard on deterministic client errors like 404 (missing release asset). - test_visual_typing_char_by_char.py: validate the typed glyph against the post-prompt region of the last line, so a "cmux" already in the prompt/path can't satisfy the check; robust to trailing zsh-autosuggestion glyphs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 5d4c282535a497ad263128b21f8805a34556f55d) * Address review round 2: critical walrus bug + 3 fail-loud/correctness guards - test_pane_break_swap_preserve_focus.py: fix NameError -- the lambda referenced `p` in the ternary condition before the walrus assigned it. Bind the walrus inside the condition so the predicate reads panes once and returns them. - perf-activation-session.py: a missing measurement (actual is None) is now always a blocking failure, not advisory -- absent data is a benchmark-contract violation, distinct from load-sensitive over-budget timing. - perf-activation-session.py: copy the real-scrollback measurement before storing it under snapshot_with_scrollback, so best_of_snapshot_timing's in-place mutation no longer clobbers the raw snapshot_with_real_scrollback capture. - NotificationAndMenuBarTests.swift: on the stall-timeout path, fail fast via a lock-guarded result box instead of awaiting evaluationTask.value, which could hang until the suite timeout when the hook ignores cancellation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7e75ad467bc95cef0436236d94183f43b9aaffb2) * Add reports.md: consolidated flaky-test discovery (in-session workflow) Output of a parallel dynamic workflow (20 agents over the 419 signal-bearing test files of 2080) verifying residual flakiness on this already-de-flaked branch. 101 candidate findings across 86 files, dominated by wall-clock timing asserts, sleep-as-sync, and async races -- concentrated in Packages/ unit tests the first 44-file sweep did not cover. Each entry is a candidate; the fix phase adversarially verifies every one before changing code (repo rules allow deterministic test sleeps, so those are left alone). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit e1952d01faee805298bdff215af1d0897f63abcf) * Fix 37 verified residual flaky tests (coordinated area-team workflow) Second-pass de-flake driven by the in-session discovery + team-fix workflows (see reports.md). Adversarial per-area verification confirmed genuine flakiness before any edit; 9 candidates were refuted (deterministic-allowed sleeps, factually-wrong premises, fixes needing risky production changes) and left. Highlights: - Packages/ unit tests (missed by the first 44-file sweep): AuthCoordinator and HostBrowserSignIn now drive timeouts off an injected ManualTestClock instead of real ContinuousClock; ChatConversationStore poller results are asserted with #expect instead of discarded (fail-loud); RemoteProxyBroker / RemoteCLIRelay / CommandPaletteSearchEngine timing/race tightened. - Integration tests: per-PID/UUID socket+port paths (terminal_focus_routing, ctrl_socket, ssh_remote_* disjoint port ranges, CMUXCLIErrorOutput socket) to kill cross-run collisions; order-dependent session-restore checks now scan all workspaces instead of trusting seeding order; latent NameErrors fixed (pane_resize missing `import time` / `must`); fail-open waits made loud (new_tab tmp-write now raises); surface-targeted send to avoid focus races (tab_dragging); fd/temp-file leaks closed. - App-target Swift: removed/loosened flaky wall-clock asserts while keeping the behavior assert (TerminalAndGhostty paste, ShellStartupMatrix budget), widened burst spacing (GhosttyNotificationDispatcher), and routed UITest socket/condition waits through the shared waitForControlSocketReady helper. All Python pass py_compile; all 4 touched packages pass swift build --build-tests; app-target Swift is CI-compiled. No test run against a socket. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7a374ac84c2c96d175df7656b9fd6657a788ea68) * Fix MobileCoreRPC cancelled-while-queued race deterministically (DEBUG-only seam) The remaining discovery finding: MobileCoreRPCClientTests spun a fixed 100 `Task.yield()`s to "wait" for a queued request to reach the session writer gate before cancelling it. Under scheduler load the queued task may not have reached the gate yet, so cancellation fires before `session.send` registers it and the cancelled-while-queued invariant is never exercised (false pass). That gate state (`queuedRequestIDs`) is private to the production `MobileCoreRPCSession` actor, so there is no deterministic test-only signal. Fix adds a `#if DEBUG`, read-only `debugQueuedRequestCount()` to the session and a thin client wrapper, placed in the file's existing `#if DEBUG` test-support extension (alongside `debugWithRequestTimeout`). The test now polls that real signal until the request is registered at the gate, then cancels. The accessor is `#if DEBUG` and read-only: it is compiled out of Release builds and changes no shipping-build behavior, so it needs no dogfood. `swift build --build-tests` (debug) compiles and links the package + test target cleanly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit c56f53ece49ae1c98072bb24996c55e6d15b4fa6) * Refresh Swift file-length budget for the de-flaked test files The fail-loud guards and ManualTestClock/RunLoop-poll rewrites grew five test files past their budgets (NotificationAndMenuBar +31, CommandPaletteSearchEngine +21, MultiWindowNotifications +9, BrowserPaneNavigationKeybind +3, AgentHibernation +1) and shrank two (TerminalAndGhostty -7 after dropping a flaky wall-clock assert, BrowserFixtureInteraction -6 after routing through the shared readiness helper). Budget updated for exactly those seven files. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: drop reintroduced wall-clock asserts / fixed sleeps CodeRabbit/Greptile caught spots where the second-pass fixes leaned on the very patterns this PR is removing. All test-only: - ShellStartupMatrixTests: drop the redundant `duration < 5.0` ceiling; the bootstrap already runs under a 5s `runProcess` timeout, so status==0 + !timedOut is the causal completion signal (no wall-clock assert). - MultiWindowNotificationsUITests: assert no-foreground at the causal point (right after `waitForCommandCompletionWhileBackgrounded`) instead of polling a fixed 2s window. - test_tab_dragging.py: remove two `time.sleep(1.5)` waits; the file-content poll loop right below is the readiness signal. - test_terminal_focus_routing.py: `tempfile.mktemp` -> `mkstemp` (atomic, no TOCTOU; clears Ruff S306). - test_session_restore_stress_kill_cycles.py: replace the fixed 0.1s settle with a deadline-bounded poll on the real is-selected signal. - RemoteCLIRelayServerTests: capture errno before close() so thrown bind/listen diagnostics report the real failure code. py_compile + swift build --build-tests (CmuxRemoteWorkspace) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Make CommandPalette search benchmarks advisory instead of wall-clock-gated Greptile/CodeRabbit correctly flagged that loosening these optimized-vs-reference ratio asserts (e.g. 0.80 -> 1.25) no longer validates anything: a band tight enough to prove the optimized path is faster flakes on shared CI, and a band wide enough not to flake passes even when the optimized path regresses. The engine exposes no preparation/work counter, so there is no causal (non-wall- clock) signal to assert on here. Per the repo test-time policy (no wall-clock latency asserts on shared CI) and matching the activation-session perf gate's advisory-timing approach, drop the flaky `#expect` ratio/dropped-frame assertions across all four benchmarks and keep the `BENCH ...` diagnostic prints for trend tracking. Each test still exercises both code paths end to end; real activation latency/frame-budget regressions are gated by the dedicated activation-session job. swift build --build-tests (CmuxCommandPalette) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * MobileCoreRPC: observe queue gate via @testable, not a production debug hook Per review (Aziz): test/debug extensions should not live in production source. Remove the `#if DEBUG debugQueuedRequestCount()` accessor from MobileCoreRPCClient/Session and instead widen `session` and `queuedRequestIDs` from `private` to `internal` so the cancellation test reads the writer-gate state directly through its existing `@testable import CmuxMobileRPC`. All test scaffolding now lives in the test target; production source carries only the two access-level changes (still module-internal, no shipping behavior change). swift build --build-tests (CmuxMobileRPC) clean; no debug funcs remain in Sources/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
test: cover hidden browser screenshot freshness | 3 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Eliminate flaky tests: deterministic rewrites across Swift, Python, shell (#6392) * Fix flaky tests across Swift, Python, and shell suites Replace nondeterministic test patterns with waits on real signals so the same correct code stops failing under CI/VM load. Found by a parallel static audit of ~950 test files; 61 fixes across 44 files. - sleep-as-sync: fixed sleep + single assert replaced with a bounded poll on the actual readiness signal (existing in-file wait helpers preferred). - tight-timeout: expectation/wait timeouts widened to generous bounds; only the failure path waits longer, passing runs return as soon as the signal fires. - wall-clock-assert: single-shot timing ratios replaced with best-of-N minima plus deterministic work-count assertions; perf benchmarks no longer flip on a single scheduler spike. - order-dependence: shared static recorder suite serialized; shared state reset per test. - network/port races: external endpoints swapped for local/ephemeral servers; real-network download hardened with retries, timeouts, and a transient-failure soft-skip. Deterministic test sleeps (allowed by the repo review rules) are left intact. All changed Python/shell files pass py_compile/bash -n; all five changed SPM package test targets compile via swift build --build-tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for de-flaked test files The poll helpers and best-of-N scaffolding that replaced fixed sleeps added a few lines to 6 Swift test files, crossing their per-file budgets. Update only those rows (and track HostBrowserSignInFlowTests, newly over 500). No other files touched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Deflake activation-session perf gate: best-of-N timing, advisory budgets The activation benchmark failed nondeterministically on snapshot_with_scrollback .elapsed_ms (e.g. 1827ms > 1500ms): a single wall-clock measurement of a load-sensitive ~1.3M-char snapshot, compared to a hard absolute ceiling, run on shared/Depot GUI runners. A loaded runner trivially adds 20%+, so the same correct code flipped red across unrelated PRs. Two changes, matching the de-flake philosophy used elsewhere in this PR (performance is measured, not gated on a fixed wall-clock ceiling): - best_of_snapshot_timing: re-measure each snapshot timing N times (--budget-snapshot-samples, default 3; persist=False so reruns have no session side effects) and keep the MINIMUM elapsed_ms. A real regression slows every sample so the min still regresses; transient contention only inflates some samples, so the min is stable. Shape/char counts are preserved. - Wall-clock timing budgets (launch/restore socket-ready, snapshot elapsed_ms) are now advisory: reported as budget_warnings, non-blocking. Deterministic structural budgets (min scrollback chars, min terminal surfaces, restored workspace/terminal counts) stay blocking. Pass --fail-on-timing-budget to hard-gate timing if a dedicated, controlled perf runner is used. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Rewrite remaining flaky tests to logic/signal/virtual-clock based Eliminates the 37 residual time-dependent test primitives the determinism checker flags (27 sleep-then-assert, 10 assert-on-duration) across 17 files, applying the two principles: - Assert on causality, not latency: duration assertions (elapsed < N, Date().timeIntervalSince < N) replaced with assertions on the logical outcome the timing proxied (operation completed / right value / right state / right order), or a generous deadline-bounded wait on the real completion signal where a "did not hang" guard is genuinely needed. - Invert the time dependency: sleep-then-assert replaced with a wait on the real signal: an injected virtual clock advanced by hand (e.g. ManualTestClock via existing makeHarness(clock:) seam), a completion signal awaited directly (semaphore/continuation/main-queue drain), or a deadline-bounded poll of the real state predicate where the system exposes no completion event. No sleeps or timeouts were merely widened; no assertions were weakened. The determinism checker now reports 0 active findings. All four changed SPM package test targets compile via swift build --build-tests; Python files pass py_compile. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for rewritten test files The deterministic rewrites (virtual clocks, poll helpers, signal waits) added lines to 7 Swift test files. Bump only those rows. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Fix offscreen terminal helper PTY startup (#4233) * test: cover offscreen terminal helper startup * fix: start offscreen terminal helpers before attach Co-authored-by: austinpower1258 <austinwang115@gmail.com> * fix: isolate offscreen startup helpers * Keep headless terminal bootstrap out of health window state * test: cover direct AppKit terminal view startup * fix: keep plain terminal hosted views surface-bound * fix: remove dead main-thread branch in force refresh * fix: annotate headless window helpers main actor * fix: ignore headless window in force refresh --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> Co-authored-by: austinpower1258 <austinwang115@gmail.com> | 4 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Make cmux.json the canonical settings file (#3409) * feat: make cmux.json the settings file * fix: address cmux.json review feedback * fix: preserve legacy settings schema URL --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Fix stable IDs in mirror workspace CLI output (#8437) * test(cli): cover stable mirror workspace IDs * fix(cli): retain stable workspace inspection IDs * test(cli): assert IDs for every workspace row * fix(cli): preserve only workspace inspection UUIDs * test(cli): cover non-workspace top identifiers * fix(cli): preserve only exact workspace handles * fix(cli): simplify exact workspace ref check * test(cli): cover plural workspace identifiers --------- Co-authored-by: cmux reload-cloud <cmux-reload-cloud@users.noreply.github.com> | 2 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Fix background new-workspace commands (#4137) * test: cover background layout commands * fix: start background workspace commands * test: stabilize background workspace CLI regressions * Narrow background startup for https://github.com/manaflow-ai/cmux/issues/4090 * Preserve explicit background surface starts --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 4 个月前 | |
Flaky-test follow-up: review fixes + second-pass team sweep (#6452) * Address review: fail-loud guards + scoped soft-skips in de-flaked tests - test_browser_goto_split.py: _wait_url_loaded now raises on timeout instead of silently returning (false pass), and the local HTTP server is managed via a "with" block so the thread/socket can't leak on setup exceptions. - test_surface_move_reorder_api.py: assert the cleanup re-select of ws0 actually converges instead of ignoring the _wait() result (state leak across runs). - test_homebrew_sha.sh: only soft-skip transient transport failures (000/408/429/5xx); fail hard on deterministic client errors like 404 (missing release asset). - test_visual_typing_char_by_char.py: validate the typed glyph against the post-prompt region of the last line, so a "cmux" already in the prompt/path can't satisfy the check; robust to trailing zsh-autosuggestion glyphs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 5d4c282535a497ad263128b21f8805a34556f55d) * Address review round 2: critical walrus bug + 3 fail-loud/correctness guards - test_pane_break_swap_preserve_focus.py: fix NameError -- the lambda referenced `p` in the ternary condition before the walrus assigned it. Bind the walrus inside the condition so the predicate reads panes once and returns them. - perf-activation-session.py: a missing measurement (actual is None) is now always a blocking failure, not advisory -- absent data is a benchmark-contract violation, distinct from load-sensitive over-budget timing. - perf-activation-session.py: copy the real-scrollback measurement before storing it under snapshot_with_scrollback, so best_of_snapshot_timing's in-place mutation no longer clobbers the raw snapshot_with_real_scrollback capture. - NotificationAndMenuBarTests.swift: on the stall-timeout path, fail fast via a lock-guarded result box instead of awaiting evaluationTask.value, which could hang until the suite timeout when the hook ignores cancellation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7e75ad467bc95cef0436236d94183f43b9aaffb2) * Add reports.md: consolidated flaky-test discovery (in-session workflow) Output of a parallel dynamic workflow (20 agents over the 419 signal-bearing test files of 2080) verifying residual flakiness on this already-de-flaked branch. 101 candidate findings across 86 files, dominated by wall-clock timing asserts, sleep-as-sync, and async races -- concentrated in Packages/ unit tests the first 44-file sweep did not cover. Each entry is a candidate; the fix phase adversarially verifies every one before changing code (repo rules allow deterministic test sleeps, so those are left alone). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit e1952d01faee805298bdff215af1d0897f63abcf) * Fix 37 verified residual flaky tests (coordinated area-team workflow) Second-pass de-flake driven by the in-session discovery + team-fix workflows (see reports.md). Adversarial per-area verification confirmed genuine flakiness before any edit; 9 candidates were refuted (deterministic-allowed sleeps, factually-wrong premises, fixes needing risky production changes) and left. Highlights: - Packages/ unit tests (missed by the first 44-file sweep): AuthCoordinator and HostBrowserSignIn now drive timeouts off an injected ManualTestClock instead of real ContinuousClock; ChatConversationStore poller results are asserted with #expect instead of discarded (fail-loud); RemoteProxyBroker / RemoteCLIRelay / CommandPaletteSearchEngine timing/race tightened. - Integration tests: per-PID/UUID socket+port paths (terminal_focus_routing, ctrl_socket, ssh_remote_* disjoint port ranges, CMUXCLIErrorOutput socket) to kill cross-run collisions; order-dependent session-restore checks now scan all workspaces instead of trusting seeding order; latent NameErrors fixed (pane_resize missing `import time` / `must`); fail-open waits made loud (new_tab tmp-write now raises); surface-targeted send to avoid focus races (tab_dragging); fd/temp-file leaks closed. - App-target Swift: removed/loosened flaky wall-clock asserts while keeping the behavior assert (TerminalAndGhostty paste, ShellStartupMatrix budget), widened burst spacing (GhosttyNotificationDispatcher), and routed UITest socket/condition waits through the shared waitForControlSocketReady helper. All Python pass py_compile; all 4 touched packages pass swift build --build-tests; app-target Swift is CI-compiled. No test run against a socket. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7a374ac84c2c96d175df7656b9fd6657a788ea68) * Fix MobileCoreRPC cancelled-while-queued race deterministically (DEBUG-only seam) The remaining discovery finding: MobileCoreRPCClientTests spun a fixed 100 `Task.yield()`s to "wait" for a queued request to reach the session writer gate before cancelling it. Under scheduler load the queued task may not have reached the gate yet, so cancellation fires before `session.send` registers it and the cancelled-while-queued invariant is never exercised (false pass). That gate state (`queuedRequestIDs`) is private to the production `MobileCoreRPCSession` actor, so there is no deterministic test-only signal. Fix adds a `#if DEBUG`, read-only `debugQueuedRequestCount()` to the session and a thin client wrapper, placed in the file's existing `#if DEBUG` test-support extension (alongside `debugWithRequestTimeout`). The test now polls that real signal until the request is registered at the gate, then cancels. The accessor is `#if DEBUG` and read-only: it is compiled out of Release builds and changes no shipping-build behavior, so it needs no dogfood. `swift build --build-tests` (debug) compiles and links the package + test target cleanly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit c56f53ece49ae1c98072bb24996c55e6d15b4fa6) * Refresh Swift file-length budget for the de-flaked test files The fail-loud guards and ManualTestClock/RunLoop-poll rewrites grew five test files past their budgets (NotificationAndMenuBar +31, CommandPaletteSearchEngine +21, MultiWindowNotifications +9, BrowserPaneNavigationKeybind +3, AgentHibernation +1) and shrank two (TerminalAndGhostty -7 after dropping a flaky wall-clock assert, BrowserFixtureInteraction -6 after routing through the shared readiness helper). Budget updated for exactly those seven files. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: drop reintroduced wall-clock asserts / fixed sleeps CodeRabbit/Greptile caught spots where the second-pass fixes leaned on the very patterns this PR is removing. All test-only: - ShellStartupMatrixTests: drop the redundant `duration < 5.0` ceiling; the bootstrap already runs under a 5s `runProcess` timeout, so status==0 + !timedOut is the causal completion signal (no wall-clock assert). - MultiWindowNotificationsUITests: assert no-foreground at the causal point (right after `waitForCommandCompletionWhileBackgrounded`) instead of polling a fixed 2s window. - test_tab_dragging.py: remove two `time.sleep(1.5)` waits; the file-content poll loop right below is the readiness signal. - test_terminal_focus_routing.py: `tempfile.mktemp` -> `mkstemp` (atomic, no TOCTOU; clears Ruff S306). - test_session_restore_stress_kill_cycles.py: replace the fixed 0.1s settle with a deadline-bounded poll on the real is-selected signal. - RemoteCLIRelayServerTests: capture errno before close() so thrown bind/listen diagnostics report the real failure code. py_compile + swift build --build-tests (CmuxRemoteWorkspace) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Make CommandPalette search benchmarks advisory instead of wall-clock-gated Greptile/CodeRabbit correctly flagged that loosening these optimized-vs-reference ratio asserts (e.g. 0.80 -> 1.25) no longer validates anything: a band tight enough to prove the optimized path is faster flakes on shared CI, and a band wide enough not to flake passes even when the optimized path regresses. The engine exposes no preparation/work counter, so there is no causal (non-wall- clock) signal to assert on here. Per the repo test-time policy (no wall-clock latency asserts on shared CI) and matching the activation-session perf gate's advisory-timing approach, drop the flaky `#expect` ratio/dropped-frame assertions across all four benchmarks and keep the `BENCH ...` diagnostic prints for trend tracking. Each test still exercises both code paths end to end; real activation latency/frame-budget regressions are gated by the dedicated activation-session job. swift build --build-tests (CmuxCommandPalette) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * MobileCoreRPC: observe queue gate via @testable, not a production debug hook Per review (Aziz): test/debug extensions should not live in production source. Remove the `#if DEBUG debugQueuedRequestCount()` accessor from MobileCoreRPCClient/Session and instead widen `session` and `queuedRequestIDs` from `private` to `internal` so the cancellation test reads the writer-gate state directly through its existing `@testable import CmuxMobileRPC`. All test scaffolding now lives in the test target; production source carries only the two access-level changes (still module-internal, no shipping behavior change). swift build --build-tests (CmuxMobileRPC) clean; no debug funcs remain in Sources/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Fix background new-workspace commands (#4137) * test: cover background layout commands * fix: start background workspace commands * test: stabilize background workspace CLI regressions * Narrow background startup for https://github.com/manaflow-ai/cmux/issues/4090 * Preserve explicit background surface starts --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 4 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Expose set-status priority Summary: - Document and localize cmux set-status --priority <n> in CLI help. - Include --priority in top-level CLI usage and API examples. - Extend sidebar metadata CLI coverage to verify priority forwarding and ordering. Verification: - CI required checks green on PR #4064. | 4 个月前 | |
Eliminate flaky tests: deterministic rewrites across Swift, Python, shell (#6392) * Fix flaky tests across Swift, Python, and shell suites Replace nondeterministic test patterns with waits on real signals so the same correct code stops failing under CI/VM load. Found by a parallel static audit of ~950 test files; 61 fixes across 44 files. - sleep-as-sync: fixed sleep + single assert replaced with a bounded poll on the actual readiness signal (existing in-file wait helpers preferred). - tight-timeout: expectation/wait timeouts widened to generous bounds; only the failure path waits longer, passing runs return as soon as the signal fires. - wall-clock-assert: single-shot timing ratios replaced with best-of-N minima plus deterministic work-count assertions; perf benchmarks no longer flip on a single scheduler spike. - order-dependence: shared static recorder suite serialized; shared state reset per test. - network/port races: external endpoints swapped for local/ephemeral servers; real-network download hardened with retries, timeouts, and a transient-failure soft-skip. Deterministic test sleeps (allowed by the repo review rules) are left intact. All changed Python/shell files pass py_compile/bash -n; all five changed SPM package test targets compile via swift build --build-tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for de-flaked test files The poll helpers and best-of-N scaffolding that replaced fixed sleeps added a few lines to 6 Swift test files, crossing their per-file budgets. Update only those rows (and track HostBrowserSignInFlowTests, newly over 500). No other files touched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Deflake activation-session perf gate: best-of-N timing, advisory budgets The activation benchmark failed nondeterministically on snapshot_with_scrollback .elapsed_ms (e.g. 1827ms > 1500ms): a single wall-clock measurement of a load-sensitive ~1.3M-char snapshot, compared to a hard absolute ceiling, run on shared/Depot GUI runners. A loaded runner trivially adds 20%+, so the same correct code flipped red across unrelated PRs. Two changes, matching the de-flake philosophy used elsewhere in this PR (performance is measured, not gated on a fixed wall-clock ceiling): - best_of_snapshot_timing: re-measure each snapshot timing N times (--budget-snapshot-samples, default 3; persist=False so reruns have no session side effects) and keep the MINIMUM elapsed_ms. A real regression slows every sample so the min still regresses; transient contention only inflates some samples, so the min is stable. Shape/char counts are preserved. - Wall-clock timing budgets (launch/restore socket-ready, snapshot elapsed_ms) are now advisory: reported as budget_warnings, non-blocking. Deterministic structural budgets (min scrollback chars, min terminal surfaces, restored workspace/terminal counts) stay blocking. Pass --fail-on-timing-budget to hard-gate timing if a dedicated, controlled perf runner is used. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Rewrite remaining flaky tests to logic/signal/virtual-clock based Eliminates the 37 residual time-dependent test primitives the determinism checker flags (27 sleep-then-assert, 10 assert-on-duration) across 17 files, applying the two principles: - Assert on causality, not latency: duration assertions (elapsed < N, Date().timeIntervalSince < N) replaced with assertions on the logical outcome the timing proxied (operation completed / right value / right state / right order), or a generous deadline-bounded wait on the real completion signal where a "did not hang" guard is genuinely needed. - Invert the time dependency: sleep-then-assert replaced with a wait on the real signal: an injected virtual clock advanced by hand (e.g. ManualTestClock via existing makeHarness(clock:) seam), a completion signal awaited directly (semaphore/continuation/main-queue drain), or a deadline-bounded poll of the real state predicate where the system exposes no completion event. No sleeps or timeouts were merely widened; no assertions were weakened. The determinism checker now reports 0 active findings. All four changed SPM package test targets compile via swift build --build-tests; Python files pass py_compile. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for rewritten test files The deterministic rewrites (virtual clocks, poll helpers, signal waits) added lines to 7 Swift test files. Bump only those rows. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Eliminate flaky tests: deterministic rewrites across Swift, Python, shell (#6392) * Fix flaky tests across Swift, Python, and shell suites Replace nondeterministic test patterns with waits on real signals so the same correct code stops failing under CI/VM load. Found by a parallel static audit of ~950 test files; 61 fixes across 44 files. - sleep-as-sync: fixed sleep + single assert replaced with a bounded poll on the actual readiness signal (existing in-file wait helpers preferred). - tight-timeout: expectation/wait timeouts widened to generous bounds; only the failure path waits longer, passing runs return as soon as the signal fires. - wall-clock-assert: single-shot timing ratios replaced with best-of-N minima plus deterministic work-count assertions; perf benchmarks no longer flip on a single scheduler spike. - order-dependence: shared static recorder suite serialized; shared state reset per test. - network/port races: external endpoints swapped for local/ephemeral servers; real-network download hardened with retries, timeouts, and a transient-failure soft-skip. Deterministic test sleeps (allowed by the repo review rules) are left intact. All changed Python/shell files pass py_compile/bash -n; all five changed SPM package test targets compile via swift build --build-tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for de-flaked test files The poll helpers and best-of-N scaffolding that replaced fixed sleeps added a few lines to 6 Swift test files, crossing their per-file budgets. Update only those rows (and track HostBrowserSignInFlowTests, newly over 500). No other files touched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Deflake activation-session perf gate: best-of-N timing, advisory budgets The activation benchmark failed nondeterministically on snapshot_with_scrollback .elapsed_ms (e.g. 1827ms > 1500ms): a single wall-clock measurement of a load-sensitive ~1.3M-char snapshot, compared to a hard absolute ceiling, run on shared/Depot GUI runners. A loaded runner trivially adds 20%+, so the same correct code flipped red across unrelated PRs. Two changes, matching the de-flake philosophy used elsewhere in this PR (performance is measured, not gated on a fixed wall-clock ceiling): - best_of_snapshot_timing: re-measure each snapshot timing N times (--budget-snapshot-samples, default 3; persist=False so reruns have no session side effects) and keep the MINIMUM elapsed_ms. A real regression slows every sample so the min still regresses; transient contention only inflates some samples, so the min is stable. Shape/char counts are preserved. - Wall-clock timing budgets (launch/restore socket-ready, snapshot elapsed_ms) are now advisory: reported as budget_warnings, non-blocking. Deterministic structural budgets (min scrollback chars, min terminal surfaces, restored workspace/terminal counts) stay blocking. Pass --fail-on-timing-budget to hard-gate timing if a dedicated, controlled perf runner is used. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Rewrite remaining flaky tests to logic/signal/virtual-clock based Eliminates the 37 residual time-dependent test primitives the determinism checker flags (27 sleep-then-assert, 10 assert-on-duration) across 17 files, applying the two principles: - Assert on causality, not latency: duration assertions (elapsed < N, Date().timeIntervalSince < N) replaced with assertions on the logical outcome the timing proxied (operation completed / right value / right state / right order), or a generous deadline-bounded wait on the real completion signal where a "did not hang" guard is genuinely needed. - Invert the time dependency: sleep-then-assert replaced with a wait on the real signal: an injected virtual clock advanced by hand (e.g. ManualTestClock via existing makeHarness(clock:) seam), a completion signal awaited directly (semaphore/continuation/main-queue drain), or a deadline-bounded poll of the real state predicate where the system exposes no completion event. No sleeps or timeouts were merely widened; no assertions were weakened. The determinism checker now reports 0 active findings. All four changed SPM package test targets compile via swift build --build-tests; Python files pass py_compile. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for rewritten test files The deterministic rewrites (virtual clocks, poll helpers, signal waits) added lines to 7 Swift test files. Bump only those rows. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Eliminate flaky tests: deterministic rewrites across Swift, Python, shell (#6392) * Fix flaky tests across Swift, Python, and shell suites Replace nondeterministic test patterns with waits on real signals so the same correct code stops failing under CI/VM load. Found by a parallel static audit of ~950 test files; 61 fixes across 44 files. - sleep-as-sync: fixed sleep + single assert replaced with a bounded poll on the actual readiness signal (existing in-file wait helpers preferred). - tight-timeout: expectation/wait timeouts widened to generous bounds; only the failure path waits longer, passing runs return as soon as the signal fires. - wall-clock-assert: single-shot timing ratios replaced with best-of-N minima plus deterministic work-count assertions; perf benchmarks no longer flip on a single scheduler spike. - order-dependence: shared static recorder suite serialized; shared state reset per test. - network/port races: external endpoints swapped for local/ephemeral servers; real-network download hardened with retries, timeouts, and a transient-failure soft-skip. Deterministic test sleeps (allowed by the repo review rules) are left intact. All changed Python/shell files pass py_compile/bash -n; all five changed SPM package test targets compile via swift build --build-tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for de-flaked test files The poll helpers and best-of-N scaffolding that replaced fixed sleeps added a few lines to 6 Swift test files, crossing their per-file budgets. Update only those rows (and track HostBrowserSignInFlowTests, newly over 500). No other files touched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Deflake activation-session perf gate: best-of-N timing, advisory budgets The activation benchmark failed nondeterministically on snapshot_with_scrollback .elapsed_ms (e.g. 1827ms > 1500ms): a single wall-clock measurement of a load-sensitive ~1.3M-char snapshot, compared to a hard absolute ceiling, run on shared/Depot GUI runners. A loaded runner trivially adds 20%+, so the same correct code flipped red across unrelated PRs. Two changes, matching the de-flake philosophy used elsewhere in this PR (performance is measured, not gated on a fixed wall-clock ceiling): - best_of_snapshot_timing: re-measure each snapshot timing N times (--budget-snapshot-samples, default 3; persist=False so reruns have no session side effects) and keep the MINIMUM elapsed_ms. A real regression slows every sample so the min still regresses; transient contention only inflates some samples, so the min is stable. Shape/char counts are preserved. - Wall-clock timing budgets (launch/restore socket-ready, snapshot elapsed_ms) are now advisory: reported as budget_warnings, non-blocking. Deterministic structural budgets (min scrollback chars, min terminal surfaces, restored workspace/terminal counts) stay blocking. Pass --fail-on-timing-budget to hard-gate timing if a dedicated, controlled perf runner is used. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Rewrite remaining flaky tests to logic/signal/virtual-clock based Eliminates the 37 residual time-dependent test primitives the determinism checker flags (27 sleep-then-assert, 10 assert-on-duration) across 17 files, applying the two principles: - Assert on causality, not latency: duration assertions (elapsed < N, Date().timeIntervalSince < N) replaced with assertions on the logical outcome the timing proxied (operation completed / right value / right state / right order), or a generous deadline-bounded wait on the real completion signal where a "did not hang" guard is genuinely needed. - Invert the time dependency: sleep-then-assert replaced with a wait on the real signal: an injected virtual clock advanced by hand (e.g. ManualTestClock via existing makeHarness(clock:) seam), a completion signal awaited directly (semaphore/continuation/main-queue drain), or a deadline-bounded poll of the real state predicate where the system exposes no completion event. No sleeps or timeouts were merely widened; no assertions were weakened. The determinism checker now reports 0 active findings. All four changed SPM package test targets compile via swift build --build-tests; Python files pass py_compile. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for rewritten test files The deterministic rewrites (virtual clocks, poll helpers, signal waits) added lines to 7 Swift test files. Bump only those rows. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Flaky-test follow-up: review fixes + second-pass team sweep (#6452) * Address review: fail-loud guards + scoped soft-skips in de-flaked tests - test_browser_goto_split.py: _wait_url_loaded now raises on timeout instead of silently returning (false pass), and the local HTTP server is managed via a "with" block so the thread/socket can't leak on setup exceptions. - test_surface_move_reorder_api.py: assert the cleanup re-select of ws0 actually converges instead of ignoring the _wait() result (state leak across runs). - test_homebrew_sha.sh: only soft-skip transient transport failures (000/408/429/5xx); fail hard on deterministic client errors like 404 (missing release asset). - test_visual_typing_char_by_char.py: validate the typed glyph against the post-prompt region of the last line, so a "cmux" already in the prompt/path can't satisfy the check; robust to trailing zsh-autosuggestion glyphs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 5d4c282535a497ad263128b21f8805a34556f55d) * Address review round 2: critical walrus bug + 3 fail-loud/correctness guards - test_pane_break_swap_preserve_focus.py: fix NameError -- the lambda referenced `p` in the ternary condition before the walrus assigned it. Bind the walrus inside the condition so the predicate reads panes once and returns them. - perf-activation-session.py: a missing measurement (actual is None) is now always a blocking failure, not advisory -- absent data is a benchmark-contract violation, distinct from load-sensitive over-budget timing. - perf-activation-session.py: copy the real-scrollback measurement before storing it under snapshot_with_scrollback, so best_of_snapshot_timing's in-place mutation no longer clobbers the raw snapshot_with_real_scrollback capture. - NotificationAndMenuBarTests.swift: on the stall-timeout path, fail fast via a lock-guarded result box instead of awaiting evaluationTask.value, which could hang until the suite timeout when the hook ignores cancellation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7e75ad467bc95cef0436236d94183f43b9aaffb2) * Add reports.md: consolidated flaky-test discovery (in-session workflow) Output of a parallel dynamic workflow (20 agents over the 419 signal-bearing test files of 2080) verifying residual flakiness on this already-de-flaked branch. 101 candidate findings across 86 files, dominated by wall-clock timing asserts, sleep-as-sync, and async races -- concentrated in Packages/ unit tests the first 44-file sweep did not cover. Each entry is a candidate; the fix phase adversarially verifies every one before changing code (repo rules allow deterministic test sleeps, so those are left alone). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit e1952d01faee805298bdff215af1d0897f63abcf) * Fix 37 verified residual flaky tests (coordinated area-team workflow) Second-pass de-flake driven by the in-session discovery + team-fix workflows (see reports.md). Adversarial per-area verification confirmed genuine flakiness before any edit; 9 candidates were refuted (deterministic-allowed sleeps, factually-wrong premises, fixes needing risky production changes) and left. Highlights: - Packages/ unit tests (missed by the first 44-file sweep): AuthCoordinator and HostBrowserSignIn now drive timeouts off an injected ManualTestClock instead of real ContinuousClock; ChatConversationStore poller results are asserted with #expect instead of discarded (fail-loud); RemoteProxyBroker / RemoteCLIRelay / CommandPaletteSearchEngine timing/race tightened. - Integration tests: per-PID/UUID socket+port paths (terminal_focus_routing, ctrl_socket, ssh_remote_* disjoint port ranges, CMUXCLIErrorOutput socket) to kill cross-run collisions; order-dependent session-restore checks now scan all workspaces instead of trusting seeding order; latent NameErrors fixed (pane_resize missing `import time` / `must`); fail-open waits made loud (new_tab tmp-write now raises); surface-targeted send to avoid focus races (tab_dragging); fd/temp-file leaks closed. - App-target Swift: removed/loosened flaky wall-clock asserts while keeping the behavior assert (TerminalAndGhostty paste, ShellStartupMatrix budget), widened burst spacing (GhosttyNotificationDispatcher), and routed UITest socket/condition waits through the shared waitForControlSocketReady helper. All Python pass py_compile; all 4 touched packages pass swift build --build-tests; app-target Swift is CI-compiled. No test run against a socket. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7a374ac84c2c96d175df7656b9fd6657a788ea68) * Fix MobileCoreRPC cancelled-while-queued race deterministically (DEBUG-only seam) The remaining discovery finding: MobileCoreRPCClientTests spun a fixed 100 `Task.yield()`s to "wait" for a queued request to reach the session writer gate before cancelling it. Under scheduler load the queued task may not have reached the gate yet, so cancellation fires before `session.send` registers it and the cancelled-while-queued invariant is never exercised (false pass). That gate state (`queuedRequestIDs`) is private to the production `MobileCoreRPCSession` actor, so there is no deterministic test-only signal. Fix adds a `#if DEBUG`, read-only `debugQueuedRequestCount()` to the session and a thin client wrapper, placed in the file's existing `#if DEBUG` test-support extension (alongside `debugWithRequestTimeout`). The test now polls that real signal until the request is registered at the gate, then cancels. The accessor is `#if DEBUG` and read-only: it is compiled out of Release builds and changes no shipping-build behavior, so it needs no dogfood. `swift build --build-tests` (debug) compiles and links the package + test target cleanly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit c56f53ece49ae1c98072bb24996c55e6d15b4fa6) * Refresh Swift file-length budget for the de-flaked test files The fail-loud guards and ManualTestClock/RunLoop-poll rewrites grew five test files past their budgets (NotificationAndMenuBar +31, CommandPaletteSearchEngine +21, MultiWindowNotifications +9, BrowserPaneNavigationKeybind +3, AgentHibernation +1) and shrank two (TerminalAndGhostty -7 after dropping a flaky wall-clock assert, BrowserFixtureInteraction -6 after routing through the shared readiness helper). Budget updated for exactly those seven files. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: drop reintroduced wall-clock asserts / fixed sleeps CodeRabbit/Greptile caught spots where the second-pass fixes leaned on the very patterns this PR is removing. All test-only: - ShellStartupMatrixTests: drop the redundant `duration < 5.0` ceiling; the bootstrap already runs under a 5s `runProcess` timeout, so status==0 + !timedOut is the causal completion signal (no wall-clock assert). - MultiWindowNotificationsUITests: assert no-foreground at the causal point (right after `waitForCommandCompletionWhileBackgrounded`) instead of polling a fixed 2s window. - test_tab_dragging.py: remove two `time.sleep(1.5)` waits; the file-content poll loop right below is the readiness signal. - test_terminal_focus_routing.py: `tempfile.mktemp` -> `mkstemp` (atomic, no TOCTOU; clears Ruff S306). - test_session_restore_stress_kill_cycles.py: replace the fixed 0.1s settle with a deadline-bounded poll on the real is-selected signal. - RemoteCLIRelayServerTests: capture errno before close() so thrown bind/listen diagnostics report the real failure code. py_compile + swift build --build-tests (CmuxRemoteWorkspace) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Make CommandPalette search benchmarks advisory instead of wall-clock-gated Greptile/CodeRabbit correctly flagged that loosening these optimized-vs-reference ratio asserts (e.g. 0.80 -> 1.25) no longer validates anything: a band tight enough to prove the optimized path is faster flakes on shared CI, and a band wide enough not to flake passes even when the optimized path regresses. The engine exposes no preparation/work counter, so there is no causal (non-wall- clock) signal to assert on here. Per the repo test-time policy (no wall-clock latency asserts on shared CI) and matching the activation-session perf gate's advisory-timing approach, drop the flaky `#expect` ratio/dropped-frame assertions across all four benchmarks and keep the `BENCH ...` diagnostic prints for trend tracking. Each test still exercises both code paths end to end; real activation latency/frame-budget regressions are gated by the dedicated activation-session job. swift build --build-tests (CmuxCommandPalette) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * MobileCoreRPC: observe queue gate via @testable, not a production debug hook Per review (Aziz): test/debug extensions should not live in production source. Remove the `#if DEBUG debugQueuedRequestCount()` accessor from MobileCoreRPCClient/Session and instead widen `session` and `queuedRequestIDs` from `private` to `internal` so the cancellation test reads the writer-gate state directly through its existing `@testable import CmuxMobileRPC`. All test scaffolding now lives in the test target; production source carries only the two access-level changes (still module-internal, no shipping behavior change). swift build --build-tests (CmuxMobileRPC) clean; no debug funcs remain in Sources/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Fix config window to open active cmux Ghostty config (#3525) * Prove config window active-path regressions Add focused regression coverage for cmux Ghostty config resolution before changing the runtime path ownership. The tests capture config.ghostty winning over legacy config, nightly fallback to release, and the active-file behavior the settings window must expose. Constraint: Regression-test policy asks for a test commit before the fix commit Confidence: high Scope-risk: narrow Tested: Not run; regression-only commit is expected to fail before the implementation commit Not-tested: Local test execution per repository policy * Use the active cmux Ghostty config everywhere Move config-window and runtime helpers onto one cmux Ghostty config resolver. The resolver now chooses the same active file the app loads: config.ghostty when non-empty, legacy config when config.ghostty is empty, and release fallback for dev/nightly/staging channels without their own active file. The settings window no longer shows a standalone Ghostty tab, and Open in Editor/Reveals use the selected cmux file instead of ghostty_config_open_path. The launch crash found during verification came from settings-file shortcut parsing re-entering the shared shortcut store during static initialization. Settings-file parsing now normalizes recorder-compatible shortcuts without invoking the UI conflict lookup. Constraint: Issue 3518 requires the tagged reload script with --launch and origin-only pushes Rejected: Always open the channel-local config.ghostty | it hides the active legacy config when config.ghostty exists but is empty Rejected: Keep a standalone Ghostty tab | it displays a file that is not the cmux override path users need to edit Confidence: high Scope-risk: moderate Directive: Keep Settings, GhosttyApp.cmuxAppSupportConfigURLs, and GhosttyConfig.cmuxConfigPaths on CmuxGhosttyConfigPathResolver; do not reintroduce separate path policies Tested: git diff --check; jq empty Resources/Localizable.xcstrings; ./scripts/reload.sh --tag issue-3518-config-window --launch; running debug log loaded /Users/austinwang/Library/Application Support/com.cmuxterm.app/config Not-tested: Local test suites per repository policy * Match config button to settings picker sizing The Terminal Config action now uses the same trailing control column as the adjacent picker rows and the shorter Open Config label, so it visually aligns with controls like New Workspace Placement. Constraint: User requested this PR update after visual review Confidence: high Scope-risk: narrow Tested: git diff --check; jq empty Resources/Localizable.xcstrings; ./scripts/reload.sh --tag issue-3518-config-window --launch; visual check in Settings window Not-tested: Local test suites per repository policy * Use intrinsic size for config button Keep the requested Open Config label but remove the full-width button frame so the action renders like a normal settings button instead of a stretched picker-width control. Constraint: User visual review requested normal button padding Confidence: high Scope-risk: narrow Tested: git diff --check; jq empty Resources/Localizable.xcstrings; ./scripts/reload.sh --tag issue-3518-config-window --launch; visual check in Settings window Not-tested: Local test suites per repository policy * Address config resolver review feedback Preserve symlinked cmux Ghostty config URLs in the shared resolver, make synced-tab file actions name the active cmux config target, and move resolver regression coverage into a focused test file so the Swift file-length budget stays respected. Constraint: Dogfood launch waits until CI is green per latest user direction. Rejected: Load legacy config underneath non-empty config.ghostty | this would reintroduce the appearance-pinning bug that #3479 and issue #3518 are fixing. Confidence: high Scope-risk: narrow Tested: git diff --check; jq empty Resources/Localizable.xcstrings; plutil -lint GhosttyTabs.xcodeproj/project.pbxproj; python3 scripts/swift_file_length_budget.py --budget .github/swift-file-length-budget.tsv Not-tested: Local test suites and dogfood launch per repository policy and user direction; CI pending. * Resolve remaining review blockers Remove the duplicate settings-file shortcut normalization method left by the main merge, and collapse cmux config fallback channel checks into one helper. Constraint: Keep commits authored by austinywang only, with no co-author trailers. Rejected: Raise Swift file budgets for moved tests | the previous commit split coverage instead and the budget guard is now green. Confidence: high Scope-risk: narrow Tested: rg normalizedSettingsFileShortcut; git diff --check; jq empty Resources/Localizable.xcstrings; plutil -lint GhosttyTabs.xcodeproj/project.pbxproj; python3 scripts/swift_file_length_budget.py --budget .github/swift-file-length-budget.tsv Not-tested: Local test suites and dogfood launch per repository policy and user direction; CI pending. * Preserve symlinked config files when saving Save the cmux config through the shared config environment so the settings window follows symlink targets instead of atomically replacing the symlink path. This keeps symlinked config.ghostty files editable from the config window while preserving the link users expect. Constraint: Review feedback identified atomic writes against symlink URLs as a data-loss risk. Rejected: Disable atomic writes for all config saves | writing to the resolved target preserves atomicity without replacing the symlink. Confidence: high Scope-risk: narrow Tested: git diff --check; jq empty Resources/Localizable.xcstrings; plutil -lint GhosttyTabs.xcodeproj/project.pbxproj; python3 scripts/swift_file_length_budget.py --budget .github/swift-file-length-budget.tsv; ./scripts/reload.sh --tag issue-3518-config-window Not-tested: Local test suites per repository policy; CI pending after push. * Normalize new Xcode project IDs Pad the GhosttyConfigPathResolverTests PBXBuildFile and PBXFileReference identifiers to the standard 24-character Xcode project identifier length. The references keep the same relationship, only the malformed short IDs are corrected. Constraint: Cursor Bugbot flagged the newly-added A11EB project identifiers as 23 hex characters. Confidence: high Scope-risk: narrow Tested: plutil -lint GhosttyTabs.xcodeproj/project.pbxproj; rg confirmed the short A11EB IDs are gone Not-tested: Tagged reload is queued behind another xcodebuild lock; CI pending after push. * Address remaining config review feedback Treat symlinked standalone Ghostty configs as regular config inputs, make the active-config Finder action use standard Finder wording, and rename the shortcut normalization regression so the assertion matches the invariant it proves. Constraint: CodeRabbit posted new actionable review feedback after the symlink-save follow-up. Rejected: Move config-path logic into a package boundary in this PR | that is a broader extraction unrelated to issue 3518 and would expand review scope late in the loop. Confidence: high Scope-risk: narrow Tested: git diff --check; jq empty Resources/Localizable.xcstrings; plutil -lint GhosttyTabs.xcodeproj/project.pbxproj; python3 scripts/swift_file_length_budget.py --budget .github/swift-file-length-budget.tsv; ./scripts/reload.sh --tag issue-3518-config-window Not-tested: Local test suites per repository policy; CI pending after push. * Restore small control size for config button Keep the App settings Open Config button visually consistent with the other bordered settings controls by restoring the small control size modifier after the label-closure refactor. Constraint: Cursor Bot flagged the missing controlSize(.small) as a visual inconsistency. Confidence: high Scope-risk: narrow Tested: git diff --check; python3 scripts/swift_file_length_budget.py --budget .github/swift-file-length-budget.tsv Not-tested: Local test suites per repository policy; CI pending after push. * Make config action match settings controls The App settings row uses a regular-size picker in the same column, so the config action should keep the compact button chrome requested in review without looking visually smaller than adjacent controls. Constraint: Preserve .controlSize(.small) to satisfy the prior settings button review feedback. Confidence: high Scope-risk: narrow Tested: git diff --check Tested: python3 scripts/swift_file_length_budget.py --budget .github/swift-file-length-budget.tsv Not-tested: Full app visual pass not run before commit * Stabilize config materialization target Config materialization selected an active or editable URL, then delegated through the public writer which resolved the URL again. If another process changed config state between those operations, the returned URL could differ from the file that was created. Constraint: Keep cmux config selection dynamic for normal reads and saves, but make one materialization operation internally consistent. Rejected: Cache cmuxConfigURL globally | stale across dev/nightly/app variants and external file edits. Confidence: high Scope-risk: narrow Tested: git diff --check Tested: python3 scripts/swift_file_length_budget.py --budget .github/swift-file-length-budget.tsv Tested: swift FileManager override compile probe Not-tested: Full unit suite locally; CI will compile and run tests * Normalize config button sizing The config action should align with the neighboring settings controls without using a larger call-to-action label. Keep the small bordered button style while matching the row control column and using modest text metrics. Constraint: Preserve .controlSize(.small) from review feedback. Confidence: high Scope-risk: narrow Tested: git diff --check Tested: python3 scripts/swift_file_length_budget.py --budget .github/swift-file-length-budget.tsv Tested: ./scripts/reload.sh --tag issue-3518-config-window --launch Tested: cmux settings open via tagged CLI * Soften config button label weight The settings action should read like a standard compact control rather than a emphasized action. Keep the fixed size and control column, but use regular label weight for calmer visual balance. Constraint: Preserve .controlSize(.small) and picker-column alignment. Confidence: high Scope-risk: narrow Tested: git diff --check Tested: python3 scripts/swift_file_length_budget.py --budget .github/swift-file-length-budget.tsv Tested: ./scripts/reload.sh --tag issue-3518-config-window --launch Tested: cmux settings open via tagged CLI | 4 个月前 | |
Eliminate flaky tests: deterministic rewrites across Swift, Python, shell (#6392) * Fix flaky tests across Swift, Python, and shell suites Replace nondeterministic test patterns with waits on real signals so the same correct code stops failing under CI/VM load. Found by a parallel static audit of ~950 test files; 61 fixes across 44 files. - sleep-as-sync: fixed sleep + single assert replaced with a bounded poll on the actual readiness signal (existing in-file wait helpers preferred). - tight-timeout: expectation/wait timeouts widened to generous bounds; only the failure path waits longer, passing runs return as soon as the signal fires. - wall-clock-assert: single-shot timing ratios replaced with best-of-N minima plus deterministic work-count assertions; perf benchmarks no longer flip on a single scheduler spike. - order-dependence: shared static recorder suite serialized; shared state reset per test. - network/port races: external endpoints swapped for local/ephemeral servers; real-network download hardened with retries, timeouts, and a transient-failure soft-skip. Deterministic test sleeps (allowed by the repo review rules) are left intact. All changed Python/shell files pass py_compile/bash -n; all five changed SPM package test targets compile via swift build --build-tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for de-flaked test files The poll helpers and best-of-N scaffolding that replaced fixed sleeps added a few lines to 6 Swift test files, crossing their per-file budgets. Update only those rows (and track HostBrowserSignInFlowTests, newly over 500). No other files touched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Deflake activation-session perf gate: best-of-N timing, advisory budgets The activation benchmark failed nondeterministically on snapshot_with_scrollback .elapsed_ms (e.g. 1827ms > 1500ms): a single wall-clock measurement of a load-sensitive ~1.3M-char snapshot, compared to a hard absolute ceiling, run on shared/Depot GUI runners. A loaded runner trivially adds 20%+, so the same correct code flipped red across unrelated PRs. Two changes, matching the de-flake philosophy used elsewhere in this PR (performance is measured, not gated on a fixed wall-clock ceiling): - best_of_snapshot_timing: re-measure each snapshot timing N times (--budget-snapshot-samples, default 3; persist=False so reruns have no session side effects) and keep the MINIMUM elapsed_ms. A real regression slows every sample so the min still regresses; transient contention only inflates some samples, so the min is stable. Shape/char counts are preserved. - Wall-clock timing budgets (launch/restore socket-ready, snapshot elapsed_ms) are now advisory: reported as budget_warnings, non-blocking. Deterministic structural budgets (min scrollback chars, min terminal surfaces, restored workspace/terminal counts) stay blocking. Pass --fail-on-timing-budget to hard-gate timing if a dedicated, controlled perf runner is used. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Rewrite remaining flaky tests to logic/signal/virtual-clock based Eliminates the 37 residual time-dependent test primitives the determinism checker flags (27 sleep-then-assert, 10 assert-on-duration) across 17 files, applying the two principles: - Assert on causality, not latency: duration assertions (elapsed < N, Date().timeIntervalSince < N) replaced with assertions on the logical outcome the timing proxied (operation completed / right value / right state / right order), or a generous deadline-bounded wait on the real completion signal where a "did not hang" guard is genuinely needed. - Invert the time dependency: sleep-then-assert replaced with a wait on the real signal: an injected virtual clock advanced by hand (e.g. ManualTestClock via existing makeHarness(clock:) seam), a completion signal awaited directly (semaphore/continuation/main-queue drain), or a deadline-bounded poll of the real state predicate where the system exposes no completion event. No sleeps or timeouts were merely widened; no assertions were weakened. The determinism checker now reports 0 active findings. All four changed SPM package test targets compile via swift build --build-tests; Python files pass py_compile. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for rewritten test files The deterministic rewrites (virtual clocks, poll helpers, signal waits) added lines to 7 Swift test files. Bump only those rows. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Fix frozen terminals after split churn (#12) * Fix blank terminal after split operations and add visual tests ## Blank Terminal Fix - Add `needsRefreshAfterWindowChange` flag in GhosttyTerminalView - Force terminal refresh when view is added to window, even if size unchanged - Add `ghostty_surface_refresh()` call in attachToView for same-view reattachment - Add debug logging for surface attachment lifecycle (DEBUG builds only) ## Bonsplit Migration - Add bonsplit as local Swift package (vendor/bonsplit submodule) - Replace custom SplitTree with BonsplitController - Add Panel protocol with TerminalPanel and BrowserPanel implementations - Add SidebarTab as main tab container with BonsplitController - Remove old Splits/ directory (SplitTree, SplitView, TerminalSplitTreeView) ## Visual Screenshot Tests - Add test_visual_screenshots.py for automated visual regression testing - Uses in-app screenshot API (CGWindowListCreateImage) - no screen recording needed - Generates HTML report with before/after comparisons - Tests: splits, browser panels, focus switching, close operations, rapid cycles - Includes annotation fields for easy feedback ## Browser Shortcut (⌘⇧B) - Add keyboard shortcut to open browser panel in current pane - Add openBrowser() method to TabManager - Add shortcut configuration in KeyboardShortcutSettings ## Screenshot Command - Add 'screenshot' command to TerminalController for in-app window capture - Returns OK with screenshot ID and path ## Other - Add tests/visual_output/ and tests/visual_report.html to .gitignore * Add browser title subscription and set tab height to 30px - Subscribe to BrowserPanel.$pageTitle changes to update bonsplit tabs - Update tab titles in real-time as page navigation occurs - Clean up subscriptions when panels are removed - Set bonsplit tab bar and tab height to 30px (in submodule) * Fix socket API regressions in list_surfaces, list_bonsplit_tabs, focus_pane - list_surfaces: Remove [terminal]/[browser] suffix to keep UUID-only format that clients and tests expect for parsing - list_bonsplit_tabs --pane: Properly look up pane by UUID instead of creating a new PaneID (requires bonsplit PaneID.id to be public) - focus_pane: Accept both UUID strings and integer indices as documented * Fix browser panel stability and keyboard shortcuts - Prevent WKWebView focus lifecycle crashes during split/view reshuffles - Match bracket shortcuts via keyCode (Cmd+Shift+[ / ], Cmd+Ctrl+[ / ]) - Support Ghostty config goto_split:* keybinds when WebView is focused - Add focus_webview/is_webview_focused socket commands and regression tests - Rename SidebarTab to Workspace and update docs * Make ctrl+enter keybind test skippable Skip when the Ghostty keybind isn't configured or when osascript can't send keystrokes (no Accessibility permission), so VM runs stay green. * Auto-focus browser omnibar when blank When a browser surface is focused but no URL is loaded yet, focus the address bar instead of the WKWebView. * Stabilize socket surface indexing * Focus browser omnibar escape; add webview keybind UI tests - Escape in omnibar now returns focus to WKWebView\n- Add UI tests for Cmd+Ctrl+H pane navigation with WebKit focused (including Ghostty config)\n- Avoid flaky element screenshots in UpdatePillUITests on the UTM VM * Fix browser drag-to-split blanks and socket parsing * Fix webview-focused shortcuts and stabilize browser splits - Match ctrl/shift shortcuts by keyCode where needed (Ctrl+H, bracket keys) - Load Ghostty goto_split triggers reliably and refresh on config load - Add debug socket helpers: set_shortcut + simulate_shortcut for tests - Convert browser goto_split/keybind tests to socket-based injection (no osascript) - Bump bonsplit for drag-to-split fixes * Fix split layout collapse and harden socket pane APIs * Stabilize OSC 99 notification test timing * Fix terminal focus routing after split reparent * Support simulate_shortcut enter for focus routing test * Stabilize terminal focus routing test * Fix frozen new terminal tabs after many splits * Fix frozen new terminal tabs after splits * Fix terminal freeze on launch/new tabs * Update ghostty submodule * Fix terminal focus/render stalls after split churn * Fix nested split collapsing existing pane * Fix nested split collapse + stabilize new-surface focus * Update bonsplit submodule * Fix SIGINT test flake * Remove bonsplit tab-switch crossfade * Remove PROJECTS.md * Remove bonsplit tab selection animation * Ignore generated test reports * Middle click closes tab * Revert unintended .gitignore change * Fix build after main merge * Revert "Fix build after main merge" This reverts commit 16bf9816d0856b5385d52f886aa5eb50f3c9d9a4. * Revert "Merge remote-tracking branch 'origin/main' into fix/blank-terminal-and-visual-tests" This reverts commit 7c20fb53fd71fea7a19a3673f2dd73e5f0c783c4, reversing changes made to 0aff107d787bc9d8bbc28220090b4ca7af72e040. * Remove tab close fade animation * Use terminal.fill icon * Make terminal tab icon smaller * Match browser globe tab icon size * Bonsplit: tab min width 48 and tighter close button * Bonsplit: smaller tab title font * Show unread notification badge in bonsplit tabs and improve UI polish Sync unread notification state to bonsplit tab badges (blue dot). Improve EmptyPanelView with Terminal/Browser buttons and shortcut hints. Add tooltips to close tab button and search overlay buttons. * Fix reload.sh single-instance safety check on macOS Replace GNU-only `ps -o etimes=` with portable `ps -o etime=` and parse the dd-hh:mm:ss format manually for macOS compatibility. * Centralize keyboard shortcut definitions into Action enum Replace per-shortcut boilerplate with a single Action enum that holds the label, defaults key, and default binding for each shortcut. All call sites now use shortcut(for:). Settings UI is data-driven via ForEach(Action.allCases). Titlebar tooltips update dynamically when shortcuts are changed. Remove duplicate .keyboardShortcut() modifiers from menu items that are already handled by the event monitor. * Fix WKWebView consuming app menu shortcuts and close panel confirmation Add CmuxWebView subclass that routes key equivalents through the main menu before WebKit, so Cmd+N/Cmd+W/tab switching work when a browser pane is focused. Fix Cmd+W close-panel path: bypass Bonsplit delegate gating after the user confirms the running-process dialog by tracking forceCloseTabIds. Add unit tests (CmuxWebViewKeyEquivalentTests) and UI test scaffolding (MenuKeyEquivalentRoutingUITests) with a new cmux-unit Xcode scheme. * Update CLAUDE.md and PROJECTS.md with recent changes CLAUDE.md: enforce --tag for reload commands, add cleanup safety rules. PROJECTS.md: log notification badge, reload.sh fix, Cmd+W fix, WebView key equiv fix, and centralized shortcuts work. * Keep selection index stable on close * Add concepts page documenting terminology hierarchy New docs page explaining Window > Workspace > Pane > Surface > Panel hierarchy with aligned ASCII diagram. Updated tabs.mdx and splits.mdx to use consistent terminology (workspace instead of tab, surface instead of panel) and corrected outdated CLI command references. * Update bonsplit submodule * WIP: improve split close stability and UI regressions * Close terminal panel on child exit; hide terminal dirty dot * Fix split close/focus regressions and stabilize UI tests * Add unread Dock/Cmd+Tab badge with settings toggle * Fix browser-surface shortcuts and Cmd+L browser opening * Snapshot current workspace state before regression fixes * Update bonsplit submodule snapshot * Stabilize split-close regression capture and sidebar resize assertions * Change default Show Notifications shortcut from Cmd+Shift+I to Cmd+I * Fix update check readiness race, enable release update logging, and improve checking spinner * Restore terminal file drop, fix browser omnibar click focus, and add panel workspace ID mutation for surface moves * Add Cmd+digit workspace hints, titlebar shortcut pills, sidebar drag-reorder, and workspace placement settings * Add v2 browser automation API, surface move/reorder commands, and short-handle ref system to TerminalController * Add CLI browser command surface, --id-format flag, and move/reorder commands * Extend test clients with move/reorder APIs, ref-handle support, and increased timeouts * Harden test runner scripts with deterministic builds, retry logic, and robust socket readiness * Stabilize existing test suites with focus-wait helpers, increased timeouts, and API shape updates * Add terminal file drop e2e regression test * Add v2 browser API, CLI ref resolution, and surface move/reorder test suites * Add unit tests for shortcut hints, workspace reorder, drop planner, and update UI test stabilization * Add cmux-debug-windows skill with snapshot script and agent config * Update project docs: mark browser parity and move/reorder phases complete, add parallel agent workflow guidelines * Update bonsplit submodule: re-entrant setPosition guard, tab shortcut hints, and moveTab/reorderTab API * Add browser agent UX improvements: snapshot refs, placement reuse, diagnostics, and skill docs - Upgrade browser.snapshot to emit accessibility tree text with element refs (eN) - Add right-sibling pane reuse policy for browser.open_split placement - Add rich not_found diagnostics with retry logic for selector actions - Support --snapshot-after for post-action verification on mutating commands - Allow browser fill with empty text for clearing inputs - Default CLI --id-format to refs-first (UUIDs opt-in via --id-format uuids|both) - Format legacy new-pane/new-surface output with short surface refs - Add skills/cmuxterm-browser/ and skills/cmuxterm/ end-user skill docs - Add regression tests for placement policy, snapshot refs, diagnostics, and ID defaults * Update bonsplit submodule: keep raster favicons in color when inactive | 7 个月前 | |
Eliminate flaky tests: deterministic rewrites across Swift, Python, shell (#6392) * Fix flaky tests across Swift, Python, and shell suites Replace nondeterministic test patterns with waits on real signals so the same correct code stops failing under CI/VM load. Found by a parallel static audit of ~950 test files; 61 fixes across 44 files. - sleep-as-sync: fixed sleep + single assert replaced with a bounded poll on the actual readiness signal (existing in-file wait helpers preferred). - tight-timeout: expectation/wait timeouts widened to generous bounds; only the failure path waits longer, passing runs return as soon as the signal fires. - wall-clock-assert: single-shot timing ratios replaced with best-of-N minima plus deterministic work-count assertions; perf benchmarks no longer flip on a single scheduler spike. - order-dependence: shared static recorder suite serialized; shared state reset per test. - network/port races: external endpoints swapped for local/ephemeral servers; real-network download hardened with retries, timeouts, and a transient-failure soft-skip. Deterministic test sleeps (allowed by the repo review rules) are left intact. All changed Python/shell files pass py_compile/bash -n; all five changed SPM package test targets compile via swift build --build-tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for de-flaked test files The poll helpers and best-of-N scaffolding that replaced fixed sleeps added a few lines to 6 Swift test files, crossing their per-file budgets. Update only those rows (and track HostBrowserSignInFlowTests, newly over 500). No other files touched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Deflake activation-session perf gate: best-of-N timing, advisory budgets The activation benchmark failed nondeterministically on snapshot_with_scrollback .elapsed_ms (e.g. 1827ms > 1500ms): a single wall-clock measurement of a load-sensitive ~1.3M-char snapshot, compared to a hard absolute ceiling, run on shared/Depot GUI runners. A loaded runner trivially adds 20%+, so the same correct code flipped red across unrelated PRs. Two changes, matching the de-flake philosophy used elsewhere in this PR (performance is measured, not gated on a fixed wall-clock ceiling): - best_of_snapshot_timing: re-measure each snapshot timing N times (--budget-snapshot-samples, default 3; persist=False so reruns have no session side effects) and keep the MINIMUM elapsed_ms. A real regression slows every sample so the min still regresses; transient contention only inflates some samples, so the min is stable. Shape/char counts are preserved. - Wall-clock timing budgets (launch/restore socket-ready, snapshot elapsed_ms) are now advisory: reported as budget_warnings, non-blocking. Deterministic structural budgets (min scrollback chars, min terminal surfaces, restored workspace/terminal counts) stay blocking. Pass --fail-on-timing-budget to hard-gate timing if a dedicated, controlled perf runner is used. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Rewrite remaining flaky tests to logic/signal/virtual-clock based Eliminates the 37 residual time-dependent test primitives the determinism checker flags (27 sleep-then-assert, 10 assert-on-duration) across 17 files, applying the two principles: - Assert on causality, not latency: duration assertions (elapsed < N, Date().timeIntervalSince < N) replaced with assertions on the logical outcome the timing proxied (operation completed / right value / right state / right order), or a generous deadline-bounded wait on the real completion signal where a "did not hang" guard is genuinely needed. - Invert the time dependency: sleep-then-assert replaced with a wait on the real signal: an injected virtual clock advanced by hand (e.g. ManualTestClock via existing makeHarness(clock:) seam), a completion signal awaited directly (semaphore/continuation/main-queue drain), or a deadline-bounded poll of the real state predicate where the system exposes no completion event. No sleeps or timeouts were merely widened; no assertions were weakened. The determinism checker now reports 0 active findings. All four changed SPM package test targets compile via swift build --build-tests; Python files pass py_compile. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for rewritten test files The deterministic rewrites (virtual clocks, poll helpers, signal waits) added lines to 7 Swift test files. Bump only those rows. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Fix frozen terminals after split churn (#12) * Fix blank terminal after split operations and add visual tests ## Blank Terminal Fix - Add `needsRefreshAfterWindowChange` flag in GhosttyTerminalView - Force terminal refresh when view is added to window, even if size unchanged - Add `ghostty_surface_refresh()` call in attachToView for same-view reattachment - Add debug logging for surface attachment lifecycle (DEBUG builds only) ## Bonsplit Migration - Add bonsplit as local Swift package (vendor/bonsplit submodule) - Replace custom SplitTree with BonsplitController - Add Panel protocol with TerminalPanel and BrowserPanel implementations - Add SidebarTab as main tab container with BonsplitController - Remove old Splits/ directory (SplitTree, SplitView, TerminalSplitTreeView) ## Visual Screenshot Tests - Add test_visual_screenshots.py for automated visual regression testing - Uses in-app screenshot API (CGWindowListCreateImage) - no screen recording needed - Generates HTML report with before/after comparisons - Tests: splits, browser panels, focus switching, close operations, rapid cycles - Includes annotation fields for easy feedback ## Browser Shortcut (⌘⇧B) - Add keyboard shortcut to open browser panel in current pane - Add openBrowser() method to TabManager - Add shortcut configuration in KeyboardShortcutSettings ## Screenshot Command - Add 'screenshot' command to TerminalController for in-app window capture - Returns OK with screenshot ID and path ## Other - Add tests/visual_output/ and tests/visual_report.html to .gitignore * Add browser title subscription and set tab height to 30px - Subscribe to BrowserPanel.$pageTitle changes to update bonsplit tabs - Update tab titles in real-time as page navigation occurs - Clean up subscriptions when panels are removed - Set bonsplit tab bar and tab height to 30px (in submodule) * Fix socket API regressions in list_surfaces, list_bonsplit_tabs, focus_pane - list_surfaces: Remove [terminal]/[browser] suffix to keep UUID-only format that clients and tests expect for parsing - list_bonsplit_tabs --pane: Properly look up pane by UUID instead of creating a new PaneID (requires bonsplit PaneID.id to be public) - focus_pane: Accept both UUID strings and integer indices as documented * Fix browser panel stability and keyboard shortcuts - Prevent WKWebView focus lifecycle crashes during split/view reshuffles - Match bracket shortcuts via keyCode (Cmd+Shift+[ / ], Cmd+Ctrl+[ / ]) - Support Ghostty config goto_split:* keybinds when WebView is focused - Add focus_webview/is_webview_focused socket commands and regression tests - Rename SidebarTab to Workspace and update docs * Make ctrl+enter keybind test skippable Skip when the Ghostty keybind isn't configured or when osascript can't send keystrokes (no Accessibility permission), so VM runs stay green. * Auto-focus browser omnibar when blank When a browser surface is focused but no URL is loaded yet, focus the address bar instead of the WKWebView. * Stabilize socket surface indexing * Focus browser omnibar escape; add webview keybind UI tests - Escape in omnibar now returns focus to WKWebView\n- Add UI tests for Cmd+Ctrl+H pane navigation with WebKit focused (including Ghostty config)\n- Avoid flaky element screenshots in UpdatePillUITests on the UTM VM * Fix browser drag-to-split blanks and socket parsing * Fix webview-focused shortcuts and stabilize browser splits - Match ctrl/shift shortcuts by keyCode where needed (Ctrl+H, bracket keys) - Load Ghostty goto_split triggers reliably and refresh on config load - Add debug socket helpers: set_shortcut + simulate_shortcut for tests - Convert browser goto_split/keybind tests to socket-based injection (no osascript) - Bump bonsplit for drag-to-split fixes * Fix split layout collapse and harden socket pane APIs * Stabilize OSC 99 notification test timing * Fix terminal focus routing after split reparent * Support simulate_shortcut enter for focus routing test * Stabilize terminal focus routing test * Fix frozen new terminal tabs after many splits * Fix frozen new terminal tabs after splits * Fix terminal freeze on launch/new tabs * Update ghostty submodule * Fix terminal focus/render stalls after split churn * Fix nested split collapsing existing pane * Fix nested split collapse + stabilize new-surface focus * Update bonsplit submodule * Fix SIGINT test flake * Remove bonsplit tab-switch crossfade * Remove PROJECTS.md * Remove bonsplit tab selection animation * Ignore generated test reports * Middle click closes tab * Revert unintended .gitignore change * Fix build after main merge * Revert "Fix build after main merge" This reverts commit 16bf9816d0856b5385d52f886aa5eb50f3c9d9a4. * Revert "Merge remote-tracking branch 'origin/main' into fix/blank-terminal-and-visual-tests" This reverts commit 7c20fb53fd71fea7a19a3673f2dd73e5f0c783c4, reversing changes made to 0aff107d787bc9d8bbc28220090b4ca7af72e040. * Remove tab close fade animation * Use terminal.fill icon * Make terminal tab icon smaller * Match browser globe tab icon size * Bonsplit: tab min width 48 and tighter close button * Bonsplit: smaller tab title font * Show unread notification badge in bonsplit tabs and improve UI polish Sync unread notification state to bonsplit tab badges (blue dot). Improve EmptyPanelView with Terminal/Browser buttons and shortcut hints. Add tooltips to close tab button and search overlay buttons. * Fix reload.sh single-instance safety check on macOS Replace GNU-only `ps -o etimes=` with portable `ps -o etime=` and parse the dd-hh:mm:ss format manually for macOS compatibility. * Centralize keyboard shortcut definitions into Action enum Replace per-shortcut boilerplate with a single Action enum that holds the label, defaults key, and default binding for each shortcut. All call sites now use shortcut(for:). Settings UI is data-driven via ForEach(Action.allCases). Titlebar tooltips update dynamically when shortcuts are changed. Remove duplicate .keyboardShortcut() modifiers from menu items that are already handled by the event monitor. * Fix WKWebView consuming app menu shortcuts and close panel confirmation Add CmuxWebView subclass that routes key equivalents through the main menu before WebKit, so Cmd+N/Cmd+W/tab switching work when a browser pane is focused. Fix Cmd+W close-panel path: bypass Bonsplit delegate gating after the user confirms the running-process dialog by tracking forceCloseTabIds. Add unit tests (CmuxWebViewKeyEquivalentTests) and UI test scaffolding (MenuKeyEquivalentRoutingUITests) with a new cmux-unit Xcode scheme. * Update CLAUDE.md and PROJECTS.md with recent changes CLAUDE.md: enforce --tag for reload commands, add cleanup safety rules. PROJECTS.md: log notification badge, reload.sh fix, Cmd+W fix, WebView key equiv fix, and centralized shortcuts work. * Keep selection index stable on close * Add concepts page documenting terminology hierarchy New docs page explaining Window > Workspace > Pane > Surface > Panel hierarchy with aligned ASCII diagram. Updated tabs.mdx and splits.mdx to use consistent terminology (workspace instead of tab, surface instead of panel) and corrected outdated CLI command references. * Update bonsplit submodule * WIP: improve split close stability and UI regressions * Close terminal panel on child exit; hide terminal dirty dot * Fix split close/focus regressions and stabilize UI tests * Add unread Dock/Cmd+Tab badge with settings toggle * Fix browser-surface shortcuts and Cmd+L browser opening * Snapshot current workspace state before regression fixes * Update bonsplit submodule snapshot * Stabilize split-close regression capture and sidebar resize assertions * Change default Show Notifications shortcut from Cmd+Shift+I to Cmd+I * Fix update check readiness race, enable release update logging, and improve checking spinner * Restore terminal file drop, fix browser omnibar click focus, and add panel workspace ID mutation for surface moves * Add Cmd+digit workspace hints, titlebar shortcut pills, sidebar drag-reorder, and workspace placement settings * Add v2 browser automation API, surface move/reorder commands, and short-handle ref system to TerminalController * Add CLI browser command surface, --id-format flag, and move/reorder commands * Extend test clients with move/reorder APIs, ref-handle support, and increased timeouts * Harden test runner scripts with deterministic builds, retry logic, and robust socket readiness * Stabilize existing test suites with focus-wait helpers, increased timeouts, and API shape updates * Add terminal file drop e2e regression test * Add v2 browser API, CLI ref resolution, and surface move/reorder test suites * Add unit tests for shortcut hints, workspace reorder, drop planner, and update UI test stabilization * Add cmux-debug-windows skill with snapshot script and agent config * Update project docs: mark browser parity and move/reorder phases complete, add parallel agent workflow guidelines * Update bonsplit submodule: re-entrant setPosition guard, tab shortcut hints, and moveTab/reorderTab API * Add browser agent UX improvements: snapshot refs, placement reuse, diagnostics, and skill docs - Upgrade browser.snapshot to emit accessibility tree text with element refs (eN) - Add right-sibling pane reuse policy for browser.open_split placement - Add rich not_found diagnostics with retry logic for selector actions - Support --snapshot-after for post-action verification on mutating commands - Allow browser fill with empty text for clearing inputs - Default CLI --id-format to refs-first (UUIDs opt-in via --id-format uuids|both) - Format legacy new-pane/new-surface output with short surface refs - Add skills/cmuxterm-browser/ and skills/cmuxterm/ end-user skill docs - Add regression tests for placement policy, snapshot refs, diagnostics, and ID defaults * Update bonsplit submodule: keep raster favicons in color when inactive | 7 个月前 | |
Fix frozen terminals after split churn (#12) * Fix blank terminal after split operations and add visual tests ## Blank Terminal Fix - Add `needsRefreshAfterWindowChange` flag in GhosttyTerminalView - Force terminal refresh when view is added to window, even if size unchanged - Add `ghostty_surface_refresh()` call in attachToView for same-view reattachment - Add debug logging for surface attachment lifecycle (DEBUG builds only) ## Bonsplit Migration - Add bonsplit as local Swift package (vendor/bonsplit submodule) - Replace custom SplitTree with BonsplitController - Add Panel protocol with TerminalPanel and BrowserPanel implementations - Add SidebarTab as main tab container with BonsplitController - Remove old Splits/ directory (SplitTree, SplitView, TerminalSplitTreeView) ## Visual Screenshot Tests - Add test_visual_screenshots.py for automated visual regression testing - Uses in-app screenshot API (CGWindowListCreateImage) - no screen recording needed - Generates HTML report with before/after comparisons - Tests: splits, browser panels, focus switching, close operations, rapid cycles - Includes annotation fields for easy feedback ## Browser Shortcut (⌘⇧B) - Add keyboard shortcut to open browser panel in current pane - Add openBrowser() method to TabManager - Add shortcut configuration in KeyboardShortcutSettings ## Screenshot Command - Add 'screenshot' command to TerminalController for in-app window capture - Returns OK with screenshot ID and path ## Other - Add tests/visual_output/ and tests/visual_report.html to .gitignore * Add browser title subscription and set tab height to 30px - Subscribe to BrowserPanel.$pageTitle changes to update bonsplit tabs - Update tab titles in real-time as page navigation occurs - Clean up subscriptions when panels are removed - Set bonsplit tab bar and tab height to 30px (in submodule) * Fix socket API regressions in list_surfaces, list_bonsplit_tabs, focus_pane - list_surfaces: Remove [terminal]/[browser] suffix to keep UUID-only format that clients and tests expect for parsing - list_bonsplit_tabs --pane: Properly look up pane by UUID instead of creating a new PaneID (requires bonsplit PaneID.id to be public) - focus_pane: Accept both UUID strings and integer indices as documented * Fix browser panel stability and keyboard shortcuts - Prevent WKWebView focus lifecycle crashes during split/view reshuffles - Match bracket shortcuts via keyCode (Cmd+Shift+[ / ], Cmd+Ctrl+[ / ]) - Support Ghostty config goto_split:* keybinds when WebView is focused - Add focus_webview/is_webview_focused socket commands and regression tests - Rename SidebarTab to Workspace and update docs * Make ctrl+enter keybind test skippable Skip when the Ghostty keybind isn't configured or when osascript can't send keystrokes (no Accessibility permission), so VM runs stay green. * Auto-focus browser omnibar when blank When a browser surface is focused but no URL is loaded yet, focus the address bar instead of the WKWebView. * Stabilize socket surface indexing * Focus browser omnibar escape; add webview keybind UI tests - Escape in omnibar now returns focus to WKWebView\n- Add UI tests for Cmd+Ctrl+H pane navigation with WebKit focused (including Ghostty config)\n- Avoid flaky element screenshots in UpdatePillUITests on the UTM VM * Fix browser drag-to-split blanks and socket parsing * Fix webview-focused shortcuts and stabilize browser splits - Match ctrl/shift shortcuts by keyCode where needed (Ctrl+H, bracket keys) - Load Ghostty goto_split triggers reliably and refresh on config load - Add debug socket helpers: set_shortcut + simulate_shortcut for tests - Convert browser goto_split/keybind tests to socket-based injection (no osascript) - Bump bonsplit for drag-to-split fixes * Fix split layout collapse and harden socket pane APIs * Stabilize OSC 99 notification test timing * Fix terminal focus routing after split reparent * Support simulate_shortcut enter for focus routing test * Stabilize terminal focus routing test * Fix frozen new terminal tabs after many splits * Fix frozen new terminal tabs after splits * Fix terminal freeze on launch/new tabs * Update ghostty submodule * Fix terminal focus/render stalls after split churn * Fix nested split collapsing existing pane * Fix nested split collapse + stabilize new-surface focus * Update bonsplit submodule * Fix SIGINT test flake * Remove bonsplit tab-switch crossfade * Remove PROJECTS.md * Remove bonsplit tab selection animation * Ignore generated test reports * Middle click closes tab * Revert unintended .gitignore change * Fix build after main merge * Revert "Fix build after main merge" This reverts commit 16bf9816d0856b5385d52f886aa5eb50f3c9d9a4. * Revert "Merge remote-tracking branch 'origin/main' into fix/blank-terminal-and-visual-tests" This reverts commit 7c20fb53fd71fea7a19a3673f2dd73e5f0c783c4, reversing changes made to 0aff107d787bc9d8bbc28220090b4ca7af72e040. * Remove tab close fade animation * Use terminal.fill icon * Make terminal tab icon smaller * Match browser globe tab icon size * Bonsplit: tab min width 48 and tighter close button * Bonsplit: smaller tab title font * Show unread notification badge in bonsplit tabs and improve UI polish Sync unread notification state to bonsplit tab badges (blue dot). Improve EmptyPanelView with Terminal/Browser buttons and shortcut hints. Add tooltips to close tab button and search overlay buttons. * Fix reload.sh single-instance safety check on macOS Replace GNU-only `ps -o etimes=` with portable `ps -o etime=` and parse the dd-hh:mm:ss format manually for macOS compatibility. * Centralize keyboard shortcut definitions into Action enum Replace per-shortcut boilerplate with a single Action enum that holds the label, defaults key, and default binding for each shortcut. All call sites now use shortcut(for:). Settings UI is data-driven via ForEach(Action.allCases). Titlebar tooltips update dynamically when shortcuts are changed. Remove duplicate .keyboardShortcut() modifiers from menu items that are already handled by the event monitor. * Fix WKWebView consuming app menu shortcuts and close panel confirmation Add CmuxWebView subclass that routes key equivalents through the main menu before WebKit, so Cmd+N/Cmd+W/tab switching work when a browser pane is focused. Fix Cmd+W close-panel path: bypass Bonsplit delegate gating after the user confirms the running-process dialog by tracking forceCloseTabIds. Add unit tests (CmuxWebViewKeyEquivalentTests) and UI test scaffolding (MenuKeyEquivalentRoutingUITests) with a new cmux-unit Xcode scheme. * Update CLAUDE.md and PROJECTS.md with recent changes CLAUDE.md: enforce --tag for reload commands, add cleanup safety rules. PROJECTS.md: log notification badge, reload.sh fix, Cmd+W fix, WebView key equiv fix, and centralized shortcuts work. * Keep selection index stable on close * Add concepts page documenting terminology hierarchy New docs page explaining Window > Workspace > Pane > Surface > Panel hierarchy with aligned ASCII diagram. Updated tabs.mdx and splits.mdx to use consistent terminology (workspace instead of tab, surface instead of panel) and corrected outdated CLI command references. * Update bonsplit submodule * WIP: improve split close stability and UI regressions * Close terminal panel on child exit; hide terminal dirty dot * Fix split close/focus regressions and stabilize UI tests * Add unread Dock/Cmd+Tab badge with settings toggle * Fix browser-surface shortcuts and Cmd+L browser opening * Snapshot current workspace state before regression fixes * Update bonsplit submodule snapshot * Stabilize split-close regression capture and sidebar resize assertions * Change default Show Notifications shortcut from Cmd+Shift+I to Cmd+I * Fix update check readiness race, enable release update logging, and improve checking spinner * Restore terminal file drop, fix browser omnibar click focus, and add panel workspace ID mutation for surface moves * Add Cmd+digit workspace hints, titlebar shortcut pills, sidebar drag-reorder, and workspace placement settings * Add v2 browser automation API, surface move/reorder commands, and short-handle ref system to TerminalController * Add CLI browser command surface, --id-format flag, and move/reorder commands * Extend test clients with move/reorder APIs, ref-handle support, and increased timeouts * Harden test runner scripts with deterministic builds, retry logic, and robust socket readiness * Stabilize existing test suites with focus-wait helpers, increased timeouts, and API shape updates * Add terminal file drop e2e regression test * Add v2 browser API, CLI ref resolution, and surface move/reorder test suites * Add unit tests for shortcut hints, workspace reorder, drop planner, and update UI test stabilization * Add cmux-debug-windows skill with snapshot script and agent config * Update project docs: mark browser parity and move/reorder phases complete, add parallel agent workflow guidelines * Update bonsplit submodule: re-entrant setPosition guard, tab shortcut hints, and moveTab/reorderTab API * Add browser agent UX improvements: snapshot refs, placement reuse, diagnostics, and skill docs - Upgrade browser.snapshot to emit accessibility tree text with element refs (eN) - Add right-sibling pane reuse policy for browser.open_split placement - Add rich not_found diagnostics with retry logic for selector actions - Support --snapshot-after for post-action verification on mutating commands - Allow browser fill with empty text for clearing inputs - Default CLI --id-format to refs-first (UUIDs opt-in via --id-format uuids|both) - Format legacy new-pane/new-surface output with short surface refs - Add skills/cmuxterm-browser/ and skills/cmuxterm/ end-user skill docs - Add regression tests for placement policy, snapshot refs, diagnostics, and ID defaults * Update bonsplit submodule: keep raster favicons in color when inactive | 7 个月前 | |
Revert "Fix notification unread persistence when workspaces regain focus (#971)" (#992) This reverts commit 5f43a3fc32045cf63cd4bab97befe8f0756c6c16. | 6 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Saved split layouts: capture, store, and reopen named workspace layouts (#7414) * Add saved split layouts: capture, store, and reopen named workspace layouts Users can save the current workspace's split layout (pane tree, split directions, divider ratios, and per-pane surface types: terminal with cwd, browser with URL, project) as a named template, then open new workspaces from it later. Layouts persist in ~/.config/cmux/layouts.json using the existing CmuxLayoutNode/CmuxWorkspaceDefinition schema, so a saved layout is copy-pasteable into cmux.json workspace commands and compatible with workspace.create --layout. Three entrypoints share one action path (TabManager.openWorkspace( fromSavedLayout:)): command palette ("Save Layout as Template…" plus one dynamic "New Workspace from Layout: <name>" entry per saved layout), a new cmux layout CLI namespace (save/list/get/open/delete), and layout.* debug socket verbs. Capture is the inverse of applyCustomLayout: a walk of the live bonsplit tree emitting the declarative schema, with unsupported panel kinds preserved as placeholder terminals. Implements the core of https://github.com/manaflow-ai/cmux/issues/1055; related: https://github.com/manaflow-ai/cmux/issues/3448. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Surface saved layouts: ⌘⇧S save shortcut + plus-button layout submenu Adds a saveLayoutTemplate keyboard shortcut (default cmd+shift+S, free slot; cmd+S stays with saveFilePreview) wired per shortcut policy: both ShortcutAction enums, Settings-visible, rebindable, automatic shortcuts.bindings.saveLayoutTemplate support in cmux.json, and a docs row in web/data/cmux-shortcuts.ts (EN+JA). Dispatch posts a window-scoped savedLayoutSaveRequested notification; the focused window's ContentView presents the existing save-name dialog, so the shortcut, palette command, and future callers share one path. The palette entry now shows the current binding as its shortcut hint. The tab-bar plus button's right-click menu gains a "New Workspace from Layout" submenu (one item per saved layout, hidden when none exist), modeled on the move-surface submenu pattern; items re-fetch the layout by name and call the shared TabManager.openWorkspace(fromSavedLayout:). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address review findings and file-length budget for saved layouts Review fixes: capture now throws on an unrecognized bonsplit orientation string instead of silently defaulting to horizontal, and counts unmapped tabs as unsupported placeholders; palette open handlers re-resolve the layout by name at invocation (no stale snapshots); the save command dismisses the palette before presenting its name prompt; the saved-layout store's mtime cache is shared across instances (keyed by file path); layout CLI subcommands reject unknown flags before consuming the name; layout.save/layout.list nil-context fallbacks return unavailable-style errors instead of misleading not-found/empty results; user-facing alerts no longer surface raw decoder or system error text (CLI/socket keep full detail); the e2e browser fixture uses about:blank instead of a live URL. Budget gate: cmux layout help text moved into CLI/cmux_layout.swift, the ControlLayoutContext test defaults into ControlLayoutContextTestStubs .swift, the shortcut alignment test into KeyboardShortcutSavedLayoutTemplateTests.swift, and the shortcut matcher body into AppDelegate+SavedLayoutMenu.swift, restoring those files to their existing budgets. Only Sources/AppDelegate.swift (+2) and Sources/KeyboardShortcutSettings.swift (+4) budgets grew, covering the compiler-enforced shortcut registry entries. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Widen two AppDelegate shortcut helpers to internal for the extension file Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Beep when saved-layout menu open returns nil, matching sibling failure paths Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Capture project pane paths so saved layouts can restore project surfaces The canonical review caught that .project surfaces were serialized with no cwd/url while the apply side rebuilds them from url ?? cwd, so any saved layout with a project pane reopened as an empty placeholder. Capture now stores ProjectPanel.projectURL.path in cwd (mirroring session persistence) and falls back to a counted placeholder terminal when the path is missing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Fix corrupt-file cache so deleting layouts.json recovers to empty state The canonical review caught that load() checked the corrupt short-circuit before file existence: a nil-mtime cache entry from a pre-corruption read matched the nil mtime after the user deleted the bad file, wedging the store on corruptFile until restart. Nonexistence now resets the corrupt state first, with a regression test covering the delete-to-recover path. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Make saved layouts relocatable, move default binding off cmd+shift+S, reject unresolvable workspace refs Review findings from the canonical helper: terminal cwds and project paths under the workspace root now capture as relative paths so layout open --cwd re-roots them (apply-side project paths resolve via resolveCwd like terminals; paths outside the root stay absolute deliberately); the default saveLayoutTemplate binding moves to ctrl+cmd+S because cmd+shift+S is a common save-as shortcut inside browser panes (both enums, docs row, and alignment test updated); layout.save now errors with not_found when a workspace selector is present but unresolvable instead of silently capturing the focused workspace. Verified: socket package build + 195 tests, tagged build, e2e suite, and a live relocation proof (nested terminal saved from /tmp/laytest/a reopened rooted under --cwd /tmp/laytest/b). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Scope layout.save workspace lookup to an explicit window selector An explicit window_id must win over a workspace selector per the control routing precedence; previously a workspace in another window escaped the requested window's scope and could overwrite a saved layout with the wrong workspace contents. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Document the v1 tab-selection scope cut at the capture site Per-pane selected tabs are not representable in the declarative layout schema; extending it is tracked in https://github.com/manaflow-ai/cmux/issues/7444. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> | 2 个月前 | |
Rename GhosttyTabs project to cmux (#4205) * Rename GhosttyTabs project to cmux * Use tagged reload in debug windows skill * Update command palette test project path * Fix debug windows skill list numbering --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 4 个月前 | |
iOS: show workspaces from all Mac windows, not just the front one (#5565) v2MobileWorkspaceList resolved a single window's TabManager (key window via v2ResolveTabManager) and returned only that window's tabs, so the iOS app saw workspaces from one window. When no specific workspace_id/window_id is named, iterate all registered main windows (listMainWindowSummaries + tabManagerFor) and flatten every window's tabs into one list, deduped. is_selected is computed from the frontmost window only. Scoped requests (workspace_id/window_id named) keep today's single-window behavior, so create/input/select routing through v2ResolveTabManager is unaffected. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> | 3 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Flaky-test follow-up: review fixes + second-pass team sweep (#6452) * Address review: fail-loud guards + scoped soft-skips in de-flaked tests - test_browser_goto_split.py: _wait_url_loaded now raises on timeout instead of silently returning (false pass), and the local HTTP server is managed via a "with" block so the thread/socket can't leak on setup exceptions. - test_surface_move_reorder_api.py: assert the cleanup re-select of ws0 actually converges instead of ignoring the _wait() result (state leak across runs). - test_homebrew_sha.sh: only soft-skip transient transport failures (000/408/429/5xx); fail hard on deterministic client errors like 404 (missing release asset). - test_visual_typing_char_by_char.py: validate the typed glyph against the post-prompt region of the last line, so a "cmux" already in the prompt/path can't satisfy the check; robust to trailing zsh-autosuggestion glyphs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 5d4c282535a497ad263128b21f8805a34556f55d) * Address review round 2: critical walrus bug + 3 fail-loud/correctness guards - test_pane_break_swap_preserve_focus.py: fix NameError -- the lambda referenced `p` in the ternary condition before the walrus assigned it. Bind the walrus inside the condition so the predicate reads panes once and returns them. - perf-activation-session.py: a missing measurement (actual is None) is now always a blocking failure, not advisory -- absent data is a benchmark-contract violation, distinct from load-sensitive over-budget timing. - perf-activation-session.py: copy the real-scrollback measurement before storing it under snapshot_with_scrollback, so best_of_snapshot_timing's in-place mutation no longer clobbers the raw snapshot_with_real_scrollback capture. - NotificationAndMenuBarTests.swift: on the stall-timeout path, fail fast via a lock-guarded result box instead of awaiting evaluationTask.value, which could hang until the suite timeout when the hook ignores cancellation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7e75ad467bc95cef0436236d94183f43b9aaffb2) * Add reports.md: consolidated flaky-test discovery (in-session workflow) Output of a parallel dynamic workflow (20 agents over the 419 signal-bearing test files of 2080) verifying residual flakiness on this already-de-flaked branch. 101 candidate findings across 86 files, dominated by wall-clock timing asserts, sleep-as-sync, and async races -- concentrated in Packages/ unit tests the first 44-file sweep did not cover. Each entry is a candidate; the fix phase adversarially verifies every one before changing code (repo rules allow deterministic test sleeps, so those are left alone). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit e1952d01faee805298bdff215af1d0897f63abcf) * Fix 37 verified residual flaky tests (coordinated area-team workflow) Second-pass de-flake driven by the in-session discovery + team-fix workflows (see reports.md). Adversarial per-area verification confirmed genuine flakiness before any edit; 9 candidates were refuted (deterministic-allowed sleeps, factually-wrong premises, fixes needing risky production changes) and left. Highlights: - Packages/ unit tests (missed by the first 44-file sweep): AuthCoordinator and HostBrowserSignIn now drive timeouts off an injected ManualTestClock instead of real ContinuousClock; ChatConversationStore poller results are asserted with #expect instead of discarded (fail-loud); RemoteProxyBroker / RemoteCLIRelay / CommandPaletteSearchEngine timing/race tightened. - Integration tests: per-PID/UUID socket+port paths (terminal_focus_routing, ctrl_socket, ssh_remote_* disjoint port ranges, CMUXCLIErrorOutput socket) to kill cross-run collisions; order-dependent session-restore checks now scan all workspaces instead of trusting seeding order; latent NameErrors fixed (pane_resize missing `import time` / `must`); fail-open waits made loud (new_tab tmp-write now raises); surface-targeted send to avoid focus races (tab_dragging); fd/temp-file leaks closed. - App-target Swift: removed/loosened flaky wall-clock asserts while keeping the behavior assert (TerminalAndGhostty paste, ShellStartupMatrix budget), widened burst spacing (GhosttyNotificationDispatcher), and routed UITest socket/condition waits through the shared waitForControlSocketReady helper. All Python pass py_compile; all 4 touched packages pass swift build --build-tests; app-target Swift is CI-compiled. No test run against a socket. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7a374ac84c2c96d175df7656b9fd6657a788ea68) * Fix MobileCoreRPC cancelled-while-queued race deterministically (DEBUG-only seam) The remaining discovery finding: MobileCoreRPCClientTests spun a fixed 100 `Task.yield()`s to "wait" for a queued request to reach the session writer gate before cancelling it. Under scheduler load the queued task may not have reached the gate yet, so cancellation fires before `session.send` registers it and the cancelled-while-queued invariant is never exercised (false pass). That gate state (`queuedRequestIDs`) is private to the production `MobileCoreRPCSession` actor, so there is no deterministic test-only signal. Fix adds a `#if DEBUG`, read-only `debugQueuedRequestCount()` to the session and a thin client wrapper, placed in the file's existing `#if DEBUG` test-support extension (alongside `debugWithRequestTimeout`). The test now polls that real signal until the request is registered at the gate, then cancels. The accessor is `#if DEBUG` and read-only: it is compiled out of Release builds and changes no shipping-build behavior, so it needs no dogfood. `swift build --build-tests` (debug) compiles and links the package + test target cleanly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit c56f53ece49ae1c98072bb24996c55e6d15b4fa6) * Refresh Swift file-length budget for the de-flaked test files The fail-loud guards and ManualTestClock/RunLoop-poll rewrites grew five test files past their budgets (NotificationAndMenuBar +31, CommandPaletteSearchEngine +21, MultiWindowNotifications +9, BrowserPaneNavigationKeybind +3, AgentHibernation +1) and shrank two (TerminalAndGhostty -7 after dropping a flaky wall-clock assert, BrowserFixtureInteraction -6 after routing through the shared readiness helper). Budget updated for exactly those seven files. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: drop reintroduced wall-clock asserts / fixed sleeps CodeRabbit/Greptile caught spots where the second-pass fixes leaned on the very patterns this PR is removing. All test-only: - ShellStartupMatrixTests: drop the redundant `duration < 5.0` ceiling; the bootstrap already runs under a 5s `runProcess` timeout, so status==0 + !timedOut is the causal completion signal (no wall-clock assert). - MultiWindowNotificationsUITests: assert no-foreground at the causal point (right after `waitForCommandCompletionWhileBackgrounded`) instead of polling a fixed 2s window. - test_tab_dragging.py: remove two `time.sleep(1.5)` waits; the file-content poll loop right below is the readiness signal. - test_terminal_focus_routing.py: `tempfile.mktemp` -> `mkstemp` (atomic, no TOCTOU; clears Ruff S306). - test_session_restore_stress_kill_cycles.py: replace the fixed 0.1s settle with a deadline-bounded poll on the real is-selected signal. - RemoteCLIRelayServerTests: capture errno before close() so thrown bind/listen diagnostics report the real failure code. py_compile + swift build --build-tests (CmuxRemoteWorkspace) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Make CommandPalette search benchmarks advisory instead of wall-clock-gated Greptile/CodeRabbit correctly flagged that loosening these optimized-vs-reference ratio asserts (e.g. 0.80 -> 1.25) no longer validates anything: a band tight enough to prove the optimized path is faster flakes on shared CI, and a band wide enough not to flake passes even when the optimized path regresses. The engine exposes no preparation/work counter, so there is no causal (non-wall- clock) signal to assert on here. Per the repo test-time policy (no wall-clock latency asserts on shared CI) and matching the activation-session perf gate's advisory-timing approach, drop the flaky `#expect` ratio/dropped-frame assertions across all four benchmarks and keep the `BENCH ...` diagnostic prints for trend tracking. Each test still exercises both code paths end to end; real activation latency/frame-budget regressions are gated by the dedicated activation-session job. swift build --build-tests (CmuxCommandPalette) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * MobileCoreRPC: observe queue gate via @testable, not a production debug hook Per review (Aziz): test/debug extensions should not live in production source. Remove the `#if DEBUG debugQueuedRequestCount()` accessor from MobileCoreRPCClient/Session and instead widen `session` and `queuedRequestIDs` from `private` to `internal` so the cancellation test reads the writer-gate state directly through its existing `@testable import CmuxMobileRPC`. All test scaffolding now lives in the test target; production source carries only the two access-level changes (still module-internal, no shipping behavior change). swift build --build-tests (CmuxMobileRPC) clean; no debug funcs remain in Sources/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Revert "Fix notification unread persistence when workspaces regain focus (#971)" (#992) This reverts commit 5f43a3fc32045cf63cd4bab97befe8f0756c6c16. | 6 个月前 | |
Flaky-test follow-up: review fixes + second-pass team sweep (#6452) * Address review: fail-loud guards + scoped soft-skips in de-flaked tests - test_browser_goto_split.py: _wait_url_loaded now raises on timeout instead of silently returning (false pass), and the local HTTP server is managed via a "with" block so the thread/socket can't leak on setup exceptions. - test_surface_move_reorder_api.py: assert the cleanup re-select of ws0 actually converges instead of ignoring the _wait() result (state leak across runs). - test_homebrew_sha.sh: only soft-skip transient transport failures (000/408/429/5xx); fail hard on deterministic client errors like 404 (missing release asset). - test_visual_typing_char_by_char.py: validate the typed glyph against the post-prompt region of the last line, so a "cmux" already in the prompt/path can't satisfy the check; robust to trailing zsh-autosuggestion glyphs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 5d4c282535a497ad263128b21f8805a34556f55d) * Address review round 2: critical walrus bug + 3 fail-loud/correctness guards - test_pane_break_swap_preserve_focus.py: fix NameError -- the lambda referenced `p` in the ternary condition before the walrus assigned it. Bind the walrus inside the condition so the predicate reads panes once and returns them. - perf-activation-session.py: a missing measurement (actual is None) is now always a blocking failure, not advisory -- absent data is a benchmark-contract violation, distinct from load-sensitive over-budget timing. - perf-activation-session.py: copy the real-scrollback measurement before storing it under snapshot_with_scrollback, so best_of_snapshot_timing's in-place mutation no longer clobbers the raw snapshot_with_real_scrollback capture. - NotificationAndMenuBarTests.swift: on the stall-timeout path, fail fast via a lock-guarded result box instead of awaiting evaluationTask.value, which could hang until the suite timeout when the hook ignores cancellation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7e75ad467bc95cef0436236d94183f43b9aaffb2) * Add reports.md: consolidated flaky-test discovery (in-session workflow) Output of a parallel dynamic workflow (20 agents over the 419 signal-bearing test files of 2080) verifying residual flakiness on this already-de-flaked branch. 101 candidate findings across 86 files, dominated by wall-clock timing asserts, sleep-as-sync, and async races -- concentrated in Packages/ unit tests the first 44-file sweep did not cover. Each entry is a candidate; the fix phase adversarially verifies every one before changing code (repo rules allow deterministic test sleeps, so those are left alone). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit e1952d01faee805298bdff215af1d0897f63abcf) * Fix 37 verified residual flaky tests (coordinated area-team workflow) Second-pass de-flake driven by the in-session discovery + team-fix workflows (see reports.md). Adversarial per-area verification confirmed genuine flakiness before any edit; 9 candidates were refuted (deterministic-allowed sleeps, factually-wrong premises, fixes needing risky production changes) and left. Highlights: - Packages/ unit tests (missed by the first 44-file sweep): AuthCoordinator and HostBrowserSignIn now drive timeouts off an injected ManualTestClock instead of real ContinuousClock; ChatConversationStore poller results are asserted with #expect instead of discarded (fail-loud); RemoteProxyBroker / RemoteCLIRelay / CommandPaletteSearchEngine timing/race tightened. - Integration tests: per-PID/UUID socket+port paths (terminal_focus_routing, ctrl_socket, ssh_remote_* disjoint port ranges, CMUXCLIErrorOutput socket) to kill cross-run collisions; order-dependent session-restore checks now scan all workspaces instead of trusting seeding order; latent NameErrors fixed (pane_resize missing `import time` / `must`); fail-open waits made loud (new_tab tmp-write now raises); surface-targeted send to avoid focus races (tab_dragging); fd/temp-file leaks closed. - App-target Swift: removed/loosened flaky wall-clock asserts while keeping the behavior assert (TerminalAndGhostty paste, ShellStartupMatrix budget), widened burst spacing (GhosttyNotificationDispatcher), and routed UITest socket/condition waits through the shared waitForControlSocketReady helper. All Python pass py_compile; all 4 touched packages pass swift build --build-tests; app-target Swift is CI-compiled. No test run against a socket. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7a374ac84c2c96d175df7656b9fd6657a788ea68) * Fix MobileCoreRPC cancelled-while-queued race deterministically (DEBUG-only seam) The remaining discovery finding: MobileCoreRPCClientTests spun a fixed 100 `Task.yield()`s to "wait" for a queued request to reach the session writer gate before cancelling it. Under scheduler load the queued task may not have reached the gate yet, so cancellation fires before `session.send` registers it and the cancelled-while-queued invariant is never exercised (false pass). That gate state (`queuedRequestIDs`) is private to the production `MobileCoreRPCSession` actor, so there is no deterministic test-only signal. Fix adds a `#if DEBUG`, read-only `debugQueuedRequestCount()` to the session and a thin client wrapper, placed in the file's existing `#if DEBUG` test-support extension (alongside `debugWithRequestTimeout`). The test now polls that real signal until the request is registered at the gate, then cancels. The accessor is `#if DEBUG` and read-only: it is compiled out of Release builds and changes no shipping-build behavior, so it needs no dogfood. `swift build --build-tests` (debug) compiles and links the package + test target cleanly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit c56f53ece49ae1c98072bb24996c55e6d15b4fa6) * Refresh Swift file-length budget for the de-flaked test files The fail-loud guards and ManualTestClock/RunLoop-poll rewrites grew five test files past their budgets (NotificationAndMenuBar +31, CommandPaletteSearchEngine +21, MultiWindowNotifications +9, BrowserPaneNavigationKeybind +3, AgentHibernation +1) and shrank two (TerminalAndGhostty -7 after dropping a flaky wall-clock assert, BrowserFixtureInteraction -6 after routing through the shared readiness helper). Budget updated for exactly those seven files. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: drop reintroduced wall-clock asserts / fixed sleeps CodeRabbit/Greptile caught spots where the second-pass fixes leaned on the very patterns this PR is removing. All test-only: - ShellStartupMatrixTests: drop the redundant `duration < 5.0` ceiling; the bootstrap already runs under a 5s `runProcess` timeout, so status==0 + !timedOut is the causal completion signal (no wall-clock assert). - MultiWindowNotificationsUITests: assert no-foreground at the causal point (right after `waitForCommandCompletionWhileBackgrounded`) instead of polling a fixed 2s window. - test_tab_dragging.py: remove two `time.sleep(1.5)` waits; the file-content poll loop right below is the readiness signal. - test_terminal_focus_routing.py: `tempfile.mktemp` -> `mkstemp` (atomic, no TOCTOU; clears Ruff S306). - test_session_restore_stress_kill_cycles.py: replace the fixed 0.1s settle with a deadline-bounded poll on the real is-selected signal. - RemoteCLIRelayServerTests: capture errno before close() so thrown bind/listen diagnostics report the real failure code. py_compile + swift build --build-tests (CmuxRemoteWorkspace) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Make CommandPalette search benchmarks advisory instead of wall-clock-gated Greptile/CodeRabbit correctly flagged that loosening these optimized-vs-reference ratio asserts (e.g. 0.80 -> 1.25) no longer validates anything: a band tight enough to prove the optimized path is faster flakes on shared CI, and a band wide enough not to flake passes even when the optimized path regresses. The engine exposes no preparation/work counter, so there is no causal (non-wall- clock) signal to assert on here. Per the repo test-time policy (no wall-clock latency asserts on shared CI) and matching the activation-session perf gate's advisory-timing approach, drop the flaky `#expect` ratio/dropped-frame assertions across all four benchmarks and keep the `BENCH ...` diagnostic prints for trend tracking. Each test still exercises both code paths end to end; real activation latency/frame-budget regressions are gated by the dedicated activation-session job. swift build --build-tests (CmuxCommandPalette) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * MobileCoreRPC: observe queue gate via @testable, not a production debug hook Per review (Aziz): test/debug extensions should not live in production source. Remove the `#if DEBUG debugQueuedRequestCount()` accessor from MobileCoreRPCClient/Session and instead widen `session` and `queuedRequestIDs` from `private` to `internal` so the cancellation test reads the writer-gate state directly through its existing `@testable import CmuxMobileRPC`. All test scaffolding now lives in the test target; production source carries only the two access-level changes (still module-internal, no shipping behavior change). swift build --build-tests (CmuxMobileRPC) clean; no debug funcs remain in Sources/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Flaky-test follow-up: review fixes + second-pass team sweep (#6452) * Address review: fail-loud guards + scoped soft-skips in de-flaked tests - test_browser_goto_split.py: _wait_url_loaded now raises on timeout instead of silently returning (false pass), and the local HTTP server is managed via a "with" block so the thread/socket can't leak on setup exceptions. - test_surface_move_reorder_api.py: assert the cleanup re-select of ws0 actually converges instead of ignoring the _wait() result (state leak across runs). - test_homebrew_sha.sh: only soft-skip transient transport failures (000/408/429/5xx); fail hard on deterministic client errors like 404 (missing release asset). - test_visual_typing_char_by_char.py: validate the typed glyph against the post-prompt region of the last line, so a "cmux" already in the prompt/path can't satisfy the check; robust to trailing zsh-autosuggestion glyphs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 5d4c282535a497ad263128b21f8805a34556f55d) * Address review round 2: critical walrus bug + 3 fail-loud/correctness guards - test_pane_break_swap_preserve_focus.py: fix NameError -- the lambda referenced `p` in the ternary condition before the walrus assigned it. Bind the walrus inside the condition so the predicate reads panes once and returns them. - perf-activation-session.py: a missing measurement (actual is None) is now always a blocking failure, not advisory -- absent data is a benchmark-contract violation, distinct from load-sensitive over-budget timing. - perf-activation-session.py: copy the real-scrollback measurement before storing it under snapshot_with_scrollback, so best_of_snapshot_timing's in-place mutation no longer clobbers the raw snapshot_with_real_scrollback capture. - NotificationAndMenuBarTests.swift: on the stall-timeout path, fail fast via a lock-guarded result box instead of awaiting evaluationTask.value, which could hang until the suite timeout when the hook ignores cancellation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7e75ad467bc95cef0436236d94183f43b9aaffb2) * Add reports.md: consolidated flaky-test discovery (in-session workflow) Output of a parallel dynamic workflow (20 agents over the 419 signal-bearing test files of 2080) verifying residual flakiness on this already-de-flaked branch. 101 candidate findings across 86 files, dominated by wall-clock timing asserts, sleep-as-sync, and async races -- concentrated in Packages/ unit tests the first 44-file sweep did not cover. Each entry is a candidate; the fix phase adversarially verifies every one before changing code (repo rules allow deterministic test sleeps, so those are left alone). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit e1952d01faee805298bdff215af1d0897f63abcf) * Fix 37 verified residual flaky tests (coordinated area-team workflow) Second-pass de-flake driven by the in-session discovery + team-fix workflows (see reports.md). Adversarial per-area verification confirmed genuine flakiness before any edit; 9 candidates were refuted (deterministic-allowed sleeps, factually-wrong premises, fixes needing risky production changes) and left. Highlights: - Packages/ unit tests (missed by the first 44-file sweep): AuthCoordinator and HostBrowserSignIn now drive timeouts off an injected ManualTestClock instead of real ContinuousClock; ChatConversationStore poller results are asserted with #expect instead of discarded (fail-loud); RemoteProxyBroker / RemoteCLIRelay / CommandPaletteSearchEngine timing/race tightened. - Integration tests: per-PID/UUID socket+port paths (terminal_focus_routing, ctrl_socket, ssh_remote_* disjoint port ranges, CMUXCLIErrorOutput socket) to kill cross-run collisions; order-dependent session-restore checks now scan all workspaces instead of trusting seeding order; latent NameErrors fixed (pane_resize missing `import time` / `must`); fail-open waits made loud (new_tab tmp-write now raises); surface-targeted send to avoid focus races (tab_dragging); fd/temp-file leaks closed. - App-target Swift: removed/loosened flaky wall-clock asserts while keeping the behavior assert (TerminalAndGhostty paste, ShellStartupMatrix budget), widened burst spacing (GhosttyNotificationDispatcher), and routed UITest socket/condition waits through the shared waitForControlSocketReady helper. All Python pass py_compile; all 4 touched packages pass swift build --build-tests; app-target Swift is CI-compiled. No test run against a socket. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7a374ac84c2c96d175df7656b9fd6657a788ea68) * Fix MobileCoreRPC cancelled-while-queued race deterministically (DEBUG-only seam) The remaining discovery finding: MobileCoreRPCClientTests spun a fixed 100 `Task.yield()`s to "wait" for a queued request to reach the session writer gate before cancelling it. Under scheduler load the queued task may not have reached the gate yet, so cancellation fires before `session.send` registers it and the cancelled-while-queued invariant is never exercised (false pass). That gate state (`queuedRequestIDs`) is private to the production `MobileCoreRPCSession` actor, so there is no deterministic test-only signal. Fix adds a `#if DEBUG`, read-only `debugQueuedRequestCount()` to the session and a thin client wrapper, placed in the file's existing `#if DEBUG` test-support extension (alongside `debugWithRequestTimeout`). The test now polls that real signal until the request is registered at the gate, then cancels. The accessor is `#if DEBUG` and read-only: it is compiled out of Release builds and changes no shipping-build behavior, so it needs no dogfood. `swift build --build-tests` (debug) compiles and links the package + test target cleanly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit c56f53ece49ae1c98072bb24996c55e6d15b4fa6) * Refresh Swift file-length budget for the de-flaked test files The fail-loud guards and ManualTestClock/RunLoop-poll rewrites grew five test files past their budgets (NotificationAndMenuBar +31, CommandPaletteSearchEngine +21, MultiWindowNotifications +9, BrowserPaneNavigationKeybind +3, AgentHibernation +1) and shrank two (TerminalAndGhostty -7 after dropping a flaky wall-clock assert, BrowserFixtureInteraction -6 after routing through the shared readiness helper). Budget updated for exactly those seven files. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: drop reintroduced wall-clock asserts / fixed sleeps CodeRabbit/Greptile caught spots where the second-pass fixes leaned on the very patterns this PR is removing. All test-only: - ShellStartupMatrixTests: drop the redundant `duration < 5.0` ceiling; the bootstrap already runs under a 5s `runProcess` timeout, so status==0 + !timedOut is the causal completion signal (no wall-clock assert). - MultiWindowNotificationsUITests: assert no-foreground at the causal point (right after `waitForCommandCompletionWhileBackgrounded`) instead of polling a fixed 2s window. - test_tab_dragging.py: remove two `time.sleep(1.5)` waits; the file-content poll loop right below is the readiness signal. - test_terminal_focus_routing.py: `tempfile.mktemp` -> `mkstemp` (atomic, no TOCTOU; clears Ruff S306). - test_session_restore_stress_kill_cycles.py: replace the fixed 0.1s settle with a deadline-bounded poll on the real is-selected signal. - RemoteCLIRelayServerTests: capture errno before close() so thrown bind/listen diagnostics report the real failure code. py_compile + swift build --build-tests (CmuxRemoteWorkspace) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Make CommandPalette search benchmarks advisory instead of wall-clock-gated Greptile/CodeRabbit correctly flagged that loosening these optimized-vs-reference ratio asserts (e.g. 0.80 -> 1.25) no longer validates anything: a band tight enough to prove the optimized path is faster flakes on shared CI, and a band wide enough not to flake passes even when the optimized path regresses. The engine exposes no preparation/work counter, so there is no causal (non-wall- clock) signal to assert on here. Per the repo test-time policy (no wall-clock latency asserts on shared CI) and matching the activation-session perf gate's advisory-timing approach, drop the flaky `#expect` ratio/dropped-frame assertions across all four benchmarks and keep the `BENCH ...` diagnostic prints for trend tracking. Each test still exercises both code paths end to end; real activation latency/frame-budget regressions are gated by the dedicated activation-session job. swift build --build-tests (CmuxCommandPalette) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * MobileCoreRPC: observe queue gate via @testable, not a production debug hook Per review (Aziz): test/debug extensions should not live in production source. Remove the `#if DEBUG debugQueuedRequestCount()` accessor from MobileCoreRPCClient/Session and instead widen `session` and `queuedRequestIDs` from `private` to `internal` so the cancellation test reads the writer-gate state directly through its existing `@testable import CmuxMobileRPC`. All test scaffolding now lives in the test target; production source carries only the two access-level changes (still module-internal, no shipping behavior change). swift build --build-tests (CmuxMobileRPC) clean; no debug funcs remain in Sources/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Flaky-test follow-up: review fixes + second-pass team sweep (#6452) * Address review: fail-loud guards + scoped soft-skips in de-flaked tests - test_browser_goto_split.py: _wait_url_loaded now raises on timeout instead of silently returning (false pass), and the local HTTP server is managed via a "with" block so the thread/socket can't leak on setup exceptions. - test_surface_move_reorder_api.py: assert the cleanup re-select of ws0 actually converges instead of ignoring the _wait() result (state leak across runs). - test_homebrew_sha.sh: only soft-skip transient transport failures (000/408/429/5xx); fail hard on deterministic client errors like 404 (missing release asset). - test_visual_typing_char_by_char.py: validate the typed glyph against the post-prompt region of the last line, so a "cmux" already in the prompt/path can't satisfy the check; robust to trailing zsh-autosuggestion glyphs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 5d4c282535a497ad263128b21f8805a34556f55d) * Address review round 2: critical walrus bug + 3 fail-loud/correctness guards - test_pane_break_swap_preserve_focus.py: fix NameError -- the lambda referenced `p` in the ternary condition before the walrus assigned it. Bind the walrus inside the condition so the predicate reads panes once and returns them. - perf-activation-session.py: a missing measurement (actual is None) is now always a blocking failure, not advisory -- absent data is a benchmark-contract violation, distinct from load-sensitive over-budget timing. - perf-activation-session.py: copy the real-scrollback measurement before storing it under snapshot_with_scrollback, so best_of_snapshot_timing's in-place mutation no longer clobbers the raw snapshot_with_real_scrollback capture. - NotificationAndMenuBarTests.swift: on the stall-timeout path, fail fast via a lock-guarded result box instead of awaiting evaluationTask.value, which could hang until the suite timeout when the hook ignores cancellation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7e75ad467bc95cef0436236d94183f43b9aaffb2) * Add reports.md: consolidated flaky-test discovery (in-session workflow) Output of a parallel dynamic workflow (20 agents over the 419 signal-bearing test files of 2080) verifying residual flakiness on this already-de-flaked branch. 101 candidate findings across 86 files, dominated by wall-clock timing asserts, sleep-as-sync, and async races -- concentrated in Packages/ unit tests the first 44-file sweep did not cover. Each entry is a candidate; the fix phase adversarially verifies every one before changing code (repo rules allow deterministic test sleeps, so those are left alone). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit e1952d01faee805298bdff215af1d0897f63abcf) * Fix 37 verified residual flaky tests (coordinated area-team workflow) Second-pass de-flake driven by the in-session discovery + team-fix workflows (see reports.md). Adversarial per-area verification confirmed genuine flakiness before any edit; 9 candidates were refuted (deterministic-allowed sleeps, factually-wrong premises, fixes needing risky production changes) and left. Highlights: - Packages/ unit tests (missed by the first 44-file sweep): AuthCoordinator and HostBrowserSignIn now drive timeouts off an injected ManualTestClock instead of real ContinuousClock; ChatConversationStore poller results are asserted with #expect instead of discarded (fail-loud); RemoteProxyBroker / RemoteCLIRelay / CommandPaletteSearchEngine timing/race tightened. - Integration tests: per-PID/UUID socket+port paths (terminal_focus_routing, ctrl_socket, ssh_remote_* disjoint port ranges, CMUXCLIErrorOutput socket) to kill cross-run collisions; order-dependent session-restore checks now scan all workspaces instead of trusting seeding order; latent NameErrors fixed (pane_resize missing `import time` / `must`); fail-open waits made loud (new_tab tmp-write now raises); surface-targeted send to avoid focus races (tab_dragging); fd/temp-file leaks closed. - App-target Swift: removed/loosened flaky wall-clock asserts while keeping the behavior assert (TerminalAndGhostty paste, ShellStartupMatrix budget), widened burst spacing (GhosttyNotificationDispatcher), and routed UITest socket/condition waits through the shared waitForControlSocketReady helper. All Python pass py_compile; all 4 touched packages pass swift build --build-tests; app-target Swift is CI-compiled. No test run against a socket. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7a374ac84c2c96d175df7656b9fd6657a788ea68) * Fix MobileCoreRPC cancelled-while-queued race deterministically (DEBUG-only seam) The remaining discovery finding: MobileCoreRPCClientTests spun a fixed 100 `Task.yield()`s to "wait" for a queued request to reach the session writer gate before cancelling it. Under scheduler load the queued task may not have reached the gate yet, so cancellation fires before `session.send` registers it and the cancelled-while-queued invariant is never exercised (false pass). That gate state (`queuedRequestIDs`) is private to the production `MobileCoreRPCSession` actor, so there is no deterministic test-only signal. Fix adds a `#if DEBUG`, read-only `debugQueuedRequestCount()` to the session and a thin client wrapper, placed in the file's existing `#if DEBUG` test-support extension (alongside `debugWithRequestTimeout`). The test now polls that real signal until the request is registered at the gate, then cancels. The accessor is `#if DEBUG` and read-only: it is compiled out of Release builds and changes no shipping-build behavior, so it needs no dogfood. `swift build --build-tests` (debug) compiles and links the package + test target cleanly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit c56f53ece49ae1c98072bb24996c55e6d15b4fa6) * Refresh Swift file-length budget for the de-flaked test files The fail-loud guards and ManualTestClock/RunLoop-poll rewrites grew five test files past their budgets (NotificationAndMenuBar +31, CommandPaletteSearchEngine +21, MultiWindowNotifications +9, BrowserPaneNavigationKeybind +3, AgentHibernation +1) and shrank two (TerminalAndGhostty -7 after dropping a flaky wall-clock assert, BrowserFixtureInteraction -6 after routing through the shared readiness helper). Budget updated for exactly those seven files. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: drop reintroduced wall-clock asserts / fixed sleeps CodeRabbit/Greptile caught spots where the second-pass fixes leaned on the very patterns this PR is removing. All test-only: - ShellStartupMatrixTests: drop the redundant `duration < 5.0` ceiling; the bootstrap already runs under a 5s `runProcess` timeout, so status==0 + !timedOut is the causal completion signal (no wall-clock assert). - MultiWindowNotificationsUITests: assert no-foreground at the causal point (right after `waitForCommandCompletionWhileBackgrounded`) instead of polling a fixed 2s window. - test_tab_dragging.py: remove two `time.sleep(1.5)` waits; the file-content poll loop right below is the readiness signal. - test_terminal_focus_routing.py: `tempfile.mktemp` -> `mkstemp` (atomic, no TOCTOU; clears Ruff S306). - test_session_restore_stress_kill_cycles.py: replace the fixed 0.1s settle with a deadline-bounded poll on the real is-selected signal. - RemoteCLIRelayServerTests: capture errno before close() so thrown bind/listen diagnostics report the real failure code. py_compile + swift build --build-tests (CmuxRemoteWorkspace) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Make CommandPalette search benchmarks advisory instead of wall-clock-gated Greptile/CodeRabbit correctly flagged that loosening these optimized-vs-reference ratio asserts (e.g. 0.80 -> 1.25) no longer validates anything: a band tight enough to prove the optimized path is faster flakes on shared CI, and a band wide enough not to flake passes even when the optimized path regresses. The engine exposes no preparation/work counter, so there is no causal (non-wall- clock) signal to assert on here. Per the repo test-time policy (no wall-clock latency asserts on shared CI) and matching the activation-session perf gate's advisory-timing approach, drop the flaky `#expect` ratio/dropped-frame assertions across all four benchmarks and keep the `BENCH ...` diagnostic prints for trend tracking. Each test still exercises both code paths end to end; real activation latency/frame-budget regressions are gated by the dedicated activation-session job. swift build --build-tests (CmuxCommandPalette) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * MobileCoreRPC: observe queue gate via @testable, not a production debug hook Per review (Aziz): test/debug extensions should not live in production source. Remove the `#if DEBUG debugQueuedRequestCount()` accessor from MobileCoreRPCClient/Session and instead widen `session` and `queuedRequestIDs` from `private` to `internal` so the cancellation test reads the writer-gate state directly through its existing `@testable import CmuxMobileRPC`. All test scaffolding now lives in the test target; production source carries only the two access-level changes (still module-internal, no shipping behavior change). swift build --build-tests (CmuxMobileRPC) clean; no debug funcs remain in Sources/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Fix frozen terminals after split churn (#12) * Fix blank terminal after split operations and add visual tests ## Blank Terminal Fix - Add `needsRefreshAfterWindowChange` flag in GhosttyTerminalView - Force terminal refresh when view is added to window, even if size unchanged - Add `ghostty_surface_refresh()` call in attachToView for same-view reattachment - Add debug logging for surface attachment lifecycle (DEBUG builds only) ## Bonsplit Migration - Add bonsplit as local Swift package (vendor/bonsplit submodule) - Replace custom SplitTree with BonsplitController - Add Panel protocol with TerminalPanel and BrowserPanel implementations - Add SidebarTab as main tab container with BonsplitController - Remove old Splits/ directory (SplitTree, SplitView, TerminalSplitTreeView) ## Visual Screenshot Tests - Add test_visual_screenshots.py for automated visual regression testing - Uses in-app screenshot API (CGWindowListCreateImage) - no screen recording needed - Generates HTML report with before/after comparisons - Tests: splits, browser panels, focus switching, close operations, rapid cycles - Includes annotation fields for easy feedback ## Browser Shortcut (⌘⇧B) - Add keyboard shortcut to open browser panel in current pane - Add openBrowser() method to TabManager - Add shortcut configuration in KeyboardShortcutSettings ## Screenshot Command - Add 'screenshot' command to TerminalController for in-app window capture - Returns OK with screenshot ID and path ## Other - Add tests/visual_output/ and tests/visual_report.html to .gitignore * Add browser title subscription and set tab height to 30px - Subscribe to BrowserPanel.$pageTitle changes to update bonsplit tabs - Update tab titles in real-time as page navigation occurs - Clean up subscriptions when panels are removed - Set bonsplit tab bar and tab height to 30px (in submodule) * Fix socket API regressions in list_surfaces, list_bonsplit_tabs, focus_pane - list_surfaces: Remove [terminal]/[browser] suffix to keep UUID-only format that clients and tests expect for parsing - list_bonsplit_tabs --pane: Properly look up pane by UUID instead of creating a new PaneID (requires bonsplit PaneID.id to be public) - focus_pane: Accept both UUID strings and integer indices as documented * Fix browser panel stability and keyboard shortcuts - Prevent WKWebView focus lifecycle crashes during split/view reshuffles - Match bracket shortcuts via keyCode (Cmd+Shift+[ / ], Cmd+Ctrl+[ / ]) - Support Ghostty config goto_split:* keybinds when WebView is focused - Add focus_webview/is_webview_focused socket commands and regression tests - Rename SidebarTab to Workspace and update docs * Make ctrl+enter keybind test skippable Skip when the Ghostty keybind isn't configured or when osascript can't send keystrokes (no Accessibility permission), so VM runs stay green. * Auto-focus browser omnibar when blank When a browser surface is focused but no URL is loaded yet, focus the address bar instead of the WKWebView. * Stabilize socket surface indexing * Focus browser omnibar escape; add webview keybind UI tests - Escape in omnibar now returns focus to WKWebView\n- Add UI tests for Cmd+Ctrl+H pane navigation with WebKit focused (including Ghostty config)\n- Avoid flaky element screenshots in UpdatePillUITests on the UTM VM * Fix browser drag-to-split blanks and socket parsing * Fix webview-focused shortcuts and stabilize browser splits - Match ctrl/shift shortcuts by keyCode where needed (Ctrl+H, bracket keys) - Load Ghostty goto_split triggers reliably and refresh on config load - Add debug socket helpers: set_shortcut + simulate_shortcut for tests - Convert browser goto_split/keybind tests to socket-based injection (no osascript) - Bump bonsplit for drag-to-split fixes * Fix split layout collapse and harden socket pane APIs * Stabilize OSC 99 notification test timing * Fix terminal focus routing after split reparent * Support simulate_shortcut enter for focus routing test * Stabilize terminal focus routing test * Fix frozen new terminal tabs after many splits * Fix frozen new terminal tabs after splits * Fix terminal freeze on launch/new tabs * Update ghostty submodule * Fix terminal focus/render stalls after split churn * Fix nested split collapsing existing pane * Fix nested split collapse + stabilize new-surface focus * Update bonsplit submodule * Fix SIGINT test flake * Remove bonsplit tab-switch crossfade * Remove PROJECTS.md * Remove bonsplit tab selection animation * Ignore generated test reports * Middle click closes tab * Revert unintended .gitignore change * Fix build after main merge * Revert "Fix build after main merge" This reverts commit 16bf9816d0856b5385d52f886aa5eb50f3c9d9a4. * Revert "Merge remote-tracking branch 'origin/main' into fix/blank-terminal-and-visual-tests" This reverts commit 7c20fb53fd71fea7a19a3673f2dd73e5f0c783c4, reversing changes made to 0aff107d787bc9d8bbc28220090b4ca7af72e040. * Remove tab close fade animation * Use terminal.fill icon * Make terminal tab icon smaller * Match browser globe tab icon size * Bonsplit: tab min width 48 and tighter close button * Bonsplit: smaller tab title font * Show unread notification badge in bonsplit tabs and improve UI polish Sync unread notification state to bonsplit tab badges (blue dot). Improve EmptyPanelView with Terminal/Browser buttons and shortcut hints. Add tooltips to close tab button and search overlay buttons. * Fix reload.sh single-instance safety check on macOS Replace GNU-only `ps -o etimes=` with portable `ps -o etime=` and parse the dd-hh:mm:ss format manually for macOS compatibility. * Centralize keyboard shortcut definitions into Action enum Replace per-shortcut boilerplate with a single Action enum that holds the label, defaults key, and default binding for each shortcut. All call sites now use shortcut(for:). Settings UI is data-driven via ForEach(Action.allCases). Titlebar tooltips update dynamically when shortcuts are changed. Remove duplicate .keyboardShortcut() modifiers from menu items that are already handled by the event monitor. * Fix WKWebView consuming app menu shortcuts and close panel confirmation Add CmuxWebView subclass that routes key equivalents through the main menu before WebKit, so Cmd+N/Cmd+W/tab switching work when a browser pane is focused. Fix Cmd+W close-panel path: bypass Bonsplit delegate gating after the user confirms the running-process dialog by tracking forceCloseTabIds. Add unit tests (CmuxWebViewKeyEquivalentTests) and UI test scaffolding (MenuKeyEquivalentRoutingUITests) with a new cmux-unit Xcode scheme. * Update CLAUDE.md and PROJECTS.md with recent changes CLAUDE.md: enforce --tag for reload commands, add cleanup safety rules. PROJECTS.md: log notification badge, reload.sh fix, Cmd+W fix, WebView key equiv fix, and centralized shortcuts work. * Keep selection index stable on close * Add concepts page documenting terminology hierarchy New docs page explaining Window > Workspace > Pane > Surface > Panel hierarchy with aligned ASCII diagram. Updated tabs.mdx and splits.mdx to use consistent terminology (workspace instead of tab, surface instead of panel) and corrected outdated CLI command references. * Update bonsplit submodule * WIP: improve split close stability and UI regressions * Close terminal panel on child exit; hide terminal dirty dot * Fix split close/focus regressions and stabilize UI tests * Add unread Dock/Cmd+Tab badge with settings toggle * Fix browser-surface shortcuts and Cmd+L browser opening * Snapshot current workspace state before regression fixes * Update bonsplit submodule snapshot * Stabilize split-close regression capture and sidebar resize assertions * Change default Show Notifications shortcut from Cmd+Shift+I to Cmd+I * Fix update check readiness race, enable release update logging, and improve checking spinner * Restore terminal file drop, fix browser omnibar click focus, and add panel workspace ID mutation for surface moves * Add Cmd+digit workspace hints, titlebar shortcut pills, sidebar drag-reorder, and workspace placement settings * Add v2 browser automation API, surface move/reorder commands, and short-handle ref system to TerminalController * Add CLI browser command surface, --id-format flag, and move/reorder commands * Extend test clients with move/reorder APIs, ref-handle support, and increased timeouts * Harden test runner scripts with deterministic builds, retry logic, and robust socket readiness * Stabilize existing test suites with focus-wait helpers, increased timeouts, and API shape updates * Add terminal file drop e2e regression test * Add v2 browser API, CLI ref resolution, and surface move/reorder test suites * Add unit tests for shortcut hints, workspace reorder, drop planner, and update UI test stabilization * Add cmux-debug-windows skill with snapshot script and agent config * Update project docs: mark browser parity and move/reorder phases complete, add parallel agent workflow guidelines * Update bonsplit submodule: re-entrant setPosition guard, tab shortcut hints, and moveTab/reorderTab API * Add browser agent UX improvements: snapshot refs, placement reuse, diagnostics, and skill docs - Upgrade browser.snapshot to emit accessibility tree text with element refs (eN) - Add right-sibling pane reuse policy for browser.open_split placement - Add rich not_found diagnostics with retry logic for selector actions - Support --snapshot-after for post-action verification on mutating commands - Allow browser fill with empty text for clearing inputs - Default CLI --id-format to refs-first (UUIDs opt-in via --id-format uuids|both) - Format legacy new-pane/new-surface output with short surface refs - Add skills/cmuxterm-browser/ and skills/cmuxterm/ end-user skill docs - Add regression tests for placement policy, snapshot refs, diagnostics, and ID defaults * Update bonsplit submodule: keep raster favicons in color when inactive | 7 个月前 | |
Add native iPhone and iPad Simulator panes (#7857) * tighten simulator shortcut and mutation paths * Harden simulator operation commit boundaries * Close simulator routing and deadline races * Reconcile simulator mutations during teardown * Preserve simulator caller routing context * Route ios commands from caller pane * Preserve explicit simulator routing and input recovery * Serialize simulator text recovery * Align simulator deadlines and selection rollback * Close simulator lifecycle commit races * Reject stale simulator discovery results * Bound simulator re-enable and batch capture * Select the integrated simulator display * Close simulator transition races * Gate simulator automation on readiness * Finish simulator selection generation checks * Bound simulator context work end to end * Verify simulator descendant identities * Validate simulator process ancestry * Bound simulator process shutdown * Supervise simulator command groups * Supervise simulator pane sessions * Contain every simulator subprocess * Test simulator routing and feature flag regressions * Fail closed on stale simulator state * Test simulator lifecycle review regressions * Bound simulator control lifecycles * Test simulator cross-pane ownership regressions * Serialize simulator cross-pane mutations * Test Web Inspector cross-worker ownership * Lease Web Inspector targets across workers * Test Web Inspector stale occupancy handoff * Refresh Web Inspector occupancy during handoff * Test Web Inspector occupancy timeout * Fail closed on incomplete Inspector occupancy * Test Inspector census and release regressions * Close Inspector refresh and release races * Test stale location route teardown * Preserve location route ownership through teardown * Test Simulator launch environment privacy * Test route commit during device switch * Test feature flag telemetry consent * Own Simulator mutations through teardown * Harden Simulator isolation and ownership * Persist Simulator mutation ownership across processes * Honor telemetry consent for remote flags * Preserve failed Simulator cleanup ownership * Propagate Simulator CLI routing errors * Fix Simulator operation task syntax * Inject Simulator ownership and keep context read-only * Resolve restored Simulator context without booting * Test late Simulator display identity publication * Attach Simulator callbacks before display discovery * Test Simulator mutation ownership boundaries * Harden Simulator mutation ownership semantics * Test current Simulator default-screen contract * Support current Simulator default-screen metadata * Isolate Simulator camera transport tests * Test forwarded Simulator screen metadata * Support forwarded Simulator screen metadata * Align Simulator overlay lifecycle test with visibility * Bound Simulator recovery assertion by time * Make Simulator input recovery test deterministic * Test Simulator landscape framebuffer direction * Correct Simulator landscape presentation direction * Bound Simulator frame publication test wait * Register Simulator replay state before delivery * Test Simulator review boundary cases * Test iOS screenshot surface identity errors * Normalize Simulator control boundaries * Isolate Simulator CLI contract environment * Test native Simulator orientation dialects * Unify native Simulator orientation semantics * Test landscape Simulator digitizer coordinates * Map landscape input into native digitizer space * Test stable Simulator app switcher hold * Hold app switcher gesture stationary * Test Simulator app switcher button timing * Open Simulator app switcher with double Home * Test installed iPad DeviceKit chrome fallback * Test bounded Simulator frame publication * Scale Simulator frames to pane geometry * Test Simulator frame ring replacement cleanup * Release obsolete Simulator frame rings * Route IndexNow jobs through configured runner * Remove wall-clock assertions from RPC event tests * Test frame ring adoption race * Retire Simulator frame rings after host adoption * Keep frame completion on MainActor * Snapshot Simulator activity log before lazy layout * Regenerate webview assets after main merge * Test Simulator core readiness ordering * Stream Simulator before optional capability probes * Test control-socket Simulator selection ownership * Exclude active Simulator control action from teardown * Test natural DeviceKit chrome cap geometry * Preserve native DeviceKit chrome artwork geometry * Fix design mode test payload shadowing * Make Simulator replay tests signal-driven * Fix Simulator pane test client conformance * Test immediate Simulator context discovery * Discover Simulator before context reads * Test Simulator production edge cases * Close Simulator production review findings * Fix Simulator integration test fixtures * Fix Simulator CLI routing call site * Fix Simulator focus test fixtures * Fix Simulator app test fixtures * Test Simulator capability hydration readiness * Wait for Simulator capability hydration * Test Simulator review edge cases * Close Simulator production review gaps * Avoid recursive Simulator picker comparison * Isolate Simulator picker observation * Test Simulator application row snapshots * Snapshot Simulator application picker rows * Test static Simulator frame presentation * Drive Simulator frames without display callbacks * Test Simulator visibility remounts * Keep Simulator frames through host remounts * Isolate Simulator visibility regression suite * Test Simulator frame pacing under input load * Pace Simulator framebuffer readback * Localize project surface labels * Test Camera Injector header cache identity * Invalidate Camera Injector cache for headers * Test Simulator RPC capability discovery * Advertise Simulator RPC capabilities * test(simulator): cover interactive frame priority * fix(simulator): prioritize frames after pointer input * test(simulator): cover native tap hold duration * fix(simulator): hold synthetic taps long enough for iPadOS * Test immediate Simulator frame presentation * Present Simulator frames without an extra tick * Test failed Simulator ownership publication * Reject unsafe Simulator ownership claims * test(simulator): cover framebuffer lifecycle bounds * fix(simulator): bound framebuffer lifecycle work * test(simulator): cover review lifecycle regressions * fix(simulator): close review lifecycle gaps * test(simulator): cover routing pacing and consent * fix(simulator): scope routing pacing and consent * test(simulator): cover screenshot and cached log readiness * fix(simulator): prepare capture without eager lifecycle work * test(simulator): report capture-ready live state * fix(simulator): report live capture-ready state * test(simulator): cover control-plane and frame wakeups * fix(simulator): signal frames and preserve errored flag cache * test(simulator): cover publication wakeup races * fix(simulator): bound publication wakeups * test(simulator): cover tool editor shortcut focus * fix(simulator): preserve tool editor focus ownership * test(simulator): cover flag omission and runner injection * fix(simulator): inject async owned command execution * test(simulator): cover file drop proposal readiness * fix(simulator): validate file drop proposals * test(simulator): cover compound inspector cleanup failure * fix(simulator): preserve failed inspector cleanup state * test(simulator): cover device-scoped tool state * fix(simulator): scope tool state to selected device * test(simulator): cover final release blockers * fix(simulator): close final release blockers * test(simulator): make frame scheduling assertions deterministic * test: update Ghostty surface config ABI lock * fix(simulator): clear final build gates * fix(build): import Dock lifecycle workspace types * test(web): isolate feedback route environment * fix(build): disambiguate Dock snapshot types * test(simulator): cover review ownership blockers * fix(simulator): scope camera cleanup ownership * fix(build): make resume policy returns explicit * test(simulator): cover camera cleanup ownership retries * fix(simulator): preserve camera cleanup ownership * fix(build): clear latest main gates * test(simulator): cover final review blockers * fix(simulator): await quit cleanup and index targets * refactor(simulator): satisfy production policy * test(simulator): cover camera cleanup on device switch * fix(simulator): clean camera state before device switch * test(simulator): retain quit cleanup after panel removal * fix(simulator): make quit await durable rollback * test(simulator): cover retained cleanup recovery * fix(simulator): recover retained camera cleanup * test(simulator): cover cleanup side effects * fix(simulator): restore external cleanup state * refactor(simulator): split camera authorization record * test(web): respect delegated discovery order * test(ios): assert transition math deterministically * fix(ci): repair current-main Swift integration * fix(ci): return closed workspace restore result * test(simulator): cover final review regressions * fix(simulator): close final lifecycle gaps * test(simulator): cover review lifecycle findings * fix(simulator): resolve review lifecycle findings * test(simulator): cover durable recovery journals * fix(simulator): persist mutation recovery journals * refactor(simulator): inject durable recovery paths * test(simulator): cover durable journal transitions * fix(simulator): make recovery transitions crash consistent * test(simulator): cover recovery compatibility gaps * fix(simulator): preserve recovery compatibility * test(simulator): cover recovery ownership handoff * fix(simulator): gate recovery ownership handoff * test(simulator): cover duplicate camera journals * fix(simulator): suppress stale durable camera journals * test(simulator): cover journal reconciliation races * fix(simulator): serialize journal reconciliation * test(simulator): cover identical journal paths * fix(simulator): normalize camera journal paths * test(simulator): cover journal URL hints * fix(simulator): compare normalized journal paths * refactor(simulator): split legacy route fixture * test(simulator): observe journal lock contention | 2 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Flaky-test follow-up: review fixes + second-pass team sweep (#6452) * Address review: fail-loud guards + scoped soft-skips in de-flaked tests - test_browser_goto_split.py: _wait_url_loaded now raises on timeout instead of silently returning (false pass), and the local HTTP server is managed via a "with" block so the thread/socket can't leak on setup exceptions. - test_surface_move_reorder_api.py: assert the cleanup re-select of ws0 actually converges instead of ignoring the _wait() result (state leak across runs). - test_homebrew_sha.sh: only soft-skip transient transport failures (000/408/429/5xx); fail hard on deterministic client errors like 404 (missing release asset). - test_visual_typing_char_by_char.py: validate the typed glyph against the post-prompt region of the last line, so a "cmux" already in the prompt/path can't satisfy the check; robust to trailing zsh-autosuggestion glyphs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 5d4c282535a497ad263128b21f8805a34556f55d) * Address review round 2: critical walrus bug + 3 fail-loud/correctness guards - test_pane_break_swap_preserve_focus.py: fix NameError -- the lambda referenced `p` in the ternary condition before the walrus assigned it. Bind the walrus inside the condition so the predicate reads panes once and returns them. - perf-activation-session.py: a missing measurement (actual is None) is now always a blocking failure, not advisory -- absent data is a benchmark-contract violation, distinct from load-sensitive over-budget timing. - perf-activation-session.py: copy the real-scrollback measurement before storing it under snapshot_with_scrollback, so best_of_snapshot_timing's in-place mutation no longer clobbers the raw snapshot_with_real_scrollback capture. - NotificationAndMenuBarTests.swift: on the stall-timeout path, fail fast via a lock-guarded result box instead of awaiting evaluationTask.value, which could hang until the suite timeout when the hook ignores cancellation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7e75ad467bc95cef0436236d94183f43b9aaffb2) * Add reports.md: consolidated flaky-test discovery (in-session workflow) Output of a parallel dynamic workflow (20 agents over the 419 signal-bearing test files of 2080) verifying residual flakiness on this already-de-flaked branch. 101 candidate findings across 86 files, dominated by wall-clock timing asserts, sleep-as-sync, and async races -- concentrated in Packages/ unit tests the first 44-file sweep did not cover. Each entry is a candidate; the fix phase adversarially verifies every one before changing code (repo rules allow deterministic test sleeps, so those are left alone). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit e1952d01faee805298bdff215af1d0897f63abcf) * Fix 37 verified residual flaky tests (coordinated area-team workflow) Second-pass de-flake driven by the in-session discovery + team-fix workflows (see reports.md). Adversarial per-area verification confirmed genuine flakiness before any edit; 9 candidates were refuted (deterministic-allowed sleeps, factually-wrong premises, fixes needing risky production changes) and left. Highlights: - Packages/ unit tests (missed by the first 44-file sweep): AuthCoordinator and HostBrowserSignIn now drive timeouts off an injected ManualTestClock instead of real ContinuousClock; ChatConversationStore poller results are asserted with #expect instead of discarded (fail-loud); RemoteProxyBroker / RemoteCLIRelay / CommandPaletteSearchEngine timing/race tightened. - Integration tests: per-PID/UUID socket+port paths (terminal_focus_routing, ctrl_socket, ssh_remote_* disjoint port ranges, CMUXCLIErrorOutput socket) to kill cross-run collisions; order-dependent session-restore checks now scan all workspaces instead of trusting seeding order; latent NameErrors fixed (pane_resize missing `import time` / `must`); fail-open waits made loud (new_tab tmp-write now raises); surface-targeted send to avoid focus races (tab_dragging); fd/temp-file leaks closed. - App-target Swift: removed/loosened flaky wall-clock asserts while keeping the behavior assert (TerminalAndGhostty paste, ShellStartupMatrix budget), widened burst spacing (GhosttyNotificationDispatcher), and routed UITest socket/condition waits through the shared waitForControlSocketReady helper. All Python pass py_compile; all 4 touched packages pass swift build --build-tests; app-target Swift is CI-compiled. No test run against a socket. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7a374ac84c2c96d175df7656b9fd6657a788ea68) * Fix MobileCoreRPC cancelled-while-queued race deterministically (DEBUG-only seam) The remaining discovery finding: MobileCoreRPCClientTests spun a fixed 100 `Task.yield()`s to "wait" for a queued request to reach the session writer gate before cancelling it. Under scheduler load the queued task may not have reached the gate yet, so cancellation fires before `session.send` registers it and the cancelled-while-queued invariant is never exercised (false pass). That gate state (`queuedRequestIDs`) is private to the production `MobileCoreRPCSession` actor, so there is no deterministic test-only signal. Fix adds a `#if DEBUG`, read-only `debugQueuedRequestCount()` to the session and a thin client wrapper, placed in the file's existing `#if DEBUG` test-support extension (alongside `debugWithRequestTimeout`). The test now polls that real signal until the request is registered at the gate, then cancels. The accessor is `#if DEBUG` and read-only: it is compiled out of Release builds and changes no shipping-build behavior, so it needs no dogfood. `swift build --build-tests` (debug) compiles and links the package + test target cleanly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit c56f53ece49ae1c98072bb24996c55e6d15b4fa6) * Refresh Swift file-length budget for the de-flaked test files The fail-loud guards and ManualTestClock/RunLoop-poll rewrites grew five test files past their budgets (NotificationAndMenuBar +31, CommandPaletteSearchEngine +21, MultiWindowNotifications +9, BrowserPaneNavigationKeybind +3, AgentHibernation +1) and shrank two (TerminalAndGhostty -7 after dropping a flaky wall-clock assert, BrowserFixtureInteraction -6 after routing through the shared readiness helper). Budget updated for exactly those seven files. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: drop reintroduced wall-clock asserts / fixed sleeps CodeRabbit/Greptile caught spots where the second-pass fixes leaned on the very patterns this PR is removing. All test-only: - ShellStartupMatrixTests: drop the redundant `duration < 5.0` ceiling; the bootstrap already runs under a 5s `runProcess` timeout, so status==0 + !timedOut is the causal completion signal (no wall-clock assert). - MultiWindowNotificationsUITests: assert no-foreground at the causal point (right after `waitForCommandCompletionWhileBackgrounded`) instead of polling a fixed 2s window. - test_tab_dragging.py: remove two `time.sleep(1.5)` waits; the file-content poll loop right below is the readiness signal. - test_terminal_focus_routing.py: `tempfile.mktemp` -> `mkstemp` (atomic, no TOCTOU; clears Ruff S306). - test_session_restore_stress_kill_cycles.py: replace the fixed 0.1s settle with a deadline-bounded poll on the real is-selected signal. - RemoteCLIRelayServerTests: capture errno before close() so thrown bind/listen diagnostics report the real failure code. py_compile + swift build --build-tests (CmuxRemoteWorkspace) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Make CommandPalette search benchmarks advisory instead of wall-clock-gated Greptile/CodeRabbit correctly flagged that loosening these optimized-vs-reference ratio asserts (e.g. 0.80 -> 1.25) no longer validates anything: a band tight enough to prove the optimized path is faster flakes on shared CI, and a band wide enough not to flake passes even when the optimized path regresses. The engine exposes no preparation/work counter, so there is no causal (non-wall- clock) signal to assert on here. Per the repo test-time policy (no wall-clock latency asserts on shared CI) and matching the activation-session perf gate's advisory-timing approach, drop the flaky `#expect` ratio/dropped-frame assertions across all four benchmarks and keep the `BENCH ...` diagnostic prints for trend tracking. Each test still exercises both code paths end to end; real activation latency/frame-budget regressions are gated by the dedicated activation-session job. swift build --build-tests (CmuxCommandPalette) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * MobileCoreRPC: observe queue gate via @testable, not a production debug hook Per review (Aziz): test/debug extensions should not live in production source. Remove the `#if DEBUG debugQueuedRequestCount()` accessor from MobileCoreRPCClient/Session and instead widen `session` and `queuedRequestIDs` from `private` to `internal` so the cancellation test reads the writer-gate state directly through its existing `@testable import CmuxMobileRPC`. All test scaffolding now lives in the test target; production source carries only the two access-level changes (still module-internal, no shipping behavior change). swift build --build-tests (CmuxMobileRPC) clean; no debug funcs remain in Sources/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Flaky-test follow-up: review fixes + second-pass team sweep (#6452) * Address review: fail-loud guards + scoped soft-skips in de-flaked tests - test_browser_goto_split.py: _wait_url_loaded now raises on timeout instead of silently returning (false pass), and the local HTTP server is managed via a "with" block so the thread/socket can't leak on setup exceptions. - test_surface_move_reorder_api.py: assert the cleanup re-select of ws0 actually converges instead of ignoring the _wait() result (state leak across runs). - test_homebrew_sha.sh: only soft-skip transient transport failures (000/408/429/5xx); fail hard on deterministic client errors like 404 (missing release asset). - test_visual_typing_char_by_char.py: validate the typed glyph against the post-prompt region of the last line, so a "cmux" already in the prompt/path can't satisfy the check; robust to trailing zsh-autosuggestion glyphs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 5d4c282535a497ad263128b21f8805a34556f55d) * Address review round 2: critical walrus bug + 3 fail-loud/correctness guards - test_pane_break_swap_preserve_focus.py: fix NameError -- the lambda referenced `p` in the ternary condition before the walrus assigned it. Bind the walrus inside the condition so the predicate reads panes once and returns them. - perf-activation-session.py: a missing measurement (actual is None) is now always a blocking failure, not advisory -- absent data is a benchmark-contract violation, distinct from load-sensitive over-budget timing. - perf-activation-session.py: copy the real-scrollback measurement before storing it under snapshot_with_scrollback, so best_of_snapshot_timing's in-place mutation no longer clobbers the raw snapshot_with_real_scrollback capture. - NotificationAndMenuBarTests.swift: on the stall-timeout path, fail fast via a lock-guarded result box instead of awaiting evaluationTask.value, which could hang until the suite timeout when the hook ignores cancellation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7e75ad467bc95cef0436236d94183f43b9aaffb2) * Add reports.md: consolidated flaky-test discovery (in-session workflow) Output of a parallel dynamic workflow (20 agents over the 419 signal-bearing test files of 2080) verifying residual flakiness on this already-de-flaked branch. 101 candidate findings across 86 files, dominated by wall-clock timing asserts, sleep-as-sync, and async races -- concentrated in Packages/ unit tests the first 44-file sweep did not cover. Each entry is a candidate; the fix phase adversarially verifies every one before changing code (repo rules allow deterministic test sleeps, so those are left alone). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit e1952d01faee805298bdff215af1d0897f63abcf) * Fix 37 verified residual flaky tests (coordinated area-team workflow) Second-pass de-flake driven by the in-session discovery + team-fix workflows (see reports.md). Adversarial per-area verification confirmed genuine flakiness before any edit; 9 candidates were refuted (deterministic-allowed sleeps, factually-wrong premises, fixes needing risky production changes) and left. Highlights: - Packages/ unit tests (missed by the first 44-file sweep): AuthCoordinator and HostBrowserSignIn now drive timeouts off an injected ManualTestClock instead of real ContinuousClock; ChatConversationStore poller results are asserted with #expect instead of discarded (fail-loud); RemoteProxyBroker / RemoteCLIRelay / CommandPaletteSearchEngine timing/race tightened. - Integration tests: per-PID/UUID socket+port paths (terminal_focus_routing, ctrl_socket, ssh_remote_* disjoint port ranges, CMUXCLIErrorOutput socket) to kill cross-run collisions; order-dependent session-restore checks now scan all workspaces instead of trusting seeding order; latent NameErrors fixed (pane_resize missing `import time` / `must`); fail-open waits made loud (new_tab tmp-write now raises); surface-targeted send to avoid focus races (tab_dragging); fd/temp-file leaks closed. - App-target Swift: removed/loosened flaky wall-clock asserts while keeping the behavior assert (TerminalAndGhostty paste, ShellStartupMatrix budget), widened burst spacing (GhosttyNotificationDispatcher), and routed UITest socket/condition waits through the shared waitForControlSocketReady helper. All Python pass py_compile; all 4 touched packages pass swift build --build-tests; app-target Swift is CI-compiled. No test run against a socket. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7a374ac84c2c96d175df7656b9fd6657a788ea68) * Fix MobileCoreRPC cancelled-while-queued race deterministically (DEBUG-only seam) The remaining discovery finding: MobileCoreRPCClientTests spun a fixed 100 `Task.yield()`s to "wait" for a queued request to reach the session writer gate before cancelling it. Under scheduler load the queued task may not have reached the gate yet, so cancellation fires before `session.send` registers it and the cancelled-while-queued invariant is never exercised (false pass). That gate state (`queuedRequestIDs`) is private to the production `MobileCoreRPCSession` actor, so there is no deterministic test-only signal. Fix adds a `#if DEBUG`, read-only `debugQueuedRequestCount()` to the session and a thin client wrapper, placed in the file's existing `#if DEBUG` test-support extension (alongside `debugWithRequestTimeout`). The test now polls that real signal until the request is registered at the gate, then cancels. The accessor is `#if DEBUG` and read-only: it is compiled out of Release builds and changes no shipping-build behavior, so it needs no dogfood. `swift build --build-tests` (debug) compiles and links the package + test target cleanly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit c56f53ece49ae1c98072bb24996c55e6d15b4fa6) * Refresh Swift file-length budget for the de-flaked test files The fail-loud guards and ManualTestClock/RunLoop-poll rewrites grew five test files past their budgets (NotificationAndMenuBar +31, CommandPaletteSearchEngine +21, MultiWindowNotifications +9, BrowserPaneNavigationKeybind +3, AgentHibernation +1) and shrank two (TerminalAndGhostty -7 after dropping a flaky wall-clock assert, BrowserFixtureInteraction -6 after routing through the shared readiness helper). Budget updated for exactly those seven files. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: drop reintroduced wall-clock asserts / fixed sleeps CodeRabbit/Greptile caught spots where the second-pass fixes leaned on the very patterns this PR is removing. All test-only: - ShellStartupMatrixTests: drop the redundant `duration < 5.0` ceiling; the bootstrap already runs under a 5s `runProcess` timeout, so status==0 + !timedOut is the causal completion signal (no wall-clock assert). - MultiWindowNotificationsUITests: assert no-foreground at the causal point (right after `waitForCommandCompletionWhileBackgrounded`) instead of polling a fixed 2s window. - test_tab_dragging.py: remove two `time.sleep(1.5)` waits; the file-content poll loop right below is the readiness signal. - test_terminal_focus_routing.py: `tempfile.mktemp` -> `mkstemp` (atomic, no TOCTOU; clears Ruff S306). - test_session_restore_stress_kill_cycles.py: replace the fixed 0.1s settle with a deadline-bounded poll on the real is-selected signal. - RemoteCLIRelayServerTests: capture errno before close() so thrown bind/listen diagnostics report the real failure code. py_compile + swift build --build-tests (CmuxRemoteWorkspace) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Make CommandPalette search benchmarks advisory instead of wall-clock-gated Greptile/CodeRabbit correctly flagged that loosening these optimized-vs-reference ratio asserts (e.g. 0.80 -> 1.25) no longer validates anything: a band tight enough to prove the optimized path is faster flakes on shared CI, and a band wide enough not to flake passes even when the optimized path regresses. The engine exposes no preparation/work counter, so there is no causal (non-wall- clock) signal to assert on here. Per the repo test-time policy (no wall-clock latency asserts on shared CI) and matching the activation-session perf gate's advisory-timing approach, drop the flaky `#expect` ratio/dropped-frame assertions across all four benchmarks and keep the `BENCH ...` diagnostic prints for trend tracking. Each test still exercises both code paths end to end; real activation latency/frame-budget regressions are gated by the dedicated activation-session job. swift build --build-tests (CmuxCommandPalette) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * MobileCoreRPC: observe queue gate via @testable, not a production debug hook Per review (Aziz): test/debug extensions should not live in production source. Remove the `#if DEBUG debugQueuedRequestCount()` accessor from MobileCoreRPCClient/Session and instead widen `session` and `queuedRequestIDs` from `private` to `internal` so the cancellation test reads the writer-gate state directly through its existing `@testable import CmuxMobileRPC`. All test scaffolding now lives in the test target; production source carries only the two access-level changes (still module-internal, no shipping behavior change). swift build --build-tests (CmuxMobileRPC) clean; no debug funcs remain in Sources/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
CLI: allow cmux ssh to run an initial remote command (#8439) * Add initial command to cmux ssh * Test initial command with fallback shells * Run initial command portably for fallback shells * Test private remote command staging * Stage remote initial command privately * Test initial command generations sharing a relay port * Key initial command state by launch generation * Test concurrent initial command staging * Make initial command staging race-safe * Test fallback command shell state * Preserve fallback shell command state * Test unknown remote shell fallback * Retain unknown remote shell fallback * Test exec initial command cleanup * Run initial command without decoded files * Test unknown shell interactive command state * Run fallback commands in interactive shells * Test Nushell initial command startup * Run Nushell initial commands interactively * Verify Nushell enters its initial REPL * Document Nushell execute contract * Make mobile liveness tests deterministic * Test remote initial command failure recovery * Retry remote initial commands safely * Test private remote command retry staging * Stage remote initial commands privately * Update SSH bootstrap metadata assertion * Fix Swift Testing diagnostic comments * Test zsh initial command login ordering * Run zsh initial command after login setup | 2 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Eliminate flaky tests: deterministic rewrites across Swift, Python, shell (#6392) * Fix flaky tests across Swift, Python, and shell suites Replace nondeterministic test patterns with waits on real signals so the same correct code stops failing under CI/VM load. Found by a parallel static audit of ~950 test files; 61 fixes across 44 files. - sleep-as-sync: fixed sleep + single assert replaced with a bounded poll on the actual readiness signal (existing in-file wait helpers preferred). - tight-timeout: expectation/wait timeouts widened to generous bounds; only the failure path waits longer, passing runs return as soon as the signal fires. - wall-clock-assert: single-shot timing ratios replaced with best-of-N minima plus deterministic work-count assertions; perf benchmarks no longer flip on a single scheduler spike. - order-dependence: shared static recorder suite serialized; shared state reset per test. - network/port races: external endpoints swapped for local/ephemeral servers; real-network download hardened with retries, timeouts, and a transient-failure soft-skip. Deterministic test sleeps (allowed by the repo review rules) are left intact. All changed Python/shell files pass py_compile/bash -n; all five changed SPM package test targets compile via swift build --build-tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for de-flaked test files The poll helpers and best-of-N scaffolding that replaced fixed sleeps added a few lines to 6 Swift test files, crossing their per-file budgets. Update only those rows (and track HostBrowserSignInFlowTests, newly over 500). No other files touched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Deflake activation-session perf gate: best-of-N timing, advisory budgets The activation benchmark failed nondeterministically on snapshot_with_scrollback .elapsed_ms (e.g. 1827ms > 1500ms): a single wall-clock measurement of a load-sensitive ~1.3M-char snapshot, compared to a hard absolute ceiling, run on shared/Depot GUI runners. A loaded runner trivially adds 20%+, so the same correct code flipped red across unrelated PRs. Two changes, matching the de-flake philosophy used elsewhere in this PR (performance is measured, not gated on a fixed wall-clock ceiling): - best_of_snapshot_timing: re-measure each snapshot timing N times (--budget-snapshot-samples, default 3; persist=False so reruns have no session side effects) and keep the MINIMUM elapsed_ms. A real regression slows every sample so the min still regresses; transient contention only inflates some samples, so the min is stable. Shape/char counts are preserved. - Wall-clock timing budgets (launch/restore socket-ready, snapshot elapsed_ms) are now advisory: reported as budget_warnings, non-blocking. Deterministic structural budgets (min scrollback chars, min terminal surfaces, restored workspace/terminal counts) stay blocking. Pass --fail-on-timing-budget to hard-gate timing if a dedicated, controlled perf runner is used. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Rewrite remaining flaky tests to logic/signal/virtual-clock based Eliminates the 37 residual time-dependent test primitives the determinism checker flags (27 sleep-then-assert, 10 assert-on-duration) across 17 files, applying the two principles: - Assert on causality, not latency: duration assertions (elapsed < N, Date().timeIntervalSince < N) replaced with assertions on the logical outcome the timing proxied (operation completed / right value / right state / right order), or a generous deadline-bounded wait on the real completion signal where a "did not hang" guard is genuinely needed. - Invert the time dependency: sleep-then-assert replaced with a wait on the real signal: an injected virtual clock advanced by hand (e.g. ManualTestClock via existing makeHarness(clock:) seam), a completion signal awaited directly (semaphore/continuation/main-queue drain), or a deadline-bounded poll of the real state predicate where the system exposes no completion event. No sleeps or timeouts were merely widened; no assertions were weakened. The determinism checker now reports 0 active findings. All four changed SPM package test targets compile via swift build --build-tests; Python files pass py_compile. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for rewritten test files The deterministic rewrites (virtual clocks, poll helpers, signal waits) added lines to 7 Swift test files. Bump only those rows. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Add detachable SSH PTY daemon persistence (#4807) * Add detachable SSH PTY daemon persistence * Fix persistent SSH PTY review findings * Bound persistent daemon auth handshake * Fix persistent SSH daemon review findings * Enforce persistent daemon slot invariant * Fix persistent SSH restore control paths * Preserve persistent daemon sessions on auth rotation * Fix SSH fork and daemon cleanup regressions * test: cover persistent SSH PTY relaunch snapshots * fix: persist SSH PTY session IDs through relaunch * fix: preserve workspace id for SSH PTY fallback * fix: preserve SSH options in reconnect restore * fix: remap restored SSH relay context * fix: use unique daemon token temp files * fix: address ssh pty review findings * fix: tighten ssh pty review feedback * fix: keep relay alias rewrites on relay queue * fix: clean stale persistent ssh proxies * fix: scope persistent ssh cleanup by slot * fix: gate persistent ssh restore on socket * fix: preserve generic ssh cleanup with slots * fix: bound persistent daemon slot lifecycle * fix: parse persistent daemon cleanup slots by token * fix: simplify restored relay token guards * fix: avoid reattaching ended ssh ptys * fix: match escaped ssh daemon slot cleanup * fix: harden persistent daemon socket dir * test: cover ssh pty reattach command restore * fix: execute persisted ssh pty attach scripts * fix: keep only persistent ssh pty exits visible * fix: harden ssh pty restore fallbacks * fix: restore ssh pty pane bootstrap * fix: remap restored ssh relay anchors * fix: guard restored ssh shell placeholders * fix: remap restored ssh panel aliases * fix: cancel stale ssh relay forwards * fix: preserve ssh relay aliases across reconnect * fix: keep ssh pty identity on relay refresh * fix: restore moved ssh pty relay aliases * fix: preserve existing terminfo environment * fix: clear failed persistent ssh attach state * ci: retrigger stuck PR checks * fix: reset persistent ssh sessions on relay changes * test: keep relay refresh on stable port * test: cover stale ssh relay listener cleanup * fix: clean stale ssh relay listeners * fix: default persistent ssh daemon slot * fix: preserve relay command line endings * fix: harden persistent ssh relay cleanup * fix: suppress remote restore scaffold attaches * test: use workspaceId in SSH PTY snapshot test * fix: address remote relay review feedback * refactor: split remote shell bootstrap builder * fix: preserve bash login startup precedence * fix: address detachable pty review and ci * test: harden websocket pty close assertion * ci: prefer homebrew zig for macos release build * ci: build release artifacts on macos 15 * ci: extend lag regression cold-build timeout * fix: address persistent ssh restore review * fix: preserve failed persistent pty split panes * fix: suppress ended persistent pty restore startup * test: avoid ready fd reuse in daemon tests * fix: wire browser tab mute context action * fix: preserve persistent attach failures * fix: preserve local cwd on remote restore * fix: reconnect restored remote browser workspaces * test: preserve daemon events during rpc calls * fix: wait for restored pty foreground auth --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 4 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Flaky-test follow-up: review fixes + second-pass team sweep (#6452) * Address review: fail-loud guards + scoped soft-skips in de-flaked tests - test_browser_goto_split.py: _wait_url_loaded now raises on timeout instead of silently returning (false pass), and the local HTTP server is managed via a "with" block so the thread/socket can't leak on setup exceptions. - test_surface_move_reorder_api.py: assert the cleanup re-select of ws0 actually converges instead of ignoring the _wait() result (state leak across runs). - test_homebrew_sha.sh: only soft-skip transient transport failures (000/408/429/5xx); fail hard on deterministic client errors like 404 (missing release asset). - test_visual_typing_char_by_char.py: validate the typed glyph against the post-prompt region of the last line, so a "cmux" already in the prompt/path can't satisfy the check; robust to trailing zsh-autosuggestion glyphs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 5d4c282535a497ad263128b21f8805a34556f55d) * Address review round 2: critical walrus bug + 3 fail-loud/correctness guards - test_pane_break_swap_preserve_focus.py: fix NameError -- the lambda referenced `p` in the ternary condition before the walrus assigned it. Bind the walrus inside the condition so the predicate reads panes once and returns them. - perf-activation-session.py: a missing measurement (actual is None) is now always a blocking failure, not advisory -- absent data is a benchmark-contract violation, distinct from load-sensitive over-budget timing. - perf-activation-session.py: copy the real-scrollback measurement before storing it under snapshot_with_scrollback, so best_of_snapshot_timing's in-place mutation no longer clobbers the raw snapshot_with_real_scrollback capture. - NotificationAndMenuBarTests.swift: on the stall-timeout path, fail fast via a lock-guarded result box instead of awaiting evaluationTask.value, which could hang until the suite timeout when the hook ignores cancellation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7e75ad467bc95cef0436236d94183f43b9aaffb2) * Add reports.md: consolidated flaky-test discovery (in-session workflow) Output of a parallel dynamic workflow (20 agents over the 419 signal-bearing test files of 2080) verifying residual flakiness on this already-de-flaked branch. 101 candidate findings across 86 files, dominated by wall-clock timing asserts, sleep-as-sync, and async races -- concentrated in Packages/ unit tests the first 44-file sweep did not cover. Each entry is a candidate; the fix phase adversarially verifies every one before changing code (repo rules allow deterministic test sleeps, so those are left alone). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit e1952d01faee805298bdff215af1d0897f63abcf) * Fix 37 verified residual flaky tests (coordinated area-team workflow) Second-pass de-flake driven by the in-session discovery + team-fix workflows (see reports.md). Adversarial per-area verification confirmed genuine flakiness before any edit; 9 candidates were refuted (deterministic-allowed sleeps, factually-wrong premises, fixes needing risky production changes) and left. Highlights: - Packages/ unit tests (missed by the first 44-file sweep): AuthCoordinator and HostBrowserSignIn now drive timeouts off an injected ManualTestClock instead of real ContinuousClock; ChatConversationStore poller results are asserted with #expect instead of discarded (fail-loud); RemoteProxyBroker / RemoteCLIRelay / CommandPaletteSearchEngine timing/race tightened. - Integration tests: per-PID/UUID socket+port paths (terminal_focus_routing, ctrl_socket, ssh_remote_* disjoint port ranges, CMUXCLIErrorOutput socket) to kill cross-run collisions; order-dependent session-restore checks now scan all workspaces instead of trusting seeding order; latent NameErrors fixed (pane_resize missing `import time` / `must`); fail-open waits made loud (new_tab tmp-write now raises); surface-targeted send to avoid focus races (tab_dragging); fd/temp-file leaks closed. - App-target Swift: removed/loosened flaky wall-clock asserts while keeping the behavior assert (TerminalAndGhostty paste, ShellStartupMatrix budget), widened burst spacing (GhosttyNotificationDispatcher), and routed UITest socket/condition waits through the shared waitForControlSocketReady helper. All Python pass py_compile; all 4 touched packages pass swift build --build-tests; app-target Swift is CI-compiled. No test run against a socket. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7a374ac84c2c96d175df7656b9fd6657a788ea68) * Fix MobileCoreRPC cancelled-while-queued race deterministically (DEBUG-only seam) The remaining discovery finding: MobileCoreRPCClientTests spun a fixed 100 `Task.yield()`s to "wait" for a queued request to reach the session writer gate before cancelling it. Under scheduler load the queued task may not have reached the gate yet, so cancellation fires before `session.send` registers it and the cancelled-while-queued invariant is never exercised (false pass). That gate state (`queuedRequestIDs`) is private to the production `MobileCoreRPCSession` actor, so there is no deterministic test-only signal. Fix adds a `#if DEBUG`, read-only `debugQueuedRequestCount()` to the session and a thin client wrapper, placed in the file's existing `#if DEBUG` test-support extension (alongside `debugWithRequestTimeout`). The test now polls that real signal until the request is registered at the gate, then cancels. The accessor is `#if DEBUG` and read-only: it is compiled out of Release builds and changes no shipping-build behavior, so it needs no dogfood. `swift build --build-tests` (debug) compiles and links the package + test target cleanly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit c56f53ece49ae1c98072bb24996c55e6d15b4fa6) * Refresh Swift file-length budget for the de-flaked test files The fail-loud guards and ManualTestClock/RunLoop-poll rewrites grew five test files past their budgets (NotificationAndMenuBar +31, CommandPaletteSearchEngine +21, MultiWindowNotifications +9, BrowserPaneNavigationKeybind +3, AgentHibernation +1) and shrank two (TerminalAndGhostty -7 after dropping a flaky wall-clock assert, BrowserFixtureInteraction -6 after routing through the shared readiness helper). Budget updated for exactly those seven files. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: drop reintroduced wall-clock asserts / fixed sleeps CodeRabbit/Greptile caught spots where the second-pass fixes leaned on the very patterns this PR is removing. All test-only: - ShellStartupMatrixTests: drop the redundant `duration < 5.0` ceiling; the bootstrap already runs under a 5s `runProcess` timeout, so status==0 + !timedOut is the causal completion signal (no wall-clock assert). - MultiWindowNotificationsUITests: assert no-foreground at the causal point (right after `waitForCommandCompletionWhileBackgrounded`) instead of polling a fixed 2s window. - test_tab_dragging.py: remove two `time.sleep(1.5)` waits; the file-content poll loop right below is the readiness signal. - test_terminal_focus_routing.py: `tempfile.mktemp` -> `mkstemp` (atomic, no TOCTOU; clears Ruff S306). - test_session_restore_stress_kill_cycles.py: replace the fixed 0.1s settle with a deadline-bounded poll on the real is-selected signal. - RemoteCLIRelayServerTests: capture errno before close() so thrown bind/listen diagnostics report the real failure code. py_compile + swift build --build-tests (CmuxRemoteWorkspace) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Make CommandPalette search benchmarks advisory instead of wall-clock-gated Greptile/CodeRabbit correctly flagged that loosening these optimized-vs-reference ratio asserts (e.g. 0.80 -> 1.25) no longer validates anything: a band tight enough to prove the optimized path is faster flakes on shared CI, and a band wide enough not to flake passes even when the optimized path regresses. The engine exposes no preparation/work counter, so there is no causal (non-wall- clock) signal to assert on here. Per the repo test-time policy (no wall-clock latency asserts on shared CI) and matching the activation-session perf gate's advisory-timing approach, drop the flaky `#expect` ratio/dropped-frame assertions across all four benchmarks and keep the `BENCH ...` diagnostic prints for trend tracking. Each test still exercises both code paths end to end; real activation latency/frame-budget regressions are gated by the dedicated activation-session job. swift build --build-tests (CmuxCommandPalette) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * MobileCoreRPC: observe queue gate via @testable, not a production debug hook Per review (Aziz): test/debug extensions should not live in production source. Remove the `#if DEBUG debugQueuedRequestCount()` accessor from MobileCoreRPCClient/Session and instead widen `session` and `queuedRequestIDs` from `private` to `internal` so the cancellation test reads the writer-gate state directly through its existing `@testable import CmuxMobileRPC`. All test scaffolding now lives in the test target; production source carries only the two access-level changes (still module-internal, no shipping behavior change). swift build --build-tests (CmuxMobileRPC) clean; no debug funcs remain in Sources/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Fix multi-image terminal drops (#3769) * Add multi-image drop regressions * Fix multi-image terminal drops * Add CLI stderr crash regression * Fix CLI stderr crash on closed pipes * Add Claude image drop paste regression * Use one paste transaction for local image drops * Remove unused local image type helpers * Copy pasteboard item before fallback rendering * Add Claude multi-image paste regression * Fix Claude multi-image drops --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 4 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Flaky-test follow-up: review fixes + second-pass team sweep (#6452) * Address review: fail-loud guards + scoped soft-skips in de-flaked tests - test_browser_goto_split.py: _wait_url_loaded now raises on timeout instead of silently returning (false pass), and the local HTTP server is managed via a "with" block so the thread/socket can't leak on setup exceptions. - test_surface_move_reorder_api.py: assert the cleanup re-select of ws0 actually converges instead of ignoring the _wait() result (state leak across runs). - test_homebrew_sha.sh: only soft-skip transient transport failures (000/408/429/5xx); fail hard on deterministic client errors like 404 (missing release asset). - test_visual_typing_char_by_char.py: validate the typed glyph against the post-prompt region of the last line, so a "cmux" already in the prompt/path can't satisfy the check; robust to trailing zsh-autosuggestion glyphs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 5d4c282535a497ad263128b21f8805a34556f55d) * Address review round 2: critical walrus bug + 3 fail-loud/correctness guards - test_pane_break_swap_preserve_focus.py: fix NameError -- the lambda referenced `p` in the ternary condition before the walrus assigned it. Bind the walrus inside the condition so the predicate reads panes once and returns them. - perf-activation-session.py: a missing measurement (actual is None) is now always a blocking failure, not advisory -- absent data is a benchmark-contract violation, distinct from load-sensitive over-budget timing. - perf-activation-session.py: copy the real-scrollback measurement before storing it under snapshot_with_scrollback, so best_of_snapshot_timing's in-place mutation no longer clobbers the raw snapshot_with_real_scrollback capture. - NotificationAndMenuBarTests.swift: on the stall-timeout path, fail fast via a lock-guarded result box instead of awaiting evaluationTask.value, which could hang until the suite timeout when the hook ignores cancellation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7e75ad467bc95cef0436236d94183f43b9aaffb2) * Add reports.md: consolidated flaky-test discovery (in-session workflow) Output of a parallel dynamic workflow (20 agents over the 419 signal-bearing test files of 2080) verifying residual flakiness on this already-de-flaked branch. 101 candidate findings across 86 files, dominated by wall-clock timing asserts, sleep-as-sync, and async races -- concentrated in Packages/ unit tests the first 44-file sweep did not cover. Each entry is a candidate; the fix phase adversarially verifies every one before changing code (repo rules allow deterministic test sleeps, so those are left alone). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit e1952d01faee805298bdff215af1d0897f63abcf) * Fix 37 verified residual flaky tests (coordinated area-team workflow) Second-pass de-flake driven by the in-session discovery + team-fix workflows (see reports.md). Adversarial per-area verification confirmed genuine flakiness before any edit; 9 candidates were refuted (deterministic-allowed sleeps, factually-wrong premises, fixes needing risky production changes) and left. Highlights: - Packages/ unit tests (missed by the first 44-file sweep): AuthCoordinator and HostBrowserSignIn now drive timeouts off an injected ManualTestClock instead of real ContinuousClock; ChatConversationStore poller results are asserted with #expect instead of discarded (fail-loud); RemoteProxyBroker / RemoteCLIRelay / CommandPaletteSearchEngine timing/race tightened. - Integration tests: per-PID/UUID socket+port paths (terminal_focus_routing, ctrl_socket, ssh_remote_* disjoint port ranges, CMUXCLIErrorOutput socket) to kill cross-run collisions; order-dependent session-restore checks now scan all workspaces instead of trusting seeding order; latent NameErrors fixed (pane_resize missing `import time` / `must`); fail-open waits made loud (new_tab tmp-write now raises); surface-targeted send to avoid focus races (tab_dragging); fd/temp-file leaks closed. - App-target Swift: removed/loosened flaky wall-clock asserts while keeping the behavior assert (TerminalAndGhostty paste, ShellStartupMatrix budget), widened burst spacing (GhosttyNotificationDispatcher), and routed UITest socket/condition waits through the shared waitForControlSocketReady helper. All Python pass py_compile; all 4 touched packages pass swift build --build-tests; app-target Swift is CI-compiled. No test run against a socket. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7a374ac84c2c96d175df7656b9fd6657a788ea68) * Fix MobileCoreRPC cancelled-while-queued race deterministically (DEBUG-only seam) The remaining discovery finding: MobileCoreRPCClientTests spun a fixed 100 `Task.yield()`s to "wait" for a queued request to reach the session writer gate before cancelling it. Under scheduler load the queued task may not have reached the gate yet, so cancellation fires before `session.send` registers it and the cancelled-while-queued invariant is never exercised (false pass). That gate state (`queuedRequestIDs`) is private to the production `MobileCoreRPCSession` actor, so there is no deterministic test-only signal. Fix adds a `#if DEBUG`, read-only `debugQueuedRequestCount()` to the session and a thin client wrapper, placed in the file's existing `#if DEBUG` test-support extension (alongside `debugWithRequestTimeout`). The test now polls that real signal until the request is registered at the gate, then cancels. The accessor is `#if DEBUG` and read-only: it is compiled out of Release builds and changes no shipping-build behavior, so it needs no dogfood. `swift build --build-tests` (debug) compiles and links the package + test target cleanly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit c56f53ece49ae1c98072bb24996c55e6d15b4fa6) * Refresh Swift file-length budget for the de-flaked test files The fail-loud guards and ManualTestClock/RunLoop-poll rewrites grew five test files past their budgets (NotificationAndMenuBar +31, CommandPaletteSearchEngine +21, MultiWindowNotifications +9, BrowserPaneNavigationKeybind +3, AgentHibernation +1) and shrank two (TerminalAndGhostty -7 after dropping a flaky wall-clock assert, BrowserFixtureInteraction -6 after routing through the shared readiness helper). Budget updated for exactly those seven files. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: drop reintroduced wall-clock asserts / fixed sleeps CodeRabbit/Greptile caught spots where the second-pass fixes leaned on the very patterns this PR is removing. All test-only: - ShellStartupMatrixTests: drop the redundant `duration < 5.0` ceiling; the bootstrap already runs under a 5s `runProcess` timeout, so status==0 + !timedOut is the causal completion signal (no wall-clock assert). - MultiWindowNotificationsUITests: assert no-foreground at the causal point (right after `waitForCommandCompletionWhileBackgrounded`) instead of polling a fixed 2s window. - test_tab_dragging.py: remove two `time.sleep(1.5)` waits; the file-content poll loop right below is the readiness signal. - test_terminal_focus_routing.py: `tempfile.mktemp` -> `mkstemp` (atomic, no TOCTOU; clears Ruff S306). - test_session_restore_stress_kill_cycles.py: replace the fixed 0.1s settle with a deadline-bounded poll on the real is-selected signal. - RemoteCLIRelayServerTests: capture errno before close() so thrown bind/listen diagnostics report the real failure code. py_compile + swift build --build-tests (CmuxRemoteWorkspace) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Make CommandPalette search benchmarks advisory instead of wall-clock-gated Greptile/CodeRabbit correctly flagged that loosening these optimized-vs-reference ratio asserts (e.g. 0.80 -> 1.25) no longer validates anything: a band tight enough to prove the optimized path is faster flakes on shared CI, and a band wide enough not to flake passes even when the optimized path regresses. The engine exposes no preparation/work counter, so there is no causal (non-wall- clock) signal to assert on here. Per the repo test-time policy (no wall-clock latency asserts on shared CI) and matching the activation-session perf gate's advisory-timing approach, drop the flaky `#expect` ratio/dropped-frame assertions across all four benchmarks and keep the `BENCH ...` diagnostic prints for trend tracking. Each test still exercises both code paths end to end; real activation latency/frame-budget regressions are gated by the dedicated activation-session job. swift build --build-tests (CmuxCommandPalette) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * MobileCoreRPC: observe queue gate via @testable, not a production debug hook Per review (Aziz): test/debug extensions should not live in production source. Remove the `#if DEBUG debugQueuedRequestCount()` accessor from MobileCoreRPCClient/Session and instead widen `session` and `queuedRequestIDs` from `private` to `internal` so the cancellation test reads the writer-gate state directly through its existing `@testable import CmuxMobileRPC`. All test scaffolding now lives in the test target; production source carries only the two access-level changes (still module-internal, no shipping behavior change). swift build --build-tests (CmuxMobileRPC) clean; no debug funcs remain in Sources/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Eliminate flaky tests: deterministic rewrites across Swift, Python, shell (#6392) * Fix flaky tests across Swift, Python, and shell suites Replace nondeterministic test patterns with waits on real signals so the same correct code stops failing under CI/VM load. Found by a parallel static audit of ~950 test files; 61 fixes across 44 files. - sleep-as-sync: fixed sleep + single assert replaced with a bounded poll on the actual readiness signal (existing in-file wait helpers preferred). - tight-timeout: expectation/wait timeouts widened to generous bounds; only the failure path waits longer, passing runs return as soon as the signal fires. - wall-clock-assert: single-shot timing ratios replaced with best-of-N minima plus deterministic work-count assertions; perf benchmarks no longer flip on a single scheduler spike. - order-dependence: shared static recorder suite serialized; shared state reset per test. - network/port races: external endpoints swapped for local/ephemeral servers; real-network download hardened with retries, timeouts, and a transient-failure soft-skip. Deterministic test sleeps (allowed by the repo review rules) are left intact. All changed Python/shell files pass py_compile/bash -n; all five changed SPM package test targets compile via swift build --build-tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for de-flaked test files The poll helpers and best-of-N scaffolding that replaced fixed sleeps added a few lines to 6 Swift test files, crossing their per-file budgets. Update only those rows (and track HostBrowserSignInFlowTests, newly over 500). No other files touched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Deflake activation-session perf gate: best-of-N timing, advisory budgets The activation benchmark failed nondeterministically on snapshot_with_scrollback .elapsed_ms (e.g. 1827ms > 1500ms): a single wall-clock measurement of a load-sensitive ~1.3M-char snapshot, compared to a hard absolute ceiling, run on shared/Depot GUI runners. A loaded runner trivially adds 20%+, so the same correct code flipped red across unrelated PRs. Two changes, matching the de-flake philosophy used elsewhere in this PR (performance is measured, not gated on a fixed wall-clock ceiling): - best_of_snapshot_timing: re-measure each snapshot timing N times (--budget-snapshot-samples, default 3; persist=False so reruns have no session side effects) and keep the MINIMUM elapsed_ms. A real regression slows every sample so the min still regresses; transient contention only inflates some samples, so the min is stable. Shape/char counts are preserved. - Wall-clock timing budgets (launch/restore socket-ready, snapshot elapsed_ms) are now advisory: reported as budget_warnings, non-blocking. Deterministic structural budgets (min scrollback chars, min terminal surfaces, restored workspace/terminal counts) stay blocking. Pass --fail-on-timing-budget to hard-gate timing if a dedicated, controlled perf runner is used. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Rewrite remaining flaky tests to logic/signal/virtual-clock based Eliminates the 37 residual time-dependent test primitives the determinism checker flags (27 sleep-then-assert, 10 assert-on-duration) across 17 files, applying the two principles: - Assert on causality, not latency: duration assertions (elapsed < N, Date().timeIntervalSince < N) replaced with assertions on the logical outcome the timing proxied (operation completed / right value / right state / right order), or a generous deadline-bounded wait on the real completion signal where a "did not hang" guard is genuinely needed. - Invert the time dependency: sleep-then-assert replaced with a wait on the real signal: an injected virtual clock advanced by hand (e.g. ManualTestClock via existing makeHarness(clock:) seam), a completion signal awaited directly (semaphore/continuation/main-queue drain), or a deadline-bounded poll of the real state predicate where the system exposes no completion event. No sleeps or timeouts were merely widened; no assertions were weakened. The determinism checker now reports 0 active findings. All four changed SPM package test targets compile via swift build --build-tests; Python files pass py_compile. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for rewritten test files The deterministic rewrites (virtual clocks, poll helpers, signal waits) added lines to 7 Swift test files. Bump only those rows. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Flaky-test follow-up: review fixes + second-pass team sweep (#6452) * Address review: fail-loud guards + scoped soft-skips in de-flaked tests - test_browser_goto_split.py: _wait_url_loaded now raises on timeout instead of silently returning (false pass), and the local HTTP server is managed via a "with" block so the thread/socket can't leak on setup exceptions. - test_surface_move_reorder_api.py: assert the cleanup re-select of ws0 actually converges instead of ignoring the _wait() result (state leak across runs). - test_homebrew_sha.sh: only soft-skip transient transport failures (000/408/429/5xx); fail hard on deterministic client errors like 404 (missing release asset). - test_visual_typing_char_by_char.py: validate the typed glyph against the post-prompt region of the last line, so a "cmux" already in the prompt/path can't satisfy the check; robust to trailing zsh-autosuggestion glyphs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 5d4c282535a497ad263128b21f8805a34556f55d) * Address review round 2: critical walrus bug + 3 fail-loud/correctness guards - test_pane_break_swap_preserve_focus.py: fix NameError -- the lambda referenced `p` in the ternary condition before the walrus assigned it. Bind the walrus inside the condition so the predicate reads panes once and returns them. - perf-activation-session.py: a missing measurement (actual is None) is now always a blocking failure, not advisory -- absent data is a benchmark-contract violation, distinct from load-sensitive over-budget timing. - perf-activation-session.py: copy the real-scrollback measurement before storing it under snapshot_with_scrollback, so best_of_snapshot_timing's in-place mutation no longer clobbers the raw snapshot_with_real_scrollback capture. - NotificationAndMenuBarTests.swift: on the stall-timeout path, fail fast via a lock-guarded result box instead of awaiting evaluationTask.value, which could hang until the suite timeout when the hook ignores cancellation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7e75ad467bc95cef0436236d94183f43b9aaffb2) * Add reports.md: consolidated flaky-test discovery (in-session workflow) Output of a parallel dynamic workflow (20 agents over the 419 signal-bearing test files of 2080) verifying residual flakiness on this already-de-flaked branch. 101 candidate findings across 86 files, dominated by wall-clock timing asserts, sleep-as-sync, and async races -- concentrated in Packages/ unit tests the first 44-file sweep did not cover. Each entry is a candidate; the fix phase adversarially verifies every one before changing code (repo rules allow deterministic test sleeps, so those are left alone). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit e1952d01faee805298bdff215af1d0897f63abcf) * Fix 37 verified residual flaky tests (coordinated area-team workflow) Second-pass de-flake driven by the in-session discovery + team-fix workflows (see reports.md). Adversarial per-area verification confirmed genuine flakiness before any edit; 9 candidates were refuted (deterministic-allowed sleeps, factually-wrong premises, fixes needing risky production changes) and left. Highlights: - Packages/ unit tests (missed by the first 44-file sweep): AuthCoordinator and HostBrowserSignIn now drive timeouts off an injected ManualTestClock instead of real ContinuousClock; ChatConversationStore poller results are asserted with #expect instead of discarded (fail-loud); RemoteProxyBroker / RemoteCLIRelay / CommandPaletteSearchEngine timing/race tightened. - Integration tests: per-PID/UUID socket+port paths (terminal_focus_routing, ctrl_socket, ssh_remote_* disjoint port ranges, CMUXCLIErrorOutput socket) to kill cross-run collisions; order-dependent session-restore checks now scan all workspaces instead of trusting seeding order; latent NameErrors fixed (pane_resize missing `import time` / `must`); fail-open waits made loud (new_tab tmp-write now raises); surface-targeted send to avoid focus races (tab_dragging); fd/temp-file leaks closed. - App-target Swift: removed/loosened flaky wall-clock asserts while keeping the behavior assert (TerminalAndGhostty paste, ShellStartupMatrix budget), widened burst spacing (GhosttyNotificationDispatcher), and routed UITest socket/condition waits through the shared waitForControlSocketReady helper. All Python pass py_compile; all 4 touched packages pass swift build --build-tests; app-target Swift is CI-compiled. No test run against a socket. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7a374ac84c2c96d175df7656b9fd6657a788ea68) * Fix MobileCoreRPC cancelled-while-queued race deterministically (DEBUG-only seam) The remaining discovery finding: MobileCoreRPCClientTests spun a fixed 100 `Task.yield()`s to "wait" for a queued request to reach the session writer gate before cancelling it. Under scheduler load the queued task may not have reached the gate yet, so cancellation fires before `session.send` registers it and the cancelled-while-queued invariant is never exercised (false pass). That gate state (`queuedRequestIDs`) is private to the production `MobileCoreRPCSession` actor, so there is no deterministic test-only signal. Fix adds a `#if DEBUG`, read-only `debugQueuedRequestCount()` to the session and a thin client wrapper, placed in the file's existing `#if DEBUG` test-support extension (alongside `debugWithRequestTimeout`). The test now polls that real signal until the request is registered at the gate, then cancels. The accessor is `#if DEBUG` and read-only: it is compiled out of Release builds and changes no shipping-build behavior, so it needs no dogfood. `swift build --build-tests` (debug) compiles and links the package + test target cleanly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit c56f53ece49ae1c98072bb24996c55e6d15b4fa6) * Refresh Swift file-length budget for the de-flaked test files The fail-loud guards and ManualTestClock/RunLoop-poll rewrites grew five test files past their budgets (NotificationAndMenuBar +31, CommandPaletteSearchEngine +21, MultiWindowNotifications +9, BrowserPaneNavigationKeybind +3, AgentHibernation +1) and shrank two (TerminalAndGhostty -7 after dropping a flaky wall-clock assert, BrowserFixtureInteraction -6 after routing through the shared readiness helper). Budget updated for exactly those seven files. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: drop reintroduced wall-clock asserts / fixed sleeps CodeRabbit/Greptile caught spots where the second-pass fixes leaned on the very patterns this PR is removing. All test-only: - ShellStartupMatrixTests: drop the redundant `duration < 5.0` ceiling; the bootstrap already runs under a 5s `runProcess` timeout, so status==0 + !timedOut is the causal completion signal (no wall-clock assert). - MultiWindowNotificationsUITests: assert no-foreground at the causal point (right after `waitForCommandCompletionWhileBackgrounded`) instead of polling a fixed 2s window. - test_tab_dragging.py: remove two `time.sleep(1.5)` waits; the file-content poll loop right below is the readiness signal. - test_terminal_focus_routing.py: `tempfile.mktemp` -> `mkstemp` (atomic, no TOCTOU; clears Ruff S306). - test_session_restore_stress_kill_cycles.py: replace the fixed 0.1s settle with a deadline-bounded poll on the real is-selected signal. - RemoteCLIRelayServerTests: capture errno before close() so thrown bind/listen diagnostics report the real failure code. py_compile + swift build --build-tests (CmuxRemoteWorkspace) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Make CommandPalette search benchmarks advisory instead of wall-clock-gated Greptile/CodeRabbit correctly flagged that loosening these optimized-vs-reference ratio asserts (e.g. 0.80 -> 1.25) no longer validates anything: a band tight enough to prove the optimized path is faster flakes on shared CI, and a band wide enough not to flake passes even when the optimized path regresses. The engine exposes no preparation/work counter, so there is no causal (non-wall- clock) signal to assert on here. Per the repo test-time policy (no wall-clock latency asserts on shared CI) and matching the activation-session perf gate's advisory-timing approach, drop the flaky `#expect` ratio/dropped-frame assertions across all four benchmarks and keep the `BENCH ...` diagnostic prints for trend tracking. Each test still exercises both code paths end to end; real activation latency/frame-budget regressions are gated by the dedicated activation-session job. swift build --build-tests (CmuxCommandPalette) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * MobileCoreRPC: observe queue gate via @testable, not a production debug hook Per review (Aziz): test/debug extensions should not live in production source. Remove the `#if DEBUG debugQueuedRequestCount()` accessor from MobileCoreRPCClient/Session and instead widen `session` and `queuedRequestIDs` from `private` to `internal` so the cancellation test reads the writer-gate state directly through its existing `@testable import CmuxMobileRPC`. All test scaffolding now lives in the test target; production source carries only the two access-level changes (still module-internal, no shipping behavior change). swift build --build-tests (CmuxMobileRPC) clean; no debug funcs remain in Sources/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Flaky-test follow-up: review fixes + second-pass team sweep (#6452) * Address review: fail-loud guards + scoped soft-skips in de-flaked tests - test_browser_goto_split.py: _wait_url_loaded now raises on timeout instead of silently returning (false pass), and the local HTTP server is managed via a "with" block so the thread/socket can't leak on setup exceptions. - test_surface_move_reorder_api.py: assert the cleanup re-select of ws0 actually converges instead of ignoring the _wait() result (state leak across runs). - test_homebrew_sha.sh: only soft-skip transient transport failures (000/408/429/5xx); fail hard on deterministic client errors like 404 (missing release asset). - test_visual_typing_char_by_char.py: validate the typed glyph against the post-prompt region of the last line, so a "cmux" already in the prompt/path can't satisfy the check; robust to trailing zsh-autosuggestion glyphs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 5d4c282535a497ad263128b21f8805a34556f55d) * Address review round 2: critical walrus bug + 3 fail-loud/correctness guards - test_pane_break_swap_preserve_focus.py: fix NameError -- the lambda referenced `p` in the ternary condition before the walrus assigned it. Bind the walrus inside the condition so the predicate reads panes once and returns them. - perf-activation-session.py: a missing measurement (actual is None) is now always a blocking failure, not advisory -- absent data is a benchmark-contract violation, distinct from load-sensitive over-budget timing. - perf-activation-session.py: copy the real-scrollback measurement before storing it under snapshot_with_scrollback, so best_of_snapshot_timing's in-place mutation no longer clobbers the raw snapshot_with_real_scrollback capture. - NotificationAndMenuBarTests.swift: on the stall-timeout path, fail fast via a lock-guarded result box instead of awaiting evaluationTask.value, which could hang until the suite timeout when the hook ignores cancellation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7e75ad467bc95cef0436236d94183f43b9aaffb2) * Add reports.md: consolidated flaky-test discovery (in-session workflow) Output of a parallel dynamic workflow (20 agents over the 419 signal-bearing test files of 2080) verifying residual flakiness on this already-de-flaked branch. 101 candidate findings across 86 files, dominated by wall-clock timing asserts, sleep-as-sync, and async races -- concentrated in Packages/ unit tests the first 44-file sweep did not cover. Each entry is a candidate; the fix phase adversarially verifies every one before changing code (repo rules allow deterministic test sleeps, so those are left alone). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit e1952d01faee805298bdff215af1d0897f63abcf) * Fix 37 verified residual flaky tests (coordinated area-team workflow) Second-pass de-flake driven by the in-session discovery + team-fix workflows (see reports.md). Adversarial per-area verification confirmed genuine flakiness before any edit; 9 candidates were refuted (deterministic-allowed sleeps, factually-wrong premises, fixes needing risky production changes) and left. Highlights: - Packages/ unit tests (missed by the first 44-file sweep): AuthCoordinator and HostBrowserSignIn now drive timeouts off an injected ManualTestClock instead of real ContinuousClock; ChatConversationStore poller results are asserted with #expect instead of discarded (fail-loud); RemoteProxyBroker / RemoteCLIRelay / CommandPaletteSearchEngine timing/race tightened. - Integration tests: per-PID/UUID socket+port paths (terminal_focus_routing, ctrl_socket, ssh_remote_* disjoint port ranges, CMUXCLIErrorOutput socket) to kill cross-run collisions; order-dependent session-restore checks now scan all workspaces instead of trusting seeding order; latent NameErrors fixed (pane_resize missing `import time` / `must`); fail-open waits made loud (new_tab tmp-write now raises); surface-targeted send to avoid focus races (tab_dragging); fd/temp-file leaks closed. - App-target Swift: removed/loosened flaky wall-clock asserts while keeping the behavior assert (TerminalAndGhostty paste, ShellStartupMatrix budget), widened burst spacing (GhosttyNotificationDispatcher), and routed UITest socket/condition waits through the shared waitForControlSocketReady helper. All Python pass py_compile; all 4 touched packages pass swift build --build-tests; app-target Swift is CI-compiled. No test run against a socket. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7a374ac84c2c96d175df7656b9fd6657a788ea68) * Fix MobileCoreRPC cancelled-while-queued race deterministically (DEBUG-only seam) The remaining discovery finding: MobileCoreRPCClientTests spun a fixed 100 `Task.yield()`s to "wait" for a queued request to reach the session writer gate before cancelling it. Under scheduler load the queued task may not have reached the gate yet, so cancellation fires before `session.send` registers it and the cancelled-while-queued invariant is never exercised (false pass). That gate state (`queuedRequestIDs`) is private to the production `MobileCoreRPCSession` actor, so there is no deterministic test-only signal. Fix adds a `#if DEBUG`, read-only `debugQueuedRequestCount()` to the session and a thin client wrapper, placed in the file's existing `#if DEBUG` test-support extension (alongside `debugWithRequestTimeout`). The test now polls that real signal until the request is registered at the gate, then cancels. The accessor is `#if DEBUG` and read-only: it is compiled out of Release builds and changes no shipping-build behavior, so it needs no dogfood. `swift build --build-tests` (debug) compiles and links the package + test target cleanly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit c56f53ece49ae1c98072bb24996c55e6d15b4fa6) * Refresh Swift file-length budget for the de-flaked test files The fail-loud guards and ManualTestClock/RunLoop-poll rewrites grew five test files past their budgets (NotificationAndMenuBar +31, CommandPaletteSearchEngine +21, MultiWindowNotifications +9, BrowserPaneNavigationKeybind +3, AgentHibernation +1) and shrank two (TerminalAndGhostty -7 after dropping a flaky wall-clock assert, BrowserFixtureInteraction -6 after routing through the shared readiness helper). Budget updated for exactly those seven files. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: drop reintroduced wall-clock asserts / fixed sleeps CodeRabbit/Greptile caught spots where the second-pass fixes leaned on the very patterns this PR is removing. All test-only: - ShellStartupMatrixTests: drop the redundant `duration < 5.0` ceiling; the bootstrap already runs under a 5s `runProcess` timeout, so status==0 + !timedOut is the causal completion signal (no wall-clock assert). - MultiWindowNotificationsUITests: assert no-foreground at the causal point (right after `waitForCommandCompletionWhileBackgrounded`) instead of polling a fixed 2s window. - test_tab_dragging.py: remove two `time.sleep(1.5)` waits; the file-content poll loop right below is the readiness signal. - test_terminal_focus_routing.py: `tempfile.mktemp` -> `mkstemp` (atomic, no TOCTOU; clears Ruff S306). - test_session_restore_stress_kill_cycles.py: replace the fixed 0.1s settle with a deadline-bounded poll on the real is-selected signal. - RemoteCLIRelayServerTests: capture errno before close() so thrown bind/listen diagnostics report the real failure code. py_compile + swift build --build-tests (CmuxRemoteWorkspace) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Make CommandPalette search benchmarks advisory instead of wall-clock-gated Greptile/CodeRabbit correctly flagged that loosening these optimized-vs-reference ratio asserts (e.g. 0.80 -> 1.25) no longer validates anything: a band tight enough to prove the optimized path is faster flakes on shared CI, and a band wide enough not to flake passes even when the optimized path regresses. The engine exposes no preparation/work counter, so there is no causal (non-wall- clock) signal to assert on here. Per the repo test-time policy (no wall-clock latency asserts on shared CI) and matching the activation-session perf gate's advisory-timing approach, drop the flaky `#expect` ratio/dropped-frame assertions across all four benchmarks and keep the `BENCH ...` diagnostic prints for trend tracking. Each test still exercises both code paths end to end; real activation latency/frame-budget regressions are gated by the dedicated activation-session job. swift build --build-tests (CmuxCommandPalette) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * MobileCoreRPC: observe queue gate via @testable, not a production debug hook Per review (Aziz): test/debug extensions should not live in production source. Remove the `#if DEBUG debugQueuedRequestCount()` accessor from MobileCoreRPCClient/Session and instead widen `session` and `queuedRequestIDs` from `private` to `internal` so the cancellation test reads the writer-gate state directly through its existing `@testable import CmuxMobileRPC`. All test scaffolding now lives in the test target; production source carries only the two access-level changes (still module-internal, no shipping behavior change). swift build --build-tests (CmuxMobileRPC) clean; no debug funcs remain in Sources/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Flaky-test follow-up: review fixes + second-pass team sweep (#6452) * Address review: fail-loud guards + scoped soft-skips in de-flaked tests - test_browser_goto_split.py: _wait_url_loaded now raises on timeout instead of silently returning (false pass), and the local HTTP server is managed via a "with" block so the thread/socket can't leak on setup exceptions. - test_surface_move_reorder_api.py: assert the cleanup re-select of ws0 actually converges instead of ignoring the _wait() result (state leak across runs). - test_homebrew_sha.sh: only soft-skip transient transport failures (000/408/429/5xx); fail hard on deterministic client errors like 404 (missing release asset). - test_visual_typing_char_by_char.py: validate the typed glyph against the post-prompt region of the last line, so a "cmux" already in the prompt/path can't satisfy the check; robust to trailing zsh-autosuggestion glyphs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 5d4c282535a497ad263128b21f8805a34556f55d) * Address review round 2: critical walrus bug + 3 fail-loud/correctness guards - test_pane_break_swap_preserve_focus.py: fix NameError -- the lambda referenced `p` in the ternary condition before the walrus assigned it. Bind the walrus inside the condition so the predicate reads panes once and returns them. - perf-activation-session.py: a missing measurement (actual is None) is now always a blocking failure, not advisory -- absent data is a benchmark-contract violation, distinct from load-sensitive over-budget timing. - perf-activation-session.py: copy the real-scrollback measurement before storing it under snapshot_with_scrollback, so best_of_snapshot_timing's in-place mutation no longer clobbers the raw snapshot_with_real_scrollback capture. - NotificationAndMenuBarTests.swift: on the stall-timeout path, fail fast via a lock-guarded result box instead of awaiting evaluationTask.value, which could hang until the suite timeout when the hook ignores cancellation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7e75ad467bc95cef0436236d94183f43b9aaffb2) * Add reports.md: consolidated flaky-test discovery (in-session workflow) Output of a parallel dynamic workflow (20 agents over the 419 signal-bearing test files of 2080) verifying residual flakiness on this already-de-flaked branch. 101 candidate findings across 86 files, dominated by wall-clock timing asserts, sleep-as-sync, and async races -- concentrated in Packages/ unit tests the first 44-file sweep did not cover. Each entry is a candidate; the fix phase adversarially verifies every one before changing code (repo rules allow deterministic test sleeps, so those are left alone). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit e1952d01faee805298bdff215af1d0897f63abcf) * Fix 37 verified residual flaky tests (coordinated area-team workflow) Second-pass de-flake driven by the in-session discovery + team-fix workflows (see reports.md). Adversarial per-area verification confirmed genuine flakiness before any edit; 9 candidates were refuted (deterministic-allowed sleeps, factually-wrong premises, fixes needing risky production changes) and left. Highlights: - Packages/ unit tests (missed by the first 44-file sweep): AuthCoordinator and HostBrowserSignIn now drive timeouts off an injected ManualTestClock instead of real ContinuousClock; ChatConversationStore poller results are asserted with #expect instead of discarded (fail-loud); RemoteProxyBroker / RemoteCLIRelay / CommandPaletteSearchEngine timing/race tightened. - Integration tests: per-PID/UUID socket+port paths (terminal_focus_routing, ctrl_socket, ssh_remote_* disjoint port ranges, CMUXCLIErrorOutput socket) to kill cross-run collisions; order-dependent session-restore checks now scan all workspaces instead of trusting seeding order; latent NameErrors fixed (pane_resize missing `import time` / `must`); fail-open waits made loud (new_tab tmp-write now raises); surface-targeted send to avoid focus races (tab_dragging); fd/temp-file leaks closed. - App-target Swift: removed/loosened flaky wall-clock asserts while keeping the behavior assert (TerminalAndGhostty paste, ShellStartupMatrix budget), widened burst spacing (GhosttyNotificationDispatcher), and routed UITest socket/condition waits through the shared waitForControlSocketReady helper. All Python pass py_compile; all 4 touched packages pass swift build --build-tests; app-target Swift is CI-compiled. No test run against a socket. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7a374ac84c2c96d175df7656b9fd6657a788ea68) * Fix MobileCoreRPC cancelled-while-queued race deterministically (DEBUG-only seam) The remaining discovery finding: MobileCoreRPCClientTests spun a fixed 100 `Task.yield()`s to "wait" for a queued request to reach the session writer gate before cancelling it. Under scheduler load the queued task may not have reached the gate yet, so cancellation fires before `session.send` registers it and the cancelled-while-queued invariant is never exercised (false pass). That gate state (`queuedRequestIDs`) is private to the production `MobileCoreRPCSession` actor, so there is no deterministic test-only signal. Fix adds a `#if DEBUG`, read-only `debugQueuedRequestCount()` to the session and a thin client wrapper, placed in the file's existing `#if DEBUG` test-support extension (alongside `debugWithRequestTimeout`). The test now polls that real signal until the request is registered at the gate, then cancels. The accessor is `#if DEBUG` and read-only: it is compiled out of Release builds and changes no shipping-build behavior, so it needs no dogfood. `swift build --build-tests` (debug) compiles and links the package + test target cleanly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit c56f53ece49ae1c98072bb24996c55e6d15b4fa6) * Refresh Swift file-length budget for the de-flaked test files The fail-loud guards and ManualTestClock/RunLoop-poll rewrites grew five test files past their budgets (NotificationAndMenuBar +31, CommandPaletteSearchEngine +21, MultiWindowNotifications +9, BrowserPaneNavigationKeybind +3, AgentHibernation +1) and shrank two (TerminalAndGhostty -7 after dropping a flaky wall-clock assert, BrowserFixtureInteraction -6 after routing through the shared readiness helper). Budget updated for exactly those seven files. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: drop reintroduced wall-clock asserts / fixed sleeps CodeRabbit/Greptile caught spots where the second-pass fixes leaned on the very patterns this PR is removing. All test-only: - ShellStartupMatrixTests: drop the redundant `duration < 5.0` ceiling; the bootstrap already runs under a 5s `runProcess` timeout, so status==0 + !timedOut is the causal completion signal (no wall-clock assert). - MultiWindowNotificationsUITests: assert no-foreground at the causal point (right after `waitForCommandCompletionWhileBackgrounded`) instead of polling a fixed 2s window. - test_tab_dragging.py: remove two `time.sleep(1.5)` waits; the file-content poll loop right below is the readiness signal. - test_terminal_focus_routing.py: `tempfile.mktemp` -> `mkstemp` (atomic, no TOCTOU; clears Ruff S306). - test_session_restore_stress_kill_cycles.py: replace the fixed 0.1s settle with a deadline-bounded poll on the real is-selected signal. - RemoteCLIRelayServerTests: capture errno before close() so thrown bind/listen diagnostics report the real failure code. py_compile + swift build --build-tests (CmuxRemoteWorkspace) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Make CommandPalette search benchmarks advisory instead of wall-clock-gated Greptile/CodeRabbit correctly flagged that loosening these optimized-vs-reference ratio asserts (e.g. 0.80 -> 1.25) no longer validates anything: a band tight enough to prove the optimized path is faster flakes on shared CI, and a band wide enough not to flake passes even when the optimized path regresses. The engine exposes no preparation/work counter, so there is no causal (non-wall- clock) signal to assert on here. Per the repo test-time policy (no wall-clock latency asserts on shared CI) and matching the activation-session perf gate's advisory-timing approach, drop the flaky `#expect` ratio/dropped-frame assertions across all four benchmarks and keep the `BENCH ...` diagnostic prints for trend tracking. Each test still exercises both code paths end to end; real activation latency/frame-budget regressions are gated by the dedicated activation-session job. swift build --build-tests (CmuxCommandPalette) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * MobileCoreRPC: observe queue gate via @testable, not a production debug hook Per review (Aziz): test/debug extensions should not live in production source. Remove the `#if DEBUG debugQueuedRequestCount()` accessor from MobileCoreRPCClient/Session and instead widen `session` and `queuedRequestIDs` from `private` to `internal` so the cancellation test reads the writer-gate state directly through its existing `@testable import CmuxMobileRPC`. All test scaffolding now lives in the test target; production source carries only the two access-level changes (still module-internal, no shipping behavior change). swift build --build-tests (CmuxMobileRPC) clean; no debug funcs remain in Sources/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Eliminate flaky tests: deterministic rewrites across Swift, Python, shell (#6392) * Fix flaky tests across Swift, Python, and shell suites Replace nondeterministic test patterns with waits on real signals so the same correct code stops failing under CI/VM load. Found by a parallel static audit of ~950 test files; 61 fixes across 44 files. - sleep-as-sync: fixed sleep + single assert replaced with a bounded poll on the actual readiness signal (existing in-file wait helpers preferred). - tight-timeout: expectation/wait timeouts widened to generous bounds; only the failure path waits longer, passing runs return as soon as the signal fires. - wall-clock-assert: single-shot timing ratios replaced with best-of-N minima plus deterministic work-count assertions; perf benchmarks no longer flip on a single scheduler spike. - order-dependence: shared static recorder suite serialized; shared state reset per test. - network/port races: external endpoints swapped for local/ephemeral servers; real-network download hardened with retries, timeouts, and a transient-failure soft-skip. Deterministic test sleeps (allowed by the repo review rules) are left intact. All changed Python/shell files pass py_compile/bash -n; all five changed SPM package test targets compile via swift build --build-tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for de-flaked test files The poll helpers and best-of-N scaffolding that replaced fixed sleeps added a few lines to 6 Swift test files, crossing their per-file budgets. Update only those rows (and track HostBrowserSignInFlowTests, newly over 500). No other files touched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Deflake activation-session perf gate: best-of-N timing, advisory budgets The activation benchmark failed nondeterministically on snapshot_with_scrollback .elapsed_ms (e.g. 1827ms > 1500ms): a single wall-clock measurement of a load-sensitive ~1.3M-char snapshot, compared to a hard absolute ceiling, run on shared/Depot GUI runners. A loaded runner trivially adds 20%+, so the same correct code flipped red across unrelated PRs. Two changes, matching the de-flake philosophy used elsewhere in this PR (performance is measured, not gated on a fixed wall-clock ceiling): - best_of_snapshot_timing: re-measure each snapshot timing N times (--budget-snapshot-samples, default 3; persist=False so reruns have no session side effects) and keep the MINIMUM elapsed_ms. A real regression slows every sample so the min still regresses; transient contention only inflates some samples, so the min is stable. Shape/char counts are preserved. - Wall-clock timing budgets (launch/restore socket-ready, snapshot elapsed_ms) are now advisory: reported as budget_warnings, non-blocking. Deterministic structural budgets (min scrollback chars, min terminal surfaces, restored workspace/terminal counts) stay blocking. Pass --fail-on-timing-budget to hard-gate timing if a dedicated, controlled perf runner is used. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Rewrite remaining flaky tests to logic/signal/virtual-clock based Eliminates the 37 residual time-dependent test primitives the determinism checker flags (27 sleep-then-assert, 10 assert-on-duration) across 17 files, applying the two principles: - Assert on causality, not latency: duration assertions (elapsed < N, Date().timeIntervalSince < N) replaced with assertions on the logical outcome the timing proxied (operation completed / right value / right state / right order), or a generous deadline-bounded wait on the real completion signal where a "did not hang" guard is genuinely needed. - Invert the time dependency: sleep-then-assert replaced with a wait on the real signal: an injected virtual clock advanced by hand (e.g. ManualTestClock via existing makeHarness(clock:) seam), a completion signal awaited directly (semaphore/continuation/main-queue drain), or a deadline-bounded poll of the real state predicate where the system exposes no completion event. No sleeps or timeouts were merely widened; no assertions were weakened. The determinism checker now reports 0 active findings. All four changed SPM package test targets compile via swift build --build-tests; Python files pass py_compile. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for rewritten test files The deterministic rewrites (virtual clocks, poll helpers, signal waits) added lines to 7 Swift test files. Bump only those rows. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Eliminate flaky tests: deterministic rewrites across Swift, Python, shell (#6392) * Fix flaky tests across Swift, Python, and shell suites Replace nondeterministic test patterns with waits on real signals so the same correct code stops failing under CI/VM load. Found by a parallel static audit of ~950 test files; 61 fixes across 44 files. - sleep-as-sync: fixed sleep + single assert replaced with a bounded poll on the actual readiness signal (existing in-file wait helpers preferred). - tight-timeout: expectation/wait timeouts widened to generous bounds; only the failure path waits longer, passing runs return as soon as the signal fires. - wall-clock-assert: single-shot timing ratios replaced with best-of-N minima plus deterministic work-count assertions; perf benchmarks no longer flip on a single scheduler spike. - order-dependence: shared static recorder suite serialized; shared state reset per test. - network/port races: external endpoints swapped for local/ephemeral servers; real-network download hardened with retries, timeouts, and a transient-failure soft-skip. Deterministic test sleeps (allowed by the repo review rules) are left intact. All changed Python/shell files pass py_compile/bash -n; all five changed SPM package test targets compile via swift build --build-tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for de-flaked test files The poll helpers and best-of-N scaffolding that replaced fixed sleeps added a few lines to 6 Swift test files, crossing their per-file budgets. Update only those rows (and track HostBrowserSignInFlowTests, newly over 500). No other files touched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Deflake activation-session perf gate: best-of-N timing, advisory budgets The activation benchmark failed nondeterministically on snapshot_with_scrollback .elapsed_ms (e.g. 1827ms > 1500ms): a single wall-clock measurement of a load-sensitive ~1.3M-char snapshot, compared to a hard absolute ceiling, run on shared/Depot GUI runners. A loaded runner trivially adds 20%+, so the same correct code flipped red across unrelated PRs. Two changes, matching the de-flake philosophy used elsewhere in this PR (performance is measured, not gated on a fixed wall-clock ceiling): - best_of_snapshot_timing: re-measure each snapshot timing N times (--budget-snapshot-samples, default 3; persist=False so reruns have no session side effects) and keep the MINIMUM elapsed_ms. A real regression slows every sample so the min still regresses; transient contention only inflates some samples, so the min is stable. Shape/char counts are preserved. - Wall-clock timing budgets (launch/restore socket-ready, snapshot elapsed_ms) are now advisory: reported as budget_warnings, non-blocking. Deterministic structural budgets (min scrollback chars, min terminal surfaces, restored workspace/terminal counts) stay blocking. Pass --fail-on-timing-budget to hard-gate timing if a dedicated, controlled perf runner is used. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Rewrite remaining flaky tests to logic/signal/virtual-clock based Eliminates the 37 residual time-dependent test primitives the determinism checker flags (27 sleep-then-assert, 10 assert-on-duration) across 17 files, applying the two principles: - Assert on causality, not latency: duration assertions (elapsed < N, Date().timeIntervalSince < N) replaced with assertions on the logical outcome the timing proxied (operation completed / right value / right state / right order), or a generous deadline-bounded wait on the real completion signal where a "did not hang" guard is genuinely needed. - Invert the time dependency: sleep-then-assert replaced with a wait on the real signal: an injected virtual clock advanced by hand (e.g. ManualTestClock via existing makeHarness(clock:) seam), a completion signal awaited directly (semaphore/continuation/main-queue drain), or a deadline-bounded poll of the real state predicate where the system exposes no completion event. No sleeps or timeouts were merely widened; no assertions were weakened. The determinism checker now reports 0 active findings. All four changed SPM package test targets compile via swift build --build-tests; Python files pass py_compile. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for rewritten test files The deterministic rewrites (virtual clocks, poll helpers, signal waits) added lines to 7 Swift test files. Bump only those rows. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Flaky-test follow-up: review fixes + second-pass team sweep (#6452) * Address review: fail-loud guards + scoped soft-skips in de-flaked tests - test_browser_goto_split.py: _wait_url_loaded now raises on timeout instead of silently returning (false pass), and the local HTTP server is managed via a "with" block so the thread/socket can't leak on setup exceptions. - test_surface_move_reorder_api.py: assert the cleanup re-select of ws0 actually converges instead of ignoring the _wait() result (state leak across runs). - test_homebrew_sha.sh: only soft-skip transient transport failures (000/408/429/5xx); fail hard on deterministic client errors like 404 (missing release asset). - test_visual_typing_char_by_char.py: validate the typed glyph against the post-prompt region of the last line, so a "cmux" already in the prompt/path can't satisfy the check; robust to trailing zsh-autosuggestion glyphs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 5d4c282535a497ad263128b21f8805a34556f55d) * Address review round 2: critical walrus bug + 3 fail-loud/correctness guards - test_pane_break_swap_preserve_focus.py: fix NameError -- the lambda referenced `p` in the ternary condition before the walrus assigned it. Bind the walrus inside the condition so the predicate reads panes once and returns them. - perf-activation-session.py: a missing measurement (actual is None) is now always a blocking failure, not advisory -- absent data is a benchmark-contract violation, distinct from load-sensitive over-budget timing. - perf-activation-session.py: copy the real-scrollback measurement before storing it under snapshot_with_scrollback, so best_of_snapshot_timing's in-place mutation no longer clobbers the raw snapshot_with_real_scrollback capture. - NotificationAndMenuBarTests.swift: on the stall-timeout path, fail fast via a lock-guarded result box instead of awaiting evaluationTask.value, which could hang until the suite timeout when the hook ignores cancellation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7e75ad467bc95cef0436236d94183f43b9aaffb2) * Add reports.md: consolidated flaky-test discovery (in-session workflow) Output of a parallel dynamic workflow (20 agents over the 419 signal-bearing test files of 2080) verifying residual flakiness on this already-de-flaked branch. 101 candidate findings across 86 files, dominated by wall-clock timing asserts, sleep-as-sync, and async races -- concentrated in Packages/ unit tests the first 44-file sweep did not cover. Each entry is a candidate; the fix phase adversarially verifies every one before changing code (repo rules allow deterministic test sleeps, so those are left alone). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit e1952d01faee805298bdff215af1d0897f63abcf) * Fix 37 verified residual flaky tests (coordinated area-team workflow) Second-pass de-flake driven by the in-session discovery + team-fix workflows (see reports.md). Adversarial per-area verification confirmed genuine flakiness before any edit; 9 candidates were refuted (deterministic-allowed sleeps, factually-wrong premises, fixes needing risky production changes) and left. Highlights: - Packages/ unit tests (missed by the first 44-file sweep): AuthCoordinator and HostBrowserSignIn now drive timeouts off an injected ManualTestClock instead of real ContinuousClock; ChatConversationStore poller results are asserted with #expect instead of discarded (fail-loud); RemoteProxyBroker / RemoteCLIRelay / CommandPaletteSearchEngine timing/race tightened. - Integration tests: per-PID/UUID socket+port paths (terminal_focus_routing, ctrl_socket, ssh_remote_* disjoint port ranges, CMUXCLIErrorOutput socket) to kill cross-run collisions; order-dependent session-restore checks now scan all workspaces instead of trusting seeding order; latent NameErrors fixed (pane_resize missing `import time` / `must`); fail-open waits made loud (new_tab tmp-write now raises); surface-targeted send to avoid focus races (tab_dragging); fd/temp-file leaks closed. - App-target Swift: removed/loosened flaky wall-clock asserts while keeping the behavior assert (TerminalAndGhostty paste, ShellStartupMatrix budget), widened burst spacing (GhosttyNotificationDispatcher), and routed UITest socket/condition waits through the shared waitForControlSocketReady helper. All Python pass py_compile; all 4 touched packages pass swift build --build-tests; app-target Swift is CI-compiled. No test run against a socket. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit 7a374ac84c2c96d175df7656b9fd6657a788ea68) * Fix MobileCoreRPC cancelled-while-queued race deterministically (DEBUG-only seam) The remaining discovery finding: MobileCoreRPCClientTests spun a fixed 100 `Task.yield()`s to "wait" for a queued request to reach the session writer gate before cancelling it. Under scheduler load the queued task may not have reached the gate yet, so cancellation fires before `session.send` registers it and the cancelled-while-queued invariant is never exercised (false pass). That gate state (`queuedRequestIDs`) is private to the production `MobileCoreRPCSession` actor, so there is no deterministic test-only signal. Fix adds a `#if DEBUG`, read-only `debugQueuedRequestCount()` to the session and a thin client wrapper, placed in the file's existing `#if DEBUG` test-support extension (alongside `debugWithRequestTimeout`). The test now polls that real signal until the request is registered at the gate, then cancels. The accessor is `#if DEBUG` and read-only: it is compiled out of Release builds and changes no shipping-build behavior, so it needs no dogfood. `swift build --build-tests` (debug) compiles and links the package + test target cleanly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> (cherry picked from commit c56f53ece49ae1c98072bb24996c55e6d15b4fa6) * Refresh Swift file-length budget for the de-flaked test files The fail-loud guards and ManualTestClock/RunLoop-poll rewrites grew five test files past their budgets (NotificationAndMenuBar +31, CommandPaletteSearchEngine +21, MultiWindowNotifications +9, BrowserPaneNavigationKeybind +3, AgentHibernation +1) and shrank two (TerminalAndGhostty -7 after dropping a flaky wall-clock assert, BrowserFixtureInteraction -6 after routing through the shared readiness helper). Budget updated for exactly those seven files. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Address review: drop reintroduced wall-clock asserts / fixed sleeps CodeRabbit/Greptile caught spots where the second-pass fixes leaned on the very patterns this PR is removing. All test-only: - ShellStartupMatrixTests: drop the redundant `duration < 5.0` ceiling; the bootstrap already runs under a 5s `runProcess` timeout, so status==0 + !timedOut is the causal completion signal (no wall-clock assert). - MultiWindowNotificationsUITests: assert no-foreground at the causal point (right after `waitForCommandCompletionWhileBackgrounded`) instead of polling a fixed 2s window. - test_tab_dragging.py: remove two `time.sleep(1.5)` waits; the file-content poll loop right below is the readiness signal. - test_terminal_focus_routing.py: `tempfile.mktemp` -> `mkstemp` (atomic, no TOCTOU; clears Ruff S306). - test_session_restore_stress_kill_cycles.py: replace the fixed 0.1s settle with a deadline-bounded poll on the real is-selected signal. - RemoteCLIRelayServerTests: capture errno before close() so thrown bind/listen diagnostics report the real failure code. py_compile + swift build --build-tests (CmuxRemoteWorkspace) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Make CommandPalette search benchmarks advisory instead of wall-clock-gated Greptile/CodeRabbit correctly flagged that loosening these optimized-vs-reference ratio asserts (e.g. 0.80 -> 1.25) no longer validates anything: a band tight enough to prove the optimized path is faster flakes on shared CI, and a band wide enough not to flake passes even when the optimized path regresses. The engine exposes no preparation/work counter, so there is no causal (non-wall- clock) signal to assert on here. Per the repo test-time policy (no wall-clock latency asserts on shared CI) and matching the activation-session perf gate's advisory-timing approach, drop the flaky `#expect` ratio/dropped-frame assertions across all four benchmarks and keep the `BENCH ...` diagnostic prints for trend tracking. Each test still exercises both code paths end to end; real activation latency/frame-budget regressions are gated by the dedicated activation-session job. swift build --build-tests (CmuxCommandPalette) clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * MobileCoreRPC: observe queue gate via @testable, not a production debug hook Per review (Aziz): test/debug extensions should not live in production source. Remove the `#if DEBUG debugQueuedRequestCount()` accessor from MobileCoreRPCClient/Session and instead widen `session` and `queuedRequestIDs` from `private` to `internal` so the cancellation test reads the writer-gate state directly through its existing `@testable import CmuxMobileRPC`. All test scaffolding now lives in the test target; production source carries only the two access-level changes (still module-internal, no shipping behavior change). swift build --build-tests (CmuxMobileRPC) clean; no debug funcs remain in Sources/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Eliminate flaky tests: deterministic rewrites across Swift, Python, shell (#6392) * Fix flaky tests across Swift, Python, and shell suites Replace nondeterministic test patterns with waits on real signals so the same correct code stops failing under CI/VM load. Found by a parallel static audit of ~950 test files; 61 fixes across 44 files. - sleep-as-sync: fixed sleep + single assert replaced with a bounded poll on the actual readiness signal (existing in-file wait helpers preferred). - tight-timeout: expectation/wait timeouts widened to generous bounds; only the failure path waits longer, passing runs return as soon as the signal fires. - wall-clock-assert: single-shot timing ratios replaced with best-of-N minima plus deterministic work-count assertions; perf benchmarks no longer flip on a single scheduler spike. - order-dependence: shared static recorder suite serialized; shared state reset per test. - network/port races: external endpoints swapped for local/ephemeral servers; real-network download hardened with retries, timeouts, and a transient-failure soft-skip. Deterministic test sleeps (allowed by the repo review rules) are left intact. All changed Python/shell files pass py_compile/bash -n; all five changed SPM package test targets compile via swift build --build-tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for de-flaked test files The poll helpers and best-of-N scaffolding that replaced fixed sleeps added a few lines to 6 Swift test files, crossing their per-file budgets. Update only those rows (and track HostBrowserSignInFlowTests, newly over 500). No other files touched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Deflake activation-session perf gate: best-of-N timing, advisory budgets The activation benchmark failed nondeterministically on snapshot_with_scrollback .elapsed_ms (e.g. 1827ms > 1500ms): a single wall-clock measurement of a load-sensitive ~1.3M-char snapshot, compared to a hard absolute ceiling, run on shared/Depot GUI runners. A loaded runner trivially adds 20%+, so the same correct code flipped red across unrelated PRs. Two changes, matching the de-flake philosophy used elsewhere in this PR (performance is measured, not gated on a fixed wall-clock ceiling): - best_of_snapshot_timing: re-measure each snapshot timing N times (--budget-snapshot-samples, default 3; persist=False so reruns have no session side effects) and keep the MINIMUM elapsed_ms. A real regression slows every sample so the min still regresses; transient contention only inflates some samples, so the min is stable. Shape/char counts are preserved. - Wall-clock timing budgets (launch/restore socket-ready, snapshot elapsed_ms) are now advisory: reported as budget_warnings, non-blocking. Deterministic structural budgets (min scrollback chars, min terminal surfaces, restored workspace/terminal counts) stay blocking. Pass --fail-on-timing-budget to hard-gate timing if a dedicated, controlled perf runner is used. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Rewrite remaining flaky tests to logic/signal/virtual-clock based Eliminates the 37 residual time-dependent test primitives the determinism checker flags (27 sleep-then-assert, 10 assert-on-duration) across 17 files, applying the two principles: - Assert on causality, not latency: duration assertions (elapsed < N, Date().timeIntervalSince < N) replaced with assertions on the logical outcome the timing proxied (operation completed / right value / right state / right order), or a generous deadline-bounded wait on the real completion signal where a "did not hang" guard is genuinely needed. - Invert the time dependency: sleep-then-assert replaced with a wait on the real signal: an injected virtual clock advanced by hand (e.g. ManualTestClock via existing makeHarness(clock:) seam), a completion signal awaited directly (semaphore/continuation/main-queue drain), or a deadline-bounded poll of the real state predicate where the system exposes no completion event. No sleeps or timeouts were merely widened; no assertions were weakened. The determinism checker now reports 0 active findings. All four changed SPM package test targets compile via swift build --build-tests; Python files pass py_compile. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for rewritten test files The deterministic rewrites (virtual clocks, poll helpers, signal waits) added lines to 7 Swift test files. Bump only those rows. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Fix frozen terminals after split churn (#12) * Fix blank terminal after split operations and add visual tests ## Blank Terminal Fix - Add `needsRefreshAfterWindowChange` flag in GhosttyTerminalView - Force terminal refresh when view is added to window, even if size unchanged - Add `ghostty_surface_refresh()` call in attachToView for same-view reattachment - Add debug logging for surface attachment lifecycle (DEBUG builds only) ## Bonsplit Migration - Add bonsplit as local Swift package (vendor/bonsplit submodule) - Replace custom SplitTree with BonsplitController - Add Panel protocol with TerminalPanel and BrowserPanel implementations - Add SidebarTab as main tab container with BonsplitController - Remove old Splits/ directory (SplitTree, SplitView, TerminalSplitTreeView) ## Visual Screenshot Tests - Add test_visual_screenshots.py for automated visual regression testing - Uses in-app screenshot API (CGWindowListCreateImage) - no screen recording needed - Generates HTML report with before/after comparisons - Tests: splits, browser panels, focus switching, close operations, rapid cycles - Includes annotation fields for easy feedback ## Browser Shortcut (⌘⇧B) - Add keyboard shortcut to open browser panel in current pane - Add openBrowser() method to TabManager - Add shortcut configuration in KeyboardShortcutSettings ## Screenshot Command - Add 'screenshot' command to TerminalController for in-app window capture - Returns OK with screenshot ID and path ## Other - Add tests/visual_output/ and tests/visual_report.html to .gitignore * Add browser title subscription and set tab height to 30px - Subscribe to BrowserPanel.$pageTitle changes to update bonsplit tabs - Update tab titles in real-time as page navigation occurs - Clean up subscriptions when panels are removed - Set bonsplit tab bar and tab height to 30px (in submodule) * Fix socket API regressions in list_surfaces, list_bonsplit_tabs, focus_pane - list_surfaces: Remove [terminal]/[browser] suffix to keep UUID-only format that clients and tests expect for parsing - list_bonsplit_tabs --pane: Properly look up pane by UUID instead of creating a new PaneID (requires bonsplit PaneID.id to be public) - focus_pane: Accept both UUID strings and integer indices as documented * Fix browser panel stability and keyboard shortcuts - Prevent WKWebView focus lifecycle crashes during split/view reshuffles - Match bracket shortcuts via keyCode (Cmd+Shift+[ / ], Cmd+Ctrl+[ / ]) - Support Ghostty config goto_split:* keybinds when WebView is focused - Add focus_webview/is_webview_focused socket commands and regression tests - Rename SidebarTab to Workspace and update docs * Make ctrl+enter keybind test skippable Skip when the Ghostty keybind isn't configured or when osascript can't send keystrokes (no Accessibility permission), so VM runs stay green. * Auto-focus browser omnibar when blank When a browser surface is focused but no URL is loaded yet, focus the address bar instead of the WKWebView. * Stabilize socket surface indexing * Focus browser omnibar escape; add webview keybind UI tests - Escape in omnibar now returns focus to WKWebView\n- Add UI tests for Cmd+Ctrl+H pane navigation with WebKit focused (including Ghostty config)\n- Avoid flaky element screenshots in UpdatePillUITests on the UTM VM * Fix browser drag-to-split blanks and socket parsing * Fix webview-focused shortcuts and stabilize browser splits - Match ctrl/shift shortcuts by keyCode where needed (Ctrl+H, bracket keys) - Load Ghostty goto_split triggers reliably and refresh on config load - Add debug socket helpers: set_shortcut + simulate_shortcut for tests - Convert browser goto_split/keybind tests to socket-based injection (no osascript) - Bump bonsplit for drag-to-split fixes * Fix split layout collapse and harden socket pane APIs * Stabilize OSC 99 notification test timing * Fix terminal focus routing after split reparent * Support simulate_shortcut enter for focus routing test * Stabilize terminal focus routing test * Fix frozen new terminal tabs after many splits * Fix frozen new terminal tabs after splits * Fix terminal freeze on launch/new tabs * Update ghostty submodule * Fix terminal focus/render stalls after split churn * Fix nested split collapsing existing pane * Fix nested split collapse + stabilize new-surface focus * Update bonsplit submodule * Fix SIGINT test flake * Remove bonsplit tab-switch crossfade * Remove PROJECTS.md * Remove bonsplit tab selection animation * Ignore generated test reports * Middle click closes tab * Revert unintended .gitignore change * Fix build after main merge * Revert "Fix build after main merge" This reverts commit 16bf9816d0856b5385d52f886aa5eb50f3c9d9a4. * Revert "Merge remote-tracking branch 'origin/main' into fix/blank-terminal-and-visual-tests" This reverts commit 7c20fb53fd71fea7a19a3673f2dd73e5f0c783c4, reversing changes made to 0aff107d787bc9d8bbc28220090b4ca7af72e040. * Remove tab close fade animation * Use terminal.fill icon * Make terminal tab icon smaller * Match browser globe tab icon size * Bonsplit: tab min width 48 and tighter close button * Bonsplit: smaller tab title font * Show unread notification badge in bonsplit tabs and improve UI polish Sync unread notification state to bonsplit tab badges (blue dot). Improve EmptyPanelView with Terminal/Browser buttons and shortcut hints. Add tooltips to close tab button and search overlay buttons. * Fix reload.sh single-instance safety check on macOS Replace GNU-only `ps -o etimes=` with portable `ps -o etime=` and parse the dd-hh:mm:ss format manually for macOS compatibility. * Centralize keyboard shortcut definitions into Action enum Replace per-shortcut boilerplate with a single Action enum that holds the label, defaults key, and default binding for each shortcut. All call sites now use shortcut(for:). Settings UI is data-driven via ForEach(Action.allCases). Titlebar tooltips update dynamically when shortcuts are changed. Remove duplicate .keyboardShortcut() modifiers from menu items that are already handled by the event monitor. * Fix WKWebView consuming app menu shortcuts and close panel confirmation Add CmuxWebView subclass that routes key equivalents through the main menu before WebKit, so Cmd+N/Cmd+W/tab switching work when a browser pane is focused. Fix Cmd+W close-panel path: bypass Bonsplit delegate gating after the user confirms the running-process dialog by tracking forceCloseTabIds. Add unit tests (CmuxWebViewKeyEquivalentTests) and UI test scaffolding (MenuKeyEquivalentRoutingUITests) with a new cmux-unit Xcode scheme. * Update CLAUDE.md and PROJECTS.md with recent changes CLAUDE.md: enforce --tag for reload commands, add cleanup safety rules. PROJECTS.md: log notification badge, reload.sh fix, Cmd+W fix, WebView key equiv fix, and centralized shortcuts work. * Keep selection index stable on close * Add concepts page documenting terminology hierarchy New docs page explaining Window > Workspace > Pane > Surface > Panel hierarchy with aligned ASCII diagram. Updated tabs.mdx and splits.mdx to use consistent terminology (workspace instead of tab, surface instead of panel) and corrected outdated CLI command references. * Update bonsplit submodule * WIP: improve split close stability and UI regressions * Close terminal panel on child exit; hide terminal dirty dot * Fix split close/focus regressions and stabilize UI tests * Add unread Dock/Cmd+Tab badge with settings toggle * Fix browser-surface shortcuts and Cmd+L browser opening * Snapshot current workspace state before regression fixes * Update bonsplit submodule snapshot * Stabilize split-close regression capture and sidebar resize assertions * Change default Show Notifications shortcut from Cmd+Shift+I to Cmd+I * Fix update check readiness race, enable release update logging, and improve checking spinner * Restore terminal file drop, fix browser omnibar click focus, and add panel workspace ID mutation for surface moves * Add Cmd+digit workspace hints, titlebar shortcut pills, sidebar drag-reorder, and workspace placement settings * Add v2 browser automation API, surface move/reorder commands, and short-handle ref system to TerminalController * Add CLI browser command surface, --id-format flag, and move/reorder commands * Extend test clients with move/reorder APIs, ref-handle support, and increased timeouts * Harden test runner scripts with deterministic builds, retry logic, and robust socket readiness * Stabilize existing test suites with focus-wait helpers, increased timeouts, and API shape updates * Add terminal file drop e2e regression test * Add v2 browser API, CLI ref resolution, and surface move/reorder test suites * Add unit tests for shortcut hints, workspace reorder, drop planner, and update UI test stabilization * Add cmux-debug-windows skill with snapshot script and agent config * Update project docs: mark browser parity and move/reorder phases complete, add parallel agent workflow guidelines * Update bonsplit submodule: re-entrant setPosition guard, tab shortcut hints, and moveTab/reorderTab API * Add browser agent UX improvements: snapshot refs, placement reuse, diagnostics, and skill docs - Upgrade browser.snapshot to emit accessibility tree text with element refs (eN) - Add right-sibling pane reuse policy for browser.open_split placement - Add rich not_found diagnostics with retry logic for selector actions - Support --snapshot-after for post-action verification on mutating commands - Allow browser fill with empty text for clearing inputs - Default CLI --id-format to refs-first (UUIDs opt-in via --id-format uuids|both) - Format legacy new-pane/new-surface output with short surface refs - Add skills/cmuxterm-browser/ and skills/cmuxterm/ end-user skill docs - Add regression tests for placement policy, snapshot refs, diagnostics, and ID defaults * Update bonsplit submodule: keep raster favicons in color when inactive | 7 个月前 | |
Eliminate flaky tests: deterministic rewrites across Swift, Python, shell (#6392) * Fix flaky tests across Swift, Python, and shell suites Replace nondeterministic test patterns with waits on real signals so the same correct code stops failing under CI/VM load. Found by a parallel static audit of ~950 test files; 61 fixes across 44 files. - sleep-as-sync: fixed sleep + single assert replaced with a bounded poll on the actual readiness signal (existing in-file wait helpers preferred). - tight-timeout: expectation/wait timeouts widened to generous bounds; only the failure path waits longer, passing runs return as soon as the signal fires. - wall-clock-assert: single-shot timing ratios replaced with best-of-N minima plus deterministic work-count assertions; perf benchmarks no longer flip on a single scheduler spike. - order-dependence: shared static recorder suite serialized; shared state reset per test. - network/port races: external endpoints swapped for local/ephemeral servers; real-network download hardened with retries, timeouts, and a transient-failure soft-skip. Deterministic test sleeps (allowed by the repo review rules) are left intact. All changed Python/shell files pass py_compile/bash -n; all five changed SPM package test targets compile via swift build --build-tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for de-flaked test files The poll helpers and best-of-N scaffolding that replaced fixed sleeps added a few lines to 6 Swift test files, crossing their per-file budgets. Update only those rows (and track HostBrowserSignInFlowTests, newly over 500). No other files touched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Deflake activation-session perf gate: best-of-N timing, advisory budgets The activation benchmark failed nondeterministically on snapshot_with_scrollback .elapsed_ms (e.g. 1827ms > 1500ms): a single wall-clock measurement of a load-sensitive ~1.3M-char snapshot, compared to a hard absolute ceiling, run on shared/Depot GUI runners. A loaded runner trivially adds 20%+, so the same correct code flipped red across unrelated PRs. Two changes, matching the de-flake philosophy used elsewhere in this PR (performance is measured, not gated on a fixed wall-clock ceiling): - best_of_snapshot_timing: re-measure each snapshot timing N times (--budget-snapshot-samples, default 3; persist=False so reruns have no session side effects) and keep the MINIMUM elapsed_ms. A real regression slows every sample so the min still regresses; transient contention only inflates some samples, so the min is stable. Shape/char counts are preserved. - Wall-clock timing budgets (launch/restore socket-ready, snapshot elapsed_ms) are now advisory: reported as budget_warnings, non-blocking. Deterministic structural budgets (min scrollback chars, min terminal surfaces, restored workspace/terminal counts) stay blocking. Pass --fail-on-timing-budget to hard-gate timing if a dedicated, controlled perf runner is used. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Rewrite remaining flaky tests to logic/signal/virtual-clock based Eliminates the 37 residual time-dependent test primitives the determinism checker flags (27 sleep-then-assert, 10 assert-on-duration) across 17 files, applying the two principles: - Assert on causality, not latency: duration assertions (elapsed < N, Date().timeIntervalSince < N) replaced with assertions on the logical outcome the timing proxied (operation completed / right value / right state / right order), or a generous deadline-bounded wait on the real completion signal where a "did not hang" guard is genuinely needed. - Invert the time dependency: sleep-then-assert replaced with a wait on the real signal: an injected virtual clock advanced by hand (e.g. ManualTestClock via existing makeHarness(clock:) seam), a completion signal awaited directly (semaphore/continuation/main-queue drain), or a deadline-bounded poll of the real state predicate where the system exposes no completion event. No sleeps or timeouts were merely widened; no assertions were weakened. The determinism checker now reports 0 active findings. All four changed SPM package test targets compile via swift build --build-tests; Python files pass py_compile. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for rewritten test files The deterministic rewrites (virtual clocks, poll helpers, signal waits) added lines to 7 Swift test files. Bump only those rows. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Gate visual screenshot harness failures (#4915) Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 4 个月前 | |
Eliminate flaky tests: deterministic rewrites across Swift, Python, shell (#6392) * Fix flaky tests across Swift, Python, and shell suites Replace nondeterministic test patterns with waits on real signals so the same correct code stops failing under CI/VM load. Found by a parallel static audit of ~950 test files; 61 fixes across 44 files. - sleep-as-sync: fixed sleep + single assert replaced with a bounded poll on the actual readiness signal (existing in-file wait helpers preferred). - tight-timeout: expectation/wait timeouts widened to generous bounds; only the failure path waits longer, passing runs return as soon as the signal fires. - wall-clock-assert: single-shot timing ratios replaced with best-of-N minima plus deterministic work-count assertions; perf benchmarks no longer flip on a single scheduler spike. - order-dependence: shared static recorder suite serialized; shared state reset per test. - network/port races: external endpoints swapped for local/ephemeral servers; real-network download hardened with retries, timeouts, and a transient-failure soft-skip. Deterministic test sleeps (allowed by the repo review rules) are left intact. All changed Python/shell files pass py_compile/bash -n; all five changed SPM package test targets compile via swift build --build-tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for de-flaked test files The poll helpers and best-of-N scaffolding that replaced fixed sleeps added a few lines to 6 Swift test files, crossing their per-file budgets. Update only those rows (and track HostBrowserSignInFlowTests, newly over 500). No other files touched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Deflake activation-session perf gate: best-of-N timing, advisory budgets The activation benchmark failed nondeterministically on snapshot_with_scrollback .elapsed_ms (e.g. 1827ms > 1500ms): a single wall-clock measurement of a load-sensitive ~1.3M-char snapshot, compared to a hard absolute ceiling, run on shared/Depot GUI runners. A loaded runner trivially adds 20%+, so the same correct code flipped red across unrelated PRs. Two changes, matching the de-flake philosophy used elsewhere in this PR (performance is measured, not gated on a fixed wall-clock ceiling): - best_of_snapshot_timing: re-measure each snapshot timing N times (--budget-snapshot-samples, default 3; persist=False so reruns have no session side effects) and keep the MINIMUM elapsed_ms. A real regression slows every sample so the min still regresses; transient contention only inflates some samples, so the min is stable. Shape/char counts are preserved. - Wall-clock timing budgets (launch/restore socket-ready, snapshot elapsed_ms) are now advisory: reported as budget_warnings, non-blocking. Deterministic structural budgets (min scrollback chars, min terminal surfaces, restored workspace/terminal counts) stay blocking. Pass --fail-on-timing-budget to hard-gate timing if a dedicated, controlled perf runner is used. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Rewrite remaining flaky tests to logic/signal/virtual-clock based Eliminates the 37 residual time-dependent test primitives the determinism checker flags (27 sleep-then-assert, 10 assert-on-duration) across 17 files, applying the two principles: - Assert on causality, not latency: duration assertions (elapsed < N, Date().timeIntervalSince < N) replaced with assertions on the logical outcome the timing proxied (operation completed / right value / right state / right order), or a generous deadline-bounded wait on the real completion signal where a "did not hang" guard is genuinely needed. - Invert the time dependency: sleep-then-assert replaced with a wait on the real signal: an injected virtual clock advanced by hand (e.g. ManualTestClock via existing makeHarness(clock:) seam), a completion signal awaited directly (semaphore/continuation/main-queue drain), or a deadline-bounded poll of the real state predicate where the system exposes no completion event. No sleeps or timeouts were merely widened; no assertions were weakened. The determinism checker now reports 0 active findings. All four changed SPM package test targets compile via swift build --build-tests; Python files pass py_compile. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Bump Swift file-length budget for rewritten test files The deterministic rewrites (virtual clocks, poll helpers, signal waits) added lines to 7 Swift test files. Bump only those rows. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> | 3 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Avoid idle background terminal surface priming (#4184) * Add regression for idle background surface priming * Avoid idle background terminal surface priming * Simplify background prime guards * test: surface background workspace cleanup failures * Fix background template startup detection * Fix deferred startup predicate return | 4 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 | |
Prefer CMUX_SOCKET_PATH in socket tooling (#3364) * Prefer CMUX_SOCKET_PATH in socket tooling * Fix socket env conflict handling * Reduce socket env helper growth * Clear inherited legacy socket aliases * Tighten socket alias compatibility * Keep CLI help socket-independent * Keep no-socket CLI commands conflict-free * Clear legacy socket from child envs * Keep welcome socket-independent --------- Co-authored-by: Lawrence Chen <lawrencecchen@users.noreply.github.com> | 5 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 4 个月前 | ||
| 3 个月前 | ||
| 6 个月前 | ||
| 4 个月前 | ||
| 4 个月前 | ||
| 5 个月前 | ||
| 3 个月前 | ||
| 5 个月前 | ||
| 2 个月前 | ||
| 2 个月前 | ||
| 5 个月前 | ||
| 3 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 3 个月前 | ||
| 3 个月前 | ||
| 5 个月前 | ||
| 3 个月前 | ||
| 4 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 2 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 4 个月前 | ||
| 3 个月前 | ||
| 4 个月前 | ||
| 5 个月前 | ||
| 4 个月前 | ||
| 3 个月前 | ||
| 3 个月前 | ||
| 5 个月前 | ||
| 3 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 3 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 4 个月前 | ||
| 3 个月前 | ||
| 7 个月前 | ||
| 3 个月前 | ||
| 7 个月前 | ||
| 7 个月前 | ||
| 6 个月前 | ||
| 5 个月前 | ||
| 2 个月前 | ||
| 4 个月前 | ||
| 3 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 3 个月前 | ||
| 5 个月前 | ||
| 6 个月前 | ||
| 3 个月前 | ||
| 3 个月前 | ||
| 3 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 7 个月前 | ||
| 2 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 3 个月前 | ||
| 3 个月前 | ||
| 2 个月前 | ||
| 5 个月前 | ||
| 3 个月前 | ||
| 4 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 3 个月前 | ||
| 4 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 3 个月前 | ||
| 5 个月前 | ||
| 3 个月前 | ||
| 3 个月前 | ||
| 3 个月前 | ||
| 5 个月前 | ||
| 3 个月前 | ||
| 5 个月前 | ||
| 3 个月前 | ||
| 5 个月前 | ||
| 3 个月前 | ||
| 3 个月前 | ||
| 3 个月前 | ||
| 7 个月前 | ||
| 3 个月前 | ||
| 4 个月前 | ||
| 3 个月前 | ||
| 5 个月前 | ||
| 4 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 |