| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
harden(backend,app): client-compatibility replay contract (C10) (#14372) * harden(dev-harness): define V1 live-session and V8 spine contracts Add strict Python/Dart pending contracts and introducing-commit protection. Reserve the live-session CLI, wire, ownership/attachment seams and failing builder acceptance tests. Add live attribution before freezing evidence v1. Verification: - bash scripts/dev-harness/run-tests.sh: 287 passed, 7 existing skips, 40 strict xfailed (24 pending V1 markers), Python 3.11.15. - make preflight: 16 checks passed, using supported PYTHON selection. - make mobile-verify fast --paths harness/session contract inputs with an absolute evidence directory: 20 executed, 20 passed, no failures/skips. - bash app/test.sh: 2014 passed, 5 skipped, one unchanged UTC+7 search_rank_grouping_test failure; affected baseline inputs match main. - Real module CLI refusals exit 2 without traceback; regression covered. - Dart SDK probe proves ext. prefix required; marker-removal formatter probe proves no incidental Dart diff. Runtime is intentionally unimplemented. App UI must repair current invalid VM extension names before live acceptance; unsupported journeys block. Simulator/emulator latency and detached-process acceptance remain V1 work. make setup installed hooks but full backend wheel sync stalled; isolated Python 3.11 harness dependencies ran the owner checks without any skip flag. Failure-Class: none * harden(dev-harness): scope the spine-contracts check to spine paths The check protects files under the spine roots and their registry, so it only needs to run when those change. A repo-wide trigger would run it, and its full-history requirement, on every unrelated contributor PR. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * docs(dev-harness): clarify the live journey adapter boundary V1 must refuse live journey execution until a shared executable adapter exists; B1 addressability alone cannot attach Dart integration tests. Pure admission contracts remain valid with injected supported IDs. Verification: harness 287 passed, 7 skipped, 40 strict xfailed; make preflight passed 20 checks. * harden(dev-harness): pin reviewed spine revisions and align B0 capabilities Coordinator turn-4 review authorizes the V1 fixture correction. Exact before/after digests pin the revised oracle; immutable revision records do not permit builder assertion edits. Active tests reject missing/tampered records and prove the fake machine response. Verification: full harness 289 passed, 7 skipped, 40 strict xfailed; UTC app 2020 passed; spine 24 pending/7 protected; make preflight 16 passed. No live device acceptance claimed. * harden(dev-harness): preserve marker retirement across spine revisions Track retained marker occurrences rather than counts so retiring a second test cannot restore the first test's retired marker. The active regression covers this swap and legitimate continued retirement. Verification: full harness 290 passed, 7 skipped, 40 strict xfailed; spine 24 pending/7 protected; make preflight 16 passed. Dart sources unchanged after the earlier UTC app 2020-pass run. * harden(dev-harness): retain corrected spine oracles through squash merges A squash can introduce a corrected contract and its review records in the same commit. Absorb only the revision prefix ending at that introducing file's exact digest, preserving immutable records and later marker-only edits. Active temp-repo regression covers multiple corrections in one squash and rejects restoring old or altered assertions. Verification: harness 291 passed, 7 skipped, 40 strict xfailed; B1 make preflight 20 passed; final B1 journeys 20/20. Dart unchanged after UTC app 41 opt-in and 2051 ordinary passed. * test(dev-harness): revise live session acceptance after builder review * test(dev-harness): bind machine fixture to actual child identity * test(contracts): define released mobile client compatibility replay C10 pins adopted consumer shapes and release-source provenance, reserves frozen-decoder replay, and leaves capture plus real backend acceptance visibly pending. Build tags are not release evidence. Verified app 2064, harness 309 passed/7 skipped/61 xfailed, spine 41 markers/11 protected, preflight 31 checks, journeys 3/3. Backend E2E stopped at missing fake_firestore prerequisite; no setup or live contact. * harden(dev-harness): scope compatibility replay to supported captures Keep the two-build bootstrap oracle historical while ongoing replay checks exact current router responses for supported cases. Archived captures remain immutable; validated server retirement is explicitly out of scope. Permit append-only coverage bundles for an existing build. Verified: UTC app suite 2064 passed (Dart unchanged); full harness 310 passed, 7 skipped, 61 strict xfailed; spine 41 pending, 11 protected; make preflight 31 passed; fast journeys 3/3. Three historical backend commits pass the adoption-scoped check. Backend E2E remains blocked at its dependency prerequisite; no real-router execution claimed. * harden(dev-harness): separate oracle revisions from implementations Use the actual PR base and immutable exact scaffolding snapshots to reject mixed revision/implementation diffs. Exercise a real squash and child merge. Correct local_dev wire spelling from the simulator RPC and add adversarial V1 conformance oracles after reviewing the first broker. Verification: UTC app 2064 passed; full harness 303 passed, 7 skipped, 55 strict xfailed; spine 37 pending/9 protected; preflight 16 checks passed. Isolated reviewed-broker probes failed 9/9 as documented in the conformance report. No real app/device launch. * fix(ci): repair the checks-manifest merge of main into the client-compat contract branch The coordinator's conflict resolution left a truncated duplicate spine-contracts entry; the pre-push manifest validation caught it. Keep one entry, with this branch's extra trigger path. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * harden(dev-harness): retain pinned corrections after older parent merges * harden(dev-harness): pin oracle content across shared runners and merge orders Keep exact oracle digests and immutable revision records. Traverse full reachable history and match revision payloads by digest. Shared runner metadata requires direct fail-fast suite invocation instead of pinning shared file bytes. Both new regressions fail on the old checker. UTC app 2064 passed; final harness 306 passed, 7 skipped, 59 strict xfailed; spine 39 pending/10 protected; preflight 16 passed. No protected oracle bytes or legacy scope pins changed. * harden(dev-harness): separate replay engine from evidenced release admission Bind distribution receipts to captured identities, pin handwritten decoder sources and exact candidate requests, require original-Flutter extraction equivalence, and execute a real Dart decoder oracle. Preserve unadopted endpoint freedom and report missing OpenAPI reference closure. Verification: app2064; harness326 passed/7 skipped/78 strict xfailed; spine52/15; preflight31 passed. No real release receipts or backend replay claimed. * harden(dev-harness): inventory both literal pending marker quote styles Single-quoted Python contracts executed under pytest but were invisible to inventory and could not retire their markers. Recognize paired quote styles in both languages; reject nonliteral builder markers while allowing active MECHANISM-owned API self-tests. Preserve all oracle bytes and digest pins. Verification: harness328 passed/7 skipped/79 strict xfailed; spine55 pending/16 protected; preflight31 passed; unchanged Dart suite2064 passed. * test(dev-harness): require replay loader to preserve frozen request vectors Compare every admitted Case to its pinned method, ordered query/headers, body, status, decoder and observations. A HEAD-shaped request with the right endpoint name cannot pass. Real capture admission remains pending after synthetic engine delivery. Verification: harness328 passed/7 skipped/79 strict xfailed; spine55/16; preflight31 passed; Dart2064 passed. --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> | 5 天前 | |
feat(dev-harness): opt-in make lane-backend (fast wheel install) for the mobile harness and pre-push gates (#14349) * fix(app): refuse remote APIs and non-dev pairings in hermetic tests app/test.sh skipped writing .dev.env when generated files already existed, so a leftover API_BASE_URL or mobile_beta/prod pairing could silently survive. Fail closed at the checker, test.sh, and doctor without rewriting the file. Failure-Class: new Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dev-harness): add lane-bootstrap that skips the full backend lock A linked worktree can run cheap pre-push gates and app/test.sh after one idempotent command: pinned Flutter, shared PUB_CACHE, Python 3.11 via uv, and yaml+dotenv only. An incomplete .venv directory is no longer treated as ready. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): keep fixture venvs passing the cheap-gate import probe Manifest-contract stubs only accepted `import yaml`. The resolver now probes `import dotenv, yaml` so an empty .venv is not treated as ready, and those stubs were skipped. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dev-harness): make setup-backend the fast wheel install, not pylock sync uv pip sync of pylock.macos.toml lists hashed sdist+wheel for av/llvmlite/scipy/pyarrow and stalled ~20 minutes; uv pip install -r requirements.txt takes index wheels (69s here) and covers uvicorn/pyright/yaml/dotenv/google.auth for session start and the pre-push typecheck. Lock-faithful sync stays in sync-python-deps.sh. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): keep setup-backend as the lock sync; add opt-in lane-backend make setup-backend is the contributor lock-synced entrypoint; pointing it at unlocked requirements.txt broke scripts/test-make-setup.sh. The fast wheel install is make lane-backend. generate-app-env runs build_runner without the flags that deleted tracked manifest.g.dart. Co-authored-by: Cursor <cursoragent@cursor.com> * docs(backend): keep AGENTS.md under the lean budget while naming lane-backend Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> | 6 天前 | |
harden(backend,app): client-compatibility replay contract (C10) (#14372) * harden(dev-harness): define V1 live-session and V8 spine contracts Add strict Python/Dart pending contracts and introducing-commit protection. Reserve the live-session CLI, wire, ownership/attachment seams and failing builder acceptance tests. Add live attribution before freezing evidence v1. Verification: - bash scripts/dev-harness/run-tests.sh: 287 passed, 7 existing skips, 40 strict xfailed (24 pending V1 markers), Python 3.11.15. - make preflight: 16 checks passed, using supported PYTHON selection. - make mobile-verify fast --paths harness/session contract inputs with an absolute evidence directory: 20 executed, 20 passed, no failures/skips. - bash app/test.sh: 2014 passed, 5 skipped, one unchanged UTC+7 search_rank_grouping_test failure; affected baseline inputs match main. - Real module CLI refusals exit 2 without traceback; regression covered. - Dart SDK probe proves ext. prefix required; marker-removal formatter probe proves no incidental Dart diff. Runtime is intentionally unimplemented. App UI must repair current invalid VM extension names before live acceptance; unsupported journeys block. Simulator/emulator latency and detached-process acceptance remain V1 work. make setup installed hooks but full backend wheel sync stalled; isolated Python 3.11 harness dependencies ran the owner checks without any skip flag. Failure-Class: none * harden(dev-harness): scope the spine-contracts check to spine paths The check protects files under the spine roots and their registry, so it only needs to run when those change. A repo-wide trigger would run it, and its full-history requirement, on every unrelated contributor PR. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * docs(dev-harness): clarify the live journey adapter boundary V1 must refuse live journey execution until a shared executable adapter exists; B1 addressability alone cannot attach Dart integration tests. Pure admission contracts remain valid with injected supported IDs. Verification: harness 287 passed, 7 skipped, 40 strict xfailed; make preflight passed 20 checks. * harden(dev-harness): pin reviewed spine revisions and align B0 capabilities Coordinator turn-4 review authorizes the V1 fixture correction. Exact before/after digests pin the revised oracle; immutable revision records do not permit builder assertion edits. Active tests reject missing/tampered records and prove the fake machine response. Verification: full harness 289 passed, 7 skipped, 40 strict xfailed; UTC app 2020 passed; spine 24 pending/7 protected; make preflight 16 passed. No live device acceptance claimed. * harden(dev-harness): preserve marker retirement across spine revisions Track retained marker occurrences rather than counts so retiring a second test cannot restore the first test's retired marker. The active regression covers this swap and legitimate continued retirement. Verification: full harness 290 passed, 7 skipped, 40 strict xfailed; spine 24 pending/7 protected; make preflight 16 passed. Dart sources unchanged after the earlier UTC app 2020-pass run. * harden(dev-harness): retain corrected spine oracles through squash merges A squash can introduce a corrected contract and its review records in the same commit. Absorb only the revision prefix ending at that introducing file's exact digest, preserving immutable records and later marker-only edits. Active temp-repo regression covers multiple corrections in one squash and rejects restoring old or altered assertions. Verification: harness 291 passed, 7 skipped, 40 strict xfailed; B1 make preflight 20 passed; final B1 journeys 20/20. Dart unchanged after UTC app 41 opt-in and 2051 ordinary passed. * test(dev-harness): revise live session acceptance after builder review * test(dev-harness): bind machine fixture to actual child identity * test(contracts): define released mobile client compatibility replay C10 pins adopted consumer shapes and release-source provenance, reserves frozen-decoder replay, and leaves capture plus real backend acceptance visibly pending. Build tags are not release evidence. Verified app 2064, harness 309 passed/7 skipped/61 xfailed, spine 41 markers/11 protected, preflight 31 checks, journeys 3/3. Backend E2E stopped at missing fake_firestore prerequisite; no setup or live contact. * harden(dev-harness): scope compatibility replay to supported captures Keep the two-build bootstrap oracle historical while ongoing replay checks exact current router responses for supported cases. Archived captures remain immutable; validated server retirement is explicitly out of scope. Permit append-only coverage bundles for an existing build. Verified: UTC app suite 2064 passed (Dart unchanged); full harness 310 passed, 7 skipped, 61 strict xfailed; spine 41 pending, 11 protected; make preflight 31 passed; fast journeys 3/3. Three historical backend commits pass the adoption-scoped check. Backend E2E remains blocked at its dependency prerequisite; no real-router execution claimed. * harden(dev-harness): separate oracle revisions from implementations Use the actual PR base and immutable exact scaffolding snapshots to reject mixed revision/implementation diffs. Exercise a real squash and child merge. Correct local_dev wire spelling from the simulator RPC and add adversarial V1 conformance oracles after reviewing the first broker. Verification: UTC app 2064 passed; full harness 303 passed, 7 skipped, 55 strict xfailed; spine 37 pending/9 protected; preflight 16 checks passed. Isolated reviewed-broker probes failed 9/9 as documented in the conformance report. No real app/device launch. * fix(ci): repair the checks-manifest merge of main into the client-compat contract branch The coordinator's conflict resolution left a truncated duplicate spine-contracts entry; the pre-push manifest validation caught it. Keep one entry, with this branch's extra trigger path. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * harden(dev-harness): retain pinned corrections after older parent merges * harden(dev-harness): pin oracle content across shared runners and merge orders Keep exact oracle digests and immutable revision records. Traverse full reachable history and match revision payloads by digest. Shared runner metadata requires direct fail-fast suite invocation instead of pinning shared file bytes. Both new regressions fail on the old checker. UTC app 2064 passed; final harness 306 passed, 7 skipped, 59 strict xfailed; spine 39 pending/10 protected; preflight 16 passed. No protected oracle bytes or legacy scope pins changed. * harden(dev-harness): separate replay engine from evidenced release admission Bind distribution receipts to captured identities, pin handwritten decoder sources and exact candidate requests, require original-Flutter extraction equivalence, and execute a real Dart decoder oracle. Preserve unadopted endpoint freedom and report missing OpenAPI reference closure. Verification: app2064; harness326 passed/7 skipped/78 strict xfailed; spine52/15; preflight31 passed. No real release receipts or backend replay claimed. * harden(dev-harness): inventory both literal pending marker quote styles Single-quoted Python contracts executed under pytest but were invisible to inventory and could not retire their markers. Recognize paired quote styles in both languages; reject nonliteral builder markers while allowing active MECHANISM-owned API self-tests. Preserve all oracle bytes and digest pins. Verification: harness328 passed/7 skipped/79 strict xfailed; spine55 pending/16 protected; preflight31 passed; unchanged Dart suite2064 passed. * test(dev-harness): require replay loader to preserve frozen request vectors Compare every admitted Case to its pinned method, ordered query/headers, body, status, decoder and observations. A HEAD-shaped request with the right endpoint name cannot pass. Real capture admission remains pending after synthetic engine delivery. Verification: harness328 passed/7 skipped/79 strict xfailed; spine55/16; preflight31 passed; Dart2064 passed. --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> | 5 天前 | |
fix(dev-harness): keep local desktop auth-scoped calls off production The local dev harness desktop profile set OMI_PYTHON_API_URL and OMI_DESKTOP_API_URL but not OMI_AUTH_API_URL, so the macOS app's DesktopBackendEnvironment.authBaseURL() fell back to the production API for auth-scoped calls (GET /v2/desktop/prompts, GET /v1/csat/config). A local named bundle sent its Firebase Auth emulator token to production, got 401, and AuthSessionCoordinator treated the post-refresh 401 as a dead session (invalidateSession reason=postRefreshHTTP401), signing every local desktop profile out about a second after signing in. Point OMI_AUTH_API_URL at the local python backend, same as OMI_PYTHON_API_URL, so auth-scoped calls stay on the local backend. Verified locally: with the fix the bundle stays signed in and those calls return 200 from the local backend; scripts/dev-harness/tests/test_desktop_profile.py passes (5 tests). Co-Authored-By: ZCode (GLM-5.3) on behalf of Claude Code session | 20 小时前 | |
SCA-486: isolated mobile sessions, local-dev auth, capture replay, and seeded journeys (#14213) * mobile sessions: freeze session-evidence-v1 receipt contract Freeze the minimal versioned session/evidence schema shared by the mobile development-foundation consumers (C2 journeys, C3 capture replay, C4 verification, C5 devices) and its executable validator. - contracts/session/session-evidence-v1.schema.json: closed v1 object binding source SHA + dirty digest, built artifact identity, loopback-only endpoints, fixture/runner versions, real timestamps, status/blocked reason and exact execution counts. No credential-shaped field exists anywhere. - dev_harness/session_evidence.py: builder + validator enforcing the cross-field semantics a JSON Schema cannot express: ready/running require an artifact whose git_sha matches the source (a stale build cannot be reported ready), production-family profiles are rejected as session targets, counts must account exactly, zero-execution receipts are only valid pre-run, and credential-shaped keys are refused at any depth. - Egress guards: validate_local_http_url/validate_local_host_port reject non-loopback or non-plain-HTTP endpoints (api.omi.me / api.omiapi.com by name) before any request is attempted. Evidence: scripts/dev-harness/run-tests.sh (test_session_evidence.py, 20 contract tests incl. stale-artifact, egress, credential and accounting rejections). * mobile sessions: deterministic synthetic auth fixture v1 Seed a synthetic Auth-emulator user through the local-dev custom-token endpoint contract from open PR #11784 (feat/app+backend local-development sign-in without OAuth) — reuse, not a competing endpoint. The PR stays with its owner; scripts/dev-harness/MOBILE_SESSIONS.md records the integration plan and provenance. - fixtures/mobile/v1.json: one deterministic user (omi-fixture-v1-user-1@local.test); RFC-reserved domain so fixture identities can never collide with a real account. No real Google/Apple user, provider key, or copied token involved. - dev_harness/mobile_fixtures.py: fail-closed seeding client — the backend URL is validated as loopback plain-HTTP before any request, production hosts are denied by name, and the persisted receipt records identity and outcome only: token_minted/token_retained, never the token itself. Evidence: test_mobile_fixtures.py (16 tests) — determinism, reserved-domain enforcement, pre-request egress refusal, 404/unreachable/wrong-uid fail-closed paths, credential-free receipts. * mobile sessions: structured doctor for the session lanes Every readiness failure classifies exactly one of ready / agent-remediable (with the exact resumption command) / operator-action-needed (privileged install, license, host capacity), per lane (backend, android, ios). - Flutter version is read from the mobile CI pin in .github/workflows/mobile-app-checks.yml — never 'latest'; inconsistent pins refuse rather than guess. - Backend lane: python3.11 (venv must be 3.11, ambient 3.14 must not select the runtime), JDK 21 for the firebase emulators, firebase-tools, redis/typesense via native binary or a responding docker daemon. - Android lane: ANDROID_HOME + adb + emulator engine + system image, each with the exact sdkmanager remedy and a capacity-gated download note. - iOS lane: Xcode + simctl runtime; missing runtime is an operator action. - Capacity: <12GiB free on the shared Data/scratch container is an operator gate for emulator/build lanes (agents never free space themselves); contract/unit lanes skip it via --skip-capacity. - Egress: an ambient production OMI_LOCAL_API_BASE_URL override is reported as a blocking misconfiguration. Evidence: test_mobile_doctor.py (13 tests) over an injected runner — lane filtering, ready/degraded/blocked classification, pin parsing, capacity and operator-gate behavior; live run on m1-mac-studio via 'make mobile-session ARGS="doctor --platform android --platform ios"' reports backend+ios ready, android emulator engine agent-remediable. * mobile sessions: isolated session lifecycle CLI behind one entrypoint 'make mobile-session ARGS="…"' (scripts/dev-harness/mobile-session.sh) owns a uniquely-leased local mobile session: doctor / acquire / start / seed / reset / status / evidence / stop / recover / release. A session is an existing dev-harness instance + port offset + device lease + seed receipt + evidence receipt — the harness lifecycle is reused in-process under OMI_LOCAL_INSTANCE/OMI_HARNESS_PORT_OFFSET, not duplicated. Ownership is fail-closed: - leases are created atomically (O_EXCL) with owner host/user/pid and a harness-standard sentinel; a live foreign owner or another local user's session is never touched; cross-host takeover is an operator decision; recover bumps the generation for same-host/same-user takeovers. - ports come from a claimed offset registry; a foreign process occupying a port is refused (never killed) and the allocator skips that offset; release frees the claim only when it belongs to the session. - start gates device attach on doctor readiness (precise blocked reason, not a crash); ios-simulator devices are created/booted/deleted session-owned via simctl. - seed/reset/stop/release are idempotent; reset only touches the session's own harness instance (sentinel-validated underneath). - evidence emits session-evidence-v1 receipts; ready/running refuse without a bound artifact and refuse when the source moved since acquire. app/setup.sh (separate commit): OMI_IOS_DEVICE_ID pins non-interactive device selection; OMI_DEVICE_SUFFIX overrides hostname identity. Evidence: test_mobile_session.py (19 tests) + wrapper tests — exclusivity, disjoint ports, dead-owner/live-foreign/different-user/cross-host refusals, foreign-port refusal with a real live listener, idempotent release, artifact binding, stale-source refusal, harness env handoff. Live CLI run: acquire/list/evidence/seed-fail-closed/stop/release with exit codes 0/2 on m1-mac-studio. * app/setup.sh: non-interactive device pin and per-session device suffix - OMI_IOS_DEVICE_ID: when set, select_ios_device uses exactly that device id, failing precisely (with the available device list) when absent, instead of enumerating and prompting — the mobile-session harness, CI and nested agents cannot answer an interactive prompt, and the current no-TTY path errors out whenever more than one iOS destination exists. - OMI_DEVICE_SUFFIX: let a session harness (or a second checkout on one host) inject a unique device-identity suffix instead of the hostname, which collides across concurrent sessions on the same machine. Unset behavior is unchanged. Verified by sourcing the function with a stubbed flutter devices --machine: pinned-present emits the id; pinned-absent fails with the list; unpinned multi-device no-TTY keeps the existing enumeration failure. * mobile sessions: apply repo python formatter to the new modules black 26.5.1, --line-length 120 --skip-string-normalization via scripts/backend-python-format; behavior unchanged, dev-harness lane re-run green (214 passed; 1 pre-existing environmental failure — the host's global git worktree guard blocks pytest-tmp linked worktrees). * test: place linked-worktree pytest fixtures under OMI_WORKTREES The managed git wrapper correctly refuses worktrees in /private/tmp. Keep that guard and put the fixture where task worktrees are allowed. Co-authored-by: Cursor <cursoragent@cursor.com> * docs: treat session-evidence-v1 as proposed until consumers review it C1 shipped the schema; freeze it only after C2/C3/C4 agree, not from a single worker declaration. Co-authored-by: Cursor <cursoragent@cursor.com> * feat: reuse PR 11784 local-dev custom-token auth on current main Copy the reviewed emulator-gated sign-in path onto this integration branch so synthetic seed talks to real local services. Leave the original PR open and unmerged. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: boot isolated sessions on real CoreSimulator IDs and offline STT Use the installed iPhone 17 Pro / iOS 26.5 identifiers, pin PROVIDER_MODE=offline, and drop soniox from the offline STT chain so the local backend can start without a paid key. Co-authored-by: Cursor <cursoragent@cursor.com> * test: isolate provider-secret fixtures from ambient PROVIDER_MODE A previous offline session left PROVIDER_MODE in the shell and made the secret-injection tests read ambient offline instead of the fixture file. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(app): injectable capture seams for deterministic recovery replay CaptureController and the phone WAL resolved clock, timers, connectivity, auth, mic, socket and upload policy through global singletons, so the capture -> WAL -> recovery path could not be replayed deterministically. Add narrow constructor seams (capture_seams.dart) with production-identical defaults: CaptureScheduling, CaptureAuthBoundary, CaptureConnectivityBoundary, plus wal/phoneMic/clock/scheduler injection on CaptureController; clock/periodic/job-status injection on LocalWalSyncImpl threaded through WalSyncs/WalService; periodic-timer injection on the NativeMicRecorderService watchdogs. The in-progress-conversation loader seam now covers the socket-connect path too, and streamRecording honors the microphone permission requester like the batch path already did. No behavior change with default construction; every seam is optional. Evidence: bash app/test.sh (1990 passed, 5 pre-existing skips); analyze ratchet green. * test(app): deterministic capture-recovery replay schedules (SCA-489/C3) Replay the REAL production capture pipeline (CaptureController, NativeMicRecorderService, TranscriptSegmentSocketService, WalService, RecordingTransferCoordinator) against controlled external I/O: virtual clock, manual bounded scheduler, scripted transport/upload boundary, fake native host. Restart evidence destroys and reconstructs the object graph from real temp files (torn wals.json -> backup recovery, missing audio -> terminal corruption, process kill -> disk reload and re-upload). Six schedules with invariant oracles: network loss/reconnect mid-capture (exact frame identity in the stored WAL, single upload), stale native events after stop/new session (session-identity gate, no double teardown), interruption/resumption (live + batch, bounded stall escalation), partial/torn persistence plus reconstruction, failed upload with bounded backoff and persisted/enqueued/server-acknowledged distinctions, and ownership transition (signed-out reconnect cancellation, bounded 4001 token refresh). Also publishes the C2/C4 adapter (capture_scenario.dart: catalog + result contract) and the C5 native-event vector schema (phone-mic-native-events/v1) mirroring the Pigeon PhoneMicFlutterApi contract without touching Pigeon. Falsification evidence: removing the NativeMicRecorderService session gate flips the stale-idle schedule to failure (record->stop); removing the finalizeCurrentSession unsynced-retention guard drops the WAL and fails the network-loss schedule. Evidence: flutter test test/unit/capture_recovery_replay_scenarios_test.dart (17 passed); bash app/test.sh full suite green. * chore(app): allowlist SCA-489 replay contract libs in the dead-code ratchet The scenario catalog/result contract and the native-event vector schema are library-only by design until the C2/C4 and C5 lanes import them; the ratchet demands an explicit allowlist entry with a reason for exactly this case. * test: pin conversation-window capture session id across sequential phone-mic lives activeCaptureSessionId is WAL/conversation-scoped so a late ConversationEvent can still stamp WALs. C2 must use activeRecordingId as the live recording identity. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: wait for unawaited capture-upload retries before asserting and teardown CI failed the bounded-backoff replay because cooldown wakes are unawaited and settle used wall-clock sleeps that missed the drain under load, then deleted the temp WAL dir mid-write. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(app): typed debug semantic controls for the local journey lane Extends the existing debug Marionette surface (same debug VM-service transport, no new server/framework) with product-semantic controls: versioned capabilities (semantic-controls/v1), privacy-safe state (route, principal, capture lifecycle with activeRecordingId as the authoritative recording identity), bounded wait_ready, production-path navigation, and named journey faults. Fail-closed eligibility: kDebugMode AND local_dev profile AND OMI_DEV_CONTROLS=1 dart-define. Ineligible builds (including production-flavor debug) install nothing and the HTTP fault chokepoint is a pure pass-through — pinned by semantic_controls_guard_test.dart. Narrow seams added for the hermetic journey lane: - AuthService.installLocalHarnessTokenGateway (debug+local_dev gated Firebase token I/O boundary; isSignedIn routes through the gateway) - PlatformManager.initializeForLocalHarness (header fields only) - CrashlyticsManager report paths tolerate a missing Firebase app the same way main.dart's zone handler already does, so host-lane errors surface instead of being masked by [core/no-app] Verified: flutter test test/unit/semantic_controls_guard_test.dart (10 passed); auth regression suites (34 passed); C3 capture replay (17 passed); dead-code ratchet at baseline. * test(app): five strict seeded acceptance journeys with negative fault variants Canonical executable definitions (one per behavior) under app/integration_test/journeys/, runnable hermetically (flutter-tester + loopback fixture backend) or on a simulator via run_journeys.sh: j1 seeded conversation detail — real provider fetch + real detail page, exact synthetic identity; negative: wrong-owner session refused. j2 chat send -> distinct assistant reply — real input/send-button keys (omi.chat.input / omi.chat.send), request observed server-side, server-minted ai-role reply distinct from the prompt, rendered; negatives: suppress-send, suppress-assistant-reply, wrong-owner-session. j3 memory create/edit surviving reload — production provider path, server-minted id required after reload; negative: drop-memory-save. j4 expired session — transient failure re-mints via the real custom-token endpoint; terminal failure emits expiry and blocks requests; negative: production-family profiles never silently re-mint. j5 capture interruption/reconnect — C3 capture-scenario/v1 adapter: real temp files, process reconstruction, drain exactly once; negative: fail-capture-recovery. Each negative arms exactly one named fault and must fail with the invariant named. Evidence receipts follow session-evidence-v1 accounting; zero-execution runs never pass. Verified: bash integration_test/journeys/run_journeys.sh (5/5 pass); repeated deterministic vertical: bash integration_test/journeys/run_journeys.sh --filter j2 --runs 5. * fix: stamp journey evidence finished_at at write time Receipts were recording construction time as the end timestamp, so duration could not be distinguished from start. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(app): clear new analyzer-ratchet regressions in the integrated journeys The integrated checkpoint (b217eb1a9d) fails app/scripts/analyze_ratchet.sh with 7 new occurrences: 3 unused imports plus a bogus 'show WalStatus' in j5, an unused-looking nested import that actually provides SingleChildWidget in hermetic_boot, a missing const in j4, and two depend_on_referenced_packages for test-only platform interfaces. Declares path_provider_platform_interface and nested as direct dev dependencies (same pattern as the existing web_socket_channel dev deps) and removes the dead imports. Mechanical lint repairs only — j4/j5 hermetic journeys re-run green after the change. * feat: unified mobile verify lanes, mechanical journey selection, and CI/contributor path (SCA-490/C4) One canonical verification entrypoint over the proven lanes: make mobile-verify select|doctor|fast|smoke|physical (scripts/dev-harness/mobile-verify.sh -> dev_harness.mobile_verify). It never adds a second runner: journeys delegate to the C2 canonical runner and session infrastructure to the C1 session CLI. Selection is mechanical and fail-closed: journeys are glob-discovered (runner --list contract test), changed paths map through a per-seam rule table, unknown app/lib impact falls back to the full suite, and an empty selection is drift (exit 65), never PASS(0) — run_journeys.sh now fails closed the same way. Receipts are validated against session-evidence-v1 accounting (honest counts, zero-execution never passes) and every lane writes a verify-receipt.json binding source SHA + dirty digest, runner versions, outcomes, and the exact rerun command. smoke is fail-closed (exit 2 + remedy, never CI), physical is a separately reported admission lane. CI runs the same command in a new journeys-hermetic job in the existing mobile-app-checks.yml when has_app_journeys fires (journey definitions and support, C3 replay world, dev controls, non-generated app/lib Dart, evidence contract, or this entrypoint) — synthetic fixtures only, fork-safe, receipts uploaded on pass and failure. Selection is resolved by the shared pre_push_ci_prediction.py and deliberately stays out of the bounded pre-push gate. Docs reconciled around the real command: app README, app AGENTS (within the lean budget), and the e2e SKILL now point here instead of diverging on setup/auth. * chore(app): stop tracking Flutter's iOS ephemeral tree app/ios/Flutter/ephemeral/** is regenerated by flutter on every pub get and self-describes as 'Generated file. Do not edit.' It was committed by accident in a formatting sweep (dec329a84a) and has been stale ever since: the tracked SwiftPM Package.swift lists pods (in_app_review, pasteboard) that no pub dependency provides, so any flutter run rewrites it, dirties every worktree, and fails the diff-hygiene push gate on regenerated trailing whitespace. Untrack the four files and ignore the tree, mirroring the existing **/macos/Flutter/ephemeral/ rules. Xcode resolves the local package after flutter regenerates it during setup; nothing consumes a committed copy. * feat: native lifecycle seams, vector replay, and leased device qualification (SCA-491/C5) - PhoneMicController (iOS + Android) now consumes narrow, injectable environment/ports seams: event sink, engine, permission, session config, interruption source, batch pipeline, main loop. Production behavior is unchanged; all live wiring lives in PhoneMicHostApiImpl.swift (iOS) and PhoneMicControllerPorts.production (Android). - Canonical phone-mic-native-events/v1 vector fixtures (8 schedules incl. session adoption) shared by Dart guard, iOS ruby harness, Android JVM harness; Pigeon contract types extracted at iOS test time (drift-guarded). - iOS: ios/test/phone_mic_lifecycle_replay_test.rb replays all vectors through the production controller+emitter with fakes for OS I/O only. - Android: PhoneMicLifecycleReplayTest (JVM, virtual main loop + manual audio queue) replays the same vectors through the production controller. - device_lease.py: exclusive physical-device leases with qualification registry (personal-device refusal), bounded acquisition, live-lease never-stolen, stale-owner recovery with generation bump, safe release. - device_runner.py + 'mobile-session device' CLI: readiness doctor with exact operator steps, and a runner consuming C1 session manifests (install/adb-reverse/untethered launch/permission cycle/device-run evidence v1). All hermetically tested with fake devices (25 tests). - PHYSICAL_DEVICES.md: m1-mac-studio read-only inventory, operator runbook, and the external physical-test handoff template. Physical acceptance stays pending user-run evidence by design. * docs: point mobile-verify physical at the C5 device handoff C4's physical lane stays fail-closed (exit 2). After C5 landed, the admission document should name the real runner and PHYSICAL_DEVICES.md instead of implying the software path is still missing. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: expect four Flutter pins in mobile-app-checks C4 added journeys-hermetic as a fourth Flutter job on the same repository toolchain pin. The workflow-contract count of 3 was stale. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: raise Desktop Swift PR-lane suite budget to 3000s Run 35134593036 measured 2778s against 2700s on a cache-hit PR lane. The overrun was one 1500s batch ceiling plus isolation, not a slow desktop suite; this mobile PR has no desktop sources. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> | 7 天前 | |
fix(app): refuse remote APIs in hermetic tests; add a lane-bootstrap that skips the full backend lock (#14321) * fix(app): refuse remote APIs and non-dev pairings in hermetic tests app/test.sh skipped writing .dev.env when generated files already existed, so a leftover API_BASE_URL or mobile_beta/prod pairing could silently survive. Fail closed at the checker, test.sh, and doctor without rewriting the file. Failure-Class: new Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dev-harness): add lane-bootstrap that skips the full backend lock A linked worktree can run cheap pre-push gates and app/test.sh after one idempotent command: pinned Flutter, shared PUB_CACHE, Python 3.11 via uv, and yaml+dotenv only. An incomplete .venv directory is no longer treated as ready. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): keep fixture venvs passing the cheap-gate import probe Manifest-contract stubs only accepted `import yaml`. The resolver now probes `import dotenv, yaml` so an empty .venv is not treated as ready, and those stubs were skipped. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> | 6 天前 | |
SCA-486: isolated mobile sessions, local-dev auth, capture replay, and seeded journeys (#14213) * mobile sessions: freeze session-evidence-v1 receipt contract Freeze the minimal versioned session/evidence schema shared by the mobile development-foundation consumers (C2 journeys, C3 capture replay, C4 verification, C5 devices) and its executable validator. - contracts/session/session-evidence-v1.schema.json: closed v1 object binding source SHA + dirty digest, built artifact identity, loopback-only endpoints, fixture/runner versions, real timestamps, status/blocked reason and exact execution counts. No credential-shaped field exists anywhere. - dev_harness/session_evidence.py: builder + validator enforcing the cross-field semantics a JSON Schema cannot express: ready/running require an artifact whose git_sha matches the source (a stale build cannot be reported ready), production-family profiles are rejected as session targets, counts must account exactly, zero-execution receipts are only valid pre-run, and credential-shaped keys are refused at any depth. - Egress guards: validate_local_http_url/validate_local_host_port reject non-loopback or non-plain-HTTP endpoints (api.omi.me / api.omiapi.com by name) before any request is attempted. Evidence: scripts/dev-harness/run-tests.sh (test_session_evidence.py, 20 contract tests incl. stale-artifact, egress, credential and accounting rejections). * mobile sessions: deterministic synthetic auth fixture v1 Seed a synthetic Auth-emulator user through the local-dev custom-token endpoint contract from open PR #11784 (feat/app+backend local-development sign-in without OAuth) — reuse, not a competing endpoint. The PR stays with its owner; scripts/dev-harness/MOBILE_SESSIONS.md records the integration plan and provenance. - fixtures/mobile/v1.json: one deterministic user (omi-fixture-v1-user-1@local.test); RFC-reserved domain so fixture identities can never collide with a real account. No real Google/Apple user, provider key, or copied token involved. - dev_harness/mobile_fixtures.py: fail-closed seeding client — the backend URL is validated as loopback plain-HTTP before any request, production hosts are denied by name, and the persisted receipt records identity and outcome only: token_minted/token_retained, never the token itself. Evidence: test_mobile_fixtures.py (16 tests) — determinism, reserved-domain enforcement, pre-request egress refusal, 404/unreachable/wrong-uid fail-closed paths, credential-free receipts. * mobile sessions: structured doctor for the session lanes Every readiness failure classifies exactly one of ready / agent-remediable (with the exact resumption command) / operator-action-needed (privileged install, license, host capacity), per lane (backend, android, ios). - Flutter version is read from the mobile CI pin in .github/workflows/mobile-app-checks.yml — never 'latest'; inconsistent pins refuse rather than guess. - Backend lane: python3.11 (venv must be 3.11, ambient 3.14 must not select the runtime), JDK 21 for the firebase emulators, firebase-tools, redis/typesense via native binary or a responding docker daemon. - Android lane: ANDROID_HOME + adb + emulator engine + system image, each with the exact sdkmanager remedy and a capacity-gated download note. - iOS lane: Xcode + simctl runtime; missing runtime is an operator action. - Capacity: <12GiB free on the shared Data/scratch container is an operator gate for emulator/build lanes (agents never free space themselves); contract/unit lanes skip it via --skip-capacity. - Egress: an ambient production OMI_LOCAL_API_BASE_URL override is reported as a blocking misconfiguration. Evidence: test_mobile_doctor.py (13 tests) over an injected runner — lane filtering, ready/degraded/blocked classification, pin parsing, capacity and operator-gate behavior; live run on m1-mac-studio via 'make mobile-session ARGS="doctor --platform android --platform ios"' reports backend+ios ready, android emulator engine agent-remediable. * mobile sessions: isolated session lifecycle CLI behind one entrypoint 'make mobile-session ARGS="…"' (scripts/dev-harness/mobile-session.sh) owns a uniquely-leased local mobile session: doctor / acquire / start / seed / reset / status / evidence / stop / recover / release. A session is an existing dev-harness instance + port offset + device lease + seed receipt + evidence receipt — the harness lifecycle is reused in-process under OMI_LOCAL_INSTANCE/OMI_HARNESS_PORT_OFFSET, not duplicated. Ownership is fail-closed: - leases are created atomically (O_EXCL) with owner host/user/pid and a harness-standard sentinel; a live foreign owner or another local user's session is never touched; cross-host takeover is an operator decision; recover bumps the generation for same-host/same-user takeovers. - ports come from a claimed offset registry; a foreign process occupying a port is refused (never killed) and the allocator skips that offset; release frees the claim only when it belongs to the session. - start gates device attach on doctor readiness (precise blocked reason, not a crash); ios-simulator devices are created/booted/deleted session-owned via simctl. - seed/reset/stop/release are idempotent; reset only touches the session's own harness instance (sentinel-validated underneath). - evidence emits session-evidence-v1 receipts; ready/running refuse without a bound artifact and refuse when the source moved since acquire. app/setup.sh (separate commit): OMI_IOS_DEVICE_ID pins non-interactive device selection; OMI_DEVICE_SUFFIX overrides hostname identity. Evidence: test_mobile_session.py (19 tests) + wrapper tests — exclusivity, disjoint ports, dead-owner/live-foreign/different-user/cross-host refusals, foreign-port refusal with a real live listener, idempotent release, artifact binding, stale-source refusal, harness env handoff. Live CLI run: acquire/list/evidence/seed-fail-closed/stop/release with exit codes 0/2 on m1-mac-studio. * app/setup.sh: non-interactive device pin and per-session device suffix - OMI_IOS_DEVICE_ID: when set, select_ios_device uses exactly that device id, failing precisely (with the available device list) when absent, instead of enumerating and prompting — the mobile-session harness, CI and nested agents cannot answer an interactive prompt, and the current no-TTY path errors out whenever more than one iOS destination exists. - OMI_DEVICE_SUFFIX: let a session harness (or a second checkout on one host) inject a unique device-identity suffix instead of the hostname, which collides across concurrent sessions on the same machine. Unset behavior is unchanged. Verified by sourcing the function with a stubbed flutter devices --machine: pinned-present emits the id; pinned-absent fails with the list; unpinned multi-device no-TTY keeps the existing enumeration failure. * mobile sessions: apply repo python formatter to the new modules black 26.5.1, --line-length 120 --skip-string-normalization via scripts/backend-python-format; behavior unchanged, dev-harness lane re-run green (214 passed; 1 pre-existing environmental failure — the host's global git worktree guard blocks pytest-tmp linked worktrees). * test: place linked-worktree pytest fixtures under OMI_WORKTREES The managed git wrapper correctly refuses worktrees in /private/tmp. Keep that guard and put the fixture where task worktrees are allowed. Co-authored-by: Cursor <cursoragent@cursor.com> * docs: treat session-evidence-v1 as proposed until consumers review it C1 shipped the schema; freeze it only after C2/C3/C4 agree, not from a single worker declaration. Co-authored-by: Cursor <cursoragent@cursor.com> * feat: reuse PR 11784 local-dev custom-token auth on current main Copy the reviewed emulator-gated sign-in path onto this integration branch so synthetic seed talks to real local services. Leave the original PR open and unmerged. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: boot isolated sessions on real CoreSimulator IDs and offline STT Use the installed iPhone 17 Pro / iOS 26.5 identifiers, pin PROVIDER_MODE=offline, and drop soniox from the offline STT chain so the local backend can start without a paid key. Co-authored-by: Cursor <cursoragent@cursor.com> * test: isolate provider-secret fixtures from ambient PROVIDER_MODE A previous offline session left PROVIDER_MODE in the shell and made the secret-injection tests read ambient offline instead of the fixture file. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(app): injectable capture seams for deterministic recovery replay CaptureController and the phone WAL resolved clock, timers, connectivity, auth, mic, socket and upload policy through global singletons, so the capture -> WAL -> recovery path could not be replayed deterministically. Add narrow constructor seams (capture_seams.dart) with production-identical defaults: CaptureScheduling, CaptureAuthBoundary, CaptureConnectivityBoundary, plus wal/phoneMic/clock/scheduler injection on CaptureController; clock/periodic/job-status injection on LocalWalSyncImpl threaded through WalSyncs/WalService; periodic-timer injection on the NativeMicRecorderService watchdogs. The in-progress-conversation loader seam now covers the socket-connect path too, and streamRecording honors the microphone permission requester like the batch path already did. No behavior change with default construction; every seam is optional. Evidence: bash app/test.sh (1990 passed, 5 pre-existing skips); analyze ratchet green. * test(app): deterministic capture-recovery replay schedules (SCA-489/C3) Replay the REAL production capture pipeline (CaptureController, NativeMicRecorderService, TranscriptSegmentSocketService, WalService, RecordingTransferCoordinator) against controlled external I/O: virtual clock, manual bounded scheduler, scripted transport/upload boundary, fake native host. Restart evidence destroys and reconstructs the object graph from real temp files (torn wals.json -> backup recovery, missing audio -> terminal corruption, process kill -> disk reload and re-upload). Six schedules with invariant oracles: network loss/reconnect mid-capture (exact frame identity in the stored WAL, single upload), stale native events after stop/new session (session-identity gate, no double teardown), interruption/resumption (live + batch, bounded stall escalation), partial/torn persistence plus reconstruction, failed upload with bounded backoff and persisted/enqueued/server-acknowledged distinctions, and ownership transition (signed-out reconnect cancellation, bounded 4001 token refresh). Also publishes the C2/C4 adapter (capture_scenario.dart: catalog + result contract) and the C5 native-event vector schema (phone-mic-native-events/v1) mirroring the Pigeon PhoneMicFlutterApi contract without touching Pigeon. Falsification evidence: removing the NativeMicRecorderService session gate flips the stale-idle schedule to failure (record->stop); removing the finalizeCurrentSession unsynced-retention guard drops the WAL and fails the network-loss schedule. Evidence: flutter test test/unit/capture_recovery_replay_scenarios_test.dart (17 passed); bash app/test.sh full suite green. * chore(app): allowlist SCA-489 replay contract libs in the dead-code ratchet The scenario catalog/result contract and the native-event vector schema are library-only by design until the C2/C4 and C5 lanes import them; the ratchet demands an explicit allowlist entry with a reason for exactly this case. * test: pin conversation-window capture session id across sequential phone-mic lives activeCaptureSessionId is WAL/conversation-scoped so a late ConversationEvent can still stamp WALs. C2 must use activeRecordingId as the live recording identity. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: wait for unawaited capture-upload retries before asserting and teardown CI failed the bounded-backoff replay because cooldown wakes are unawaited and settle used wall-clock sleeps that missed the drain under load, then deleted the temp WAL dir mid-write. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(app): typed debug semantic controls for the local journey lane Extends the existing debug Marionette surface (same debug VM-service transport, no new server/framework) with product-semantic controls: versioned capabilities (semantic-controls/v1), privacy-safe state (route, principal, capture lifecycle with activeRecordingId as the authoritative recording identity), bounded wait_ready, production-path navigation, and named journey faults. Fail-closed eligibility: kDebugMode AND local_dev profile AND OMI_DEV_CONTROLS=1 dart-define. Ineligible builds (including production-flavor debug) install nothing and the HTTP fault chokepoint is a pure pass-through — pinned by semantic_controls_guard_test.dart. Narrow seams added for the hermetic journey lane: - AuthService.installLocalHarnessTokenGateway (debug+local_dev gated Firebase token I/O boundary; isSignedIn routes through the gateway) - PlatformManager.initializeForLocalHarness (header fields only) - CrashlyticsManager report paths tolerate a missing Firebase app the same way main.dart's zone handler already does, so host-lane errors surface instead of being masked by [core/no-app] Verified: flutter test test/unit/semantic_controls_guard_test.dart (10 passed); auth regression suites (34 passed); C3 capture replay (17 passed); dead-code ratchet at baseline. * test(app): five strict seeded acceptance journeys with negative fault variants Canonical executable definitions (one per behavior) under app/integration_test/journeys/, runnable hermetically (flutter-tester + loopback fixture backend) or on a simulator via run_journeys.sh: j1 seeded conversation detail — real provider fetch + real detail page, exact synthetic identity; negative: wrong-owner session refused. j2 chat send -> distinct assistant reply — real input/send-button keys (omi.chat.input / omi.chat.send), request observed server-side, server-minted ai-role reply distinct from the prompt, rendered; negatives: suppress-send, suppress-assistant-reply, wrong-owner-session. j3 memory create/edit surviving reload — production provider path, server-minted id required after reload; negative: drop-memory-save. j4 expired session — transient failure re-mints via the real custom-token endpoint; terminal failure emits expiry and blocks requests; negative: production-family profiles never silently re-mint. j5 capture interruption/reconnect — C3 capture-scenario/v1 adapter: real temp files, process reconstruction, drain exactly once; negative: fail-capture-recovery. Each negative arms exactly one named fault and must fail with the invariant named. Evidence receipts follow session-evidence-v1 accounting; zero-execution runs never pass. Verified: bash integration_test/journeys/run_journeys.sh (5/5 pass); repeated deterministic vertical: bash integration_test/journeys/run_journeys.sh --filter j2 --runs 5. * fix: stamp journey evidence finished_at at write time Receipts were recording construction time as the end timestamp, so duration could not be distinguished from start. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(app): clear new analyzer-ratchet regressions in the integrated journeys The integrated checkpoint (b217eb1a9d) fails app/scripts/analyze_ratchet.sh with 7 new occurrences: 3 unused imports plus a bogus 'show WalStatus' in j5, an unused-looking nested import that actually provides SingleChildWidget in hermetic_boot, a missing const in j4, and two depend_on_referenced_packages for test-only platform interfaces. Declares path_provider_platform_interface and nested as direct dev dependencies (same pattern as the existing web_socket_channel dev deps) and removes the dead imports. Mechanical lint repairs only — j4/j5 hermetic journeys re-run green after the change. * feat: unified mobile verify lanes, mechanical journey selection, and CI/contributor path (SCA-490/C4) One canonical verification entrypoint over the proven lanes: make mobile-verify select|doctor|fast|smoke|physical (scripts/dev-harness/mobile-verify.sh -> dev_harness.mobile_verify). It never adds a second runner: journeys delegate to the C2 canonical runner and session infrastructure to the C1 session CLI. Selection is mechanical and fail-closed: journeys are glob-discovered (runner --list contract test), changed paths map through a per-seam rule table, unknown app/lib impact falls back to the full suite, and an empty selection is drift (exit 65), never PASS(0) — run_journeys.sh now fails closed the same way. Receipts are validated against session-evidence-v1 accounting (honest counts, zero-execution never passes) and every lane writes a verify-receipt.json binding source SHA + dirty digest, runner versions, outcomes, and the exact rerun command. smoke is fail-closed (exit 2 + remedy, never CI), physical is a separately reported admission lane. CI runs the same command in a new journeys-hermetic job in the existing mobile-app-checks.yml when has_app_journeys fires (journey definitions and support, C3 replay world, dev controls, non-generated app/lib Dart, evidence contract, or this entrypoint) — synthetic fixtures only, fork-safe, receipts uploaded on pass and failure. Selection is resolved by the shared pre_push_ci_prediction.py and deliberately stays out of the bounded pre-push gate. Docs reconciled around the real command: app README, app AGENTS (within the lean budget), and the e2e SKILL now point here instead of diverging on setup/auth. * chore(app): stop tracking Flutter's iOS ephemeral tree app/ios/Flutter/ephemeral/** is regenerated by flutter on every pub get and self-describes as 'Generated file. Do not edit.' It was committed by accident in a formatting sweep (dec329a84a) and has been stale ever since: the tracked SwiftPM Package.swift lists pods (in_app_review, pasteboard) that no pub dependency provides, so any flutter run rewrites it, dirties every worktree, and fails the diff-hygiene push gate on regenerated trailing whitespace. Untrack the four files and ignore the tree, mirroring the existing **/macos/Flutter/ephemeral/ rules. Xcode resolves the local package after flutter regenerates it during setup; nothing consumes a committed copy. * feat: native lifecycle seams, vector replay, and leased device qualification (SCA-491/C5) - PhoneMicController (iOS + Android) now consumes narrow, injectable environment/ports seams: event sink, engine, permission, session config, interruption source, batch pipeline, main loop. Production behavior is unchanged; all live wiring lives in PhoneMicHostApiImpl.swift (iOS) and PhoneMicControllerPorts.production (Android). - Canonical phone-mic-native-events/v1 vector fixtures (8 schedules incl. session adoption) shared by Dart guard, iOS ruby harness, Android JVM harness; Pigeon contract types extracted at iOS test time (drift-guarded). - iOS: ios/test/phone_mic_lifecycle_replay_test.rb replays all vectors through the production controller+emitter with fakes for OS I/O only. - Android: PhoneMicLifecycleReplayTest (JVM, virtual main loop + manual audio queue) replays the same vectors through the production controller. - device_lease.py: exclusive physical-device leases with qualification registry (personal-device refusal), bounded acquisition, live-lease never-stolen, stale-owner recovery with generation bump, safe release. - device_runner.py + 'mobile-session device' CLI: readiness doctor with exact operator steps, and a runner consuming C1 session manifests (install/adb-reverse/untethered launch/permission cycle/device-run evidence v1). All hermetically tested with fake devices (25 tests). - PHYSICAL_DEVICES.md: m1-mac-studio read-only inventory, operator runbook, and the external physical-test handoff template. Physical acceptance stays pending user-run evidence by design. * docs: point mobile-verify physical at the C5 device handoff C4's physical lane stays fail-closed (exit 2). After C5 landed, the admission document should name the real runner and PHYSICAL_DEVICES.md instead of implying the software path is still missing. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: expect four Flutter pins in mobile-app-checks C4 added journeys-hermetic as a fourth Flutter job on the same repository toolchain pin. The workflow-contract count of 3 was stale. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: raise Desktop Swift PR-lane suite budget to 3000s Run 35134593036 measured 2778s against 2700s on a cache-hit PR lane. The overrun was one 1500s batch ceiling plus isolation, not a slow desktop suite; this mobile PR has no desktop sources. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> | 7 天前 | |
feat(mobile): add iPhone capture probes and offline replay (#15956) * feat(mobile): add offline iPhone capture replay and diagnostic harnesses * fix(dev-harness): mark Apple certificate fingerprint as non-security * fix(dev-harness): use platform certificate fingerprint tooling * fix(mobile): preserve startup ordering tripwire | 23 小时前 | |
| 5 天前 | ||
feat: mobile product telemetry and PostHog experimentation (#15883) * feat: add mobile product telemetry and PostHog experiment infrastructure * test: verify billing projection at the webhook persistence boundary * Harden PostHog scorecard completeness * harden: keep unknown feedback failure causes explicit * chore: regenerate clients for mobile feedback contract * fix: restore backend CI contracts for telemetry changes * fix: keep feedback receipt validation strict | 1 天前 | |
feat(dev-harness): pin D7 fixture-audio contract without injecting yet (#14391) A later phone-mic journey must hear the LibriSpeech release-probe WAV and name platform_mic vs in_app_fake; this pins that receipt so -no-audio and a capture-source fake cannot pose as a microphone. Co-authored-by: Cursor <cursoragent@cursor.com> | 5 天前 | |
feat(dev-harness): opt-in make lane-backend (fast wheel install) for the mobile harness and pre-push gates (#14349) * fix(app): refuse remote APIs and non-dev pairings in hermetic tests app/test.sh skipped writing .dev.env when generated files already existed, so a leftover API_BASE_URL or mobile_beta/prod pairing could silently survive. Fail closed at the checker, test.sh, and doctor without rewriting the file. Failure-Class: new Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dev-harness): add lane-bootstrap that skips the full backend lock A linked worktree can run cheap pre-push gates and app/test.sh after one idempotent command: pinned Flutter, shared PUB_CACHE, Python 3.11 via uv, and yaml+dotenv only. An incomplete .venv directory is no longer treated as ready. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): keep fixture venvs passing the cheap-gate import probe Manifest-contract stubs only accepted `import yaml`. The resolver now probes `import dotenv, yaml` so an empty .venv is not treated as ready, and those stubs were skipped. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dev-harness): make setup-backend the fast wheel install, not pylock sync uv pip sync of pylock.macos.toml lists hashed sdist+wheel for av/llvmlite/scipy/pyarrow and stalled ~20 minutes; uv pip install -r requirements.txt takes index wheels (69s here) and covers uvicorn/pyright/yaml/dotenv/google.auth for session start and the pre-push typecheck. Lock-faithful sync stays in sync-python-deps.sh. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): keep setup-backend as the lock sync; add opt-in lane-backend make setup-backend is the contributor lock-synced entrypoint; pointing it at unlocked requirements.txt broke scripts/test-make-setup.sh. The fast wheel install is make lane-backend. generate-app-env runs build_runner without the flags that deleted tracked manifest.g.dart. Co-authored-by: Cursor <cursoragent@cursor.com> * docs(backend): keep AGENTS.md under the lean budget while naming lane-backend Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> | 6 天前 | |
feat: add JIT knowledge ledger foundation and guarded adoption (#12084) * feat: add JIT knowledge ledger foundation * chore: refresh integration OpenAPI contract * fix: make trigger evaluation release-safe Failure-Class: none * fix: preserve lifecycle semantics in ledger apply Failure-Class: none * feat: adopt guarded JIT knowledge surfaces Route the agent preference writer through the intent-backed ledger, register a privacy-filtered entity timeline tool, render optional evidence on Windows, and add a base-ref-protected Gate F legacy-surface ratchet. Failure-Class: none * feat: add progressive JIT knowledge reads Register owner-scoped current-ledger search and explicit playbook hydration, with pre-limit semantic filtering and bounded outputs. Add a content-free planner/resume migration fixture without claiming canonical transaction completion.\n\nValidation: 141 focused backend tests passed; backend typecheck reported 0 errors; repository preflight passed 120 checks. * feat: render chat evidence on web Render bounded, fail-soft conversation evidence after authoritative answers in both web chat entry points. Unsupported, future, duplicate, and raw failure details remain inert.\n\nValidation: 337 web tests passed; web typecheck, oxlint, and Prettier passed; repository preflight passed 120 checks. * feat: require intent-backed ledger search results Apply the intent-backed requirement at the final merged canonical/history filter, with a passive historical-row regression case.\n\nValidation: 54 focused backend tests passed. * test: keep agent tool isolation stubs current * feat: gate JIT conversation retrieval * fix: make entity timeline scans deterministic Failure-Class: none * fix: honor rejected ledger projections Failure-Class: none * feat: render inert screen evidence on web * fix: reuse canonical review projection Failure-Class: none * fix(web): await recap context effect Failure-Class: none * test: amortize preference tool isolation load Failure-Class: none * feat(app): add knowledge ledger review surface Failure-Class: none * feat(macos): use canonical ledger prompt projection Failure-Class: none * test(memory): classify legacy surface inventory roles Failure-Class: none * fix(app): preserve ledger history completeness state Failure-Class: none * feat(macos): preserve canonical ledger mirror metadata Failure-Class: none * fix(app): match canonical ledger ordering Failure-Class: none * test(api): prove ledger client schema parity Failure-Class: none * feat(memory): expose bounded ledger history Failure-Class: none * chore(api): generate ledger history clients Failure-Class: none * feat(retrieval): add bounded card participants Failure-Class: none * feat(macos): project ledger trigger watchlist Failure-Class: none * fix(clients): fail closed on ledger authority Failure-Class: none * feat(macos): expose bounded trigger snapshot Failure-Class: none * fix(memory): keep closed history read only Failure-Class: none * feat(app): disclose partial ledger history Failure-Class: none * test(macos): cover ledger trigger bridge Failure-Class: none * fix(app): use neutral ledger accents Failure-Class: none * chore(api): declare ledger history route policy Failure-Class: none * test(macos): remove unsafe JSON fixture unwraps Failure-Class: none * fix(memory): satisfy typed history boundary Failure-Class: none * fix(macos): require prompt snapshot authority Failure-Class: none * test(memory): prove ledger migration on emulator Failure-Class: none * feat(retrieval): emit bounded screen evidence Failure-Class: none * test(memory): classify maintenance retirement readiness Failure-Class: none * feat(macos): adapt Rewind metadata for triggers Failure-Class: none * test(retrieval): align screen timestamp contract Failure-Class: none * feat(memory): correct ledger facts by amendment Failure-Class: none * test(memory): prove ledger correction on emulator Failure-Class: none * feat(macos): harden local trigger observations Failure-Class: none * feat(agent): search bounded historical facts Failure-Class: none * test(macos): cover trigger observation adapter * fix(memory): gate historical fact retrieval * feat(memory): add gated JIT retrieval strategy * test(memory): prove mixed-version JIT runtime parity * chore(memory): keep JIT gate exports type-safe * refactor(memory): isolate JIT prompt contract * fix(conversations): round-trip owner-scoped references Accept the conversation:<id> references emitted by JIT result cards while retaining strict UUID-only bare IDs and share links. Restrict machine IDs to a bounded safe alphabet so evidence suffixes and path-like values fail closed. Failure-Class: none * test(memory): join JIT citations to evidence envelope * fix(retrieval): enforce JIT conversation search budget Cap JIT summary searches per request and bound database hydration to the projection limit before reads. Preserve the legacy path when JIT is disabled. Failure-Class: FC-unbounded-user-collection-in-prompt * fix(memory): keep JIT retrieval request scoped * test(macos): prove future JIT evidence stays inert * fix(memory): keep JIT card citations request-global Failure-Class: new * fix(retrieval): separate JIT hydration from search Treat gated owner-scoped references as exact hydration without searching transcript text for the reference. Charge every JIT candidate search to the shared four-search request budget, including snippet-bearing requests, while keeping exact hydration free and preserving released JIT-off UUID/share-link behavior.\n\nVerified:\n- cd backend && ./.venv/bin/python -m pytest tests/unit/test_conversation_jit_processing.py tests/unit/test_conversation_exact_reference_search.py -q (58 passed)\n- cd backend && uvx --from pyright==1.1.403 pyright -p pyrightconfig.json --pythonpath .venv/bin/python (0 errors)\n- git diff --check\n\nFailure-Class: FC-unbounded-user-collection-in-prompt * fix(memory): keep repeated JIT cards index-safe * fix(retrieval): satisfy JIT card type contract * fix(retrieval): hydrate collected JIT cards * test(app): preserve answers during delayed evidence requests * test(app): exercise production evidence composition * feat(memories): restore superseded ledger facts * fix(memories): reconcile reverted ledger facts * feat(memories): append reverted ledger facts * feat(memories): synchronize revert client contract * fix(memory): name ledger revert identity * fix(memories): type and enlarge revert controls * fix(memories): fence revert retries and refreshes * fix(memories): fence ledger revert authority * test(memory): count ledger revert rate limit * feat: expose agent-controlled historical facts * feat: reopen standalone ledger facts * feat: add fail-closed JIT QA bundle routing * feat: add safe local JIT QA backend stack * fix: harden isolated JIT QA stack * feat: add explicit multi-source entity timeline * feat(backend): add JIT rollout authority * feat(backend): fence every proactive paid boundary * fix(backend): release proactive quota on cancellation Release the reserved proactive quota exactly once when cancellation interrupts paid-boundary refresh or a provider retry, then re-raise cancellation without emitting retry telemetry. Add deterministic regression coverage for both cancellation points. Failure-Class: FC-proactive-quota-cancellation | new * fix(backend): make proactive quota cancellation safe Detach in-flight Redis reservations on request cancellation and release only admitted slots once they settle. Move direct-provider fallback telemetry behind the fresh paid-boundary rollout check so late kill or unknown decisions cannot report false recovery.\n\nFailure-Class: FC-proactive-quota-cancellation | new * fix(backend): preserve quota compensation during shutdown Keep late Redis reservation compensators outside the ordinary cancellable background-task drain. Desktop and main application shutdown paths now wait for these critical compensators before cancelling ordinary work, with deterministic blocked-thread and lifecycle-order regressions.\n\nFailure-Class: FC-proactive-quota-cancellation | new * fix(backend): use expiring proactive quota leases * fix(backend): make quota finalization clock-safe * fix(backend): isolate jit rollout control plane * fix(backend): close jit control plane safely * fix(backend): emit retry recovery after quota commit * test(backend): keep rollout app contract fast * feat(jit): add guarded proactivity and first-open policies * chore(desktop): mark jit policy as internal * test(desktop): cover jit proactivity policy flow * feat(backend): wire durable JIT first-open processing * feat(desktop): fence JIT proactivity runtime admission * feat: activate authoritative JIT proactivity runtime * fix: harden JIT proactivity authority * fix: close proactive runtime authority gaps * fix(jit): make first-open effects resumable * fix(jit): fence outstanding first-open work * fix(jit): resume app usage receipts * fix(jit): make app usage retries no-op Failure-Class: none * fix(jit): allow completed usage after app deletion Failure-Class: none * fix(jit): register first-open folder query Failure-Class: none * Fix first-open import isolation * feat(memory): govern ledger slots and prompt winners * feat(macos): stage guarded ledger prompt adoption * feat(jit): adopt authoritative ledger prompts on macOS * fix(jit): close ledger adoption authority leaks * fix(jit): reauthorize every ledger migration write * fix(jit): fence ledger cutover publication * fix: keep ledger prompt rollback reversible * feat(jit): add guarded frame request retention contracts * fix(jit): close frame retention authority and evidence lifecycle * fix(jit): make frame retention retries and cleanup durable * fix(jit): make frame evidence recovery and retention complete * fix(jit): close frame retention recovery gaps * Harden temporary frame retention and deployment * fix: harden JIT frame retention and consumption * fix: close JIT frame lifecycle recovery gaps * fix: unify JIT frame authority and retention Failure-Class: FC-split-mutation-authority * docs: keep frame retention guidance lean * fix: retire duplicate frame flag bindings Failure-Class: FC-split-mutation-authority * fix: register frame keyframe queries Failure-Class: FC-split-mutation-authority * fix: serialize frame retention deploys Failure-Class: FC-split-mutation-authority * test: cover frame pixel deletion ordering * style: format cumulative Dart changes * fix(app): retain permanent conversation photo fetches * fix: bound frame vision retention and authority * fix: drain terminal frame request metadata * chore: record internal ledger adoption change * feat(memory): add dark daily sweep authority * feat(memory): harden daily sweep fences and runtime seam * feat(memory): reconcile existing standing triggers in sweep adapter * fix(memory): harden daily sweep recovery and source fences * fix(memory): close daily sweep source producers * fix(memory): close daily sweep review findings * Add dark daily memory sweep authority and recovery * fix(memory): harden daily sweep rejection repairs * test(listen): stub onboarding admission in bootstrap regression The daily sweep PR fences onboarding mode behind the server-owned backend admission (get_backend_onboarding_admission), so the bootstrap regression test now simulates an admitted session instead of failing closed on a real Firestore read. Verification: focused test passes in 1.64s (previously failed after a 4m27s Firestore timeout); full test_listen_runtime_regressions.py + test_onboarding_question_start.py: 26 passed; black --check clean. * fix(memory): close daily sweep rollout and retry cursors * fix(memory): isolate daily sweep lifecycle and retry fairness * Harden daily sweep admission and completed-day staging * fix daily memory sweep reliability boundaries * preserve daily sweep invocation tombstones * close daily sweep invocation lifecycle fences * fix: keep daily sweep lifecycle cleanup active * fix: acquire ledger snapshot client off event loop * fix(memory): preserve migration tier fence without legacy growth * test(memory): prove legacy adjudication race fences * fix(dev): allow bounded ADC readiness refresh * test: keep ledger prepush deterministic * test(memory): register prompt receipt control path * fix(memory): fence ledger writer transitions * feat(backend): preserve closed ledger history in export * feat(memory): define ledger query semantics * fix(backend): fence trigger snapshots on final authority * fix(backend): bypass stale coalesced JIT refreshes * feat(macos): mirror bounded memory evidence Decode generated v3 evidence into a domain mirror, persist canonical bounded JSON through the memory cache, and preserve it across compatibility sync and older-local conflicts. Invalid, future-shaped, oversized, and over-count payloads fail closed without hiding memory text or granting prompt authority. Tests: xcrun swift test --package-path Desktop --filter ServerMemoryV17DecodingTests Tests: xcrun swift test --package-path Desktop --filter MemoryLedgerMirrorTests Tests: python3 scripts/check_desktop_test_quality.py Failure-Class: none * fix(macos): fence and classify memory evidence Keep generated memory fields independent from malformed evidence, distinguish absent valid and invalid evidence states, preserve prior evidence on invalid payloads, and gate replacements on a monotonic server timestamp so stale active evidence cannot resurrect redacted rows. Cover populated-table migration upgrades. Tests: xcrun swift test --package-path Desktop --filter ServerMemoryV17DecodingTests Tests: xcrun swift test --package-path Desktop --filter MemoryLedgerMirrorTests Tests: python3 scripts/check_desktop_test_quality.py Failure-Class: none * fix(macos): preserve evidence fences and scrub redactions Advance evidence revisions for identical valid payloads, fence stale active responses after a local edit, and remove artifact/device pointers from redacted evidence before canonical persistence. Tests: xcrun swift test --package-path Desktop --filter ServerMemoryV17DecodingTests Tests: xcrun swift test --package-path Desktop --filter MemoryLedgerMirrorTests Tests: python3 scripts/check_desktop_test_quality.py Failure-Class: none * chore(macos): record ledger evidence mirror * feat(macos): deep-link local evidence cards to Rewind * fix(macos): fence Rewind frame evidence version * fix(macos): validate Rewind evidence card availability * fix(macos): bind task detail Rewind navigation to local leases * fix(macos): fence Rewind citation owner handoff * chore(macos): register Rewind evidence deep links * test(macos): cover Rewind evidence navigation * feat(desktop): evaluate JIT trigger watchlists locally * feat(desktop): wire authoritative JIT trigger runtime * feat(desktop): bind JIT claims to snapshot authority * fix(desktop): revalidate trigger authority at execution * fix(desktop): keep JIT execution leases live * test(memory): bind standalone reopen to direct-user writer * fix: make JIT QA sign-in self-contained Failure-Class: new Verification: bash desktop/macos/tests/test-jit-qa-target.sh; bash desktop/macos/tests/test-yolo-dev-backend.sh; repaired named-bundle Google sign-in reached authenticated onboarding. * feat(memory): complete JIT policy and native Windows parity * docs(backend): keep service map within context budget * test(macos): cover JIT client and staging flows * chore(backend): declare JIT mirror route policy * fix(backend): use strict Firestore boundary for JIT admission Failure-Class: FC-malformed-doc-read * chore(quality): register malformed-document guard surface * fix(backend): fail closed on malformed JIT authority Failure-Class: FC-malformed-doc-read * refactor(backend): name JIT workflow boundary results * test: repair JIT CI contracts * fix(backend): preserve ledger query exports Retain the explicit same-name re-exports consumed by tests and downstream callers while satisfying the enforced Pyright unused-import boundary after the main rebase. Failure-Class: none * test(backend): isolate gateway setup timing Failure-Class: none * style(memory): format direct-user evidence path Failure-Class: none * test(agent): isolate ACP process-group fallback Failure-Class: none * fix(dev-harness): preserve ownership markers in narrow CI * test(jit): refresh emulator fixtures for current contracts * test(jit): orchestrate local rollout dogfood * test(jit): harden local dogfood authority * fix(dev-harness): install PostHog for CI tests * fix(chat): project server JIT rollout into retrieval Resolve the backend-owned PostHog decision inside the bounded agent setup path and pass only its boolean result to prompt/tool configuration. Unknown or failed authority remains on the released legacy path, while callers cannot self-enroll through configurable input.\n\nVerification: backend/.venv/bin/python -m pytest -q backend/tests/unit/test_chat_async_offload.py backend/tests/unit/test_atomicity_lifecycle_regressions.py (41 passed)\n\nFailure-Class: new * fix(memory): preserve preference writer compatibility Select the agent preference write path from the canonical per-user writer control. Default compatibility mode retains the released MemoryService payload and receipt behavior; ledger mode keeps the retry-stable ledger write, and transition states fail closed.\n\nVerification: backend/.venv/bin/python -m pytest -q backend/tests/unit/test_chat_async_offload.py backend/tests/unit/test_atomicity_lifecycle_regressions.py (41 passed)\n\nFailure-Class: FC-split-mutation-authority * fix(jit): separate migration rollout authority Keep staged JIT chat and proactive exposure independent from legacy-row migration and writer cutover. Migration now requires its own default-off PostHog flag and still rechecks the shared kill switch at every mutation and publication boundary. Repair the isolated conversation-JIT fixture for main's chat-scope import. Verification: 217 focused JIT, chat-scope, migration, and lifecycle tests passed; 28 conversation-JIT fixture tests passed; independent Sol review accepted the split for QA-only dev rollout. Failure-Class: FC-split-mutation-authority * fix(photos): preserve retained image retrieval Treat an empty legacy inline marker as absent when permanent storage is authoritative, while malformed non-empty inline payloads still fail closed. Route live and retained thumbnails through the storage-aware image loader and preserve the conversation identity through the full-screen viewer.\n\nVerification: backend data-export tests 32 passed; Flutter photo-viewer tests 5 passed; focused Dart analysis clean; independent Sol review found and verified the viewer identity repair.\n\nFailure-Class: none * fix(memory): keep disabled daily sweep dark Resolve the backend-owned authority before inventory and require its literal true decision before any UID discovery, registry, cleanup, scheduler, model, or commit work. Missing, malformed, throwing, disabled, and kill-switched authority now exits without touching user data; enabled behavior is preserved.\n\nVerification: 60 focused daily-sweep job, scheduler, and inventory tests passed; independent Sol review accepted the fail-closed gate.\n\nFailure-Class: FC-split-mutation-authority * fix(jit): satisfy fail-closed type contracts * test(backend): admit full runtime contract checks * style(backend): format conversation bound test * test(backend): keep conversation router isolation current * test(backend): admit export boundary duration * fix(macos): persist failed chat turn notice Failure-Class: none * fix(macos): repair JIT rollout admission contracts Failure-Class: none * fix(windows): treat JIT screen evidence as untrusted Failure-Class: none * fix(backend): preserve explicit app failure contract Failure-Class: none * fix(app): finish photo viewer consolidation * fix(backend): make provider writes lock-free against the deletion gate The account-wide legal-hold deletion gate wrapped every GCS upload and Pinecone/Typesense upsert in an exclusive per-uid Firestore mutex with no lease: concurrent same-account writes hard-failed (dropped audio, lost vectors) and a crash between acquire and finish blocked the account's gated operations forever, with no janitor. Provider writes now use a lock-free fence that refuses only during account deletion or a live destructive operation; destructive kinds keep exclusive ownership, an abandoned gate self-expires after six hours, and releasing a gate on the failure path can no longer mask the original error. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(backend): issue onboarding admission at socket connect The completed-onboarding early exit returned False from an Optional[str] function; the listen runtime derives admission via 'is not None', so users who had already completed onboarding were admitted with a fabricated session id — the exact provenance forgery the admission exists to prevent. Separately, the 20-minute admission TTL was anchored to the app-launch state read, so a user reaching the speech-profile step late (or any client that never calls the state endpoint) silently lost onboarding questions and is_user tagging. The bootstrap now issues or refreshes the admission from the durable account state at connect time; completed accounts still can never re-enter, and issuing stays best-effort with the read failing closed. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(backend): keep the released proactivity lane open for legacy clients Gating /v1/desktop/proactivity/completions on the JIT cohort returned 403 to every non-admitted user — which is the entire deployed desktop fleet on deploy day, since shipped clients poll this route continuously and treat 403 as a plain error. Context-bucket extraction and the director would have died fleet-wide, dark cohort or not, and any environment without a PostHog key (local, self-host) would have lost the lane entirely. The route returns to merge-base admission semantics (tier quotas only); JIT admission remains enforced on the JIT reservation routes, and retiring this lane stays a later explicit operation after clients migrate. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(backend): withhold JIT tools and history reads outside the rollout Five new tools (search_knowledge, search_historical_facts, read_playbook, get_entity_timeline, look_at_frame) sat unconditionally in CORE_TOOLS, so every legacy chat request carried their schemas and the model burned tool budget on 'no entries found' answers. They are now filtered per request off the same resolved rollout boolean that gates the JIT prompt appendix. The memories-tab ledger-history endpoint likewise answered every user with a bounded 501-row provider scan that can only ever be empty outside the rollout; it now returns empty without the scan for non-admitted (and unknown/error) states. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(backend): bound rollout control-plane cost and confine sync resolution Synchronous callers resolved rollout flags via per-call asyncio.run against the shared provider singleton, crossing event loops: awaiting a Task attached to another loop raises, a timed-out asyncio.run strands a coalescer entry that then serves stale UNKNOWN forever, and the LRU cache was mutated from multiple threads. Sync resolution now runs on one long-lived control-loop thread with its own authority instance. Unknown snapshots gain a 5-second negative cache — UNKNOWN can never authorize work, and without it a fleet whose flags are simply absent pays one uncached PostHog call per conversation finalization. The screen-sync loop drops its force_refresh (one uncached decide per device per minute fleet-wide) and moves to its own rate bucket so two Macs' background sync can no longer starve conversation photo reads out of the shared 120/hour frame-requests bucket. The first-open policy's kill-switch telemetry label also reported str(Enum) instead of the value and could never match. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(backend): skip eager extraction under a non-compatibility writer mode A ledger-cutover user still ran the full L1 extraction model call at finalization, after which writer admission refused the compatibility write — the conflict retried, exhausted, and failed the entire finalization for every conversation, with the model spend already paid. Extraction now checks the canonical writer mode first and skips when the daily sweep owns memory formation; only a positively-read non-compatibility mode skips, so any control-state read failure preserves the legacy eager path. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(backend): export tolerates byte-less legacy photo rows A conversation photo row carrying the legacy empty inline marker and no storage reference failed the whole portability export forever, though it holds no durable image anywhere — there is nothing to omit. Such rows now export as metadata with a content-free gap reason. Frame requests in a retained state keep the fail-closed contract via an explicit require_bytes parameter. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(windows): harden JIT delivery, admission, and bootstrap boundaries Five verified defects: (1) the exclusive notification delivery slot leaked on any throw between reservation and commit — one SQLite hiccup during a JIT turn permanently silenced every proactive lane; the span is now try/finally-guarded and stale slots expire after ten minutes. (2) The ambient lane interpolated the raw window title into a tool-capable agent prompt; the turn now carries only the opaque context handle plus a sanitized executable name, framed as untrusted data like the nano-triage lane. (3) Google Calendar was fetched every ~60s before admission, so non-cohort users with Google connected paid ~1,440 reads a day for a refused feature; observation now gates calendar evidence on the cached authority. (4) Rollout-authority errors reset the cache and retried every frame (~1 req/s offline, forever); failures now back off from 30s to 10 minutes. (5) An unguarded JIT schema exec inside the shared database open could abort local storage for all features; the mirror bootstrap is now isolated, keeps the host-facing tables alive, and JIT stays inert when unavailable. Also re-checks the control-plane owner before committing the toast so an account switch mid-turn cannot show the previous owner's advice. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(macos): restore screen provenance, guard migrations, fence chat turns Four verified defects: (1) every pre-existing screen-derived task lost its 'Screen context / Open Rewind' source row because the new evidence policy dropped any provenance that is not rewind_frame.v1; the merge-base fallback row is restored for capture.v2/legacy refs (a test flipped to match the regression is restored to its merge-base assertions). (2) RewindDatabase published its pool before migrating, latching a failed migration into a permanent false-initialized state, and three unguarded ALTER TABLE memories migrations died with duplicate-column on machines that ran earlier builds of this branch; migration now precedes publication and the ALTERs/CREATEs are existence-guarded. (3) EventKit was queried on every context visit before the flags check; non-admitted owners now build no observation inputs. (4) A failed chat turn's reconstructed notice could be appended into a different conversation's transcript when the user switched sessions or cleared chat mid-flight; both transcript resets now revoke the active turn like selectApp already did. The pre-terminalized discard class (user Stop/watchdog) still drops the durable notice on relaunch — pinned by a characterization test in agent/tests/conversation-journal.test.ts with the least-invasive fix described there. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(testing): resolve firebase-tools from the checked-in dependency npx --prefix resolves the package bin against the current directory on some npm versions, and the admission runner deliberately launches from an isolated temp dir (firebase writes debug logs to cwd) — surfacing as 'sh: firebase: command not found' on hosts without brew node@22. Prefer the vendored node_modules binary when it matches the pin; npx remains the fallback. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(backend): keep one eager-extraction call site for the surface ratchet Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(backend): gate eager extraction at the public boundary The writer-mode skip moves from _extract_memories_inner to extract_memories: the replace-policy contract test pins the inner helper to exactly the canonical replacement path, and the public boundary is the better seam anyway — a sweep-owned user now skips parity capture and usage tracking along with the model call. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(listen): stub onboarding admission issuance in bootstrap regression The connect-time ensure call landed in a harness that only stubbed the read, so the bootstrap test paid an extra real-module exception path and grazed the 0.30s fast-unit CPU budget under fanout load. Stub the issuance like the read. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(listen): allowlist the bootstrap regression's CPU budget The full listen-runtime bootstrap test measures exactly at the 0.30s fast-unit CPU budget under a saturated pre-push fanout (CPU inflates ~2x there per the guard's own notes) while passing comfortably alone. It exercises deliberately heavyweight machinery; record it as an intentional exception rather than trimming the coverage. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(backend): keep list(CORE_TOOLS) literal through JIT tool gating The JIT-only tool filter replaced the list(CORE_TOOLS) assignment with an inline comprehension, which broke the prompt-cache structural invariant (test_prompt_cache_optimization.py::test_core_tools_used_in_both_functions). Restore the list(CORE_TOOLS) copy and apply the JIT-only filter as a conditional pass, preserving rollout semantics and tool order. * feat(jit): drop automatic goal updates from the JIT featureset Product decision (David, 2026-08-26): goals change only through explicit user action for JIT-admitted conversations. Goal progress is no longer a first-open obligation — the effect is removed from FIRST_OPEN_EFFECTS and the worker, and the policy plan can no longer express deferring it. Legacy obligations carrying a pending goal_progress row are normalized away and complete on the remaining two effects. Non-JIT (legacy eager) conversations keep today's automatic goal updates unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(sweep): one summary-spine agent pass per day, with folder backstop Replaces the per-conversation transcript extractor in the completed-day producer with a single two-phase agent run: the whole day's conversation summaries go in as one bounded spine (200 conversations / 120k chars — effectively unreachable, so heavy days no longer stall the cursor), and the agent may request up to 8 raw transcript excerpts (8k chars each) to verify specifics before finalizing. At most two provider calls per user per day, both inside the existing at-most-once invocation fence; the staged page carries the memory candidates AND folder assignments for the day's unopened, unfiled conversations, applied idempotently (first-open or user assignment always wins). Memories must cite their source conversations; uncited output is dropped. The cost gate becomes a worst-case ceiling checked before any call. The onboarding cold-start channel keeps per-conversation transcript extraction unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(sweep): harden the daily agent prompts from a real-data lab pass Iterated on one real heavy day (26 conversations) with strong- and weak-model stand-ins, an adversarial judge, and hand-verified transcript ground truths. Rules added, each pinned to an observed failure: actor binding in active voice with a personal-attribute gate (a discussed or recommended topic is never someone's attribute; judgments about named people are stored as assessments); decision-state basis labels binding the verb (decided/proposed/observed, discussed-no-outcome dropped); salience ordering (money, metrics, named-party intent, identity, and durable decisions before any operational fact; one fact per memory); never guessing the direction of an invitation/offer/commitment (verify or drop); and no deferring the whole answer to verification. The agent output schema gains a 'basis' field. The memories QoS call-site inventories now count the daily-sweep agent's call site (3 -> 4). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(sweep): tune the daily agent prompts against the real memories model Ran the assembled prompts against gpt-5.6-luna (the real 'memories' route model) on the same real day. Three refinements from observed behavior: the basis label no longer leaks into memory text (metrics read as metrics, not 'David observed that…'); the never-guess-direction trigger is mechanical (passive/verbless summary phrasing or 'Speaker' as the actor forces a transcript_request — luna confidently inverted 'Tim: Invited to New York' until this; with it, phase B verifies and corrects to the true direction), hedging is itself a request signal, and nothing high-salience may be silently dropped; and a rich-day yield anchor (8-16 memories for 15+ conversations) counters the model's over-pruning without inviting padding. Final real-model run: 11 true memories + 2 legitimate verification requests, zero fabrications, ~22k tokens (~2 calls) for a 26-conversation day. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(sweep): profile-maintaining slots, ledger lookups, cache-ready prompts The daily agent now sees the user's current profile (the same get_prompt_memories seam chat uses — the ledger render for migrated users), may run up to 4 owner-scoped prior-memory keyword lookups (provider fail-soft; hits re-read through the canonical store before disclosure) to dedup and supersede, and may name a slot for standing attributes — an occupied slot becomes an amend through the existing canonical occupancy check, so the daily run maintains the rendered profile with no second write path. Both phase prompts share a byte-identical prefix (pinned by a test) and pass a per-user prompt_cache_key through get_llm; measured against gpt-5.6-luna the provider cache is exact-match rather than prefix-based today, so this is future-proofing rather than present savings. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(sweep): type the memory-searcher seam for the pyright contract CI's authoritative typecheck rejected the untyped lookup seam (memories.py: list(Any or [])). The searcher is now Optional[Callable[[str], Sequence[str]]] and results are built through a typed comprehension; behavior unchanged (absent or failing searcher still degrades to an empty result block). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: repair four main-inherited CI breakages after sync origin/main is currently red on its own tip; syncing it into this PR inherits the breakage, so the fixes ride here: - subscription.py: drop the unused get_byok_keys import (pyright reportUnusedImport fails the Backend unit suite). - AppState+Transcription.swift: explicit self for alertPresenter inside the escaping showAlert completion (strict-concurrency compile error in all three Desktop Swift lanes, shipped red on main by d49f978512). - AppState+Permissions.swift: pinned swift-format drift from the same main commit (desktop-swift-format-lint). - web/app/bun.lock: add the prettier + prettier-plugin-tailwindcss entries 64db30c791 pinned in package.json without updating the lockfile (frozen install fails web-app-checks). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(sweep): close the second review round's findings Three parallel adversarial reviews over the post-takeover additions: - Clamp every model-controlled phase-B input (draft memories, request reasons, lookup queries/results) and add the clamped worst case to the pre-call cost ceiling, which previously under-estimated phase B. - Attest an empty consumed day when the staged page carries an older stage schema version instead of stalling the cursor forever on every deploy-boundary schema bump. - Make the folder backstop's unfiled check and write share one transaction so a concurrent first-open/user assignment always wins. - Let equal-rank sweep candidates amend sweep-authored slot occupants: the profile-maintenance path froze after a slot's first write. User statements still always win; slotless subject matches still dedup. - Neutralize ``` fences in summaries/excerpts/lookup results, and mark raw-transcript fallback rows '(unstructured transcript excerpt)' with a prompt rule refusing slots/personal attributes from them without transcript verification (test pins the marker to the rule). - Remove the dead first-open goal-authority threading left by the goals removal, and update the stale jit-first-open-runtime doc. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: repair three more main-inherited breakages All shipped red on main and only surfaced once earlier failures were cleared: - AppState.swift: move the alertPresenter default out of the stored property initializer — Xcode 16.4's SILGen segfaults (signal 11) emitting it, which failed all three Desktop Swift lanes even after the explicit-self fix. - test_byok_security.py: main's BYOK rewrite (d0e3a4eb3a, 1da8880175) changed request_has_llm_byok_key to per-provider enrollment checks and made partial headers fail closed, but left the tests targeting the old get_byok_keys()-based lenient contract (masked on main because pyright failed before pytest ran). The tests now assert the shipped strict contract their own docstrings already describe. - subscription.py: pinned-black formatting for the BYOK fallback expression (the Formatting lane rejects the file as main wrote it). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): stub the chat-agent gateway route pin in the chat router harness Main's a6988be309 made routers.chat import CHAT_AGENT_ROUTE_DIRECT / get_chat_agent_route from utils.llm.gateway_client, but the chat-router test harness (and test_chat_file_upload_unsupported's local override) stub utils.llm.gateway_client without those symbols, so every suite that loads the real router failed at import — masked on main because pyright fails its Backend unit suite before pytest runs. Ninth main-inherited repair in this sync. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): teach test_chat_quota's utils.byok stub the rewritten import surface utils/subscription.py now imports get_byok_uid and get_cached_byok_state (main's BYOK rewrite); the module-scoped utils.byok fake predates them, so reloading subscription under the fake raised ImportError at setup — and the polluted process took test_chat_openapi_operation_ids and test_desktop_screen_crisp down with it in CI's batched run (all three pass standalone). Tenth main-inherited repair, same pyright-masked pattern. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): update three more suites for main's BYOK/gateway import surface Same pyright-masked pattern as the harness and test_chat_quota repairs: - test_desktop_transcribe stubbed utils.llm as a non-package, so routers.chat's new utils.llm.gateway_client import could not resolve (50 failures); the submodule is now in its stub list. - test_paywall_reconnect_gate's BYOK escape-hatch tests never set the request uid context that the enrollment-verifying rewrite requires (middleware sets it in production); they now do, and teardown clears it. - test_chat_session_app_identity's enforce_chat_quota stub rejected the new required_llm_provider keyword. All three suites pass locally (69 + 35 + 6). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): enroll fingerprints in the desktop BYOK tests PR #11454 moved macOS BYOK activation to enrollment-verified fingerprints (isByokActive and usableBYOKEnvironment gate on persistEnrolledFingerprints), and its own test lanes shipped red: the tests store raw keys but never enroll them, so every key reads as inactive. Their teardowns already clear enrollment — the setups now enroll what they store, matching the production activation path. All 8 previously-failing cases (BYOKPaywallTests + the two AgentRuntimeProcessTests BYOK-environment cases) pass locally. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(deploy): enable the daily memory sweep on development The sweep's five deployment inputs were pinned off in every environment, so cohort enrolment alone could never start it -- turning it on for a dogfood account required a second PR. Development now carries the live values: - ENABLED/MODEL_ENABLED on, so the job stops exiting at its first authority gate and the model authority can budget a route. - MODEL_NAME pinned to gpt-5.6-luna, which is the declaration interlock the runner checks against get_model('memories') before any provider call. - MAX_MODEL_COST_USD 0.80, the worst-case pre-call ceiling for a maximal day including phase B's clamped draft/reason/lookup overhead. - COHORT_ENABLED on with COHORT_FLAG daily-memory-sweep-v1, so enrolment is a per-uid PostHog boolean and an unnamed cohort stays a closed rollout. Production is deliberately untouched and stays fully pinned off. The job still cannot form a memory for anyone until that flag exists and resolves true for a uid, which remains a control-plane action rather than a deployment one. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(firestore): terminate the daily-sweep occupant indexes with __name__ The six daily-sweep occupant lookups were the only declarations in the manifest without a trailing __name__ field -- 63 of 69 entries carry one, and main had none missing it. Firestore appends the terminator itself and reports the index back that way, so these six could never match the live inventory. The failure mode is not a missing index; the indexes build fine. It is that reconciliation never converges: every run reports the same six as missing, tries to create them, and fails on ALREADY_EXISTS. That takes down the Firestore schema workflow on both environments permanently, and with it the development backend deploy's readiness gate -- the same class of outage the workflow's own header records from the hourly_usage index in PR #11979. The derived specs previously appended their extra predicates to the base spec's index_fields, which would have placed them after the terminator, so the shared prefixes are now named explicitly and each spec ends with __name__. Verified against real Firestore: reconciliation reports zero missing indexes in both based-hardware and based-hardware-dev. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: close final JIT rollout and CI gaps Fence direct JIT tools and frame pixels, keep Windows account wipes safe after optional schema failures, and repair inherited CI regressions. Failure-Class: none --------- Co-authored-by: David Zhang <9387252+Git-on-my-level@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> | 27 天前 | |
feat(dev-harness): fake-backed V1 live-session broker (reload, restart, controls, evidence, owned teardown) (#14362) * harden(dev-harness): define V1 live-session and V8 spine contracts Add strict Python/Dart pending contracts and introducing-commit protection. Reserve the live-session CLI, wire, ownership/attachment seams and failing builder acceptance tests. Add live attribution before freezing evidence v1. Verification: - bash scripts/dev-harness/run-tests.sh: 287 passed, 7 existing skips, 40 strict xfailed (24 pending V1 markers), Python 3.11.15. - make preflight: 16 checks passed, using supported PYTHON selection. - make mobile-verify fast --paths harness/session contract inputs with an absolute evidence directory: 20 executed, 20 passed, no failures/skips. - bash app/test.sh: 2014 passed, 5 skipped, one unchanged UTC+7 search_rank_grouping_test failure; affected baseline inputs match main. - Real module CLI refusals exit 2 without traceback; regression covered. - Dart SDK probe proves ext. prefix required; marker-removal formatter probe proves no incidental Dart diff. Runtime is intentionally unimplemented. App UI must repair current invalid VM extension names before live acceptance; unsupported journeys block. Simulator/emulator latency and detached-process acceptance remain V1 work. make setup installed hooks but full backend wheel sync stalled; isolated Python 3.11 harness dependencies ran the owner checks without any skip flag. Failure-Class: none * harden(dev-harness): scope the spine-contracts check to spine paths The check protects files under the spine roots and their registry, so it only needs to run when those change. A repo-wide trigger would run it, and its full-history requirement, on every unrelated contributor PR. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * docs(dev-harness): clarify the live journey adapter boundary V1 must refuse live journey execution until a shared executable adapter exists; B1 addressability alone cannot attach Dart integration tests. Pure admission contracts remain valid with injected supported IDs. Verification: harness 287 passed, 7 skipped, 40 strict xfailed; make preflight passed 20 checks. * harden(dev-harness): pin reviewed spine revisions and align B0 capabilities Coordinator turn-4 review authorizes the V1 fixture correction. Exact before/after digests pin the revised oracle; immutable revision records do not permit builder assertion edits. Active tests reject missing/tampered records and prove the fake machine response. Verification: full harness 289 passed, 7 skipped, 40 strict xfailed; UTC app 2020 passed; spine 24 pending/7 protected; make preflight 16 passed. No live device acceptance claimed. * harden(dev-harness): preserve marker retirement across spine revisions Track retained marker occurrences rather than counts so retiring a second test cannot restore the first test's retired marker. The active regression covers this swap and legitimate continued retirement. Verification: full harness 290 passed, 7 skipped, 40 strict xfailed; spine 24 pending/7 protected; make preflight 16 passed. Dart sources unchanged after the earlier UTC app 2020-pass run. * harden(dev-harness): retain corrected spine oracles through squash merges A squash can introduce a corrected contract and its review records in the same commit. Absorb only the revision prefix ending at that introducing file's exact digest, preserving immutable records and later marker-only edits. Active temp-repo regression covers multiple corrections in one squash and rejects restoring old or altered assertions. Verification: harness 291 passed, 7 skipped, 40 strict xfailed; B1 make preflight 20 passed; final B1 journeys 20/20. Dart unchanged after UTC app 41 opt-in and 2051 ordinary passed. * test(dev-harness): revise live session acceptance after builder review * test(dev-harness): bind machine fixture to actual child identity * feat(dev-harness): implement the fake-backed V1 live session broker LiveSession speaks the Flutter machine wire against an injected child, records BrokerIdentity in live.json, and refuses real flutter run from CLI dispatch. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dev-harness): tear down live before stop, reset, recover, and release Lifecycle order is live teardown, then services, then device detach; recover still has generation 1 while teardown runs. Co-authored-by: Cursor <cursoragent@cursor.com> * test(dev-harness): retire V1 live-session pending markers The fake-backed broker now satisfies those spine tests; only whole pending-marker lines were deleted. Co-authored-by: Cursor <cursoragent@cursor.com> * test(dev-harness): cover lazy live-session holes with real child processes A dead flutter child mid-reload is blocked, never success; teardown will not signal a live PID whose start time, marker, or boot id does not match. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): hash iOS .app bundles as directory trees file_sha256 raised IsADirectoryError on a simulator .app; identity is now the sorted tree of regular files. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): pin live-session source, generation, and stdio The architect's thirteen probes rejected attributing a compile to source observed after launch, announcing stopped after a failed reap, and grepping one fixture secret. Keep one monotonic RPC deadline and actually validate negotiated capabilities. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): fence live publication, boolean readiness, and auth redaction _assert_generation only covered admission, so a lease roll during ext.omi.controls.state still published ok and advanced loaded identity. Recheck lease and child generation after daemon/readiness work, refuse non-boolean readiness, and redact complete Authorization values before logs are retained. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): restore live teardown after main merge and retire passed T7 markers Keep main's port-claim mobile_session.py; put back V1 live teardown on stop/reset/recover. The formatted round-7 oracles now pass, so remove only the pending-marker lines. Co-authored-by: Cursor <cursoragent@cursor.com> * chore(dev-harness): restore app sources that the merge hook reformatted The merge commit ran dart format without package:flutter_lints resolved and rewrote main's files. Restore origin/main bytes so the V1 PR does not carry unrelated UI diffs. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): fail fast when a live owner holds the session lease Name the session and holder pid immediately. The same overlap that used to look like a controls-extension miss is a held lease. Co-authored-by: Cursor <cursoragent@cursor.com> * test(dev-harness): retire V1 pending markers that main's live fences now pass Merging origin/main imported strict xfails for lease-roll, negotiation, and auth-redaction fences this branch already implements. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): fail closed on non-boolean live readiness flags Finding 2 of the adversarial review reproduced: a fake reporting readiness.signedIn as the string "false" was classified not-ready and fell through to the wait_ready recovery path, which then observed real booleans and published ok. _is_ready now raises LiveError (malformed-response) for present-but-non-boolean flags, so start blocks instead of self-healing; start/reload/restart wrap it and surface blocked/malformed-response. The stringy-false regression test's expected error_code moves unready -> malformed-response; blocked-stays-blocked is unchanged. The satisfied V1 pending marker on the boolean probe is retired. Findings 1, 3 and 4 did not reproduce at this head (probes re-run verbatim; evidence in PR comment). Failure-Class: new --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Cursor <cursoragent@cursor.com> | 5 天前 | |
harden(dev-harness): live-session V1 contract, evidence-v1 freeze, and strict pending contracts (#14310) * harden(dev-harness): define V1 live-session and V8 spine contracts Add strict Python/Dart pending contracts and introducing-commit protection. Reserve the live-session CLI, wire, ownership/attachment seams and failing builder acceptance tests. Add live attribution before freezing evidence v1. Verification: - bash scripts/dev-harness/run-tests.sh: 287 passed, 7 existing skips, 40 strict xfailed (24 pending V1 markers), Python 3.11.15. - make preflight: 16 checks passed, using supported PYTHON selection. - make mobile-verify fast --paths harness/session contract inputs with an absolute evidence directory: 20 executed, 20 passed, no failures/skips. - bash app/test.sh: 2014 passed, 5 skipped, one unchanged UTC+7 search_rank_grouping_test failure; affected baseline inputs match main. - Real module CLI refusals exit 2 without traceback; regression covered. - Dart SDK probe proves ext. prefix required; marker-removal formatter probe proves no incidental Dart diff. Runtime is intentionally unimplemented. App UI must repair current invalid VM extension names before live acceptance; unsupported journeys block. Simulator/emulator latency and detached-process acceptance remain V1 work. make setup installed hooks but full backend wheel sync stalled; isolated Python 3.11 harness dependencies ran the owner checks without any skip flag. Failure-Class: none * harden(dev-harness): scope the spine-contracts check to spine paths The check protects files under the spine roots and their registry, so it only needs to run when those change. A repo-wide trigger would run it, and its full-history requirement, on every unrelated contributor PR. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * docs(dev-harness): clarify the live journey adapter boundary V1 must refuse live journey execution until a shared executable adapter exists; B1 addressability alone cannot attach Dart integration tests. Pure admission contracts remain valid with injected supported IDs. Verification: harness 287 passed, 7 skipped, 40 strict xfailed; make preflight passed 20 checks. * harden(dev-harness): pin reviewed spine revisions and align B0 capabilities Coordinator turn-4 review authorizes the V1 fixture correction. Exact before/after digests pin the revised oracle; immutable revision records do not permit builder assertion edits. Active tests reject missing/tampered records and prove the fake machine response. Verification: full harness 289 passed, 7 skipped, 40 strict xfailed; UTC app 2020 passed; spine 24 pending/7 protected; make preflight 16 passed. No live device acceptance claimed. * harden(dev-harness): preserve marker retirement across spine revisions Track retained marker occurrences rather than counts so retiring a second test cannot restore the first test's retired marker. The active regression covers this swap and legitimate continued retirement. Verification: full harness 290 passed, 7 skipped, 40 strict xfailed; spine 24 pending/7 protected; make preflight 16 passed. Dart sources unchanged after the earlier UTC app 2020-pass run. * harden(dev-harness): retain corrected spine oracles through squash merges A squash can introduce a corrected contract and its review records in the same commit. Absorb only the revision prefix ending at that introducing file's exact digest, preserving immutable records and later marker-only edits. Active temp-repo regression covers multiple corrections in one squash and rejects restoring old or altered assertions. Verification: harness 291 passed, 7 skipped, 40 strict xfailed; B1 make preflight 20 passed; final B1 journeys 20/20. Dart unchanged after UTC app 41 opt-in and 2051 ordinary passed. * test(dev-harness): revise live session acceptance after builder review * test(dev-harness): bind machine fixture to actual child identity --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> | 6 天前 | |
Merge origin/main into PR #10171 | 1 个月前 | |
fix(dev-harness): distinguish failed Android inventory from a missing engine (#14400) sdkmanager --list_installed exiting nonzero was reported as engine present. A missing emulator binary, a successful empty listing, and a failed look are three different findings; cmdline-tools 23 slash paths still prove an image is installed. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> | 6 天前 | |
SCA-486: isolated mobile sessions, local-dev auth, capture replay, and seeded journeys (#14213) * mobile sessions: freeze session-evidence-v1 receipt contract Freeze the minimal versioned session/evidence schema shared by the mobile development-foundation consumers (C2 journeys, C3 capture replay, C4 verification, C5 devices) and its executable validator. - contracts/session/session-evidence-v1.schema.json: closed v1 object binding source SHA + dirty digest, built artifact identity, loopback-only endpoints, fixture/runner versions, real timestamps, status/blocked reason and exact execution counts. No credential-shaped field exists anywhere. - dev_harness/session_evidence.py: builder + validator enforcing the cross-field semantics a JSON Schema cannot express: ready/running require an artifact whose git_sha matches the source (a stale build cannot be reported ready), production-family profiles are rejected as session targets, counts must account exactly, zero-execution receipts are only valid pre-run, and credential-shaped keys are refused at any depth. - Egress guards: validate_local_http_url/validate_local_host_port reject non-loopback or non-plain-HTTP endpoints (api.omi.me / api.omiapi.com by name) before any request is attempted. Evidence: scripts/dev-harness/run-tests.sh (test_session_evidence.py, 20 contract tests incl. stale-artifact, egress, credential and accounting rejections). * mobile sessions: deterministic synthetic auth fixture v1 Seed a synthetic Auth-emulator user through the local-dev custom-token endpoint contract from open PR #11784 (feat/app+backend local-development sign-in without OAuth) — reuse, not a competing endpoint. The PR stays with its owner; scripts/dev-harness/MOBILE_SESSIONS.md records the integration plan and provenance. - fixtures/mobile/v1.json: one deterministic user (omi-fixture-v1-user-1@local.test); RFC-reserved domain so fixture identities can never collide with a real account. No real Google/Apple user, provider key, or copied token involved. - dev_harness/mobile_fixtures.py: fail-closed seeding client — the backend URL is validated as loopback plain-HTTP before any request, production hosts are denied by name, and the persisted receipt records identity and outcome only: token_minted/token_retained, never the token itself. Evidence: test_mobile_fixtures.py (16 tests) — determinism, reserved-domain enforcement, pre-request egress refusal, 404/unreachable/wrong-uid fail-closed paths, credential-free receipts. * mobile sessions: structured doctor for the session lanes Every readiness failure classifies exactly one of ready / agent-remediable (with the exact resumption command) / operator-action-needed (privileged install, license, host capacity), per lane (backend, android, ios). - Flutter version is read from the mobile CI pin in .github/workflows/mobile-app-checks.yml — never 'latest'; inconsistent pins refuse rather than guess. - Backend lane: python3.11 (venv must be 3.11, ambient 3.14 must not select the runtime), JDK 21 for the firebase emulators, firebase-tools, redis/typesense via native binary or a responding docker daemon. - Android lane: ANDROID_HOME + adb + emulator engine + system image, each with the exact sdkmanager remedy and a capacity-gated download note. - iOS lane: Xcode + simctl runtime; missing runtime is an operator action. - Capacity: <12GiB free on the shared Data/scratch container is an operator gate for emulator/build lanes (agents never free space themselves); contract/unit lanes skip it via --skip-capacity. - Egress: an ambient production OMI_LOCAL_API_BASE_URL override is reported as a blocking misconfiguration. Evidence: test_mobile_doctor.py (13 tests) over an injected runner — lane filtering, ready/degraded/blocked classification, pin parsing, capacity and operator-gate behavior; live run on m1-mac-studio via 'make mobile-session ARGS="doctor --platform android --platform ios"' reports backend+ios ready, android emulator engine agent-remediable. * mobile sessions: isolated session lifecycle CLI behind one entrypoint 'make mobile-session ARGS="…"' (scripts/dev-harness/mobile-session.sh) owns a uniquely-leased local mobile session: doctor / acquire / start / seed / reset / status / evidence / stop / recover / release. A session is an existing dev-harness instance + port offset + device lease + seed receipt + evidence receipt — the harness lifecycle is reused in-process under OMI_LOCAL_INSTANCE/OMI_HARNESS_PORT_OFFSET, not duplicated. Ownership is fail-closed: - leases are created atomically (O_EXCL) with owner host/user/pid and a harness-standard sentinel; a live foreign owner or another local user's session is never touched; cross-host takeover is an operator decision; recover bumps the generation for same-host/same-user takeovers. - ports come from a claimed offset registry; a foreign process occupying a port is refused (never killed) and the allocator skips that offset; release frees the claim only when it belongs to the session. - start gates device attach on doctor readiness (precise blocked reason, not a crash); ios-simulator devices are created/booted/deleted session-owned via simctl. - seed/reset/stop/release are idempotent; reset only touches the session's own harness instance (sentinel-validated underneath). - evidence emits session-evidence-v1 receipts; ready/running refuse without a bound artifact and refuse when the source moved since acquire. app/setup.sh (separate commit): OMI_IOS_DEVICE_ID pins non-interactive device selection; OMI_DEVICE_SUFFIX overrides hostname identity. Evidence: test_mobile_session.py (19 tests) + wrapper tests — exclusivity, disjoint ports, dead-owner/live-foreign/different-user/cross-host refusals, foreign-port refusal with a real live listener, idempotent release, artifact binding, stale-source refusal, harness env handoff. Live CLI run: acquire/list/evidence/seed-fail-closed/stop/release with exit codes 0/2 on m1-mac-studio. * app/setup.sh: non-interactive device pin and per-session device suffix - OMI_IOS_DEVICE_ID: when set, select_ios_device uses exactly that device id, failing precisely (with the available device list) when absent, instead of enumerating and prompting — the mobile-session harness, CI and nested agents cannot answer an interactive prompt, and the current no-TTY path errors out whenever more than one iOS destination exists. - OMI_DEVICE_SUFFIX: let a session harness (or a second checkout on one host) inject a unique device-identity suffix instead of the hostname, which collides across concurrent sessions on the same machine. Unset behavior is unchanged. Verified by sourcing the function with a stubbed flutter devices --machine: pinned-present emits the id; pinned-absent fails with the list; unpinned multi-device no-TTY keeps the existing enumeration failure. * mobile sessions: apply repo python formatter to the new modules black 26.5.1, --line-length 120 --skip-string-normalization via scripts/backend-python-format; behavior unchanged, dev-harness lane re-run green (214 passed; 1 pre-existing environmental failure — the host's global git worktree guard blocks pytest-tmp linked worktrees). * test: place linked-worktree pytest fixtures under OMI_WORKTREES The managed git wrapper correctly refuses worktrees in /private/tmp. Keep that guard and put the fixture where task worktrees are allowed. Co-authored-by: Cursor <cursoragent@cursor.com> * docs: treat session-evidence-v1 as proposed until consumers review it C1 shipped the schema; freeze it only after C2/C3/C4 agree, not from a single worker declaration. Co-authored-by: Cursor <cursoragent@cursor.com> * feat: reuse PR 11784 local-dev custom-token auth on current main Copy the reviewed emulator-gated sign-in path onto this integration branch so synthetic seed talks to real local services. Leave the original PR open and unmerged. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: boot isolated sessions on real CoreSimulator IDs and offline STT Use the installed iPhone 17 Pro / iOS 26.5 identifiers, pin PROVIDER_MODE=offline, and drop soniox from the offline STT chain so the local backend can start without a paid key. Co-authored-by: Cursor <cursoragent@cursor.com> * test: isolate provider-secret fixtures from ambient PROVIDER_MODE A previous offline session left PROVIDER_MODE in the shell and made the secret-injection tests read ambient offline instead of the fixture file. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(app): injectable capture seams for deterministic recovery replay CaptureController and the phone WAL resolved clock, timers, connectivity, auth, mic, socket and upload policy through global singletons, so the capture -> WAL -> recovery path could not be replayed deterministically. Add narrow constructor seams (capture_seams.dart) with production-identical defaults: CaptureScheduling, CaptureAuthBoundary, CaptureConnectivityBoundary, plus wal/phoneMic/clock/scheduler injection on CaptureController; clock/periodic/job-status injection on LocalWalSyncImpl threaded through WalSyncs/WalService; periodic-timer injection on the NativeMicRecorderService watchdogs. The in-progress-conversation loader seam now covers the socket-connect path too, and streamRecording honors the microphone permission requester like the batch path already did. No behavior change with default construction; every seam is optional. Evidence: bash app/test.sh (1990 passed, 5 pre-existing skips); analyze ratchet green. * test(app): deterministic capture-recovery replay schedules (SCA-489/C3) Replay the REAL production capture pipeline (CaptureController, NativeMicRecorderService, TranscriptSegmentSocketService, WalService, RecordingTransferCoordinator) against controlled external I/O: virtual clock, manual bounded scheduler, scripted transport/upload boundary, fake native host. Restart evidence destroys and reconstructs the object graph from real temp files (torn wals.json -> backup recovery, missing audio -> terminal corruption, process kill -> disk reload and re-upload). Six schedules with invariant oracles: network loss/reconnect mid-capture (exact frame identity in the stored WAL, single upload), stale native events after stop/new session (session-identity gate, no double teardown), interruption/resumption (live + batch, bounded stall escalation), partial/torn persistence plus reconstruction, failed upload with bounded backoff and persisted/enqueued/server-acknowledged distinctions, and ownership transition (signed-out reconnect cancellation, bounded 4001 token refresh). Also publishes the C2/C4 adapter (capture_scenario.dart: catalog + result contract) and the C5 native-event vector schema (phone-mic-native-events/v1) mirroring the Pigeon PhoneMicFlutterApi contract without touching Pigeon. Falsification evidence: removing the NativeMicRecorderService session gate flips the stale-idle schedule to failure (record->stop); removing the finalizeCurrentSession unsynced-retention guard drops the WAL and fails the network-loss schedule. Evidence: flutter test test/unit/capture_recovery_replay_scenarios_test.dart (17 passed); bash app/test.sh full suite green. * chore(app): allowlist SCA-489 replay contract libs in the dead-code ratchet The scenario catalog/result contract and the native-event vector schema are library-only by design until the C2/C4 and C5 lanes import them; the ratchet demands an explicit allowlist entry with a reason for exactly this case. * test: pin conversation-window capture session id across sequential phone-mic lives activeCaptureSessionId is WAL/conversation-scoped so a late ConversationEvent can still stamp WALs. C2 must use activeRecordingId as the live recording identity. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: wait for unawaited capture-upload retries before asserting and teardown CI failed the bounded-backoff replay because cooldown wakes are unawaited and settle used wall-clock sleeps that missed the drain under load, then deleted the temp WAL dir mid-write. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(app): typed debug semantic controls for the local journey lane Extends the existing debug Marionette surface (same debug VM-service transport, no new server/framework) with product-semantic controls: versioned capabilities (semantic-controls/v1), privacy-safe state (route, principal, capture lifecycle with activeRecordingId as the authoritative recording identity), bounded wait_ready, production-path navigation, and named journey faults. Fail-closed eligibility: kDebugMode AND local_dev profile AND OMI_DEV_CONTROLS=1 dart-define. Ineligible builds (including production-flavor debug) install nothing and the HTTP fault chokepoint is a pure pass-through — pinned by semantic_controls_guard_test.dart. Narrow seams added for the hermetic journey lane: - AuthService.installLocalHarnessTokenGateway (debug+local_dev gated Firebase token I/O boundary; isSignedIn routes through the gateway) - PlatformManager.initializeForLocalHarness (header fields only) - CrashlyticsManager report paths tolerate a missing Firebase app the same way main.dart's zone handler already does, so host-lane errors surface instead of being masked by [core/no-app] Verified: flutter test test/unit/semantic_controls_guard_test.dart (10 passed); auth regression suites (34 passed); C3 capture replay (17 passed); dead-code ratchet at baseline. * test(app): five strict seeded acceptance journeys with negative fault variants Canonical executable definitions (one per behavior) under app/integration_test/journeys/, runnable hermetically (flutter-tester + loopback fixture backend) or on a simulator via run_journeys.sh: j1 seeded conversation detail — real provider fetch + real detail page, exact synthetic identity; negative: wrong-owner session refused. j2 chat send -> distinct assistant reply — real input/send-button keys (omi.chat.input / omi.chat.send), request observed server-side, server-minted ai-role reply distinct from the prompt, rendered; negatives: suppress-send, suppress-assistant-reply, wrong-owner-session. j3 memory create/edit surviving reload — production provider path, server-minted id required after reload; negative: drop-memory-save. j4 expired session — transient failure re-mints via the real custom-token endpoint; terminal failure emits expiry and blocks requests; negative: production-family profiles never silently re-mint. j5 capture interruption/reconnect — C3 capture-scenario/v1 adapter: real temp files, process reconstruction, drain exactly once; negative: fail-capture-recovery. Each negative arms exactly one named fault and must fail with the invariant named. Evidence receipts follow session-evidence-v1 accounting; zero-execution runs never pass. Verified: bash integration_test/journeys/run_journeys.sh (5/5 pass); repeated deterministic vertical: bash integration_test/journeys/run_journeys.sh --filter j2 --runs 5. * fix: stamp journey evidence finished_at at write time Receipts were recording construction time as the end timestamp, so duration could not be distinguished from start. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(app): clear new analyzer-ratchet regressions in the integrated journeys The integrated checkpoint (b217eb1a9d) fails app/scripts/analyze_ratchet.sh with 7 new occurrences: 3 unused imports plus a bogus 'show WalStatus' in j5, an unused-looking nested import that actually provides SingleChildWidget in hermetic_boot, a missing const in j4, and two depend_on_referenced_packages for test-only platform interfaces. Declares path_provider_platform_interface and nested as direct dev dependencies (same pattern as the existing web_socket_channel dev deps) and removes the dead imports. Mechanical lint repairs only — j4/j5 hermetic journeys re-run green after the change. * feat: unified mobile verify lanes, mechanical journey selection, and CI/contributor path (SCA-490/C4) One canonical verification entrypoint over the proven lanes: make mobile-verify select|doctor|fast|smoke|physical (scripts/dev-harness/mobile-verify.sh -> dev_harness.mobile_verify). It never adds a second runner: journeys delegate to the C2 canonical runner and session infrastructure to the C1 session CLI. Selection is mechanical and fail-closed: journeys are glob-discovered (runner --list contract test), changed paths map through a per-seam rule table, unknown app/lib impact falls back to the full suite, and an empty selection is drift (exit 65), never PASS(0) — run_journeys.sh now fails closed the same way. Receipts are validated against session-evidence-v1 accounting (honest counts, zero-execution never passes) and every lane writes a verify-receipt.json binding source SHA + dirty digest, runner versions, outcomes, and the exact rerun command. smoke is fail-closed (exit 2 + remedy, never CI), physical is a separately reported admission lane. CI runs the same command in a new journeys-hermetic job in the existing mobile-app-checks.yml when has_app_journeys fires (journey definitions and support, C3 replay world, dev controls, non-generated app/lib Dart, evidence contract, or this entrypoint) — synthetic fixtures only, fork-safe, receipts uploaded on pass and failure. Selection is resolved by the shared pre_push_ci_prediction.py and deliberately stays out of the bounded pre-push gate. Docs reconciled around the real command: app README, app AGENTS (within the lean budget), and the e2e SKILL now point here instead of diverging on setup/auth. * chore(app): stop tracking Flutter's iOS ephemeral tree app/ios/Flutter/ephemeral/** is regenerated by flutter on every pub get and self-describes as 'Generated file. Do not edit.' It was committed by accident in a formatting sweep (dec329a84a) and has been stale ever since: the tracked SwiftPM Package.swift lists pods (in_app_review, pasteboard) that no pub dependency provides, so any flutter run rewrites it, dirties every worktree, and fails the diff-hygiene push gate on regenerated trailing whitespace. Untrack the four files and ignore the tree, mirroring the existing **/macos/Flutter/ephemeral/ rules. Xcode resolves the local package after flutter regenerates it during setup; nothing consumes a committed copy. * feat: native lifecycle seams, vector replay, and leased device qualification (SCA-491/C5) - PhoneMicController (iOS + Android) now consumes narrow, injectable environment/ports seams: event sink, engine, permission, session config, interruption source, batch pipeline, main loop. Production behavior is unchanged; all live wiring lives in PhoneMicHostApiImpl.swift (iOS) and PhoneMicControllerPorts.production (Android). - Canonical phone-mic-native-events/v1 vector fixtures (8 schedules incl. session adoption) shared by Dart guard, iOS ruby harness, Android JVM harness; Pigeon contract types extracted at iOS test time (drift-guarded). - iOS: ios/test/phone_mic_lifecycle_replay_test.rb replays all vectors through the production controller+emitter with fakes for OS I/O only. - Android: PhoneMicLifecycleReplayTest (JVM, virtual main loop + manual audio queue) replays the same vectors through the production controller. - device_lease.py: exclusive physical-device leases with qualification registry (personal-device refusal), bounded acquisition, live-lease never-stolen, stale-owner recovery with generation bump, safe release. - device_runner.py + 'mobile-session device' CLI: readiness doctor with exact operator steps, and a runner consuming C1 session manifests (install/adb-reverse/untethered launch/permission cycle/device-run evidence v1). All hermetically tested with fake devices (25 tests). - PHYSICAL_DEVICES.md: m1-mac-studio read-only inventory, operator runbook, and the external physical-test handoff template. Physical acceptance stays pending user-run evidence by design. * docs: point mobile-verify physical at the C5 device handoff C4's physical lane stays fail-closed (exit 2). After C5 landed, the admission document should name the real runner and PHYSICAL_DEVICES.md instead of implying the software path is still missing. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: expect four Flutter pins in mobile-app-checks C4 added journeys-hermetic as a fourth Flutter job on the same repository toolchain pin. The workflow-contract count of 3 was stale. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: raise Desktop Swift PR-lane suite budget to 3000s Run 35134593036 measured 2778s against 2700s on a cache-hit PR lane. The overrun was one 1500s batch ceiling plus isolation, not a slow desktop suite; this mobile PR has no desktop sources. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> | 7 天前 | |
feat(dev-harness): fake-backed V1 live-session broker (reload, restart, controls, evidence, owned teardown) (#14362) * harden(dev-harness): define V1 live-session and V8 spine contracts Add strict Python/Dart pending contracts and introducing-commit protection. Reserve the live-session CLI, wire, ownership/attachment seams and failing builder acceptance tests. Add live attribution before freezing evidence v1. Verification: - bash scripts/dev-harness/run-tests.sh: 287 passed, 7 existing skips, 40 strict xfailed (24 pending V1 markers), Python 3.11.15. - make preflight: 16 checks passed, using supported PYTHON selection. - make mobile-verify fast --paths harness/session contract inputs with an absolute evidence directory: 20 executed, 20 passed, no failures/skips. - bash app/test.sh: 2014 passed, 5 skipped, one unchanged UTC+7 search_rank_grouping_test failure; affected baseline inputs match main. - Real module CLI refusals exit 2 without traceback; regression covered. - Dart SDK probe proves ext. prefix required; marker-removal formatter probe proves no incidental Dart diff. Runtime is intentionally unimplemented. App UI must repair current invalid VM extension names before live acceptance; unsupported journeys block. Simulator/emulator latency and detached-process acceptance remain V1 work. make setup installed hooks but full backend wheel sync stalled; isolated Python 3.11 harness dependencies ran the owner checks without any skip flag. Failure-Class: none * harden(dev-harness): scope the spine-contracts check to spine paths The check protects files under the spine roots and their registry, so it only needs to run when those change. A repo-wide trigger would run it, and its full-history requirement, on every unrelated contributor PR. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * docs(dev-harness): clarify the live journey adapter boundary V1 must refuse live journey execution until a shared executable adapter exists; B1 addressability alone cannot attach Dart integration tests. Pure admission contracts remain valid with injected supported IDs. Verification: harness 287 passed, 7 skipped, 40 strict xfailed; make preflight passed 20 checks. * harden(dev-harness): pin reviewed spine revisions and align B0 capabilities Coordinator turn-4 review authorizes the V1 fixture correction. Exact before/after digests pin the revised oracle; immutable revision records do not permit builder assertion edits. Active tests reject missing/tampered records and prove the fake machine response. Verification: full harness 289 passed, 7 skipped, 40 strict xfailed; UTC app 2020 passed; spine 24 pending/7 protected; make preflight 16 passed. No live device acceptance claimed. * harden(dev-harness): preserve marker retirement across spine revisions Track retained marker occurrences rather than counts so retiring a second test cannot restore the first test's retired marker. The active regression covers this swap and legitimate continued retirement. Verification: full harness 290 passed, 7 skipped, 40 strict xfailed; spine 24 pending/7 protected; make preflight 16 passed. Dart sources unchanged after the earlier UTC app 2020-pass run. * harden(dev-harness): retain corrected spine oracles through squash merges A squash can introduce a corrected contract and its review records in the same commit. Absorb only the revision prefix ending at that introducing file's exact digest, preserving immutable records and later marker-only edits. Active temp-repo regression covers multiple corrections in one squash and rejects restoring old or altered assertions. Verification: harness 291 passed, 7 skipped, 40 strict xfailed; B1 make preflight 20 passed; final B1 journeys 20/20. Dart unchanged after UTC app 41 opt-in and 2051 ordinary passed. * test(dev-harness): revise live session acceptance after builder review * test(dev-harness): bind machine fixture to actual child identity * feat(dev-harness): implement the fake-backed V1 live session broker LiveSession speaks the Flutter machine wire against an injected child, records BrokerIdentity in live.json, and refuses real flutter run from CLI dispatch. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dev-harness): tear down live before stop, reset, recover, and release Lifecycle order is live teardown, then services, then device detach; recover still has generation 1 while teardown runs. Co-authored-by: Cursor <cursoragent@cursor.com> * test(dev-harness): retire V1 live-session pending markers The fake-backed broker now satisfies those spine tests; only whole pending-marker lines were deleted. Co-authored-by: Cursor <cursoragent@cursor.com> * test(dev-harness): cover lazy live-session holes with real child processes A dead flutter child mid-reload is blocked, never success; teardown will not signal a live PID whose start time, marker, or boot id does not match. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): hash iOS .app bundles as directory trees file_sha256 raised IsADirectoryError on a simulator .app; identity is now the sorted tree of regular files. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): pin live-session source, generation, and stdio The architect's thirteen probes rejected attributing a compile to source observed after launch, announcing stopped after a failed reap, and grepping one fixture secret. Keep one monotonic RPC deadline and actually validate negotiated capabilities. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): fence live publication, boolean readiness, and auth redaction _assert_generation only covered admission, so a lease roll during ext.omi.controls.state still published ok and advanced loaded identity. Recheck lease and child generation after daemon/readiness work, refuse non-boolean readiness, and redact complete Authorization values before logs are retained. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): restore live teardown after main merge and retire passed T7 markers Keep main's port-claim mobile_session.py; put back V1 live teardown on stop/reset/recover. The formatted round-7 oracles now pass, so remove only the pending-marker lines. Co-authored-by: Cursor <cursoragent@cursor.com> * chore(dev-harness): restore app sources that the merge hook reformatted The merge commit ran dart format without package:flutter_lints resolved and rewrote main's files. Restore origin/main bytes so the V1 PR does not carry unrelated UI diffs. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): fail fast when a live owner holds the session lease Name the session and holder pid immediately. The same overlap that used to look like a controls-extension miss is a held lease. Co-authored-by: Cursor <cursoragent@cursor.com> * test(dev-harness): retire V1 pending markers that main's live fences now pass Merging origin/main imported strict xfails for lease-roll, negotiation, and auth-redaction fences this branch already implements. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): fail closed on non-boolean live readiness flags Finding 2 of the adversarial review reproduced: a fake reporting readiness.signedIn as the string "false" was classified not-ready and fell through to the wait_ready recovery path, which then observed real booleans and published ok. _is_ready now raises LiveError (malformed-response) for present-but-non-boolean flags, so start blocks instead of self-healing; start/reload/restart wrap it and surface blocked/malformed-response. The stringy-false regression test's expected error_code moves unready -> malformed-response; blocked-stays-blocked is unchanged. The satisfied V1 pending marker on the boolean probe is retired. Findings 1, 3 and 4 did not reproduce at this head (probes re-run verbatim; evidence in PR comment). Failure-Class: new --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Cursor <cursoragent@cursor.com> | 5 天前 | |
fix(app): stop hermetic tests depending on host timezone; harden mobile-verify evidence paths and device doctor (#14313) * fix(app): stop hermetic calendar tests depending on host timezone UTC same-day .single and DateTime.now() as "today" are only true in the author's zone. Build same-day fixtures with local DateTime via localCalendarDay, default selectedDate through package:clock, and document the TZ=Pacific/{Kiritimati,Pago_Pago} entrypoint. Failure-Class: new Co-authored-by: Cursor <cursoragent@cursor.com> * harden(dev-harness): classify missing device binaries and evidence paths #14305 caught empty OMI_VERIFY_EVIDENCE_DIR and the CLI default runner. Relative --evidence-dir still wrote receipts under app/ (runner cwd) while aggregation looked at the invocation cwd; AndroidTooling._run still let FileNotFoundError escape past device_doctor. Resolve evidence paths against the invocation directory, convert missing adb/xcrun to DeviceRunnerError, and accept device doctor --json like the siblings. Co-authored-by: Cursor <cursoragent@cursor.com> * docs(app): keep the TZ second-pass contract in MOBILE_VERIFY.md app/AGENTS.md is already at its lean budget. The documented TZ=Pacific/{Kiritimati,Pago_Pago} entrypoint lives in the harness guide, which CI can call without growing agent context. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> | 6 天前 | |
feat(mobile): add iPhone capture probes and offline replay (#15956) * feat(mobile): add offline iPhone capture replay and diagnostic harnesses * fix(dev-harness): mark Apple certificate fingerprint as non-security * fix(dev-harness): use platform certificate fingerprint tooling * fix(mobile): preserve startup ordering tripwire | 23 小时前 | |
feat(mobile): add iPhone capture probes and offline replay (#15956) * feat(mobile): add offline iPhone capture replay and diagnostic harnesses * fix(dev-harness): mark Apple certificate fingerprint as non-security * fix(dev-harness): use platform certificate fingerprint tooling * fix(mobile): preserve startup ordering tripwire | 23 小时前 | |
SCA-486: isolated mobile sessions, local-dev auth, capture replay, and seeded journeys (#14213) * mobile sessions: freeze session-evidence-v1 receipt contract Freeze the minimal versioned session/evidence schema shared by the mobile development-foundation consumers (C2 journeys, C3 capture replay, C4 verification, C5 devices) and its executable validator. - contracts/session/session-evidence-v1.schema.json: closed v1 object binding source SHA + dirty digest, built artifact identity, loopback-only endpoints, fixture/runner versions, real timestamps, status/blocked reason and exact execution counts. No credential-shaped field exists anywhere. - dev_harness/session_evidence.py: builder + validator enforcing the cross-field semantics a JSON Schema cannot express: ready/running require an artifact whose git_sha matches the source (a stale build cannot be reported ready), production-family profiles are rejected as session targets, counts must account exactly, zero-execution receipts are only valid pre-run, and credential-shaped keys are refused at any depth. - Egress guards: validate_local_http_url/validate_local_host_port reject non-loopback or non-plain-HTTP endpoints (api.omi.me / api.omiapi.com by name) before any request is attempted. Evidence: scripts/dev-harness/run-tests.sh (test_session_evidence.py, 20 contract tests incl. stale-artifact, egress, credential and accounting rejections). * mobile sessions: deterministic synthetic auth fixture v1 Seed a synthetic Auth-emulator user through the local-dev custom-token endpoint contract from open PR #11784 (feat/app+backend local-development sign-in without OAuth) — reuse, not a competing endpoint. The PR stays with its owner; scripts/dev-harness/MOBILE_SESSIONS.md records the integration plan and provenance. - fixtures/mobile/v1.json: one deterministic user (omi-fixture-v1-user-1@local.test); RFC-reserved domain so fixture identities can never collide with a real account. No real Google/Apple user, provider key, or copied token involved. - dev_harness/mobile_fixtures.py: fail-closed seeding client — the backend URL is validated as loopback plain-HTTP before any request, production hosts are denied by name, and the persisted receipt records identity and outcome only: token_minted/token_retained, never the token itself. Evidence: test_mobile_fixtures.py (16 tests) — determinism, reserved-domain enforcement, pre-request egress refusal, 404/unreachable/wrong-uid fail-closed paths, credential-free receipts. * mobile sessions: structured doctor for the session lanes Every readiness failure classifies exactly one of ready / agent-remediable (with the exact resumption command) / operator-action-needed (privileged install, license, host capacity), per lane (backend, android, ios). - Flutter version is read from the mobile CI pin in .github/workflows/mobile-app-checks.yml — never 'latest'; inconsistent pins refuse rather than guess. - Backend lane: python3.11 (venv must be 3.11, ambient 3.14 must not select the runtime), JDK 21 for the firebase emulators, firebase-tools, redis/typesense via native binary or a responding docker daemon. - Android lane: ANDROID_HOME + adb + emulator engine + system image, each with the exact sdkmanager remedy and a capacity-gated download note. - iOS lane: Xcode + simctl runtime; missing runtime is an operator action. - Capacity: <12GiB free on the shared Data/scratch container is an operator gate for emulator/build lanes (agents never free space themselves); contract/unit lanes skip it via --skip-capacity. - Egress: an ambient production OMI_LOCAL_API_BASE_URL override is reported as a blocking misconfiguration. Evidence: test_mobile_doctor.py (13 tests) over an injected runner — lane filtering, ready/degraded/blocked classification, pin parsing, capacity and operator-gate behavior; live run on m1-mac-studio via 'make mobile-session ARGS="doctor --platform android --platform ios"' reports backend+ios ready, android emulator engine agent-remediable. * mobile sessions: isolated session lifecycle CLI behind one entrypoint 'make mobile-session ARGS="…"' (scripts/dev-harness/mobile-session.sh) owns a uniquely-leased local mobile session: doctor / acquire / start / seed / reset / status / evidence / stop / recover / release. A session is an existing dev-harness instance + port offset + device lease + seed receipt + evidence receipt — the harness lifecycle is reused in-process under OMI_LOCAL_INSTANCE/OMI_HARNESS_PORT_OFFSET, not duplicated. Ownership is fail-closed: - leases are created atomically (O_EXCL) with owner host/user/pid and a harness-standard sentinel; a live foreign owner or another local user's session is never touched; cross-host takeover is an operator decision; recover bumps the generation for same-host/same-user takeovers. - ports come from a claimed offset registry; a foreign process occupying a port is refused (never killed) and the allocator skips that offset; release frees the claim only when it belongs to the session. - start gates device attach on doctor readiness (precise blocked reason, not a crash); ios-simulator devices are created/booted/deleted session-owned via simctl. - seed/reset/stop/release are idempotent; reset only touches the session's own harness instance (sentinel-validated underneath). - evidence emits session-evidence-v1 receipts; ready/running refuse without a bound artifact and refuse when the source moved since acquire. app/setup.sh (separate commit): OMI_IOS_DEVICE_ID pins non-interactive device selection; OMI_DEVICE_SUFFIX overrides hostname identity. Evidence: test_mobile_session.py (19 tests) + wrapper tests — exclusivity, disjoint ports, dead-owner/live-foreign/different-user/cross-host refusals, foreign-port refusal with a real live listener, idempotent release, artifact binding, stale-source refusal, harness env handoff. Live CLI run: acquire/list/evidence/seed-fail-closed/stop/release with exit codes 0/2 on m1-mac-studio. * app/setup.sh: non-interactive device pin and per-session device suffix - OMI_IOS_DEVICE_ID: when set, select_ios_device uses exactly that device id, failing precisely (with the available device list) when absent, instead of enumerating and prompting — the mobile-session harness, CI and nested agents cannot answer an interactive prompt, and the current no-TTY path errors out whenever more than one iOS destination exists. - OMI_DEVICE_SUFFIX: let a session harness (or a second checkout on one host) inject a unique device-identity suffix instead of the hostname, which collides across concurrent sessions on the same machine. Unset behavior is unchanged. Verified by sourcing the function with a stubbed flutter devices --machine: pinned-present emits the id; pinned-absent fails with the list; unpinned multi-device no-TTY keeps the existing enumeration failure. * mobile sessions: apply repo python formatter to the new modules black 26.5.1, --line-length 120 --skip-string-normalization via scripts/backend-python-format; behavior unchanged, dev-harness lane re-run green (214 passed; 1 pre-existing environmental failure — the host's global git worktree guard blocks pytest-tmp linked worktrees). * test: place linked-worktree pytest fixtures under OMI_WORKTREES The managed git wrapper correctly refuses worktrees in /private/tmp. Keep that guard and put the fixture where task worktrees are allowed. Co-authored-by: Cursor <cursoragent@cursor.com> * docs: treat session-evidence-v1 as proposed until consumers review it C1 shipped the schema; freeze it only after C2/C3/C4 agree, not from a single worker declaration. Co-authored-by: Cursor <cursoragent@cursor.com> * feat: reuse PR 11784 local-dev custom-token auth on current main Copy the reviewed emulator-gated sign-in path onto this integration branch so synthetic seed talks to real local services. Leave the original PR open and unmerged. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: boot isolated sessions on real CoreSimulator IDs and offline STT Use the installed iPhone 17 Pro / iOS 26.5 identifiers, pin PROVIDER_MODE=offline, and drop soniox from the offline STT chain so the local backend can start without a paid key. Co-authored-by: Cursor <cursoragent@cursor.com> * test: isolate provider-secret fixtures from ambient PROVIDER_MODE A previous offline session left PROVIDER_MODE in the shell and made the secret-injection tests read ambient offline instead of the fixture file. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(app): injectable capture seams for deterministic recovery replay CaptureController and the phone WAL resolved clock, timers, connectivity, auth, mic, socket and upload policy through global singletons, so the capture -> WAL -> recovery path could not be replayed deterministically. Add narrow constructor seams (capture_seams.dart) with production-identical defaults: CaptureScheduling, CaptureAuthBoundary, CaptureConnectivityBoundary, plus wal/phoneMic/clock/scheduler injection on CaptureController; clock/periodic/job-status injection on LocalWalSyncImpl threaded through WalSyncs/WalService; periodic-timer injection on the NativeMicRecorderService watchdogs. The in-progress-conversation loader seam now covers the socket-connect path too, and streamRecording honors the microphone permission requester like the batch path already did. No behavior change with default construction; every seam is optional. Evidence: bash app/test.sh (1990 passed, 5 pre-existing skips); analyze ratchet green. * test(app): deterministic capture-recovery replay schedules (SCA-489/C3) Replay the REAL production capture pipeline (CaptureController, NativeMicRecorderService, TranscriptSegmentSocketService, WalService, RecordingTransferCoordinator) against controlled external I/O: virtual clock, manual bounded scheduler, scripted transport/upload boundary, fake native host. Restart evidence destroys and reconstructs the object graph from real temp files (torn wals.json -> backup recovery, missing audio -> terminal corruption, process kill -> disk reload and re-upload). Six schedules with invariant oracles: network loss/reconnect mid-capture (exact frame identity in the stored WAL, single upload), stale native events after stop/new session (session-identity gate, no double teardown), interruption/resumption (live + batch, bounded stall escalation), partial/torn persistence plus reconstruction, failed upload with bounded backoff and persisted/enqueued/server-acknowledged distinctions, and ownership transition (signed-out reconnect cancellation, bounded 4001 token refresh). Also publishes the C2/C4 adapter (capture_scenario.dart: catalog + result contract) and the C5 native-event vector schema (phone-mic-native-events/v1) mirroring the Pigeon PhoneMicFlutterApi contract without touching Pigeon. Falsification evidence: removing the NativeMicRecorderService session gate flips the stale-idle schedule to failure (record->stop); removing the finalizeCurrentSession unsynced-retention guard drops the WAL and fails the network-loss schedule. Evidence: flutter test test/unit/capture_recovery_replay_scenarios_test.dart (17 passed); bash app/test.sh full suite green. * chore(app): allowlist SCA-489 replay contract libs in the dead-code ratchet The scenario catalog/result contract and the native-event vector schema are library-only by design until the C2/C4 and C5 lanes import them; the ratchet demands an explicit allowlist entry with a reason for exactly this case. * test: pin conversation-window capture session id across sequential phone-mic lives activeCaptureSessionId is WAL/conversation-scoped so a late ConversationEvent can still stamp WALs. C2 must use activeRecordingId as the live recording identity. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: wait for unawaited capture-upload retries before asserting and teardown CI failed the bounded-backoff replay because cooldown wakes are unawaited and settle used wall-clock sleeps that missed the drain under load, then deleted the temp WAL dir mid-write. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(app): typed debug semantic controls for the local journey lane Extends the existing debug Marionette surface (same debug VM-service transport, no new server/framework) with product-semantic controls: versioned capabilities (semantic-controls/v1), privacy-safe state (route, principal, capture lifecycle with activeRecordingId as the authoritative recording identity), bounded wait_ready, production-path navigation, and named journey faults. Fail-closed eligibility: kDebugMode AND local_dev profile AND OMI_DEV_CONTROLS=1 dart-define. Ineligible builds (including production-flavor debug) install nothing and the HTTP fault chokepoint is a pure pass-through — pinned by semantic_controls_guard_test.dart. Narrow seams added for the hermetic journey lane: - AuthService.installLocalHarnessTokenGateway (debug+local_dev gated Firebase token I/O boundary; isSignedIn routes through the gateway) - PlatformManager.initializeForLocalHarness (header fields only) - CrashlyticsManager report paths tolerate a missing Firebase app the same way main.dart's zone handler already does, so host-lane errors surface instead of being masked by [core/no-app] Verified: flutter test test/unit/semantic_controls_guard_test.dart (10 passed); auth regression suites (34 passed); C3 capture replay (17 passed); dead-code ratchet at baseline. * test(app): five strict seeded acceptance journeys with negative fault variants Canonical executable definitions (one per behavior) under app/integration_test/journeys/, runnable hermetically (flutter-tester + loopback fixture backend) or on a simulator via run_journeys.sh: j1 seeded conversation detail — real provider fetch + real detail page, exact synthetic identity; negative: wrong-owner session refused. j2 chat send -> distinct assistant reply — real input/send-button keys (omi.chat.input / omi.chat.send), request observed server-side, server-minted ai-role reply distinct from the prompt, rendered; negatives: suppress-send, suppress-assistant-reply, wrong-owner-session. j3 memory create/edit surviving reload — production provider path, server-minted id required after reload; negative: drop-memory-save. j4 expired session — transient failure re-mints via the real custom-token endpoint; terminal failure emits expiry and blocks requests; negative: production-family profiles never silently re-mint. j5 capture interruption/reconnect — C3 capture-scenario/v1 adapter: real temp files, process reconstruction, drain exactly once; negative: fail-capture-recovery. Each negative arms exactly one named fault and must fail with the invariant named. Evidence receipts follow session-evidence-v1 accounting; zero-execution runs never pass. Verified: bash integration_test/journeys/run_journeys.sh (5/5 pass); repeated deterministic vertical: bash integration_test/journeys/run_journeys.sh --filter j2 --runs 5. * fix: stamp journey evidence finished_at at write time Receipts were recording construction time as the end timestamp, so duration could not be distinguished from start. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(app): clear new analyzer-ratchet regressions in the integrated journeys The integrated checkpoint (b217eb1a9d) fails app/scripts/analyze_ratchet.sh with 7 new occurrences: 3 unused imports plus a bogus 'show WalStatus' in j5, an unused-looking nested import that actually provides SingleChildWidget in hermetic_boot, a missing const in j4, and two depend_on_referenced_packages for test-only platform interfaces. Declares path_provider_platform_interface and nested as direct dev dependencies (same pattern as the existing web_socket_channel dev deps) and removes the dead imports. Mechanical lint repairs only — j4/j5 hermetic journeys re-run green after the change. * feat: unified mobile verify lanes, mechanical journey selection, and CI/contributor path (SCA-490/C4) One canonical verification entrypoint over the proven lanes: make mobile-verify select|doctor|fast|smoke|physical (scripts/dev-harness/mobile-verify.sh -> dev_harness.mobile_verify). It never adds a second runner: journeys delegate to the C2 canonical runner and session infrastructure to the C1 session CLI. Selection is mechanical and fail-closed: journeys are glob-discovered (runner --list contract test), changed paths map through a per-seam rule table, unknown app/lib impact falls back to the full suite, and an empty selection is drift (exit 65), never PASS(0) — run_journeys.sh now fails closed the same way. Receipts are validated against session-evidence-v1 accounting (honest counts, zero-execution never passes) and every lane writes a verify-receipt.json binding source SHA + dirty digest, runner versions, outcomes, and the exact rerun command. smoke is fail-closed (exit 2 + remedy, never CI), physical is a separately reported admission lane. CI runs the same command in a new journeys-hermetic job in the existing mobile-app-checks.yml when has_app_journeys fires (journey definitions and support, C3 replay world, dev controls, non-generated app/lib Dart, evidence contract, or this entrypoint) — synthetic fixtures only, fork-safe, receipts uploaded on pass and failure. Selection is resolved by the shared pre_push_ci_prediction.py and deliberately stays out of the bounded pre-push gate. Docs reconciled around the real command: app README, app AGENTS (within the lean budget), and the e2e SKILL now point here instead of diverging on setup/auth. * chore(app): stop tracking Flutter's iOS ephemeral tree app/ios/Flutter/ephemeral/** is regenerated by flutter on every pub get and self-describes as 'Generated file. Do not edit.' It was committed by accident in a formatting sweep (dec329a84a) and has been stale ever since: the tracked SwiftPM Package.swift lists pods (in_app_review, pasteboard) that no pub dependency provides, so any flutter run rewrites it, dirties every worktree, and fails the diff-hygiene push gate on regenerated trailing whitespace. Untrack the four files and ignore the tree, mirroring the existing **/macos/Flutter/ephemeral/ rules. Xcode resolves the local package after flutter regenerates it during setup; nothing consumes a committed copy. * feat: native lifecycle seams, vector replay, and leased device qualification (SCA-491/C5) - PhoneMicController (iOS + Android) now consumes narrow, injectable environment/ports seams: event sink, engine, permission, session config, interruption source, batch pipeline, main loop. Production behavior is unchanged; all live wiring lives in PhoneMicHostApiImpl.swift (iOS) and PhoneMicControllerPorts.production (Android). - Canonical phone-mic-native-events/v1 vector fixtures (8 schedules incl. session adoption) shared by Dart guard, iOS ruby harness, Android JVM harness; Pigeon contract types extracted at iOS test time (drift-guarded). - iOS: ios/test/phone_mic_lifecycle_replay_test.rb replays all vectors through the production controller+emitter with fakes for OS I/O only. - Android: PhoneMicLifecycleReplayTest (JVM, virtual main loop + manual audio queue) replays the same vectors through the production controller. - device_lease.py: exclusive physical-device leases with qualification registry (personal-device refusal), bounded acquisition, live-lease never-stolen, stale-owner recovery with generation bump, safe release. - device_runner.py + 'mobile-session device' CLI: readiness doctor with exact operator steps, and a runner consuming C1 session manifests (install/adb-reverse/untethered launch/permission cycle/device-run evidence v1). All hermetically tested with fake devices (25 tests). - PHYSICAL_DEVICES.md: m1-mac-studio read-only inventory, operator runbook, and the external physical-test handoff template. Physical acceptance stays pending user-run evidence by design. * docs: point mobile-verify physical at the C5 device handoff C4's physical lane stays fail-closed (exit 2). After C5 landed, the admission document should name the real runner and PHYSICAL_DEVICES.md instead of implying the software path is still missing. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: expect four Flutter pins in mobile-app-checks C4 added journeys-hermetic as a fourth Flutter job on the same repository toolchain pin. The workflow-contract count of 3 was stale. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: raise Desktop Swift PR-lane suite budget to 3000s Run 35134593036 measured 2778s against 2700s on a cache-hit PR lane. The overrun was one 1500s batch ceiling plus isolation, not a slow desktop suite; this mobile PR has no desktop sources. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> | 7 天前 | |
fix(app): refuse remote APIs in hermetic tests; add a lane-bootstrap that skips the full backend lock (#14321) * fix(app): refuse remote APIs and non-dev pairings in hermetic tests app/test.sh skipped writing .dev.env when generated files already existed, so a leftover API_BASE_URL or mobile_beta/prod pairing could silently survive. Fail closed at the checker, test.sh, and doctor without rewriting the file. Failure-Class: new Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dev-harness): add lane-bootstrap that skips the full backend lock A linked worktree can run cheap pre-push gates and app/test.sh after one idempotent command: pinned Flutter, shared PUB_CACHE, Python 3.11 via uv, and yaml+dotenv only. An incomplete .venv directory is no longer treated as ready. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): keep fixture venvs passing the cheap-gate import probe Manifest-contract stubs only accepted `import yaml`. The resolver now probes `import dotenv, yaml` so an empty .venv is not treated as ready, and those stubs were skipped. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> | 6 天前 | |
fix(dev-harness): reject loopback-lookalike hostnames (#12086) * test(dev-harness): cover loopback hostname lookalikes * fix(dev-harness): reject loopback-lookalike hostnames Failure-Class: FC-implicit-resource-selection | 1 个月前 | |
feat(dev-harness): fake-backed V1 live-session broker (reload, restart, controls, evidence, owned teardown) (#14362) * harden(dev-harness): define V1 live-session and V8 spine contracts Add strict Python/Dart pending contracts and introducing-commit protection. Reserve the live-session CLI, wire, ownership/attachment seams and failing builder acceptance tests. Add live attribution before freezing evidence v1. Verification: - bash scripts/dev-harness/run-tests.sh: 287 passed, 7 existing skips, 40 strict xfailed (24 pending V1 markers), Python 3.11.15. - make preflight: 16 checks passed, using supported PYTHON selection. - make mobile-verify fast --paths harness/session contract inputs with an absolute evidence directory: 20 executed, 20 passed, no failures/skips. - bash app/test.sh: 2014 passed, 5 skipped, one unchanged UTC+7 search_rank_grouping_test failure; affected baseline inputs match main. - Real module CLI refusals exit 2 without traceback; regression covered. - Dart SDK probe proves ext. prefix required; marker-removal formatter probe proves no incidental Dart diff. Runtime is intentionally unimplemented. App UI must repair current invalid VM extension names before live acceptance; unsupported journeys block. Simulator/emulator latency and detached-process acceptance remain V1 work. make setup installed hooks but full backend wheel sync stalled; isolated Python 3.11 harness dependencies ran the owner checks without any skip flag. Failure-Class: none * harden(dev-harness): scope the spine-contracts check to spine paths The check protects files under the spine roots and their registry, so it only needs to run when those change. A repo-wide trigger would run it, and its full-history requirement, on every unrelated contributor PR. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * docs(dev-harness): clarify the live journey adapter boundary V1 must refuse live journey execution until a shared executable adapter exists; B1 addressability alone cannot attach Dart integration tests. Pure admission contracts remain valid with injected supported IDs. Verification: harness 287 passed, 7 skipped, 40 strict xfailed; make preflight passed 20 checks. * harden(dev-harness): pin reviewed spine revisions and align B0 capabilities Coordinator turn-4 review authorizes the V1 fixture correction. Exact before/after digests pin the revised oracle; immutable revision records do not permit builder assertion edits. Active tests reject missing/tampered records and prove the fake machine response. Verification: full harness 289 passed, 7 skipped, 40 strict xfailed; UTC app 2020 passed; spine 24 pending/7 protected; make preflight 16 passed. No live device acceptance claimed. * harden(dev-harness): preserve marker retirement across spine revisions Track retained marker occurrences rather than counts so retiring a second test cannot restore the first test's retired marker. The active regression covers this swap and legitimate continued retirement. Verification: full harness 290 passed, 7 skipped, 40 strict xfailed; spine 24 pending/7 protected; make preflight 16 passed. Dart sources unchanged after the earlier UTC app 2020-pass run. * harden(dev-harness): retain corrected spine oracles through squash merges A squash can introduce a corrected contract and its review records in the same commit. Absorb only the revision prefix ending at that introducing file's exact digest, preserving immutable records and later marker-only edits. Active temp-repo regression covers multiple corrections in one squash and rejects restoring old or altered assertions. Verification: harness 291 passed, 7 skipped, 40 strict xfailed; B1 make preflight 20 passed; final B1 journeys 20/20. Dart unchanged after UTC app 41 opt-in and 2051 ordinary passed. * test(dev-harness): revise live session acceptance after builder review * test(dev-harness): bind machine fixture to actual child identity * feat(dev-harness): implement the fake-backed V1 live session broker LiveSession speaks the Flutter machine wire against an injected child, records BrokerIdentity in live.json, and refuses real flutter run from CLI dispatch. Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dev-harness): tear down live before stop, reset, recover, and release Lifecycle order is live teardown, then services, then device detach; recover still has generation 1 while teardown runs. Co-authored-by: Cursor <cursoragent@cursor.com> * test(dev-harness): retire V1 live-session pending markers The fake-backed broker now satisfies those spine tests; only whole pending-marker lines were deleted. Co-authored-by: Cursor <cursoragent@cursor.com> * test(dev-harness): cover lazy live-session holes with real child processes A dead flutter child mid-reload is blocked, never success; teardown will not signal a live PID whose start time, marker, or boot id does not match. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): hash iOS .app bundles as directory trees file_sha256 raised IsADirectoryError on a simulator .app; identity is now the sorted tree of regular files. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): pin live-session source, generation, and stdio The architect's thirteen probes rejected attributing a compile to source observed after launch, announcing stopped after a failed reap, and grepping one fixture secret. Keep one monotonic RPC deadline and actually validate negotiated capabilities. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): fence live publication, boolean readiness, and auth redaction _assert_generation only covered admission, so a lease roll during ext.omi.controls.state still published ok and advanced loaded identity. Recheck lease and child generation after daemon/readiness work, refuse non-boolean readiness, and redact complete Authorization values before logs are retained. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): restore live teardown after main merge and retire passed T7 markers Keep main's port-claim mobile_session.py; put back V1 live teardown on stop/reset/recover. The formatted round-7 oracles now pass, so remove only the pending-marker lines. Co-authored-by: Cursor <cursoragent@cursor.com> * chore(dev-harness): restore app sources that the merge hook reformatted The merge commit ran dart format without package:flutter_lints resolved and rewrote main's files. Restore origin/main bytes so the V1 PR does not carry unrelated UI diffs. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): fail fast when a live owner holds the session lease Name the session and holder pid immediately. The same overlap that used to look like a controls-extension miss is a held lease. Co-authored-by: Cursor <cursoragent@cursor.com> * test(dev-harness): retire V1 pending markers that main's live fences now pass Merging origin/main imported strict xfails for lease-roll, negotiation, and auth-redaction fences this branch already implements. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): fail closed on non-boolean live readiness flags Finding 2 of the adversarial review reproduced: a fake reporting readiness.signedIn as the string "false" was classified not-ready and fell through to the wait_ready recovery path, which then observed real booleans and published ok. _is_ready now raises LiveError (malformed-response) for present-but-non-boolean flags, so start blocks instead of self-healing; start/reload/restart wrap it and surface blocked/malformed-response. The stringy-false regression test's expected error_code moves unready -> malformed-response; blocked-stays-blocked is unchanged. The satisfied V1 pending marker on the boolean probe is retired. Findings 1, 3 and 4 did not reproduce at this head (probes re-run verbatim; evidence in PR comment). Failure-Class: new --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Cursor <cursoragent@cursor.com> | 5 天前 | |
feat(dev-harness): opt-in make lane-backend (fast wheel install) for the mobile harness and pre-push gates (#14349) * fix(app): refuse remote APIs and non-dev pairings in hermetic tests app/test.sh skipped writing .dev.env when generated files already existed, so a leftover API_BASE_URL or mobile_beta/prod pairing could silently survive. Fail closed at the checker, test.sh, and doctor without rewriting the file. Failure-Class: new Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dev-harness): add lane-bootstrap that skips the full backend lock A linked worktree can run cheap pre-push gates and app/test.sh after one idempotent command: pinned Flutter, shared PUB_CACHE, Python 3.11 via uv, and yaml+dotenv only. An incomplete .venv directory is no longer treated as ready. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): keep fixture venvs passing the cheap-gate import probe Manifest-contract stubs only accepted `import yaml`. The resolver now probes `import dotenv, yaml` so an empty .venv is not treated as ready, and those stubs were skipped. Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> * feat(dev-harness): make setup-backend the fast wheel install, not pylock sync uv pip sync of pylock.macos.toml lists hashed sdist+wheel for av/llvmlite/scipy/pyarrow and stalled ~20 minutes; uv pip install -r requirements.txt takes index wheels (69s here) and covers uvicorn/pyright/yaml/dotenv/google.auth for session start and the pre-push typecheck. Lock-faithful sync stays in sync-python-deps.sh. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dev-harness): keep setup-backend as the lock sync; add opt-in lane-backend make setup-backend is the contributor lock-synced entrypoint; pointing it at unlocked requirements.txt broke scripts/test-make-setup.sh. The fast wheel install is make lane-backend. generate-app-env runs build_runner without the flags that deleted tracked manifest.g.dart. Co-authored-by: Cursor <cursoragent@cursor.com> * docs(backend): keep AGENTS.md under the lean budget while naming lane-backend Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> | 6 天前 | |
harden(backend,app): client-compatibility replay contract (C10) (#14372) * harden(dev-harness): define V1 live-session and V8 spine contracts Add strict Python/Dart pending contracts and introducing-commit protection. Reserve the live-session CLI, wire, ownership/attachment seams and failing builder acceptance tests. Add live attribution before freezing evidence v1. Verification: - bash scripts/dev-harness/run-tests.sh: 287 passed, 7 existing skips, 40 strict xfailed (24 pending V1 markers), Python 3.11.15. - make preflight: 16 checks passed, using supported PYTHON selection. - make mobile-verify fast --paths harness/session contract inputs with an absolute evidence directory: 20 executed, 20 passed, no failures/skips. - bash app/test.sh: 2014 passed, 5 skipped, one unchanged UTC+7 search_rank_grouping_test failure; affected baseline inputs match main. - Real module CLI refusals exit 2 without traceback; regression covered. - Dart SDK probe proves ext. prefix required; marker-removal formatter probe proves no incidental Dart diff. Runtime is intentionally unimplemented. App UI must repair current invalid VM extension names before live acceptance; unsupported journeys block. Simulator/emulator latency and detached-process acceptance remain V1 work. make setup installed hooks but full backend wheel sync stalled; isolated Python 3.11 harness dependencies ran the owner checks without any skip flag. Failure-Class: none * harden(dev-harness): scope the spine-contracts check to spine paths The check protects files under the spine roots and their registry, so it only needs to run when those change. A repo-wide trigger would run it, and its full-history requirement, on every unrelated contributor PR. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * docs(dev-harness): clarify the live journey adapter boundary V1 must refuse live journey execution until a shared executable adapter exists; B1 addressability alone cannot attach Dart integration tests. Pure admission contracts remain valid with injected supported IDs. Verification: harness 287 passed, 7 skipped, 40 strict xfailed; make preflight passed 20 checks. * harden(dev-harness): pin reviewed spine revisions and align B0 capabilities Coordinator turn-4 review authorizes the V1 fixture correction. Exact before/after digests pin the revised oracle; immutable revision records do not permit builder assertion edits. Active tests reject missing/tampered records and prove the fake machine response. Verification: full harness 289 passed, 7 skipped, 40 strict xfailed; UTC app 2020 passed; spine 24 pending/7 protected; make preflight 16 passed. No live device acceptance claimed. * harden(dev-harness): preserve marker retirement across spine revisions Track retained marker occurrences rather than counts so retiring a second test cannot restore the first test's retired marker. The active regression covers this swap and legitimate continued retirement. Verification: full harness 290 passed, 7 skipped, 40 strict xfailed; spine 24 pending/7 protected; make preflight 16 passed. Dart sources unchanged after the earlier UTC app 2020-pass run. * harden(dev-harness): retain corrected spine oracles through squash merges A squash can introduce a corrected contract and its review records in the same commit. Absorb only the revision prefix ending at that introducing file's exact digest, preserving immutable records and later marker-only edits. Active temp-repo regression covers multiple corrections in one squash and rejects restoring old or altered assertions. Verification: harness 291 passed, 7 skipped, 40 strict xfailed; B1 make preflight 20 passed; final B1 journeys 20/20. Dart unchanged after UTC app 41 opt-in and 2051 ordinary passed. * test(dev-harness): revise live session acceptance after builder review * test(dev-harness): bind machine fixture to actual child identity * test(contracts): define released mobile client compatibility replay C10 pins adopted consumer shapes and release-source provenance, reserves frozen-decoder replay, and leaves capture plus real backend acceptance visibly pending. Build tags are not release evidence. Verified app 2064, harness 309 passed/7 skipped/61 xfailed, spine 41 markers/11 protected, preflight 31 checks, journeys 3/3. Backend E2E stopped at missing fake_firestore prerequisite; no setup or live contact. * harden(dev-harness): scope compatibility replay to supported captures Keep the two-build bootstrap oracle historical while ongoing replay checks exact current router responses for supported cases. Archived captures remain immutable; validated server retirement is explicitly out of scope. Permit append-only coverage bundles for an existing build. Verified: UTC app suite 2064 passed (Dart unchanged); full harness 310 passed, 7 skipped, 61 strict xfailed; spine 41 pending, 11 protected; make preflight 31 passed; fast journeys 3/3. Three historical backend commits pass the adoption-scoped check. Backend E2E remains blocked at its dependency prerequisite; no real-router execution claimed. * harden(dev-harness): separate oracle revisions from implementations Use the actual PR base and immutable exact scaffolding snapshots to reject mixed revision/implementation diffs. Exercise a real squash and child merge. Correct local_dev wire spelling from the simulator RPC and add adversarial V1 conformance oracles after reviewing the first broker. Verification: UTC app 2064 passed; full harness 303 passed, 7 skipped, 55 strict xfailed; spine 37 pending/9 protected; preflight 16 checks passed. Isolated reviewed-broker probes failed 9/9 as documented in the conformance report. No real app/device launch. * fix(ci): repair the checks-manifest merge of main into the client-compat contract branch The coordinator's conflict resolution left a truncated duplicate spine-contracts entry; the pre-push manifest validation caught it. Keep one entry, with this branch's extra trigger path. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * harden(dev-harness): retain pinned corrections after older parent merges * harden(dev-harness): pin oracle content across shared runners and merge orders Keep exact oracle digests and immutable revision records. Traverse full reachable history and match revision payloads by digest. Shared runner metadata requires direct fail-fast suite invocation instead of pinning shared file bytes. Both new regressions fail on the old checker. UTC app 2064 passed; final harness 306 passed, 7 skipped, 59 strict xfailed; spine 39 pending/10 protected; preflight 16 passed. No protected oracle bytes or legacy scope pins changed. * harden(dev-harness): separate replay engine from evidenced release admission Bind distribution receipts to captured identities, pin handwritten decoder sources and exact candidate requests, require original-Flutter extraction equivalence, and execute a real Dart decoder oracle. Preserve unadopted endpoint freedom and report missing OpenAPI reference closure. Verification: app2064; harness326 passed/7 skipped/78 strict xfailed; spine52/15; preflight31 passed. No real release receipts or backend replay claimed. * harden(dev-harness): inventory both literal pending marker quote styles Single-quoted Python contracts executed under pytest but were invisible to inventory and could not retire their markers. Recognize paired quote styles in both languages; reject nonliteral builder markers while allowing active MECHANISM-owned API self-tests. Preserve all oracle bytes and digest pins. Verification: harness328 passed/7 skipped/79 strict xfailed; spine55 pending/16 protected; preflight31 passed; unchanged Dart suite2064 passed. * test(dev-harness): require replay loader to preserve frozen request vectors Compare every admitted Case to its pinned method, ordered query/headers, body, status, decoder and observations. A HEAD-shaped request with the right endpoint name cannot pass. Real capture admission remains pending after synthetic engine delivery. Verification: harness328 passed/7 skipped/79 strict xfailed; spine55/16; preflight31 passed; Dart2064 passed. --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> | 5 天前 | |
fix(dev-harness): reject non-executable Typesense overrides (#12087) * test(dev-harness): cover non-executable Typesense overrides * fix(dev-harness): reject non-executable Typesense overrides Failure-Class: none | 1 个月前 | |
feat(mobile): add iPhone capture probes and offline replay (#15956) * feat(mobile): add offline iPhone capture replay and diagnostic harnesses * fix(dev-harness): mark Apple certificate fingerprint as non-security * fix(dev-harness): use platform certificate fingerprint tooling * fix(mobile): preserve startup ordering tripwire | 23 小时前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 5 天前 | ||
| 6 天前 | ||
| 5 天前 | ||
| 20 小时前 | ||
| 7 天前 | ||
| 6 天前 | ||
| 7 天前 | ||
| 23 小时前 | ||
| 5 天前 | ||
| 1 天前 | ||
| 5 天前 | ||
| 6 天前 | ||
| 27 天前 | ||
| 5 天前 | ||
| 6 天前 | ||
| 1 个月前 | ||
| 6 天前 | ||
| 7 天前 | ||
| 5 天前 | ||
| 6 天前 | ||
| 23 小时前 | ||
| 23 小时前 | ||
| 7 天前 | ||
| 6 天前 | ||
| 1 个月前 | ||
| 5 天前 | ||
| 6 天前 | ||
| 5 天前 | ||
| 1 个月前 | ||
| 23 小时前 |