| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
chore(desktop): allowlist the MCP transport clients for the dead-code ratchet The four mcp-* clients are reachable only from pi-mono-extension, which imports them through the built dist path the agent-package scan does not follow. | 3 天前 | |
chore: remove root/CI/docs strays, orphaned scripts, stale citations Root: build6.done, .cross-review-verify.json, design-qa.md deleted; SPEC-meeting-receipt.md archived under docs/ (it read like standing spec for a shipped fix); Makefile run-canonical-promotion target (invokes a script deleted from main); expo-file-system + ioredis from root package.json with a full 373-entry surgical lock prune (every surviving resolution path verified identical); stale sdks/swift/Package.resolved (17 months old; SPM uses the root lockfile); sdks/omi-expo husk (three .gitignores, dead 16+ months). CI/docs: dispatch-only entellegence workflows and the disabled deploy_docs cutover residue deleted (its build copied six nonexistent doc dirs; docs npm package had no other consumer); mobile-app-checks cache key no longer hashes a nonexistent app/build.yaml; omi/firmware/bootloader/deprecated removed (live bootloader0.9.0.uf2 untouched); stale path citations fixed in integrations.mdx (device connectors), firmware_release.yml, desktop_windows_release.yml, codemagic.yaml. test(ci): dead-code ratchet checks (phase R) - import-graph reachability (flutter/agent/windows) and zero-importer buckets (backend) with committed baselines seeded from this cleanup (floors: flutter 0, backend 5 pending manual review, agent 0, windows 0) and reason-bearing allowlists; wired into checks-manifest local+ci lanes with a 10-test hermetic self-suite. Validated against the hand-verified inventory (flutter 56/56, zero false positives) before landing. Verification: manifest validator 160 checks / 0 errors; test_run_checks 47/47; ratchet --check exit 0; rg zero-refs for all deletions; edited JSON/YAML parse; dangling-import scan clean across dart/python/ts. | 3 天前 | |
chore(analytics): refresh reachability baseline after rebase onto main Co-authored-by: Cursor <cursoragent@cursor.com> | 4 天前 | |
fix(release): fence native Codemagic fallback Failure-Class: none After the bounded native-tag absence proof, make one tag-bound Codemagic fallback request only behind an all-state same-tag fence. Verify its returned build identity and retain the combined evidence without touching Beta or Stable paths. | 1 个月前 | |
Add isolated QA Typesense projection proof (#12826) * feat: add isolated JIT QA cloud plane Adds a main-only development workflow with immutable API, desktop companion, drain, and sweep images; Cloud Run readback rejects inherited credentials, cache or queue bindings and keeps migration gates closed. Pins Firebase token verification to the auth project while jobs use the dev Firestore ADC data plane.\n\nVerification: actionlint .github/workflows/jit_qa_cloud_run.yml; 30 focused backend tests; python3 backend/scripts/runtime_image_contracts.py check; make preflight; no GCP mutation performed. * feat: complete isolated JIT QA operator plane * fix: isolate JIT QA database and prove serving state * guard: fence isolated QA from shared Firestore * fix: match activation QA receipt contract * fix: fence isolated QA gateway identity * test: close QA paid execution path * ci: register isolated JIT QA deployment lock * fix: satisfy QA database and inventory type contracts * Bound JIT proactivity attempts and receipts * Document proactive timing improvements * Add manual bounded JIT qualification procedure * Avoid force unwrap after JIT contract validation * fix: offload JIT budget accounting from async gateway paths * fix: harden isolated JIT QA plane admission * Type shared JIT accounting and request contracts explicitly * Update app API snapshot for JIT budget capability * Regenerate client types for JIT budget capability * fix: bind PostHog rollout authority in JIT QA * Include isolated QA in retained sweep image contract * Enable bounded provider contract only on isolated QA services * feat(jit): bound proactive tool projection and accounting join * Preserve isolated QA budget capability across main integration * fix(jit): preserve proactivity request ids on errors Failure-Class: none * test(jit): preserve temporal context in shared full-turn prompt * fix: isolate JIT QA index reconciliation * fix: load QA secret for drain verification * fix(jit): keep QA index receipts parseable * fix: accept omitted empty QA environment value * Add isolated QA Typesense projection proof * fix(jit): enforce QA qualification and retain admission clock * Harden isolated QA Typesense projection * Fix JIT gateway accounting and stream safety * Add QA projection bootstrap and image smoke * Harden JIT streaming frame boundaries * fix(jit): preserve unknown runtime cost on interrupted streams * fix(jit): preserve budget after terminal stream cancellation * Fail closed around QA projection readiness proof * refactor(jit): declare Firestore dependency at module scope * fix(jit): match real QA Cloud Run projection contracts * docs(jit): clarify projection readiness timing * ci: lock isolated Typesense projection deploys * docs(jit): document projection crash window * fix(jit): close Typesense QA deployment contracts | 1 天前 | |
feat: add JIT knowledge ledger foundation and guarded adoption (#12084) * feat: add JIT knowledge ledger foundation * chore: refresh integration OpenAPI contract * fix: make trigger evaluation release-safe Failure-Class: none * fix: preserve lifecycle semantics in ledger apply Failure-Class: none * feat: adopt guarded JIT knowledge surfaces Route the agent preference writer through the intent-backed ledger, register a privacy-filtered entity timeline tool, render optional evidence on Windows, and add a base-ref-protected Gate F legacy-surface ratchet. Failure-Class: none * feat: add progressive JIT knowledge reads Register owner-scoped current-ledger search and explicit playbook hydration, with pre-limit semantic filtering and bounded outputs. Add a content-free planner/resume migration fixture without claiming canonical transaction completion.\n\nValidation: 141 focused backend tests passed; backend typecheck reported 0 errors; repository preflight passed 120 checks. * feat: render chat evidence on web Render bounded, fail-soft conversation evidence after authoritative answers in both web chat entry points. Unsupported, future, duplicate, and raw failure details remain inert.\n\nValidation: 337 web tests passed; web typecheck, oxlint, and Prettier passed; repository preflight passed 120 checks. * feat: require intent-backed ledger search results Apply the intent-backed requirement at the final merged canonical/history filter, with a passive historical-row regression case.\n\nValidation: 54 focused backend tests passed. * test: keep agent tool isolation stubs current * feat: gate JIT conversation retrieval * fix: make entity timeline scans deterministic Failure-Class: none * fix: honor rejected ledger projections Failure-Class: none * feat: render inert screen evidence on web * fix: reuse canonical review projection Failure-Class: none * fix(web): await recap context effect Failure-Class: none * test: amortize preference tool isolation load Failure-Class: none * feat(app): add knowledge ledger review surface Failure-Class: none * feat(macos): use canonical ledger prompt projection Failure-Class: none * test(memory): classify legacy surface inventory roles Failure-Class: none * fix(app): preserve ledger history completeness state Failure-Class: none * feat(macos): preserve canonical ledger mirror metadata Failure-Class: none * fix(app): match canonical ledger ordering Failure-Class: none * test(api): prove ledger client schema parity Failure-Class: none * feat(memory): expose bounded ledger history Failure-Class: none * chore(api): generate ledger history clients Failure-Class: none * feat(retrieval): add bounded card participants Failure-Class: none * feat(macos): project ledger trigger watchlist Failure-Class: none * fix(clients): fail closed on ledger authority Failure-Class: none * feat(macos): expose bounded trigger snapshot Failure-Class: none * fix(memory): keep closed history read only Failure-Class: none * feat(app): disclose partial ledger history Failure-Class: none * test(macos): cover ledger trigger bridge Failure-Class: none * fix(app): use neutral ledger accents Failure-Class: none * chore(api): declare ledger history route policy Failure-Class: none * test(macos): remove unsafe JSON fixture unwraps Failure-Class: none * fix(memory): satisfy typed history boundary Failure-Class: none * fix(macos): require prompt snapshot authority Failure-Class: none * test(memory): prove ledger migration on emulator Failure-Class: none * feat(retrieval): emit bounded screen evidence Failure-Class: none * test(memory): classify maintenance retirement readiness Failure-Class: none * feat(macos): adapt Rewind metadata for triggers Failure-Class: none * test(retrieval): align screen timestamp contract Failure-Class: none * feat(memory): correct ledger facts by amendment Failure-Class: none * test(memory): prove ledger correction on emulator Failure-Class: none * feat(macos): harden local trigger observations Failure-Class: none * feat(agent): search bounded historical facts Failure-Class: none * test(macos): cover trigger observation adapter * fix(memory): gate historical fact retrieval * feat(memory): add gated JIT retrieval strategy * test(memory): prove mixed-version JIT runtime parity * chore(memory): keep JIT gate exports type-safe * refactor(memory): isolate JIT prompt contract * fix(conversations): round-trip owner-scoped references Accept the conversation:<id> references emitted by JIT result cards while retaining strict UUID-only bare IDs and share links. Restrict machine IDs to a bounded safe alphabet so evidence suffixes and path-like values fail closed. Failure-Class: none * test(memory): join JIT citations to evidence envelope * fix(retrieval): enforce JIT conversation search budget Cap JIT summary searches per request and bound database hydration to the projection limit before reads. Preserve the legacy path when JIT is disabled. Failure-Class: FC-unbounded-user-collection-in-prompt * fix(memory): keep JIT retrieval request scoped * test(macos): prove future JIT evidence stays inert * fix(memory): keep JIT card citations request-global Failure-Class: new * fix(retrieval): separate JIT hydration from search Treat gated owner-scoped references as exact hydration without searching transcript text for the reference. Charge every JIT candidate search to the shared four-search request budget, including snippet-bearing requests, while keeping exact hydration free and preserving released JIT-off UUID/share-link behavior.\n\nVerified:\n- cd backend && ./.venv/bin/python -m pytest tests/unit/test_conversation_jit_processing.py tests/unit/test_conversation_exact_reference_search.py -q (58 passed)\n- cd backend && uvx --from pyright==1.1.403 pyright -p pyrightconfig.json --pythonpath .venv/bin/python (0 errors)\n- git diff --check\n\nFailure-Class: FC-unbounded-user-collection-in-prompt * fix(memory): keep repeated JIT cards index-safe * fix(retrieval): satisfy JIT card type contract * fix(retrieval): hydrate collected JIT cards * test(app): preserve answers during delayed evidence requests * test(app): exercise production evidence composition * feat(memories): restore superseded ledger facts * fix(memories): reconcile reverted ledger facts * feat(memories): append reverted ledger facts * feat(memories): synchronize revert client contract * fix(memory): name ledger revert identity * fix(memories): type and enlarge revert controls * fix(memories): fence revert retries and refreshes * fix(memories): fence ledger revert authority * test(memory): count ledger revert rate limit * feat: expose agent-controlled historical facts * feat: reopen standalone ledger facts * feat: add fail-closed JIT QA bundle routing * feat: add safe local JIT QA backend stack * fix: harden isolated JIT QA stack * feat: add explicit multi-source entity timeline * feat(backend): add JIT rollout authority * feat(backend): fence every proactive paid boundary * fix(backend): release proactive quota on cancellation Release the reserved proactive quota exactly once when cancellation interrupts paid-boundary refresh or a provider retry, then re-raise cancellation without emitting retry telemetry. Add deterministic regression coverage for both cancellation points. Failure-Class: FC-proactive-quota-cancellation | new * fix(backend): make proactive quota cancellation safe Detach in-flight Redis reservations on request cancellation and release only admitted slots once they settle. Move direct-provider fallback telemetry behind the fresh paid-boundary rollout check so late kill or unknown decisions cannot report false recovery.\n\nFailure-Class: FC-proactive-quota-cancellation | new * fix(backend): preserve quota compensation during shutdown Keep late Redis reservation compensators outside the ordinary cancellable background-task drain. Desktop and main application shutdown paths now wait for these critical compensators before cancelling ordinary work, with deterministic blocked-thread and lifecycle-order regressions.\n\nFailure-Class: FC-proactive-quota-cancellation | new * fix(backend): use expiring proactive quota leases * fix(backend): make quota finalization clock-safe * fix(backend): isolate jit rollout control plane * fix(backend): close jit control plane safely * fix(backend): emit retry recovery after quota commit * test(backend): keep rollout app contract fast * feat(jit): add guarded proactivity and first-open policies * chore(desktop): mark jit policy as internal * test(desktop): cover jit proactivity policy flow * feat(backend): wire durable JIT first-open processing * feat(desktop): fence JIT proactivity runtime admission * feat: activate authoritative JIT proactivity runtime * fix: harden JIT proactivity authority * fix: close proactive runtime authority gaps * fix(jit): make first-open effects resumable * fix(jit): fence outstanding first-open work * fix(jit): resume app usage receipts * fix(jit): make app usage retries no-op Failure-Class: none * fix(jit): allow completed usage after app deletion Failure-Class: none * fix(jit): register first-open folder query Failure-Class: none * Fix first-open import isolation * feat(memory): govern ledger slots and prompt winners * feat(macos): stage guarded ledger prompt adoption * feat(jit): adopt authoritative ledger prompts on macOS * fix(jit): close ledger adoption authority leaks * fix(jit): reauthorize every ledger migration write * fix(jit): fence ledger cutover publication * fix: keep ledger prompt rollback reversible * feat(jit): add guarded frame request retention contracts * fix(jit): close frame retention authority and evidence lifecycle * fix(jit): make frame retention retries and cleanup durable * fix(jit): make frame evidence recovery and retention complete * fix(jit): close frame retention recovery gaps * Harden temporary frame retention and deployment * fix: harden JIT frame retention and consumption * fix: close JIT frame lifecycle recovery gaps * fix: unify JIT frame authority and retention Failure-Class: FC-split-mutation-authority * docs: keep frame retention guidance lean * fix: retire duplicate frame flag bindings Failure-Class: FC-split-mutation-authority * fix: register frame keyframe queries Failure-Class: FC-split-mutation-authority * fix: serialize frame retention deploys Failure-Class: FC-split-mutation-authority * test: cover frame pixel deletion ordering * style: format cumulative Dart changes * fix(app): retain permanent conversation photo fetches * fix: bound frame vision retention and authority * fix: drain terminal frame request metadata * chore: record internal ledger adoption change * feat(memory): add dark daily sweep authority * feat(memory): harden daily sweep fences and runtime seam * feat(memory): reconcile existing standing triggers in sweep adapter * fix(memory): harden daily sweep recovery and source fences * fix(memory): close daily sweep source producers * fix(memory): close daily sweep review findings * Add dark daily memory sweep authority and recovery * fix(memory): harden daily sweep rejection repairs * test(listen): stub onboarding admission in bootstrap regression The daily sweep PR fences onboarding mode behind the server-owned backend admission (get_backend_onboarding_admission), so the bootstrap regression test now simulates an admitted session instead of failing closed on a real Firestore read. Verification: focused test passes in 1.64s (previously failed after a 4m27s Firestore timeout); full test_listen_runtime_regressions.py + test_onboarding_question_start.py: 26 passed; black --check clean. * fix(memory): close daily sweep rollout and retry cursors * fix(memory): isolate daily sweep lifecycle and retry fairness * Harden daily sweep admission and completed-day staging * fix daily memory sweep reliability boundaries * preserve daily sweep invocation tombstones * close daily sweep invocation lifecycle fences * fix: keep daily sweep lifecycle cleanup active * fix: acquire ledger snapshot client off event loop * fix(memory): preserve migration tier fence without legacy growth * test(memory): prove legacy adjudication race fences * fix(dev): allow bounded ADC readiness refresh * test: keep ledger prepush deterministic * test(memory): register prompt receipt control path * fix(memory): fence ledger writer transitions * feat(backend): preserve closed ledger history in export * feat(memory): define ledger query semantics * fix(backend): fence trigger snapshots on final authority * fix(backend): bypass stale coalesced JIT refreshes * feat(macos): mirror bounded memory evidence Decode generated v3 evidence into a domain mirror, persist canonical bounded JSON through the memory cache, and preserve it across compatibility sync and older-local conflicts. Invalid, future-shaped, oversized, and over-count payloads fail closed without hiding memory text or granting prompt authority. Tests: xcrun swift test --package-path Desktop --filter ServerMemoryV17DecodingTests Tests: xcrun swift test --package-path Desktop --filter MemoryLedgerMirrorTests Tests: python3 scripts/check_desktop_test_quality.py Failure-Class: none * fix(macos): fence and classify memory evidence Keep generated memory fields independent from malformed evidence, distinguish absent valid and invalid evidence states, preserve prior evidence on invalid payloads, and gate replacements on a monotonic server timestamp so stale active evidence cannot resurrect redacted rows. Cover populated-table migration upgrades. Tests: xcrun swift test --package-path Desktop --filter ServerMemoryV17DecodingTests Tests: xcrun swift test --package-path Desktop --filter MemoryLedgerMirrorTests Tests: python3 scripts/check_desktop_test_quality.py Failure-Class: none * fix(macos): preserve evidence fences and scrub redactions Advance evidence revisions for identical valid payloads, fence stale active responses after a local edit, and remove artifact/device pointers from redacted evidence before canonical persistence. Tests: xcrun swift test --package-path Desktop --filter ServerMemoryV17DecodingTests Tests: xcrun swift test --package-path Desktop --filter MemoryLedgerMirrorTests Tests: python3 scripts/check_desktop_test_quality.py Failure-Class: none * chore(macos): record ledger evidence mirror * feat(macos): deep-link local evidence cards to Rewind * fix(macos): fence Rewind frame evidence version * fix(macos): validate Rewind evidence card availability * fix(macos): bind task detail Rewind navigation to local leases * fix(macos): fence Rewind citation owner handoff * chore(macos): register Rewind evidence deep links * test(macos): cover Rewind evidence navigation * feat(desktop): evaluate JIT trigger watchlists locally * feat(desktop): wire authoritative JIT trigger runtime * feat(desktop): bind JIT claims to snapshot authority * fix(desktop): revalidate trigger authority at execution * fix(desktop): keep JIT execution leases live * test(memory): bind standalone reopen to direct-user writer * fix: make JIT QA sign-in self-contained Failure-Class: new Verification: bash desktop/macos/tests/test-jit-qa-target.sh; bash desktop/macos/tests/test-yolo-dev-backend.sh; repaired named-bundle Google sign-in reached authenticated onboarding. * feat(memory): complete JIT policy and native Windows parity * docs(backend): keep service map within context budget * test(macos): cover JIT client and staging flows * chore(backend): declare JIT mirror route policy * fix(backend): use strict Firestore boundary for JIT admission Failure-Class: FC-malformed-doc-read * chore(quality): register malformed-document guard surface * fix(backend): fail closed on malformed JIT authority Failure-Class: FC-malformed-doc-read * refactor(backend): name JIT workflow boundary results * test: repair JIT CI contracts * fix(backend): preserve ledger query exports Retain the explicit same-name re-exports consumed by tests and downstream callers while satisfying the enforced Pyright unused-import boundary after the main rebase. Failure-Class: none * test(backend): isolate gateway setup timing Failure-Class: none * style(memory): format direct-user evidence path Failure-Class: none * test(agent): isolate ACP process-group fallback Failure-Class: none * fix(dev-harness): preserve ownership markers in narrow CI * test(jit): refresh emulator fixtures for current contracts * test(jit): orchestrate local rollout dogfood * test(jit): harden local dogfood authority * fix(dev-harness): install PostHog for CI tests * fix(chat): project server JIT rollout into retrieval Resolve the backend-owned PostHog decision inside the bounded agent setup path and pass only its boolean result to prompt/tool configuration. Unknown or failed authority remains on the released legacy path, while callers cannot self-enroll through configurable input.\n\nVerification: backend/.venv/bin/python -m pytest -q backend/tests/unit/test_chat_async_offload.py backend/tests/unit/test_atomicity_lifecycle_regressions.py (41 passed)\n\nFailure-Class: new * fix(memory): preserve preference writer compatibility Select the agent preference write path from the canonical per-user writer control. Default compatibility mode retains the released MemoryService payload and receipt behavior; ledger mode keeps the retry-stable ledger write, and transition states fail closed.\n\nVerification: backend/.venv/bin/python -m pytest -q backend/tests/unit/test_chat_async_offload.py backend/tests/unit/test_atomicity_lifecycle_regressions.py (41 passed)\n\nFailure-Class: FC-split-mutation-authority * fix(jit): separate migration rollout authority Keep staged JIT chat and proactive exposure independent from legacy-row migration and writer cutover. Migration now requires its own default-off PostHog flag and still rechecks the shared kill switch at every mutation and publication boundary. Repair the isolated conversation-JIT fixture for main's chat-scope import. Verification: 217 focused JIT, chat-scope, migration, and lifecycle tests passed; 28 conversation-JIT fixture tests passed; independent Sol review accepted the split for QA-only dev rollout. Failure-Class: FC-split-mutation-authority * fix(photos): preserve retained image retrieval Treat an empty legacy inline marker as absent when permanent storage is authoritative, while malformed non-empty inline payloads still fail closed. Route live and retained thumbnails through the storage-aware image loader and preserve the conversation identity through the full-screen viewer.\n\nVerification: backend data-export tests 32 passed; Flutter photo-viewer tests 5 passed; focused Dart analysis clean; independent Sol review found and verified the viewer identity repair.\n\nFailure-Class: none * fix(memory): keep disabled daily sweep dark Resolve the backend-owned authority before inventory and require its literal true decision before any UID discovery, registry, cleanup, scheduler, model, or commit work. Missing, malformed, throwing, disabled, and kill-switched authority now exits without touching user data; enabled behavior is preserved.\n\nVerification: 60 focused daily-sweep job, scheduler, and inventory tests passed; independent Sol review accepted the fail-closed gate.\n\nFailure-Class: FC-split-mutation-authority * fix(jit): satisfy fail-closed type contracts * test(backend): admit full runtime contract checks * style(backend): format conversation bound test * test(backend): keep conversation router isolation current * test(backend): admit export boundary duration * fix(macos): persist failed chat turn notice Failure-Class: none * fix(macos): repair JIT rollout admission contracts Failure-Class: none * fix(windows): treat JIT screen evidence as untrusted Failure-Class: none * fix(backend): preserve explicit app failure contract Failure-Class: none * fix(app): finish photo viewer consolidation * fix(backend): make provider writes lock-free against the deletion gate The account-wide legal-hold deletion gate wrapped every GCS upload and Pinecone/Typesense upsert in an exclusive per-uid Firestore mutex with no lease: concurrent same-account writes hard-failed (dropped audio, lost vectors) and a crash between acquire and finish blocked the account's gated operations forever, with no janitor. Provider writes now use a lock-free fence that refuses only during account deletion or a live destructive operation; destructive kinds keep exclusive ownership, an abandoned gate self-expires after six hours, and releasing a gate on the failure path can no longer mask the original error. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(backend): issue onboarding admission at socket connect The completed-onboarding early exit returned False from an Optional[str] function; the listen runtime derives admission via 'is not None', so users who had already completed onboarding were admitted with a fabricated session id — the exact provenance forgery the admission exists to prevent. Separately, the 20-minute admission TTL was anchored to the app-launch state read, so a user reaching the speech-profile step late (or any client that never calls the state endpoint) silently lost onboarding questions and is_user tagging. The bootstrap now issues or refreshes the admission from the durable account state at connect time; completed accounts still can never re-enter, and issuing stays best-effort with the read failing closed. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(backend): keep the released proactivity lane open for legacy clients Gating /v1/desktop/proactivity/completions on the JIT cohort returned 403 to every non-admitted user — which is the entire deployed desktop fleet on deploy day, since shipped clients poll this route continuously and treat 403 as a plain error. Context-bucket extraction and the director would have died fleet-wide, dark cohort or not, and any environment without a PostHog key (local, self-host) would have lost the lane entirely. The route returns to merge-base admission semantics (tier quotas only); JIT admission remains enforced on the JIT reservation routes, and retiring this lane stays a later explicit operation after clients migrate. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(backend): withhold JIT tools and history reads outside the rollout Five new tools (search_knowledge, search_historical_facts, read_playbook, get_entity_timeline, look_at_frame) sat unconditionally in CORE_TOOLS, so every legacy chat request carried their schemas and the model burned tool budget on 'no entries found' answers. They are now filtered per request off the same resolved rollout boolean that gates the JIT prompt appendix. The memories-tab ledger-history endpoint likewise answered every user with a bounded 501-row provider scan that can only ever be empty outside the rollout; it now returns empty without the scan for non-admitted (and unknown/error) states. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(backend): bound rollout control-plane cost and confine sync resolution Synchronous callers resolved rollout flags via per-call asyncio.run against the shared provider singleton, crossing event loops: awaiting a Task attached to another loop raises, a timed-out asyncio.run strands a coalescer entry that then serves stale UNKNOWN forever, and the LRU cache was mutated from multiple threads. Sync resolution now runs on one long-lived control-loop thread with its own authority instance. Unknown snapshots gain a 5-second negative cache — UNKNOWN can never authorize work, and without it a fleet whose flags are simply absent pays one uncached PostHog call per conversation finalization. The screen-sync loop drops its force_refresh (one uncached decide per device per minute fleet-wide) and moves to its own rate bucket so two Macs' background sync can no longer starve conversation photo reads out of the shared 120/hour frame-requests bucket. The first-open policy's kill-switch telemetry label also reported str(Enum) instead of the value and could never match. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(backend): skip eager extraction under a non-compatibility writer mode A ledger-cutover user still ran the full L1 extraction model call at finalization, after which writer admission refused the compatibility write — the conflict retried, exhausted, and failed the entire finalization for every conversation, with the model spend already paid. Extraction now checks the canonical writer mode first and skips when the daily sweep owns memory formation; only a positively-read non-compatibility mode skips, so any control-state read failure preserves the legacy eager path. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(backend): export tolerates byte-less legacy photo rows A conversation photo row carrying the legacy empty inline marker and no storage reference failed the whole portability export forever, though it holds no durable image anywhere — there is nothing to omit. Such rows now export as metadata with a content-free gap reason. Frame requests in a retained state keep the fail-closed contract via an explicit require_bytes parameter. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(windows): harden JIT delivery, admission, and bootstrap boundaries Five verified defects: (1) the exclusive notification delivery slot leaked on any throw between reservation and commit — one SQLite hiccup during a JIT turn permanently silenced every proactive lane; the span is now try/finally-guarded and stale slots expire after ten minutes. (2) The ambient lane interpolated the raw window title into a tool-capable agent prompt; the turn now carries only the opaque context handle plus a sanitized executable name, framed as untrusted data like the nano-triage lane. (3) Google Calendar was fetched every ~60s before admission, so non-cohort users with Google connected paid ~1,440 reads a day for a refused feature; observation now gates calendar evidence on the cached authority. (4) Rollout-authority errors reset the cache and retried every frame (~1 req/s offline, forever); failures now back off from 30s to 10 minutes. (5) An unguarded JIT schema exec inside the shared database open could abort local storage for all features; the mirror bootstrap is now isolated, keeps the host-facing tables alive, and JIT stays inert when unavailable. Also re-checks the control-plane owner before committing the toast so an account switch mid-turn cannot show the previous owner's advice. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(macos): restore screen provenance, guard migrations, fence chat turns Four verified defects: (1) every pre-existing screen-derived task lost its 'Screen context / Open Rewind' source row because the new evidence policy dropped any provenance that is not rewind_frame.v1; the merge-base fallback row is restored for capture.v2/legacy refs (a test flipped to match the regression is restored to its merge-base assertions). (2) RewindDatabase published its pool before migrating, latching a failed migration into a permanent false-initialized state, and three unguarded ALTER TABLE memories migrations died with duplicate-column on machines that ran earlier builds of this branch; migration now precedes publication and the ALTERs/CREATEs are existence-guarded. (3) EventKit was queried on every context visit before the flags check; non-admitted owners now build no observation inputs. (4) A failed chat turn's reconstructed notice could be appended into a different conversation's transcript when the user switched sessions or cleared chat mid-flight; both transcript resets now revoke the active turn like selectApp already did. The pre-terminalized discard class (user Stop/watchdog) still drops the durable notice on relaunch — pinned by a characterization test in agent/tests/conversation-journal.test.ts with the least-invasive fix described there. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(testing): resolve firebase-tools from the checked-in dependency npx --prefix resolves the package bin against the current directory on some npm versions, and the admission runner deliberately launches from an isolated temp dir (firebase writes debug logs to cwd) — surfacing as 'sh: firebase: command not found' on hosts without brew node@22. Prefer the vendored node_modules binary when it matches the pin; npx remains the fallback. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(backend): keep one eager-extraction call site for the surface ratchet Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(backend): gate eager extraction at the public boundary The writer-mode skip moves from _extract_memories_inner to extract_memories: the replace-policy contract test pins the inner helper to exactly the canonical replacement path, and the public boundary is the better seam anyway — a sweep-owned user now skips parity capture and usage tracking along with the model call. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(listen): stub onboarding admission issuance in bootstrap regression The connect-time ensure call landed in a harness that only stubbed the read, so the bootstrap test paid an extra real-module exception path and grazed the 0.30s fast-unit CPU budget under fanout load. Stub the issuance like the read. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(listen): allowlist the bootstrap regression's CPU budget The full listen-runtime bootstrap test measures exactly at the 0.30s fast-unit CPU budget under a saturated pre-push fanout (CPU inflates ~2x there per the guard's own notes) while passing comfortably alone. It exercises deliberately heavyweight machinery; record it as an intentional exception rather than trimming the coverage. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(backend): keep list(CORE_TOOLS) literal through JIT tool gating The JIT-only tool filter replaced the list(CORE_TOOLS) assignment with an inline comprehension, which broke the prompt-cache structural invariant (test_prompt_cache_optimization.py::test_core_tools_used_in_both_functions). Restore the list(CORE_TOOLS) copy and apply the JIT-only filter as a conditional pass, preserving rollout semantics and tool order. * feat(jit): drop automatic goal updates from the JIT featureset Product decision (David, 2026-08-26): goals change only through explicit user action for JIT-admitted conversations. Goal progress is no longer a first-open obligation — the effect is removed from FIRST_OPEN_EFFECTS and the worker, and the policy plan can no longer express deferring it. Legacy obligations carrying a pending goal_progress row are normalized away and complete on the remaining two effects. Non-JIT (legacy eager) conversations keep today's automatic goal updates unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(sweep): one summary-spine agent pass per day, with folder backstop Replaces the per-conversation transcript extractor in the completed-day producer with a single two-phase agent run: the whole day's conversation summaries go in as one bounded spine (200 conversations / 120k chars — effectively unreachable, so heavy days no longer stall the cursor), and the agent may request up to 8 raw transcript excerpts (8k chars each) to verify specifics before finalizing. At most two provider calls per user per day, both inside the existing at-most-once invocation fence; the staged page carries the memory candidates AND folder assignments for the day's unopened, unfiled conversations, applied idempotently (first-open or user assignment always wins). Memories must cite their source conversations; uncited output is dropped. The cost gate becomes a worst-case ceiling checked before any call. The onboarding cold-start channel keeps per-conversation transcript extraction unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(sweep): harden the daily agent prompts from a real-data lab pass Iterated on one real heavy day (26 conversations) with strong- and weak-model stand-ins, an adversarial judge, and hand-verified transcript ground truths. Rules added, each pinned to an observed failure: actor binding in active voice with a personal-attribute gate (a discussed or recommended topic is never someone's attribute; judgments about named people are stored as assessments); decision-state basis labels binding the verb (decided/proposed/observed, discussed-no-outcome dropped); salience ordering (money, metrics, named-party intent, identity, and durable decisions before any operational fact; one fact per memory); never guessing the direction of an invitation/offer/commitment (verify or drop); and no deferring the whole answer to verification. The agent output schema gains a 'basis' field. The memories QoS call-site inventories now count the daily-sweep agent's call site (3 -> 4). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(sweep): tune the daily agent prompts against the real memories model Ran the assembled prompts against gpt-5.6-luna (the real 'memories' route model) on the same real day. Three refinements from observed behavior: the basis label no longer leaks into memory text (metrics read as metrics, not 'David observed that…'); the never-guess-direction trigger is mechanical (passive/verbless summary phrasing or 'Speaker' as the actor forces a transcript_request — luna confidently inverted 'Tim: Invited to New York' until this; with it, phase B verifies and corrects to the true direction), hedging is itself a request signal, and nothing high-salience may be silently dropped; and a rich-day yield anchor (8-16 memories for 15+ conversations) counters the model's over-pruning without inviting padding. Final real-model run: 11 true memories + 2 legitimate verification requests, zero fabrications, ~22k tokens (~2 calls) for a 26-conversation day. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(sweep): profile-maintaining slots, ledger lookups, cache-ready prompts The daily agent now sees the user's current profile (the same get_prompt_memories seam chat uses — the ledger render for migrated users), may run up to 4 owner-scoped prior-memory keyword lookups (provider fail-soft; hits re-read through the canonical store before disclosure) to dedup and supersede, and may name a slot for standing attributes — an occupied slot becomes an amend through the existing canonical occupancy check, so the daily run maintains the rendered profile with no second write path. Both phase prompts share a byte-identical prefix (pinned by a test) and pass a per-user prompt_cache_key through get_llm; measured against gpt-5.6-luna the provider cache is exact-match rather than prefix-based today, so this is future-proofing rather than present savings. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(sweep): type the memory-searcher seam for the pyright contract CI's authoritative typecheck rejected the untyped lookup seam (memories.py: list(Any or [])). The searcher is now Optional[Callable[[str], Sequence[str]]] and results are built through a typed comprehension; behavior unchanged (absent or failing searcher still degrades to an empty result block). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: repair four main-inherited CI breakages after sync origin/main is currently red on its own tip; syncing it into this PR inherits the breakage, so the fixes ride here: - subscription.py: drop the unused get_byok_keys import (pyright reportUnusedImport fails the Backend unit suite). - AppState+Transcription.swift: explicit self for alertPresenter inside the escaping showAlert completion (strict-concurrency compile error in all three Desktop Swift lanes, shipped red on main by d49f978512). - AppState+Permissions.swift: pinned swift-format drift from the same main commit (desktop-swift-format-lint). - web/app/bun.lock: add the prettier + prettier-plugin-tailwindcss entries 64db30c791 pinned in package.json without updating the lockfile (frozen install fails web-app-checks). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(sweep): close the second review round's findings Three parallel adversarial reviews over the post-takeover additions: - Clamp every model-controlled phase-B input (draft memories, request reasons, lookup queries/results) and add the clamped worst case to the pre-call cost ceiling, which previously under-estimated phase B. - Attest an empty consumed day when the staged page carries an older stage schema version instead of stalling the cursor forever on every deploy-boundary schema bump. - Make the folder backstop's unfiled check and write share one transaction so a concurrent first-open/user assignment always wins. - Let equal-rank sweep candidates amend sweep-authored slot occupants: the profile-maintenance path froze after a slot's first write. User statements still always win; slotless subject matches still dedup. - Neutralize ``` fences in summaries/excerpts/lookup results, and mark raw-transcript fallback rows '(unstructured transcript excerpt)' with a prompt rule refusing slots/personal attributes from them without transcript verification (test pins the marker to the rule). - Remove the dead first-open goal-authority threading left by the goals removal, and update the stale jit-first-open-runtime doc. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: repair three more main-inherited breakages All shipped red on main and only surfaced once earlier failures were cleared: - AppState.swift: move the alertPresenter default out of the stored property initializer — Xcode 16.4's SILGen segfaults (signal 11) emitting it, which failed all three Desktop Swift lanes even after the explicit-self fix. - test_byok_security.py: main's BYOK rewrite (d0e3a4eb3a, 1da8880175) changed request_has_llm_byok_key to per-provider enrollment checks and made partial headers fail closed, but left the tests targeting the old get_byok_keys()-based lenient contract (masked on main because pyright failed before pytest ran). The tests now assert the shipped strict contract their own docstrings already describe. - subscription.py: pinned-black formatting for the BYOK fallback expression (the Formatting lane rejects the file as main wrote it). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): stub the chat-agent gateway route pin in the chat router harness Main's a6988be309 made routers.chat import CHAT_AGENT_ROUTE_DIRECT / get_chat_agent_route from utils.llm.gateway_client, but the chat-router test harness (and test_chat_file_upload_unsupported's local override) stub utils.llm.gateway_client without those symbols, so every suite that loads the real router failed at import — masked on main because pyright fails its Backend unit suite before pytest runs. Ninth main-inherited repair in this sync. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): teach test_chat_quota's utils.byok stub the rewritten import surface utils/subscription.py now imports get_byok_uid and get_cached_byok_state (main's BYOK rewrite); the module-scoped utils.byok fake predates them, so reloading subscription under the fake raised ImportError at setup — and the polluted process took test_chat_openapi_operation_ids and test_desktop_screen_crisp down with it in CI's batched run (all three pass standalone). Tenth main-inherited repair, same pyright-masked pattern. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): update three more suites for main's BYOK/gateway import surface Same pyright-masked pattern as the harness and test_chat_quota repairs: - test_desktop_transcribe stubbed utils.llm as a non-package, so routers.chat's new utils.llm.gateway_client import could not resolve (50 failures); the submodule is now in its stub list. - test_paywall_reconnect_gate's BYOK escape-hatch tests never set the request uid context that the enrollment-verifying rewrite requires (middleware sets it in production); they now do, and teardown clears it. - test_chat_session_app_identity's enforce_chat_quota stub rejected the new required_llm_provider keyword. All three suites pass locally (69 + 35 + 6). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): enroll fingerprints in the desktop BYOK tests PR #11454 moved macOS BYOK activation to enrollment-verified fingerprints (isByokActive and usableBYOKEnvironment gate on persistEnrolledFingerprints), and its own test lanes shipped red: the tests store raw keys but never enroll them, so every key reads as inactive. Their teardowns already clear enrollment — the setups now enroll what they store, matching the production activation path. All 8 previously-failing cases (BYOKPaywallTests + the two AgentRuntimeProcessTests BYOK-environment cases) pass locally. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(deploy): enable the daily memory sweep on development The sweep's five deployment inputs were pinned off in every environment, so cohort enrolment alone could never start it -- turning it on for a dogfood account required a second PR. Development now carries the live values: - ENABLED/MODEL_ENABLED on, so the job stops exiting at its first authority gate and the model authority can budget a route. - MODEL_NAME pinned to gpt-5.6-luna, which is the declaration interlock the runner checks against get_model('memories') before any provider call. - MAX_MODEL_COST_USD 0.80, the worst-case pre-call ceiling for a maximal day including phase B's clamped draft/reason/lookup overhead. - COHORT_ENABLED on with COHORT_FLAG daily-memory-sweep-v1, so enrolment is a per-uid PostHog boolean and an unnamed cohort stays a closed rollout. Production is deliberately untouched and stays fully pinned off. The job still cannot form a memory for anyone until that flag exists and resolves true for a uid, which remains a control-plane action rather than a deployment one. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(firestore): terminate the daily-sweep occupant indexes with __name__ The six daily-sweep occupant lookups were the only declarations in the manifest without a trailing __name__ field -- 63 of 69 entries carry one, and main had none missing it. Firestore appends the terminator itself and reports the index back that way, so these six could never match the live inventory. The failure mode is not a missing index; the indexes build fine. It is that reconciliation never converges: every run reports the same six as missing, tries to create them, and fails on ALREADY_EXISTS. That takes down the Firestore schema workflow on both environments permanently, and with it the development backend deploy's readiness gate -- the same class of outage the workflow's own header records from the hourly_usage index in PR #11979. The derived specs previously appended their extra predicates to the base spec's index_fields, which would have placed them after the terminator, so the shared prefixes are now named explicitly and each spec ends with __name__. Verified against real Firestore: reconciliation reports zero missing indexes in both based-hardware and based-hardware-dev. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: close final JIT rollout and CI gaps Fence direct JIT tools and frame pixels, keep Windows account wipes safe after optional schema failures, and repair inherited CI regressions. Failure-Class: none --------- Co-authored-by: David Zhang <9387252+Git-on-my-level@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> | 11 天前 | |
fix: keep the macOS hourly release train moving Bind each candidate to its checked source, recover squash-orphaned changelog branches, install Bun for release eligibility, and age freshness from the oldest unshipped update. Verification: 77 focused release-control tests; actionlint; Python compile; diff check. Failure-Class: new | 18 天前 | |
fix: make no-changelog-needed survive the merge boundary The PR label greened PRs and then reddened main because push runs cannot see labels. Require an in-repo kind:none fragment for internal production desktop edits so both lanes agree. Failure-Class: FC-changelog-exemption-lost-at-merge Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: multica-agent <github@multica.ai> | 20 天前 | |
feat(desktop-release): record serving backends on the stable pointer Persist the live desktop-backend and API health identities on promote and repoint, keep them off the immutable manifest, and print a read-only drift table at promotion time so lag is visible without a new block. Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: multica-agent <github@multica.ai> | 4 天前 | |
fix(ci): ensure the release-gate-failure label exists before filing audit issue (#10389) (#10966) * fix(ci): ensure the release-gate-failure label exists before filing audit issue (#10389) Both break-glass audit jobs ran gh issue create --label release-gate-failure with no guarantee the label exists in the repo. GitHub rejects issue creation with an unknown label, so the audit issue was never filed (the prod Cloud Run bypass in #10301 was preserved manually). Ensure the label exists first with an idempotent gh label create, then file the issue — the audit trail can no longer be dropped by a missing label. Failure-Class: none Co-authored-by: CommandCodeBot <noreply@commandcode.ai> * fix(ci): require the issue-create --label flag in the break-glass admission check The gateway validator required the bare `release-gate-failure` string anywhere in gcp_llm_gateway.yml. The new `gh label create release-gate-failure` step (and its comment) satisfy that weak fragment, so a workflow that dropped `--label release-gate-failure` from the audit issue would pass the check. Tighten the fragment to the exact `--label release-gate-failure` flag on the issue-create command, which is the property the check exists to pin. Co-authored-by: CommandCodeBot <noreply@commandcode.ai> * fix(ci): fail loudly when the release-gate-failure label cannot be ensured The idempotent `gh label create ... 2>/dev/null || true` swallowed every failure — permissions, rate limits, API errors — not just the tolerated "already exists" case, so a real label-create failure would still drop the audit issue silently. Capture the output and only suppress the exact "already exists" message; any other failure now fails the job loudly before the audit issue is attempted. Co-authored-by: CommandCodeBot <noreply@commandcode.ai> --------- Co-authored-by: CommandCodeBot <noreply@commandcode.ai> | 13 天前 | |
fix(infra): harden admitted GCP deploy control plane Resolve the main merge and repair deploy control-source staging, rollback, admission, and runtime configuration guards. Failure-Class: FC-workflow-control-source-identity | 1 个月前 | |
fix(monitoring): accept the Cloud Run metrics egress filter at Cloud Monitoring (#12099) * fix(monitoring): accept the Cloud Run metrics egress filter at Cloud Monitoring The dedicated Stackdriver exporter added in #11998 has never imported a single series. Cloud Monitoring rejects a filter that mixes AND with OR across resource.labels restrictions, so every descriptor query returned HTTP 400 while the exporter stayed Available, its Prometheus target stayed up, and Grafana showed empty panels that read as no traffic. Express the namespace disjunction as one_of(...), which the filter grammar defines for this case. Keep the namespace scope: dropping it would import every omi_ series from every Cloud Run service in the project. Add omi-cloud-run-metrics-egress-query-rejected, alerting on the exporter's own upstream error rather than on its liveness. Extend the exporter contract test to cover the dev values file, which was unasserted and is what the automatic post-merge rollout installs. Correct the runbook's verification step, which queried the mangled metric name that a healthy deployment no longer produces. * fix(monitoring): key the Cloud Run observer exemption on the monitored resource The production-data-plane-routing guard exempts the Stackdriver egress values files from its retired-GKE-desktop-backend rule only when they contain the literal resource.labels.namespace="desktop-backend". That pins the exemption to one spelling of a filter rather than to what makes the file a Cloud Run observer, so rewriting the disjunction as one_of(...) to satisfy Cloud Monitoring's grammar made a read-only metrics reader look like retired GKE ownership. Key the exemption on resource.labels.cluster="__run__" instead. That is Cloud Run's reserved pseudo-cluster, so together with the prometheus.googleapis.com/ prefix it identifies the monitored resource directly and survives any future edit to the namespace set. --------- Co-authored-by: r <r@r> | 15 天前 | |
docs: take operator pages off docs.omi.me Unlisted Mintlify MDX is still a public URL. Move runbooks, flags, invariants, and agent rules next to owning code, add docs/AGENTS.md as the site allow-list, and correct the live kill-switch contract after the JIT authority page leaves the site. Co-authored-by: Cursor <cursoragent@cursor.com> | 7 天前 | |
test(admin): isolate push-range fixtures from hook GIT_DIR Pre-push exports GIT_DIR into child git commands, so the disposable classifier repos must clear GIT_* and disable inherited hooks. Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: multica-agent <github@multica.ai> | 22 天前 | |
fix(admin): classify the full push range and keep the activation shim honest Scope now diffs github.event.before...github.sha and fail-closes when before is missing or zero. Admission requires live scope wiring, workflow_dispatch, and uncommented deploy keys. Partial Firestore caches no longer overlay a silent rate, and classic keeps the daily viral-metrics fallback. Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: multica-agent <github@multica.ai> | 22 天前 | |
ci: bind agent-doc claims to the code they describe check_agent_doc_references.py proves a doc's pointers resolve. Nothing proved that what a doc asserts about behavior is still true, and a doc can keep every pointer valid while describing behavior replaced years ago. An agent reading it cannot tell current from stale. The instance this would have caught is the first commit in this branch: .cursor/cloud-agent-environment.md claimed backend/test.sh "halts at the first failing file" long after it became an aggregating worker pool, and an agent acted on that by proposing to replace the per-file process isolation the suite depends on with `pytest -n auto --dist loadfile`. Each claim binds a literal in a doc to a literal in the code that makes it true, and fails if either side moves without the other. Three seeded claims, all verified on this branch, all about backend/test.sh's execution model. check_trigger_coverage asserts every path named by a claim is matched by this check's own manifest triggers -- a claim outside its triggers never runs on the change that invalidates it, which is the same drift hiding in the wiring. Not a shared primitive already: check_agent_doc_references.py answers "does this path exist", a different question with a different failure mode, and its resolver has no notion of a doc/code pair. Deliberately not semantic -- it is a tripwire that forces a re-read, not proof the prose is true. Verifying meaning needs judgment, and a check that needs judgment grows an allowlist. docs/agents/doc-maintenance.md already asks for exactly this: "When a defect ships because guidance was misread or missing, tighten the guidance in the fix PR ... or add the check that catches it." Verification: python3 .github/scripts/check_agent_doc_claims.py -> agent doc claims OK python3 .github/scripts/test_run_checks.py -> 26/26 pass removed a code anchor from a scratch copy of test.sh -> fails naming the code side reworded the doc anchor in a scratch copy -> fails naming the doc side dropped the backend triggers from this check's entry -> fails naming the gap Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> | 1 个月前 | |
docs: take operator pages off docs.omi.me Unlisted Mintlify MDX is still a public URL. Move runbooks, flags, invariants, and agent rules next to owning code, add docs/AGENTS.md as the site allow-list, and correct the live kill-switch contract after the JIT authority page leaves the site. Co-authored-by: Cursor <cursoragent@cursor.com> | 7 天前 | |
fix(tooling): scope local repo scans to the tree CI checks out (#12678) `agents-md-lean` (#12572) and `plan-catalog-contract` (#12476) both walk the working directory. On a developer machine that directory also holds gitignored siblings — `.claude/worktrees/` is the multi-worktree pattern documented in `docs/multi-worktree-dev.md` — and each of those holds a full copy of the repo. So both checks report files that are untracked, absent from the pushed diff, and absent in a CI checkout: .claude/worktrees/agent-a49a501fb54466d1b/.github/AGENTS.md: new AGENTS.md has no budget .claude/worktrees/<old-branch>/backend/charts/values.yaml: embedded Stripe price ... is absent Both pass in CI and fail only locally, which reads as a broken gate. The escapes are `git push --no-verify` and `PRE_PUSH_SKIP_PR_PREFLIGHT=1` — each of which drops every other guard in the same lane, so a cosmetic false positive costs real coverage. Both scanners now skip what `git ls-files --others --ignored --exclude-standard --directory` reports, so a local run sees the same tree CI does. When git cannot answer the scanners fall back to their existing static exclusion lists rather than failing, so neither check gets weaker. The issues each proposed adding `.claude` to an exclusion list. Deferring to git covers the same case without a list that has to be extended for the next ignored sibling — and `.claude/worktrees/` is only one of them. Verification, both reproduced before fixing and mutation-checked after: - `check_agents_md_lean.py` self-test builds a git repo with a gitignored `.claude/worktrees/agent-abc/.github/AGENTS.md`. Reverting the fix fails it with `got ['.claude/worktrees/agent-abc/.github/AGENTS.md', 'AGENTS.md']`. - `test_source_scan_skips_gitignored_sibling_worktrees` covers the catalog scan; reverting the fix fails it. 23/23 in `test_plan_catalog_contract.py`. - Planting a real `.claude/worktrees/<name>/backend/charts/values.yaml` in this checkout reproduced the reported failure verbatim on the unfixed script, and is clean on the fixed one. Scope: only the two scanners the issues name. 25 other scripts under `.github/scripts/` and `backend/scripts/` walk the tree without excluding gitignored paths and can carry the same defect; that survey belongs in its own change rather than widened into this one. Failure-Class: none | 4 天前 | |
Repair PostHog telemetry ownership and coverage (#10660) * Repair mobile device lifecycle telemetry * Restore Omi device purchase intent telemetry * Track each permissions interstitial presentation * Report backend account deletion outcomes * Add macOS device pairing telemetry * Restore Windows PostHog delivery * Guard analytics emitter reachability * Define PostHog regression alert contracts * Document load-bearing analytics events * Add macOS pairing telemetry changelog * Cover device vendor mapping in desktop flow * Use typed pairing defaults in device tests * Stub the new telemetry module in account-deletion isolation tests Both files close the utils namespace (__path__ = []), so the added utils.integration_telemetry import in account_deletion.py resolved only against sys.modules and raised ModuleNotFoundError in CI. * Stop a Swift property modifier leaking onto the next analytics method The modifier scan walked back from func to the previous brace, so a 'private var x' declared directly above 'func setX' donated its private and the method was audited as an unreachable helper. Main's integrationConnect test seam hit exactly that shape on merge. Baseline records main's integrationConnect call-site counts and its test-only seam. * fix(telemetry): avoid deleted-UID PostHog identity and pin Windows host Account-deletion completion/failure telemetry now uses a service distinct_id with $process_person_profile=false so wiped Firebase UIDs are never re-identified in PostHog. Windows PostHog host is fixed to the CSP-allowed us.i.posthog.com origin. * fix(ci): black-format deletion telemetry and ratchet line counts Format account_deletion.py for black 26.5.1 and raise product-file line-count baselines for users.py and storage.py with justifications; ratchet OmiApp.swift down to the post-repair line count. * fix(ci): classify analytics test seams and re-baseline against main The reachability tripwire audited `set*TelemetryCaptureForTests` as if it were a production emitter, so every new scoped test seam had to be hand-added to `public_orphans` — main's `setSuggestionAssistantTelemetryCaptureForTests` failed the check for exactly that reason. Installing or forwarding to a test capture is production-unreachable by construction, so `emitters()` now drops `*ForTests`/`*ForTesting` methods before the audit and the five seam entries leave the baseline. The baseline also predated main. Regenerating it records main's own drift: 25ea86a1d2 consolidated the onboarding Google-connect flow through ConnectorImportRunner, dropping the four SBOnboardingModel+Steps emit sites (integrationConnectAttempted 3->1, Succeeded/Failed 2->1 — each still requires its remaining call site), and main's live-suggestion work grew suggestionAssistantGateOutcome to 3 and suggestionAssistantDeliveryOutcome to 2. Only those ten entries moved; nothing else was absorbed. Verified: - python3 .github/scripts/check_analytics_reachability.py -> "analytics reachability static tripwire passed" - python3 .github/scripts/test_check_analytics_reachability.py -> Ran 7 tests, OK (new test_test_only_seams_are_not_audited_as_emitters) * fix(desktop): assert the pinned PostHog host at runtime, not in source text The host-pin regression scraped analytics.ts and asserted the file never contains "VITE_POSTHOG_HOST" — which the comment explaining why the override was dropped also matches, so the check failed on its own explanation. It was a static tripwire either way; it never proved the override was inert. It now stubs VITE_POSTHOG_HOST to an origin the renderer CSP does not allow, re-imports the module, and asserts fetch still goes to us.i.posthog.com. Verified in desktop/windows: - npx vitest run src/renderer/src/lib/analytics.test.ts -> 6 passed - restoring the `import.meta.env.VITE_POSTHOG_HOST ||` fallback in analytics.ts fails it with Received "https://not-in-csp.example.com/i/v0/e/", so the test fails for the reason it claims - npx prettier --check on the file -> clean --------- Co-authored-by: Max Carter 祁明思 <136312656+undivisible@users.noreply.github.com> | 1 个月前 | |
Document memory architecture and enforce package maps | 1 个月前 | |
Make automatic development backend deploys converge (#12019) * Make automatic development backend deploys converge Automatic development deploys succeeded 8 times in the 25 runs before this change. The 13 failures had three causes, and this addresses the two that are defects rather than configuration. Admission required the Release Eligibility proof SHA to still equal main's tip. Anything merging while eligibility ran therefore rejected a merged, reviewed commit -- 8 of the 13 failures, and why development sat a day behind main. The property that protects the runtime is that the commit is merged, so require ancestry instead. Production is untouched: it deploys only by explicit dispatch, which already required ancestor-of-main rather than tip-equality. Development also had no automatic Firestore migration path. An automatic index reconciliation resolves its environment to prod, and only a manual dispatch ever targeted development, so a merged manifest addition left development with a schema that no longer matched main and every deploy failed its readiness gate until somebody noticed -- 2 more failures, most recently the hourly_usage (year, month) index from #11979. Composite reconciliation is create-only and development carries no required reviewer, so it now converges on the same merge that queues the production migration. Production's approval gate is unchanged. The manual development lane failed separately, at custom_token_signing: its candidate audio gate authenticates against production Firebase, which a development deploy identity cannot sign a custom token for. The probe can now be told which account to sign as. Left unset the behaviour is identical, so this is inert until FIREBASE_PROBE_SIGNER_SERVICE_ACCOUNT is set and the deploy identity is granted token-creator on it. Not addressed here: 3 failures came from GCP_FIRESTORE_READONLY_CREDENTIALS being intermittently unavailable in the development environment. That credential is a deliberate privilege boundary -- readiness executes admitted source and must not hold deploy credentials -- so it wants a configuration fix, not a code change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Update release-vector contract for the development index lane The static migration contract counted --provision-missing across the whole workflow, which asserted 'only one lane applies indexes'. There are now two, one per environment, so count per job instead and pin the development lane to its own environment, concurrency group, and push-only trigger. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Resolve the newest proven main source instead of the triggering one gpt-5.6-sol's review found the previous approach incomplete in two ways, and both are real. Ancestry alone was not safe. Tip-equality was doing more than proving merge status -- it was also a currentness fence. Accepting any ancestor of main lets a late-scheduled run deploy older code than development already had, because Actions concurrency groups are not FIFO, and lets a run for a commit that has since been reverted redeploy the pre-revert tree. This runtime shares production Firestore, Firebase auth, and Stripe, so that is not benign. Ancestry alone was also not sufficient. The scope job green-no-ops any triggering SHA that main has moved past, before it ever inspects changed paths. So a backend commit still never deploys if an unrelated commit merges before scope runs: the backend commit no-ops for being behind, the unrelated commit no-ops on its own diff. The regression test claimed to cover this but built a later main SHA and never passed it to scope, so scope saw the backend commit as main's tip and the assertion proved nothing. Passing it reproduces the strand. Both follow from deploying the triggering commit. Admission now resolves the newest commit on main carrying a first-attempt successful Release Eligibility proof and reachable from current main, and deploys that. Concurrent runs converge on one target rather than racing, a revert is never undone by a late run for the commit it reverted, and a behind trigger still deploys because the target moves forward instead of the run being skipped. Scope's supersession decision is removed as now-redundant, which also deletes its two GitHub API proofs and their fixture -- the contract gains tripwires against reintroducing it. --trigger-is-ancestor-of-sha keeps the resolved target at least as new as the proof that triggered the run, so a stale listing cannot move development backwards from its own trigger. The proof listing is fetched with curl --fail and no error suppression: an unreadable listing refuses to deploy. The guard checkout assertion is gone rather than re-checked-out. sol was right that it had become true by construction and added no independent evidence, and the re-checkout it needed also made an in-flight run execute a newer guard script than the workflow that invoked it. Readiness now needs actions:read to list proofs. The manual lane's readiness job already had exactly that for exactly this lookup, so the contract now expects it for both rather than treating the automatic lane as more restricted. sol's P0 -- that the new development index lane writes to production -- does not hold: RUNTIME_GCP_PROJECT_ID is based-hardware-dev in the development environment and based-hardware in prod, so the two jobs target different projects and cannot race on the same index. It read the value from runtime_env.yaml's runtime_gcp_project rather than the deployed variable. That inconsistency between the checked-in contract and the deployed value is real and worth its own look, but it is not this lane writing to production. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> | 15 天前 | |
docs: take operator pages off docs.omi.me Unlisted Mintlify MDX is still a public URL. Move runbooks, flags, invariants, and agent rules next to owning code, add docs/AGENTS.md as the site allow-list, and correct the live kill-switch contract after the JIT authority page leaves the site. Co-authored-by: Cursor <cursoragent@cursor.com> | 7 天前 | |
One chat shell: every content block renders everywhere, plus chat ergonomics (#12607) * refactor(desktop): mount one chat shell for every account DesktopHomeView held the app on a "Preparing Omi…" card until a network call decided which of two shells to mount, then rendered either ChatFirstShell or a legacy sidebar + DashboardPage tree. Both were the same product with different chrome, and the legacy branch was the only reason `useLegacyHomeDesign`, `useOldestHomeDesign`, DashboardPage's inline chat, SidebarView, and the widget hub still existed. The shell now mounts immediately for everyone. The server-owned capability still resolves — same request, same analytics event, same ChatProvider projection gate — but alongside the mounted shell rather than in front of it, and it now only decides whether the capability-gated kernel features engage. Capability-off renders the same shell. `navigate help` named a "Help from Founder" page no shell had mounted for a long time: the bridge resolved a title and then timed out. It now resolves to Settings → About, where getting help from a person actually lives. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop): render every content block as an interactable component Six of the journal's block kinds — question card, task card, goal link, capture link, conversation link, memory link — were dropped on the floor by every Chat surface except one. `ContentBlockGroup.group` skipped them unless `richBlockRenderingEnabled`, and `ChatBubble.blockView` returned `EmptyView` for each of them again. A turn whose whole content was a task card therefore read as an empty assistant reply in the task panel and in the notch, and as a card you could tick off in the main window. `ChatFirstRichBlockContext` is now non-optional on `QueryShellHome`, `QueryAnswerThread`, `ChatMessagesView` and `ChatBubble`; the task panel and the floating/notch renderers bind the shell's process-wide owners through `.auxiliary`, so a card tapped in the notch summons the main window and routes the one shell. `ChatFirstRichBlockGroupView` is the single renderer all three hosts share. Capability-off degrades rather than disappears: cards render, task check-off works (it binds `TasksStore`, not the projection), links navigate, and a question card shows its options dimmed and unpressable with an explicit "Answering is unavailable right now" line, instead of a question with no visible answers. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): every proactive card says what it is `showNotification` took an optional `kind:` and `FloatingBarNotification` quietly filled it in from `assistantId`, whose default arm is `.general`. Five producers never passed one — trial messaging, onboarding permission help, both notch moments, and the whole generic proactive path — so their cards journaled a bare `notification:<uuid>` continuity key and came back in the transcript badged "Notification" with a bell, a row that says nothing about what Omi actually noticed. `kind:` is now required and never derived inside the value type. The generic proactive path derives it once at the producer edge, from the same `from(assistantId:)` call the category gate already makes three lines earlier. `.trial` and `.onboarding` are new kinds and are excluded from journaling alongside the integration nudge: billing copy and permission help are not observations. `.functional` carries the system notices that used to ride on `.general`. `.general` survives as decode-only so historical bare keys keep reading back, and its badge arm stays for exactly those rows. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(desktop): pin one shell and six live blocks, and retire the second-shell flows The tests that protected the old shape are the reason it would come back: `ChatFirstRichBlockTests` asserted that a caller without an explicit context got *nothing*, and three flows waited on `shellVariant: legacy`. - `OneChatShellRichBlockTests` builds one turn carrying prose plus all six interactable kinds and asserts the grouping keeps every one of them, in transcript order, on the same entry point the notch and task panel call. It then drives each link's typed navigation target through the real navigation owner, and pins capability-off to "options dimmed", never "options gone" — via a new `ChatFirstQuestionCardOptionsPolicy` that separates *answered* and *retired* (hide) from *capability-off* (disable). - `ProactiveNotificationKindTests` walks every assistant id a producer ships and proves none of them derives `.general`, so no producer can mint a bare `notification:<uuid>` key; historical bare keys still decode. - `check-single-chat-shell.py` (+ manifest entries, with a self-test) is the tripwire for the vocabulary that made a second shell expressible. - `home-stage.yaml` and `dashboard.yaml` described `DashboardPage` and are deleted with it; `chat-first-capability-isolation.yaml` is repurposed to the assertion that now matters — capability-off mounts the same shell. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): the notch and the main window share one navigation owner `ChatFirstRichBlockContext.auxiliary` binds `ChatFirstShellNavigation.shared` so a card tapped in the notch or the task panel routes the shell. The root was still creating its own instance, so those taps would have moved a navigation object nothing rendered — the card would appear to do nothing. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(desktop): delete the intelligence store nothing renders any more `DashboardIntelligenceStore` fetched recommendations, projected them, and kept a feedback outbox for a section that only `DashboardPage` mounted. With the page gone its 750 lines had no renderer and no caller — the grep is exact: the class was referenced only by its own file and its own tests. Its 1,117-line test file went with it, because a test for a store nothing mounts is coverage of nothing. What stays is the part other surfaces still use: `TaskNavigationRequestStore`, the exact-record handoff `QueryShellHome` and the chat-first task card give the Tasks page instead of a tab index. The file is named for it now. Three source-reading tests still pointed at `DashboardPage.swift` and failed on the missing file rather than on anything real; they move onto `QueryAnswerThread` / `QueryShellHome`, which is where Home's error card and its colour tokens actually live. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): an explicit settings section still wins over the help default `navigate help` pre-selects About because that is where getting help from a person lives. A caller that also names a section meant that section. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): let a settled chat answer be selected, copied and seen as cut off Three things a reader could not do with a reply on screen. Select it. `OmiMarkdown` disabled native text selection outright, so a date or a name in an answer could only be retyped. The reason was real — one AppKit selection overlay per `Text`, on a row that rewrites its body every streaming flush, is a non-converging layout loop (400 segments, 2 s hangs) — but it is a reason about *streaming* rows. Selection is now opt-in through `\.chatTextSelectable`, and `ChatTextSelectionPolicy` grants it to settled rows only, on every surface that shows chat prose: the transcript, the notch, the expanded floating bar, and onboarding. The three `.textSelection(.enabled)` calls outside `OmiMarkdown` in the floating surfaces were dead — the inner `.disabled` won — and are replaced rather than left as decoration. Copy it without hunting. The copy button lived only in the hover-revealed strip, and a user turn had no copy affordance at all. Every row now has a "Copy Message" context menu over the same pasteboard write, and the copy button takes ⌘C while its row's strip holds keyboard focus — not window-wide, which would take the shortcut from selected prose and the composer. See that it stopped mid-sentence. A voice barge-in persists the partial answer with a terminal failed status, and the only failure affordance was a stamp for a row with no text — so "…arrive on Saturday," rendered exactly like a finished reply. `ChatTurnFailurePresentation` decides between that stamp and a quiet trailing "Interrupted" mark, and the empty-row case is unchanged. Also: the hover strip is now `accessibilityHidden` when it is invisible (opacity and hit-testing hid it from the eye and the mouse but not VoiceOver), each row carries a You/Omi label, and a row whose whole content is a rich block reserves no metadata band — a memory card stamps its own time and has nothing to copy or rate. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): stop charging the transcript twice for the metadata band Two consecutive one-line answers sat roughly 100 device pixels apart, and a memory card floated in symmetric dead space. Measured on the real views: 44 pt between two settled replies, and a 48 pt card row inflated to 69 pt. Two causes, both double-charges. The hover strip is 28 pt of real reserved height under every settled reply. The stack then added a full 16 pt inter-exchange gap on top of it, so the separation the band already provides was paid for twice. `ChatTranscriptLayout.spacing` now asks whether the row above reserves a band and takes a hairline when it does; every other rung of the ladder is unchanged, and a reply still binds to its question more tightly than to the next exchange. `ChatOmiMarkPlacement.rowHeight` reserved 32 pt on every assistant row for a mark that only needs it when an empty streaming reply has no height of its own. On a settled row the reservation did nothing but centre short content in a box taller than itself — which is what put equal dead space above and below the memory card. It now applies while streaming, top-aligned. Measured after: 32 pt between two replies, 48 pt for the card row, and a five-row transcript 296 pt tall instead of 341. Also collapses adjacent repeats. Dedup only ran on messages over 200 characters, so three push-to-talk tries at the same ~90-character question stuttered down the transcript untouched — each press mints a distinct `voice:<uuid>` turn, so those are three legitimate journal rows and journal identity is not the place to fix it. `adjacentDuplicateIDs` collapses a short answer repeated in the row immediately below it within ten minutes, and folds a failed barge-in fragment into the answer it is a strict prefix of. It stays behind the existing expandable "Duplicate message" chip, so nothing is hidden outright, and non-adjacent, distant, or cross-sender repeats are left alone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(app): decode chat content blocks into typed models Mobile stored `content_blocks` as raw maps and, since #12015, hid any message whose blocks were only desktop chat-first chrome (goal/task/ question) because there was no renderer for them. Both halves are now wrong: the components are coming, so the schema needs a typed projection and the hide filter has to go. Add `ChatContentBlock`, a sealed model mirroring the canonical schema in `desktop/macos/agent/src/runtime/types.ts` and the Swift codec's required-field rules, decoding both the camelCase (desktop/agent) and snake_case (chat-first spec) dialects. Malformed blocks are dropped; unknown types become `UnknownContentBlock` so the message keeps its synthesized fallback text instead of losing content. The raw list stays authoritative on the wire — `toJson` is unchanged. Delete `hideFromMobileChat` / `visibleOnMobile` and their four call sites in MessageProvider, replacing the "is this body only the fallback dump?" test with `textIsStructuredFallback`, which the renderer uses to decide whether components replace the body or sit beside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(app): add l10n keys for chat content-block components Ten new keys for the block eyebrows, destination actions, unavailable state, and the conversation link's recommended-steps header, translated into all 48 non-template locales. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(app): render chat content blocks as interactable components Every block type that macOS renders as a control now has a mobile component, driven from the same wire schema (#12598): - taskCard: a live checkbox wired to the single tasks mutation path, ActionItemsProvider.updateActionItemState. The tasks API is list-only, so the card resolves against the loaded list and mirrors the macOS loading / unavailable states rather than inventing a fetch-by-id. - goalLink: mobile has no goal detail route, so the card opens a bottom sheet with the goal's title and progress resolved from GoalsProvider, and shows the unavailable state when the id is not in the list. - captureLink / conversationLink: push ConversationDetailPage through the citation preamble already shipped in chat (grouped-map hit, then fetch by id). conversationLink also lists its recommended action items as plain rows; mobile creates tasks from the tasks surface, so the block mutates nothing. - memoryLink: opens the existing memory sheet for the resolved memory. - questionCard: options send their preparedAnswer down the normal chat send path, so the runtime stays authoritative for what an answer means. A deferral option is not special — it sends its own prepared answer. Once selectedOptionId is set only the chosen option remains, disabled, so no stale chip ever looks tappable. text/thinking/toolCall/discoveryCard/citation/agentSpawn/agentCompletion and unknown types render nothing extra — the body (or its synthesized fallback) already carries them — but they never hide the message. Where the body is only that fallback, the components replace it instead of repeating it. Every interactive element carries Key('chat-block-<type>-<id>...'). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * docs: record the presentation-cohort-drops-journaled-content failure class Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): give the reader a selectable copy instead of a selectable transcript The selection change this branch shipped is reverted. `OmiMarkdown` disables native text selection again at both sites, explicitly, and the `\.chatTextSelectable` environment key, `ChatTextSelectionPolicy` and `OmiChatTextSelectability` are gone along with the host wiring in `AIResponseView`, `FloatingControlBarView` and `OnboardingChatView` — those hosts now carry no `textSelection` modifier around `OmiMarkdown` at all, since the ones that were there before this branch were dead code under the inner `.disabled`. The settled-row gate was not enough. PR #10834 made the same argument and reopened FC-selection-overlay-layout-loop in Omi Beta 0.12.146: every sampled main-thread stack sat in `SelectionOverlay`, `setFont`, intrinsic-size invalidation and AttributeGraph, and memory grew without bound. A settled row is still rebuilt by transcript loading, scrolling, window resize and parent-state updates, which is all that loop needs. `.github/scripts/check_chat_selection_boundary.py` rejects the escape hatch and names the remedy: the existing copy actions, or a separate non-live reading surface. So this adds the reading surface. "Select Text…" sits on the row's context menu next to "Copy Message" and as an ibeam button in the hover strip, on assistant and user rows alike. It opens `ChatSelectableTextPopover`: one `NSTextView` over one message's copyable text — `isEditable` false, `isSelectable` true, ⌘A and ⌘C native, Escape closes, sized to content with a 360 pt cap and internal scrolling. It is outside the transcript's layout and does not mount until the reader asks for it, so it cannot take part in the loading, scrolling and resize passes that made in-place selection unsafe. No SwiftUI `textSelection` anywhere in it — AppKit selection is what an `NSTextView` already is. The strip now carries four controls plus the timestamp (thumbs, thumbs, copy, select, info-when-present). At 24 pt each that still leaves the timestamp its own room, so both affordances stay rather than context-menu only. Everything else on this branch is unchanged: right-click Copy, focus-gated ⌘C, the metadata-band policy, the transcript rhythm, the interrupted-turn marker, adjacent duplicate collapse, and the accessibility fixes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(desktop): drop the deleted shell files from the static guards Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(desktop): repoint ptt-lifecycle covers at ChatToolExecutor RealtimeConversationToolProjection.swift was folded into ChatToolExecutor in 758cd3f1fb and the flow kept the stale path, which fails desktop-flow-lint. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): a reply the reader watched arrive stays whole when it settles `ChatBubbleTruncation` clamps any body over 500 characters, and only once `isStreaming` goes false. So an answer rendered in full while it streamed — with the transcript following it down — collapsed to its own first paragraph the instant it finished. A forty-item list became three items and a "Show more", and the document shrank by thousands of points under a reader pinned to the live edge. Watching that happen reads as "the chat stopped scrolling": everything you just followed is taken back at the end. Truncation is for restored history, where a long transcript should not be mostly one old reply. An answer that just settled here is the opposite case, so it keeps its full body. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(desktop): the assistant reads the same task list the Tasks page shows Asking by voice what was on the list answered "you don't have any tasks overdue or due today" to someone looking at thirty of them. `get_tasks` is backed by `TasksStore.loadDashboardTasks`, which narrowed the list twice in ways the Tasks page never has: - a seven-day recency window, as a lower bound on overdue due dates and as a `createdAfter` cutoff on undated rows. The page buckets on `dueAt < startOfTomorrow` alone and ages nothing out, so a backlog older than a week was invisible to the assistant and only to the assistant. - a source filter that dropped every AI-capture row. That gate was written when a capture could still land in `action_items` unreviewed. INV-TASK-2 has since made capture suggestion-only — `TaskCaptureModePolicy.usesLegacyStaging` is false for every mode, and a capture stays a Candidate until an explicit gesture accepts it — so the gate no longer separated reviewed from unreviewed. It hid the user's own backlog, and anything they created by voice, since `create_action_item` comes back stamped `conversation`. On the reporting account those two took thirty visible tasks to zero. The buckets now carry what the page carries: 82 overdue + 3 due today against the page's "Today: 85". The per-bucket cap moves 50 → 500, because the count is spoken and 50 would have understated it. `isPendingSuggestion` stays for proactive nudges, the one consumer still asking whether a capture pipeline wrote a row. `DashboardTaskLanePolicy` had no other caller and goes. The voice tool descriptions said "overdue + due today", which is where the spoken wording came from; they now describe the whole open list. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(desktop): the transcript's own follow-scroll is not the reader taking it Found while tracking down the streaming-scroll report. `UserScrollDetector` promotes an open mouse press to reader ownership as soon as the clip view moves, because a scrollbar-track click repositions the viewport without ever emitting a drag. That test cannot tell the app's own follow-scroll from the reader's, and while an answer streams the transcript re-reaches the live edge every `ChatScrollFollowThrottle.interval` — so a press still open when one lands ends follow mode for the rest of the answer. Now that every content block is something you can click, a press inside a streaming transcript is ordinary. Two lifecycle holes alongside it: a release delivered to another window — the "Select Text…" popover and context menus present in their own — was dropped by the same-window guard, leaving the press candidate open for the life of the scroll view with every later follow-scroll able to promote it; and a second press registered its bounds observer without removing the first. The transcript now records when it moves its own viewport, and movement inside that window re-baselines the press instead of promoting it. A drag and a scrollbar-track click still take ownership, which the existing harness cases pin. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(desktop): the tasks flow checks the lanes the assistant answers from The Tasks page and `tasks_snapshot` read the same rows now, so the flow can say so: S6 compares the page's Today count against overdue_count + today_count. That is the comparison the bug failed — thirty tasks on the page, zero in the lanes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(desktop): the components the agent renders reach the transcript `render_chat_blocks` had never once been offered to the model on this Mac, and on the turns where it did run its cards were deleted three seconds later. Four separate gates, one confusion between them: *which surface is this*. One shell means main Chat and the floating bar project the same conversation, so a session the bar registered carries `surface_kind = floating_chat` while main-Chat runs execute on it. Every chat-first gate admits `main_chat` only, and three of them read the session's registration instead of the run's: the adapter metadata that sets `OMI_CHAT_FIRST_UI`, the tool-capability broker that admits the call, and the surface Swift re-validates before executing. The fourth was separate — run admission built its own context snapshot and dropped the capability on the way, because the capability map lived on `KernelSessions` where run admission could not reach it. Any one of the four was enough on its own; advertised tools go 41 -> 46 with all four fixed. Then the cards died anyway. `monotonicAcceptContentBlocks` protected exactly two block kinds across terminalization, and terminalization applies the projection *Swift* assembled from the adapter stream — text and tool calls, never a block the agent appended mid-turn, because that append is a journal mutation the surface never saw. So the replace erased every task card, goal link and memory link the turn had rendered, moments after the tool reported `ok`. The protected set is now the kinds the kernel authors and the surface never does. Confirmed live: the tool executes, and three `taskCard` blocks now survive on the finished turn. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(desktop): ticking a task card shows it done, not gone Checking a box replaced the card with "Task is no longer available" — the one message that means the row is not the user's any more. Nothing had gone away. Ticking moves the task between the store's two arrays and the toggle awaits SQLite first, so `liveTask` is briefly nil. That flips the card's `hydrationKey`, SwiftUI cancels the in-flight hydration, and `TasksStore.isCurrent` folds `!Task.isCancelled` into its lease check — so `resolveCanonicalTask` then returns nil *by construction*, not because the task is missing. The old code published that nil, cleared the retained task and set `hydrationFinished`: exactly the pair that draws the unavailable placeholder. A superseded hydration now speaks for nothing, and a store that cannot vouch for a row is no longer read as the row being retired. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(chat): components replace the writing instead of doubling it Two ways the same answer was being said twice. The tool's own instructions were the first. They told the model to render components "whenever you retrieve, create, or summarize" those entities and "do not leave them as a Markdown table/list" — so a summary of yesterday stacked three conversation cards above the prose that already said it. That is backwards: reading an entity to answer a question makes it a *source*, and sources belong in citations. A component is for the entity that IS the answer — the one the user asked to see or act on, or the one this turn created or changed — and when components are rendered they are the list, so the prose above them is one lead-in sentence at most. The second was ours. A turn that carried only components still printed the blocks' own degradation text underneath them, so three cards sat above the three lines they were made from. `ChatStructuredFallbackText` mirrors the producer case for case, and a body that is only that projection is recognised as the cards talking to themselves rather than as answer text. Mobile drew six of the nine kinds the desktop transcript draws and silently skipped the rest. Discovery cards and the two agent-run blocks have components now, and a parity test fails if the desktop grows a tenth without one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(desktop): selection lives in the words, not in a box beside them "Select Text…" opened a popover that re-printed the message in raw Markdown next to the row the reader was already looking at, and only assistant rows had it — a user turn has no hover strip, so their own words were never selectable by any means. SwiftUI's selection stays barred: PR #10834 put it back on settled rows and reopened FC-selection-overlay-layout-loop in Omi Beta 0.12.146, every sampled main-thread stack in `SelectionOverlay` and `setFont` while memory climbed. But that bar is on `SelectionOverlay`, not on selecting. An `NSTextView` *is* one selection: one view owns it, a rebuild replaces a string, and nothing per-`Text` is mounted for a parent to thrash. `ChatSelectableTextPopover` said so in its own header — it just kept that view outside the transcript. So the words are the surface now. Chat prose renders through one text view per block, parsed from exactly what the SwiftUI renderer parses, and the reader drags across an answer in place with ⌘C copying what they highlighted. Citation markers become link ranges rather than buttons, which is what made the surrounding line selectable at all — a chip in a flow layout forced the prose to be chopped into per-segment views — and they still open their source and still preview it on hover. Two things had to be got right. Height is measured beside the live view rather than inside it, because a container left holding a measurement width wrapped the answer to a width the transcript never granted; and the measurement is synchronous, because publishing a height a frame late broke the transcript's follow-scroll. The gesture harness reads pixels, and `cacheDisplay` stopped seeing prose once it was drawn from a layer, so the probe composites the layer tree and can measure the row again. The boundary check now guards this file too: the AppKit path may never quietly acquire the SwiftUI one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(chat): a rendered card is the citation, and it survives the next update Two leaks left over from making components reachable. The first: terminalization was not the only replace. The streaming projection pushes the surface's own block list several times a turn, and `updateJournalTurn` took the replacement literally — so cards that survived the terminal commit died to the very next update instead, which is why they appeared on one turn and not the next. Both paths now apply a projection over the journal rather than in place of it: an id the projection carries is the projection's to define, which is how a question card's options still get retired, and an id it omits survives only when the kernel wrote it. The second is what the reader saw. When the model renders components and skips inline markers, we append a compact `Sources: [1][2][3]` rail so provenance stays discoverable. That was a sensible garnish when components were rare; now that a component turn is one lead-in line and three cards, the rail is a row of bare markers printed under the very things they point at. Sources the turn already draws are dropped from it, and anything with no component of its own keeps its marker. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(desktop): every citation marker the transcript writes is selectable and live The AppKit prose path matched `[\d{1,3}]` of its own invention. Ordinals run to four digits and the model also writes the kind beside them — `[5004]`, `[memory 5023]` — so real markers rendered as dead text in the middle of an answer. It uses `ChatCitationMarkup.numericMarkerPattern` now, which is the transcript's own definition of a marker rather than a second copy of it. Also pins what the harness proved by hand: the mounted transcript hosts a selectable text view on rows at more than one inset, so it is both senders' words that can be dragged across, not just Omi's. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(desktop): the cohesive chat flow covers the selection renderer that replaced the popover Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(ci): the selection boundary fixtures cover the AppKit surface too The remedy is an NSTextView owning its own selection; a SwiftUI rewrite of that file would put SelectionOverlay back in the transcript under a name no pattern check can see, so the fixtures pin that rule alongside the existing ones. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(agent): the component cap is not a licence to write the list out instead Asking to see your tasks and getting seventeen of them numbered in prose is the shape components exist to replace; the cap bounds how many are drawn, not whether any are. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(agent): tasks the user asked to see are rendered, not written out The render-vs-cite rule lived only on `render_chat_blocks`, where the model reads it after it has already decided what to write. Asking "what are my open tasks right now?" still returned seventeen of a hundred as a numbered list — prose that names the tasks but cannot tick one off, which is the shape components exist to replace. The rule now sits on `get_action_items` too, at the moment the tasks arrive: if the user asked to see, review, pick from or work through them, render the few that matter and say how many more there are; if a task is only evidence for a question answered in prose, cite it and render nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(agent): a rendered task list is counted, not named The `get_action_items` guidance said to render the cards and "say how many more there are". The model read *say* as *list*: it rendered three task cards and bulleted the same three titles above them, printing every task twice — once as words that cannot be ticked off and once as the card. Say the count, never the names, and say why: naming them is what doubles the answer. Also corrects the OmiMarkdown header comment, which still described the deleted selection popover. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(desktop): stop the Removed lane tombstoning live tasks Ticking a task on a chat task card replaced it with "Task is no longer available" — the reader's own completion, erased. The tombstone was local and it was fabricated. `fetchDeletedPage` asks for retired rows with a `deleted=true` query item that `GET /v1/action-items` has never had: FastAPI drops the unknown item, and that handler's stream skips soft-deleted documents outright, so the "deleted lane" answered with the owner's live first page. Every row came back stamped `.retired()` and was synced into SQLite, so each visit to Removed tombstoned another hundred live tasks. Completing one read the tombstone back — 190 ms after the local write, before the backend was called at all. This Mac was carrying 100 of them, every one with no `deletedBy` and a canonical status still `active`. Three changes, because each is load-bearing on its own: - The lane keeps the rows the response itself reports retired and drops the rest. Removed showing fewer rows is a gap; manufacturing retirement is data loss. - A migration clears the tombstones already written, and only those: a deletion the owner performed records `deletedBy`, a server-side retirement arrives as canonical status, and a row with neither witness was retired by nothing but the stamp. - The card treats the reader's tick as authoritative. A retirement discovered afterwards does not get to undo a completion the app accepted, whatever put the retirement there. Until the backend grows a real `deleted` filter, Removed can only show deletions made on this Mac. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(desktop): clearing Chat takes the daily summary with it The summary card is chrome above the thread rather than a turn — INV-CHAT-1 keeps transcript authorship in the kernel, so nothing on the Swift side may write a journal row for it. The consequence was that Clear could not reach it: the day's summary sat alone in a chat the reader had just emptied, which reads as a clear that did not work. Clearing now records which summary was on screen and withdraws the card on the same frame. Recording the id rather than a flag is what lets tomorrow's summary come back on its own, and the key is owner-scoped like the announcement's, so clearing on one account cannot blank another reader's day on a shared Mac. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(desktop): stop one test post killing the whole test host `swift test` has been exiting 1 on this machine, before this branch and after it. Not a failing test — the runner died with SIGTRAP partway through, taking every suite alphabetically after it with no message and no crash report. The two suites involved each pass alone. `JITProactivityDeliveryTests` posts `.runtimeOwnerDidChange` from an async test body, so it lands on a cooperative thread. `NotificationCenter` delivers synchronously on whichever thread posted, and the observers are `@MainActor` types whose sink closures carry an isolation check on entry. Once `InterjectWiringTests` had mounted `FloatingControlBarManager`, that check failed `dispatch_assert_queue` and trapped the process. Nothing inside the closure could have guarded it: the check runs before the first statement. The fix is the post, not the observers. `performEffectiveOwnerTransition` — the only production poster — already posts inside `await MainActor.run`, so no shipping code does what the test did. And the observers' synchronous delivery is load-bearing, not incidental: it is what lets a surface fence itself *during* an owner transition rather than a runloop later (see `IntegrationNudgeCoordinator`). Hopping delivery downstream instead was tried first and was wrong twice over — it broke `FloatingOwnerProjectionTests`, which correctly caught the previous account's pills surviving the switch. So: the test posts on the main thread as production does, and the contract is written down where the notification is declared. The suite now runs to completion for the first time: 6695 tests, no crash. Two `JITProactivityRuntimeTests` failures remain, pre-existing and unrelated — they assert an outcome that only holds when the local mirror database is absent, which is true only when they run alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(app): restore the ephemeral iOS package Flutter regenerated Running `flutter gen-l10n` and `flutter test` on this Mac rewrote `app/ios/Flutter/ephemeral/Packages/FlutterGeneratedPluginSwiftPackage/Package.swift`, emptying its dependency list because the iOS packages were never resolved here. That is build residue from the merge, not a change anyone made. Restored to origin/main's content. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * style(app): dart format the files this branch touches `message.dart` picked up its formatting drift from the merge resolution; the generated wire file is regenerated output. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Revert "style(app): dart format the generated wire file" `subscription_usage_wire.g.dart` is generated by `backend/scripts/generate_dart_models.py`, so reformatting it makes it stale against its own generator. Restored to origin/main's content; the `message.dart` formatting in the previous commit stands. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(app): regenerate subscription_usage wire models The committed file was stale against `backend/scripts/generate_dart_models.py --group subscription_usage` — the preflight contract check rejects it on origin/main's copy too, so this is drift the merge surfaced rather than anything this branch changed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(desktop): a follow-up's citation borrows the ordinal it points back at Ordinals are assigned per attempt, so a turn that retrieved nothing has no ordinals of its own — and when the model answers "pick one conversation from that day" without a tool call, the `[1]` it writes is the first conversation of the previous answer. Left unbound, it drew as plain text beside a title the reader could not open. Only ordinals the turn cannot resolve itself are borrowed, each from the nearest earlier assistant turn that persisted it, so a turn's own provenance always outranks the past. The binding runs at finalize and again over every journal projection, so restored history opens the same source the reader could open live. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop): fold a reply after two screens of text, not five lines Five hundred characters was five lines, so every real answer collapsed and the reader clicked "Show more" under nearly everything they asked. The budget is now measured in viewports of rendered text: a reply may fill two screens before the transcript offers to fold it, and one that long starts folded. Lines are estimated per source line at the column's width, so a bulleted list is measured as the lines it takes, not the characters it holds, and the cut lands on a line boundary with any open fence closed. The transcript publishes its viewport to the rows through the environment, republished only at a coarse step so a resize drag does not re-evaluate every row per frame. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop): stream as a flow, and glide the live edge instead of snapping The wire delivers an answer in bursts — a provider chunk, a paragraph the moment a tool returns — and a flush that dumped everything it had made the transcript lurch by a sentence and then sit still. Each 35 ms flush now reveals a paced slice: a small backlog drains over about five flushes, a large one is let through fast enough that the reader is never more than two lines behind, and the tail of every burst tapers. Boundary flushes (a tool starting, the turn settling) still land everything at once, so block order is unchanged. The follow scroll during a stream eases to the live edge rather than jumping a row at a time; restores and sends stay instant, and Reduce Motion disables the glide. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(agent): native components are the default answer for what Omi draws "Pick one conversation from that day" came back as a bold title with a citation number where a conversation card belonged. The render policy now says it plainly: default to a component whenever the user asks for a task, goal, memory, conversation or capture, and reserve prose for a request to read rather than open or act — a summary, a recap, an analysis, a comparison, a count, or a list too long to render. The conversation and memory retrieval tools carry the same rule with their own block shapes, including follow-ups that narrow an earlier result. Regenerated tool surfaces and fixture. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(desktop): the hermetic chat flow covers the streaming buffer The paced reveal lives in ChatStreamingBuffer.swift, which no e2e flow declared; the hermetic chat flow streams an answer through it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(desktop): the offline readiness fixture asks for a port nothing holds The harness verifies a launched bundle's /health whenever --port is bound. On a Mac with any dev bundle up on the default 47777 the offline-only readiness cases reached that check and sourced app-config.sh, which the fixture never provides, so the pre-push launcher gate failed on a test that never left the box. The fixture now probes for a closed port and passes it explicitly. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): a conversation citation opens the conversation it names The agent cites whatever its conversation tools retrieved, and most of this user's library is desktop and phone recordings. But a citation of kind .conversation routed the capture focus, and the capture archive resolves that focus through a strictly source-scoped fetch (GET v1/conversations/{id}?source=omi). For a desktop recording that request is a 404, so the hub landed on the Conversations list with nothing opened — the report's "Open Conversation just brings me to the conversation page". Citations of omi-device captures kept working, which is why only some markers failed. Citation routing now fetches the unscoped record first and lets the record's own provenance pick the route: omi captures keep the capture focus so the transcript moment still plays; every other conversation opens as the exact fetched record through open(conversation:), which presents it even when the paginated list does not contain it. A failed or mismatched fetch navigates nowhere instead of stranding the reader. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): launch no longer prints the daily summary before the chat Opening the app admitted the daily summary card while the initial history was still loading, so launch read as a summary page that then yanked away the moment the transcript landed at the live edge. The card is chrome (INV-CHAT-2), and chrome must not outrun the thread: admission now defers for the whole initial load, the loading-complete observer admits it before the live-edge restore measures geometry, and an empty thread still meets its card. The decision lives in one pure, tested place — ChatDailySummaryAdmission. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * perf(desktop): returning to Chat no longer re-parses and re-measures every prose block Chat is torn down and remounted on every route change (the shell keys the destination by route), and each mount rebuilt the whole visible transcript window from scratch: one Foundation Markdown parse plus one NSAttributedString build per prose block, and one throwaway TextKit stack per sizeThatFits query. ChatProseRenderCache memoizes both, keyed on the exact render inputs (markdown, style, fontSize, font scale in thousandths, sorted citation ordinals), so a hit is identical to a recompute. Bounded LRU (192 entries) absorbs a streaming row minting a key per 35 ms flush; per-entry measured heights are keyed on the proposed width and dropped wholesale at eight widths, which is what a window resize wants anyway. Measured by instrumenting ChatSelectableProse.attributedString and ChatSelectableProseText.height on the mounted transcript (ChatTranscriptGestureHarnessTests.Harness, 120-message journal, compact 50-row window, debug build), cache disabled vs enabled in the same binary: - before: every mount, including every return to Chat, paid 50 parses (4-8 ms) and 600 height measures (54-62 ms of TextKit layout) - after: the cold mount pays 50 parses and 50 measures once (4-17 ms, 4-14 ms); every remount pays 0 parses and 0 measures New ChatProseRenderCacheTests pin hit identity, key sensitivity to every render input, nil-produce never being cached, the LRU bound, and the per-width height memo and its drop bound. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): the follow glide runs on the run loop, so it lands in the app and in the tests The streaming follow rode a SwiftUI withAnimation transaction, which advances on the display cycle. In the mounted-transcript guard host that cycle never runs, so the follow silently did nothing there and the three live-edge drift guards failed at full stream growth: 462-484 pt of drift against a 120 pt limit, bisected to the commit that introduced the glide. A follow that only works where nothing can observe it is exactly what INV-CHAT-2's guard tests exist to catch. ChatFollowGlide moves the clip view on a 60 Hz run-loop timer with the same ease-out and the same 0.16 s duration (now one shared constant). The run loop is pumped by both the app and the harness, so the glide the reader feels is the glide the tests measure. A newer follow retargets the clock in flight, and reader input or teardown cancels it through cancelAllPendingScrolls; the snap path stays the fallback for pre-resolution frames and for Reduce Motion. Measured on the mounted transcript (60-message window, 40x35 ms stream flushes): worst drift 462-484 pt before, 48.0 pt after; all 20 ChatTranscriptGestureHarnessTests pass again. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): the streaming flush beat keeps its period when a flush overruns The paced flush re-armed from the moment the last flush finished, so the reveal's period was interval + render work, and jitter in that work shifted every later beat back. Measured with a flush workload of 0.75x the interval driving ChatStreamingBuffer directly: completion-anchored beats ran at 151-160 ms against an 80 ms interval; anchored beats hold 80 ms while the work still fits inside one. Beats now anchor to the beat before them, clamped to now so an overrun beat fires at once and resynchronizes instead of firing late. A drained backlog ends the chain, so the next delta starts a fresh beat rather than inheriting a deadline that already passed. Instrumented per-flush costs on the mounted transcript, for the record: the markdown re-parse grows 0.8 to 4.1 ms and the TextKit height measure 1.4 to 4.9 ms over a 0.25-10 KB answer (the O(n) per-flush terms; a stable-prefix incremental parse was assessed and declined - the parse is only about a quarter of the O(n) work, the rest is TextKit measure and the live view's own relayout, which incremental parsing does not touch). The storage isEqual compare is 0.01-0.08 ms and the transcript re-evaluates one bubble body per flush - both cleared as stutter candidates. The follow glide retargets from its mid-glide position with no step: worst sample-to-sample jump 11 pt over 210 samples of a 30-flush stream. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): the settle frame folds the tool trace away instead of teleporting When isStreaming flips, the row drops its tool-call groups and pre-tool commentary and swaps in the terminal answer, citations and metadata band in one layout change - measured on the mounted transcript with a three-tool trace, a 62 pt height change landing in a single frame, which reads as a jump. The fold is now animated: an .animation scoped on message.isStreaming wraps the row content, so that one frame eases out over the motion table's settle duration while every per-token streaming frame, where isStreaming did not change, stays exactly as instantaneous as before. Reduce Motion folds instantly. Final state is unchanged: the settle probe measures identical streaming and settled heights before and after (2620 / 2558), the group-drop assertions in ChatFirstRichBlockTests, OneChatShellRichBlockTests and ChatBubbleLayoutRegressionTests pass unchanged, and the mounted-transcript harness - which streams forty-flush answers through a row carrying this modifier - stays green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(desktop): the beat-period test declares its real-time wait The flush-beat test sleeps for real by design - the subject is the period the scheduler holds under genuine elapsed work - so it carries the wall-clock-wait escape annotation the desktop test-quality gate requires, keeping the counted baseline where it was. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(desktop): prune merge survivors of views this branch deletes Taking main's side of ViewExporter kept exports for the old Dashboard page and its score gauge, and main's automation registry picked up a Home-knows action family for the legacy hub - all three views and the feature behind them are deleted on this branch, so the entries and the action family go too. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(desktop): retire the knows-list flow and pin the merged test debt The knows-list rotation flow covers the HomeKnows feature this branch deletes with the legacy Home hub, and the merged suite's real source-inspection debt is 53 files / 143 sites - lower than the inherited ceiling. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): restore the memory-review fixture the knows-list prune removed The prune of the retired knows-list action family took `enum MemoryReviewFixture` down with it, but the surviving `seed_memory_review_fixture` / `memory_review_snapshot` / `memory_review_vote` actions and MemoryReviewCardTests still reference it — the branch did not compile. The enum moves to its own automation file, content unchanged from 98ddcffab9^. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): finish the merge-fallout repairs the tree still needed Two more of the same class as the fixture restore, on tracked files the merge left referencing deleted things: - `AnimatedGIFView.swift` went out with the retired onboarding wizard views, but main's own PermissionsPage and BrowserExtensionSetup still render it, and the merge kept those references. Restored verbatim from origin/main. - The ViewExporter prune meant to drop the full-dashboard export with the deleted Dashboard page but left a bare `AnyView` expression behind; the dead entry is now actually gone. The untracked Dashboard WIP files belong to a concurrent lane and are deliberately untouched. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop): decode the recap sections the dedicated page renders DailySummaryRecord gains unresolved_questions, decisions_made, and knowledge_nuggets, conversation_ids on topic highlights, and source_conversation_id on action items — the wire contract mobile's daily_summary.dart already mirrors and the backend has served since the sections shipped. Everything is decodeIfPresent, so an older backend just leaves the sections absent. Highlight and ActionItem get explicit inits with defaulted new fields so existing call sites are untouched. Tests pin the new keys verbatim and pin the older-backend case: absent sections decode to nil rather than dropping the summary. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop): the daily recap becomes a page, and Chat shows a slim pill Recaps are now clickable into a dedicated surface. DailyRecapPage presents the whole record — date and headline, overview, stat chips, highlights, tasks with their completed state, unresolved questions, decisions, the memories-learned review rows, and learnings — and every row carrying a conversation id deep-links through the hub-owned conversation detail. Rows carry no badges or buttons anywhere else: the pill in Chat and the pill on an Activity day show only title and summary, and clicking either opens this page as a sheet. A sheet, not a ChatFirstRoute: route values persist across launches and are gated to primary destinations, while a recap is a transient read over the surface that produced it. The Chat pill replaces the pinned full card. The old bar expanded in place over the transcript, and its measured inset tracked a high-water mark that could never shrink — one tall expand padded the thread forever, and a shrink left the bar lapping over the newest message. The pill never expands (the page is the expansion), so the inset now tracks the live height both ways, bounded to a couple of lines. The full-card view is retired; its pure section projections move to ChatDailySummaryPresentation next to the other recap rules. Review rows gained a daily_summary_detail source so votes from the page are counted as page votes, matching mobile's telemetry. INV-CHAT-2 admission is untouched. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop): the Activity day folds its recap into a doorway pill The day header was already the recap's toggle — thin when the day is folded, the recap inside when it is open. What it contained changes: the stored recap renders as a pill of title and summary only. The highlight chips, the Ask-about-this-day button, and the Regenerate button leave the list entirely — the full record, its badges, and its actions are the dedicated recap page the pill opens in a sheet, reusing the hub's own onOpenConversation so a recap's conversation links can never open a second detail owner beside it. Generate stays for a day with no recap: that is the way the recap comes to exist, not recap chrome. memory-review.yaml follows the rows to the page: the pill renders no rows, so the flow opens it before reading and voting, and the section now reports the daily_summary_detail source. chat-first-cohesive covers the new pill, page, and fixture files and drives pill-to-sheet. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(desktop): follow SpineDayRecapRow to its doorway signature The row no longer takes now/calendar — the follow-up question moved to the dedicated recap page with the rest of the recap actions — so the two host-view tests construct it with content and dateKey only. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): the restoring phase always resolves, even when its attempt is superseded Three dev launches hung in appState=launching / isSignedIn=false / isRestoringAuth=true for their whole session (10:31-11:05), then an identical launch came up clean. The hung logs end at the auth listener's "skipping REST validation while launch restore is in flight" with nothing after, and the phase never leaves .restoring. The deeper defect: the restoring phase's only guaranteed escape - the 5s watchdog in configure() - was gated on isSessionAttemptCurrent(attempt). Any newer session attempt defuses it, and every fenced exit in the restore flow (validateRestoredSessionNow's entry/post-refresh guards, refreshIdToken's post-network guard, the saveAuthState/commitRestoredSession commits) then returns silently, leaving .restoring stuck with no remaining resolver and no log line. The hung trail is exactly that shape; the begin itself is invisible because the flow's OMI AUTH NSLog lines are not captured in the dev log, and the closed set of beginSessionAttempt call sites includes the restore flow's own invalidation branches, which fire on a transient empty credential read - the same launch whose Firebase cached user read also came back nil. The watchdog now resolves on the phase alone: while the app still reports restoring and no user-driven sign-in owns the UI (isLoading), it lands in the same recoverable state the ungated case always did. It arms before the restore's first await, so the restore's own awaits cannot delay the arming, and it logs - visibly in the dev log - whether the launch attempt was superseded, so the next occurrence names its race. Tests: AuthRestoreWatchdogTests drives the seam with an injectable timeout. testRestoringPhaseResolvesEvenWhenTheLaunchAttemptWasSuperseded fails on the attempt-gated watchdog (phase stuck, verified) and passes on the fix; testTheWatchdogLeavesARestoringPhaseAloneWhileASignInOwnsTheUI pins the isLoading deferral; a source-wiring pin asserts the watchdog arms before the restore await and resolves on phase alone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): re-follow the live edge when the recap pill's inset moves The live-tracked inset is document space above the reader, so when the pill's measured height changes after admission — the overview arriving a beat after the headline is enough — the document shifts under a stationary viewport and the newest row ends up buried under the pill. Caught in the running bundle: the pill rendered thin, but the first transcript row slid under it. A reader following the live edge gets admission's own answer, a re-follow; a reader scrolled away is left exactly where they are. Verified visually: chat now shows the thin pill with the live edge readable beneath it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): close the activation-actions extension the knows-list prune truncated The knows-list retirement left DesktopAutomationActivationActions.swift one brace short of its extension, so a clean checkout of the branch does not parse. Closing it restores the build without touching the pruned flow. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop): the daily recap becomes a full page on the Chat-first shell The sheet is dead. `ChatFirstRoute` gains `.dailyRecap(DailyRecapRouteRef)` — a non-primary destination in the `.more` family, hosted by the shell on the glass page lane like the other full pages. The route carries identity only (record id + date), so a relaunch onto the persisted route re-fetches through the new `getDailySummary(id:)` read; a record that is gone degrades to an honest unavailable state. `ChatFirstShellNavigation.openDailyRecap` captures the opening surface's route as a transient origin, and the page's back chevron — and Escape — return there; a tab select supersedes it and an owner change clears it. The page keeps the punch-list layout: back chevron + Daily-recap eyebrow, single-line action pills (Ask primary, Regenerate secondary — `fixedSize` so no label can ever wrap), dayEmoji + full weekday date eyebrow + headline + lede in a 720 pt column, one row of equal-width single-line stat chips that scrolls rather than folds a label, and the deep-linked sections. The eyebrow's wider date format lives beside the pill's in `ChatDailySummaryPresentation` so the arithmetic stays in one place. Chat and Activity open the same route; their flows wait on `visibleChatFirstRoute: daily-recap`. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop): the chat recap becomes an in-history day boundary The pinned overlay bar, its measured-height inset, the high-water and live-tracking preference plumbing, and INV-CHAT-2's admission gating are all deleted from ChatMessagesView. A row that lives in the thread's own history cannot shift the live edge, so none of the banner constraints apply any more — the transcript simply renders `ChatDailyRecapRow` as part of its row data, anchored above the first message on or after the recap's day. When that boundary is outside the loaded window (no message reaches the day, or the day's start is at the top of the window with more history above), the thread renders nothing rather than a marker the history cannot back up. A cleared thread keeps its recap withdrawn through the coordinator's existing clear contract. The row is a quiet, centered history marker — day label with emoji, a one-line headline, two lines of overview, whole row clickable — and it opens the typed recap route. The old pinned pill file is gone; its identifier and covers follow the rename. The admission unit tests become placement tests. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop): the Activity day card grows its recap as one continuous body The open day's recap no longer floats as a second rounded surface under the day header. The header becomes the card's top half — rounded at the top, square at the bottom, no gap — and the recap renders as the card's body on the same material, square where they meet and rounded where the card ends, with the hairline only on the outer edges so no line runs across the seam. The header's content resolution is shared with the slot, so the attached shape and the slot's content can never disagree; folding the day keeps the thin header-only card, and the generate affordance for a recap-less day stays its own small surface. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop-e2e): identity seam and drive path for the recap page `daily_summary_snapshot` reports the record's opaque id (identity, not content), and a new `open_daily_recap_page` action opens the dedicated page through the same `ChatFirstShellNavigation.openDailyRecap` the recap rows call — so flows and harnesses can reach the page without a click. Launch always lands on chat (`openMainAppChat`), so the persisted recap route never survives a relaunch on screen; the route's persistence stays harmless exactly as intended. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): stat chips hold their full label on one line 106 pt folded "4h watching" into "4h… watc…" — the exact mid-word wrap the page layout exists to forbid. The equal-width slot widens to fit the longest label ("8h 5m listening", "13 conversations"); a narrower lane scrolls the row rather than folding a chip. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * style(desktop): the in-history recap row sits below the messages Review feedback: centered bordered prose read as another answer. The row is now left-aligned one type step under the bubbles, borderless on a softer fill, with the chevron trailing - a day marker in the thread, not a competing card. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * perf(desktop): a live stream pins the following viewport to the live edge The streaming follow still jittered: every 80 ms the throttle fired a 0.16 s ease-out toward the bottom as measured when the follow fired, but the paced flush grows the document every 35 ms - so each target was stale on arrival, and the viewport's velocity cycled between decelerate, stall, accelerate for the whole answer. Mid-stream re-wraps that briefly shrank the document even drove the glide upward. While the last message is streaming and the reader is following, a 60 Hz ChatLiveEdgePinner now moves the clip view to the current maximum scroll every tick - the bottom is read and taken in the same tick, so there is no target to go stale. Armed from the streaming follow path; disarmed on stream settle, reader input, mode change, conversation switch, and teardown through cancelAllPendingScrolls; every move marks the programmatic-scroll signal so the user-scroll detector keeps classifying tracking as the app's own. The glide remains for discrete jumps (button, restore settling). Measured on the mounted transcript: worst between-flush drift 0.0 pt (the replaced glide's was 48 pt); Reduce Motion applies - pinning is positional tracking, not animation. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(desktop-e2e): the activity day header answers to spine-day-recap-header A stable handle for flows (and a reviewer) to drive a day's fold without reaching into the spine's private collapse state. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): recap stat chips hug their labels The fixed 152 pt slot was the defect: short labels left dead space inside their capsule while the row's tail clipped past the column edge. Chips now size to their content with one uniform gap, and the row scrolls only when the window is genuinely too narrow. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): the activity day card stops repeating its recap's emoji and arrow The header already carries the recap's emoji and is the fold control, so the recap body showed both a second time - two emojis stacked, and the fold chevron sitting directly over the row's own arrow. The body is now headline + summary only, and its whole surface still opens the page. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): all six recap stat chips fit the column at min window width * feat(desktop): the chat recap row shows the day's stats, compactly The chat row was the only recap surface without the day's numbers. A strip of micro chips - same data as the Activity day card, bubble-row weight, content-hugging so nothing clips - sits under the overview. Badges and actions stay page-only; this is the number line, nothing more. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): the tasks list stops nesting a lazy stack inside a lazy item TaskCategorySection is itself an item of tasksListView's LazyVStack, and it laid its rows out in a nested LazyVStack. A lazy container nested inside a lazy item makes the section's measured height non-convergent — the inner stack reports estimates while the outer measures it and real heights once placed — so every layout pass mutated lazy-item phases and re-signaled prefetch, each scheduling another transaction. The flush never drained: a permanent beachball the moment the list was scrolled with enough data (sampled: 2627/2627 main-thread samples inside GraphHost.flushTransactions, dominated by LazyStack measureEstimates and LazyLayoutViewCache updateItemPhase). The row stack is now eager, so a section's measured height is deterministic and the flush drains. The outer stack stays the virtualizer. A source inspection guard pins the no-nesting contract. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * perf(desktop): the switch into Chat stops paying a per-node preference reduce and a fixed settle ladder Navigating to Chat felt delayed. Two measured causes: 1. The composer height travelled through QueryComposerHeightKey, written at the composer and read at the surface root, so SwiftUI ran a reduce over one combiner pair per node of everything mounted between them — the whole transcript — on every switch. Sampled: 40% of the mount transaction's observer work. The height now travels through onGeometryChange into the same guarded state write; the preference key is gone. 2. The initial restore re-pinned the live edge on a fixed [0.05, 0.2, 0.5, 1.0] ladder, so every switch waited out the full ladder even when the document had stopped reflowing after the first pass — 1.16s of pure waiting at the p50. ChatInitialRestoreSettle tightens the ladder and settles early on a height-stability check (two consecutive passes measuring the same laid-out document); the last pass still completes unconditionally. Also adds the opt-in OMI_SWITCH_PERF switch-span telemetry (route select, destination teardown, transcript mount, first laid-out document, settled restore, main-thread stall watchdog) that measured all of this, and the chat-first flow cover for it. Measured: restoreSettled p50 1156ms -> 241ms against memories as the source page. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * perf(desktop): the bubble identity check stops re-deriving the answer text ChatBubbleIdentity.equal compared lhs.copyableText == rhs.copyableText, but copyableText derives purely from (contentBlocks, text, isStreaming) — every one of which the same conjunction already compares. SwiftUI runs this equality for every bubble on every transcript body pass, and each evaluation re-derived and whitespace-normalized the full answer text. Sampled mid-switch: the copyableText chain was the largest app-code cost of the mount. Removing the redundant term is behavior-preserving and dropped the navigate-to-chat settle from ~1.3s to ~240ms at the p50; a regression test pins that block-only edits stay visible through the blocks term. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * perf(desktop): bubble identity compares content blocks field-by-field, not through JSON ChatBubbleIdentity.equal's blocks term went through ChatContentCodec.comparisonData, which JSON-encodes the persistence dictionary with sorted keys — two full serializations per bubble per transcript pass, for every unchanged row, ~28 times a second while an answer streams. ChatContentBlock now conforms to Equatable with hand-written field-by-field comparison (questionCard options deep-compare as NSDictionary; ToolCallInput synthesizes). The comparison is deliberately one notch stricter than the encoder was — in-flight toolCall statuses (.running/.slow/.stalled) no longer collapse together, which can only force an extra re-render mid-turn, never stale UI: the isStreaming guard already re-renders through those states. Also gives the task-chat panel the compact transcript window: it was the one ChatMessagesView host still on the 500-row default that measured 910 ms and 607 native views against 114 ms and 84 compact. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): goals cards stop showing impossible progress and echoing their own titles UI audit findings. A goal created without an explicit target keeps the creation default (1.0) while progress updates push currentValue past it, and the row rendered that raw: "Progress: 20 of 1". When the current value has overshot that unconfigured default target, the target carries no information, so the summary reports the value alone. desiredOutcome falling back to the title echoed every card's title under its title; it is hidden when it matches. Both the focused-goal card and the list rows capped at 680pt while the page header spanned the panel — the cards now fill the lane like the rest of the page. The Rewind capture control said "Capture Off" next to a switch knob in the on position, because the knob shows the setting and the label showed health, and the two can disagree. "Off" now only ever describes the setting; a dead-but- enabled capture reads "Capture Stopped". Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * perf(desktop): the task-chat panel gets the compact transcript window Missed from d27ca379ed: TaskChatPanel is the file carrying the transcriptWindowPolicy one-liner; without it the panel mounts the 500-row default (910 ms / 607 views vs 114 ms / 84 compact for a 400-message thread). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(desktop): the content-block equality lives in the codec file ChatProvider.swift is under a hard convergence limit and has one job; the ChatContentBlock: Equatable conformance belongs beside the block codec. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(desktop): regenerate tool surfaces at the merged manifest Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(desktop): the tasks list flattens to a single lazy level TaskCategorySection was one lazy item holding an eager stack of all its rows. Eager-in-lazy meant the outer list could not estimate a section without laying out every row, so all sections materialized: leaving the page tore down every task row in the profile (~330 trees, ~600 ms of main-thread layout per navigate away), the reason navigate-to-Chat from Tasks cost ~2x the same switch from Memories. The list is now one LazyVStack whose items interleave a category header with one item per task row (TasksListItem). Rows virtualize individually, nothing nests inside a lazy item — which also retires the lazy-in-lazy estimate livelock class the eager stack had traded the freeze for. Row ids keep keyboard navigation and per-row state; the header (icon, count, Today menu, top drop zone, accessibility identifiers) survives as TaskCategorySectionHeader; per-row drag-drop and inline-create composition is byte-for-byte the old one. Collapse was never wired at the only call site; the dead parameters stay compiling with a note that a real implementation must filter the items where they are built. Measured on a ~330-task profile: freeze stress clean (zero flushTransactions / propagate_dirty samples), the teardown storm gone (TaskRow ≈ 3-5 of ~2370 samples mid-switch), tasks→chat switch parity with Memories. Interactions exercised: snapshot, create, toggle, reorder (persisted), capture. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): release compile and the Dart analyzer ratchet TaskCardView's body type-checked past the release optimizer's expression budget ("unable to type-check this expression in reasonable time" at ChatFirstContentBlockViews.swift:198); the branches are now extracted into renderedCard / unavailableBlock / unavailableDescription so each sub-expression type-checks on its own. Drops the one new unused import the Dart analyzer ratchet flagged in the content-block parity test. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(desktop): the beat-anchoring contract holds on a simulated clock The beat-period test drove a real Thread.sleep and then asserted the sleep stayed inside the beat budget; on a loaded CI runner the sleep itself overshot (periods measured 125-194 ms against a 120 ms ceiling) and the test failed on machines whose streaming was fine. The contract lives in the scheduling math, so nextBeatDeadline is extracted as a pure function and the test simulates the beat chain against it: work inside the interval holds the grid exactly, an overrunning beat resynchronizes to now instead of compounding. Deterministic, and it pins more than the wall-clock version could — the old assertion tolerated any drift under a ceiling, this one requires the grid to hold exactly. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> | 2 天前 | |
fix(ci): exclude gitignored files from the dead-code ratchet (#12756) * fix(ci): exclude gitignored files from the dead-code ratchet check_dead_code.py enumerated each area with a raw filesystem walk, so any gitignored file under app/lib, backend/ or a desktop src tree was treated as production source. Those files are absent from CI's checkout and from every diff, so the gate failed only on developer machines and read as broken. Reproduced on this repo: app/lib/firebase_options_dev.dart is gitignored at .gitignore:176 and generated during Flutter setup. With it present, --area flutter reported it as newly dead and failed; removing the file turned the same check green. Filter each scan through git ls-files --others --ignored --exclude-standard, mirroring _git_ignored_paths in backend/scripts/generate_plan_catalog.py, which was added for this same failure in #12476. --directory collapses an ignored directory to a single entry, so membership is tested through ancestors. When git cannot answer the set is empty and the previous pure-filesystem behaviour stands, which keeps the checker hermetic. Verdicts are unchanged for agent, windows and backend on this checkout; only flutter moves, from a false failure to ok. * fix(ci): fail closed on an ignored entry and harden the git filter Review follow-ups on the gitignore filter. scan_flutter seeds reachability from app/lib/main.dart. Excluding an ignored entry from the scan set while still walking from it would strand every other file and report the whole area dead, so an ignored entry now errors the way a missing one does, matching what scan_ts already did when no entry remained. git ls-files output is decoded with errors="surrogateescape", and UnicodeError joins the fallback path. A filename whose bytes are invalid for the locale previously raised out of the checker instead of degrading to the filesystem scan the fallback exists to provide. Tests: the shared scan_ts path now has its own ignored-file case, since the agent and windows areas went through it untested; the flutter and backend fixtures stage their files so the "still caught when tracked" assertions describe what they actually exercise; and the ignored-entry guard has a regression test that fails without it. | 2 天前 | |
fix(preflight): decode Git content as UTF-8 on Windows (#10595) * fix(preflight): decode deployment Git output as UTF-8 * fix(preflight): decode deferred-work Git output as UTF-8 * fix(preflight): decode line-ratchet Git output as UTF-8 The product line-count ratchet read baseline paths and JSON blobs with the Windows host code page. Decode all three Git content reads explicitly as UTF-8 and lock both sharded and legacy paths with a subprocess contract test. --------- Co-authored-by: 491034170 <142008960+491034170@users.noreply.github.com> | 1 个月前 | |
fix(desktop): make auth restoration transactional Gate authenticated UI on launch-time credential validation, preserve rollback credentials across Keychain migration failures, and add a recoverable session state. Require the exact signed beta artifact to pass an in-app synthetic Keychain write/read/delete canary before publication. | 1 个月前 | |
feat(desktop-release): isolate drift self-test from hook GIT_DIR Pre-push exports GIT_DIR for the outer checkout, so the fixture repo inherited that namespace and failed to commit. Strip GIT_* and disable hooks/gpgsign in the self-test. Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: multica-agent <github@multica.ai> | 4 天前 | |
chore(ci): make traffic-regression gate robust to hook git env and python 3.14 - argparse help: escape the literal % (python 3.14 applies %-formatting to help strings and rejects a bare '% t') - _git_is_ancestor + fixtures: strip GIT_DIR/GIT_WORK_TREE/GIT_INDEX_FILE/ GIT_OBJECT_DIRECTORY/GIT_NAMESPACE so the probes and scratch repos target the current directory's repository even when run from a pre-push hook Co-authored-by: multica-agent <github@multica.ai> | 20 天前 | |
ci: unify deterministic checks behind manifest Make local preflight, pre-push, and CI resolve the same diff-scoped deterministic checks from .github/checks-manifest.yaml. Add drift guards and local enforcement for brand and deferred-work policies.\n\nVerified: python3 -m unittest discover -s .github/scripts -p 'test_*.py' (40 passed); actionlint on changed workflows; bash syntax checks; make preflight; adversarial purple, deferred-marker, workflow-drift, and lane-removal cases. | 1 个月前 | |
docs: correct stale references surfaced in review - apps-marketplace e2e: covers now lists category_section.dart (the flow's S1-S7 browsing renders CategorySection/SectionAppItemCard; AppListItem is only rendered in search-filter results, which the flow never enters) - goals-tracking e2e: GoalsWidget sits after the processing-conversations widget in conversations_page.dart (TodayTasksWidget renders in home_content.dart); with zero goals the widget hides and the home-screen Add Goal entry point is ActionItemsPage._buildGoalsRow - goals_widget.dart: same correction in the hide-when-empty comment - 00-INDEX.md: focus/stats.ts and AutoCreatedTasksStep.tsx entries updated to match the deletion note — both actionable suggestions are now marked resolved-by-deletion instead of presented as live work - plugins/README.md: document monolith (plugins/requirements.txt -> ./omi-plugin-sdk) vs per-service (own requirements.txt -> ../omi-plugin-sdk) SDK install paths separately | 3 天前 | |
fix(ci): count a landed change once in the failure-class guard ratchet (#12119) The guard ratchet counts first-parent integration changes declaring a class, then adds the `--pr-body-file` declaration so the declaring PR fails first. Both sources can describe the same change: on a main push HEAD *is* that integration change, and `git log --first-parent` already counted it. So one change is counted twice and the total depends on where the check runs. A class with one prior declaration reads as 2 on the PR (green) and as 3 on the push of that same PR (red at the default threshold), which is exactly the after-merge red the main-push metadata path exists to prevent. Locally the count is inflated the same way, so pre-push does not predict its own PR. Subtract what HEAD already contributed: a declaration the body carries counts only when HEAD's own message did not carry it. A `pull_request` run checks out the merge ref, which declares nothing, so the pending change still counts. Failure-Class: new | 14 天前 | |
harden(backend): centralize Firestore read boundaries | 1 个月前 | |
fix(repo): reject Git identities that can never be attributed (#12244) * fix(repo): reject Git identities that can never be attributed PR #12239 was authored entirely by `r <r@r>` from a stale clone-local `user.email`. This guard ran on every one of those commits — 66 over five days — and printed "OK: Git author identity is not a test fixture" each time. It was enumerating the previous incident rather than the invariant. After #11525 minted `Ratchet Test <ratchet-test@example.invalid>`, the check learned that exact display name, that exact local-part, and three reserved TLDs. `r@r` is none of those, so it passed: `email_tld("r@r")` returns `"r"`, because the domain has no dot to split on. The rule is now the property that makes an identity wrong. An address whose domain cannot resolve can never be delivered or attributed to anyone, whatever it is called. That single predicate covers both incidents and the family they belong to, and it needs no allowlist to maintain. `failures_for` already applied its predicate to clone-local config as well as to the commit range, so one change covers every lane at once: - pre-commit (`--pending`) rejects the override before a commit is minted - the local lane of `git-author-identity` rejects it again before push - the CI lane never sees a developer's `.git/config`, so the commit range is the only unbypassable place this can be caught — and it now is Real addresses, including `users.noreply.github.com`, are untouched; a name-only identity carries an empty email and is not judged on it. The report wording said "test fixture", which is exactly what let a non-fixture bad identity read as a pass. It now says unattributable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(failure-classes): add FC-guard-encodes-incident-literals The class this PR's fix belongs to, and the reason it took a second incident to notice: a guard written after an incident encoded that incident's literal values instead of the property that made it wrong, so the next instance passed while the check reported success. evidence_prs is empty because this PR is the first instance; the adding commit is the evidence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> | 12 天前 | |
fix(ci): defer guardrail ratchet annotations on macOS Failure-Class: none | 1 个月前 | |
Enforce lifecycle headers for rollout scaffolding | 1 个月前 | |
chore(ci): reduce redundant CI work | 1 个月前 | |
build(devops): add offline OpenTofu foundation guard Establish an intentionally empty OpenTofu foundation module for #9842.\n\nThe module has a GCS backend shape, separate environment placeholders, and a locked Google provider. Its guard allows only durable identity, additive IAM, secret metadata, and state-bucket resource families; it rejects release resources, data sources, and secret values. The validation workflow is credentials-free and can only validate the source plus plan a backend-free temporary copy.\n\nVerification:\n- python3 .github/scripts/test_check_opentofu_foundation.py\n- python3 .github/scripts/check_opentofu_foundation.py --plan-json /tmp/omi-opentofu.wkAqDX/foundation-clean-plan.json\n- actionlint -shellcheck '' .github/workflows/opentofu-foundation-validate.yml\n- checksum-verified OpenTofu 1.12.4 init -backend=false, validate, and offline -refresh=false empty plan with Google credentials unset | 1 个月前 | |
Document memory architecture and enforce package maps | 1 个月前 | |
ci: regression guardrails — PR scope advisor, test-discovery ratchet, sync annotation ratchet Three mechanical regression guardrails derived from an internal audit of merged PRs: 1. PR scope advisor (advisory-only, never blocks): warns at 1,500 and 3,000 changed production-source lines so reviewers calibrate skepticism to diff size. 2. Unit-test discovery ratchet: fails when any test file is not discovered by a verified runner. Allowlists only shrink. 3. Sync parameter annotation ratchet: pyright executionEnvironments enforces reportMissingParameterType=error for utils/sync/. Co-authored-by: skanderkaroui <skanderkaroui@users.noreply.github.com> | 1 个月前 | |
fix(ci): keep routing-only push gates bounded (#11361) * fix(ci): stop routing-only diffs from waking Flutter regeneration The pre-push CI-prediction selector treated any ROUTING_INPUTS edit (for example .github/checks-manifest.yaml alone) as a reason to wake flutter-codegen and flutter-l10n. A manifest-only diff then paid a full build_runner rebuild (~17 minutes cold) at push time and had to reach for PRE_PUSH_SKIP_FLUTTER_GENERATED=1. Routing metadata cannot make a committed generated file stale, so those two lanes stay owned by their real generator inputs; every other component lane the selector can influence is still woken. Verification: python3 .github/scripts/test_pre_push_ci_prediction.py -> Ran 16 tests OK python3 scripts/pre_push_ci_prediction.py --changed-files <(manifest only) -> no flutter-codegen / flutter-l10n Mutation check: re-adding flutter-codegen/flutter-l10n to the selector fan-out fails test_routing_only_diff_does_not_wake_flutter_regeneration Failure-Class: none Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(desktop): pin FluidAudio by its real tag instead of a raw revision The comment above the FluidAudio dependency claimed upstream had removed the v0.14.8+ tags, so the manifest carried a bare revision pin. That premise is false: git ls-remote --tags shows v0.14.8 and every v0.15.x tag present, and revision 19600a4 IS tag v0.15.5. A raw revision pin forces SPM to search an unlabelled commit on a cold cache; exact: 0.15.5 resolves the same commit through a normal tag lookup. Verification (cold cache: fresh --cache-path/--scratch-path/--config-path /--security-path under /tmp, removed before the run): swift package resolve -> succeeded, 8m12s Package.resolved fluidaudio pin unchanged at revision 19600a485baa4998812e4654b70d2bab8f2c9949, now also labelled 0.15.5 swift package show-dependencies -> fluidaudio<https://github.com/FluidInference/FluidAudio.git@0.15.5> Failure-Class: none Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(desktop): restore existing FluidAudio revision pin Remove the unrelated exact-tag manifest change and its inaccurate rationale while preserving the audited resolved revision. Verification: xcrun swift package dump-package; desktop manifest and Swift CI contract tests. * ci: wake Flutter generation for mobile workflow changes * ci(app): ratchet existing analyzer diagnostics * ci: wake Flutter generation for detect-changes action changes The routing exclusion treated every ROUTING_INPUTS path as metadata, but .github/actions/detect-changes/action.yml defines and forwards has_app_codegen/has_app_l10n/has_flutter_generated. A change there emitted has_flutter_generated=false and skipped the generated-files job that would exercise the rewiring. Narrow the exclusion to genuine routing metadata via FLUTTER_GENERATION_DEFINITION_INPUTS/_PREFIXES. * fix(ci): refresh pre-push hatch guidance Keep the hatch disclosure documentation aligned with the selector boundary: routing metadata stays cheap while generator inputs and workflow definitions retain generated-output checks. Failure-Class: none --------- Co-authored-by: Max Carter 祁明思 <undivisible@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> | 18 天前 | |
ci: make CI fast and reliable — every job under 20 minutes, recurring flake removed (#12247) * ci(desktop-windows): cache electron & electron-builder downloads The Linux package job failed on main (EAI_AGAIN github.com) and on PRs (read ECONNRESET) because electron-builder re-downloads its packaging tools on every run. Cache the tool and Electron archives so steady-state builds never depend on those endpoints, on both the Linux and Windows packaging jobs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(hermetic-e2e): stop minting red checks from cancelled runs; harden sync stack readiness The merge gate ran under always() and failed hard on 'cancelled' upstream results whenever cancel-in-progress superseded a run — 3 of the last 4 re-run-to-green cycles on this workflow were exactly that. Gate now goes neutral on cancellation and defers to the superseding run. The sync Cloud Tasks stack's one genuine flake: scenarios run back-to-back, a dying child of the previous stack can still answer a TCP probe, and the authoritative health check had the tightest budget (20s) while the weak port probe had the generous one. close() now drains the stack's ports and health gets 60s. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(desktop-swift): collapse macOS-runner demand that was queueing every run Desktop Swift CI's tail (p90 150 min, max 218 min) was queue time: the repo's only macos-15 consumers demanded ~4.9 concurrent runners against a cap of ~5. Cut the demand instead of the deadline: - Run the release-mode UserNotifications regression inside the Release Compile job, next to the release build it consumes, instead of paying a second from-scratch release build (~31 min) on the verify runner. The planner still requires the Release Compile check by exact name, and the Build & Tests aggregate now hard-fails on the release job's result, so release tagging and PR merges both still gate on it (#11373/#11374 protection preserved). - Cache SwiftPM dependencies and the release Desktop/.build in the Release Compile job, saved from main pushes so every ref can restore it, and drop the forced rm -rf that guaranteed a ~20-min cold build. - Run launcher script tests 3-way parallel with per-test logs (9.3 min serial p50), and widen the build-lock recovery budgets that saturated runners pushed past 4s. - blob:none the verify checkout; full history stays for diffing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(mobile,web,parakeet): kill checkout/cache tails and 45-min unschedulable burns Mobile and Web Checks' p90-to-max tail was almost entirely actions/checkout pulling all blobs of a ~1.25GB repo with fetch-depth: 0 (a 53-minute checkout-only Web run happened); history-needing jobs now use blob:none and worktree-only jobs go shallow, with timeout caps so no job can idle unbounded again. The Android compile smoke paid a ~700s cold Gradle build every run because the default cache policy never wrote from main; main now seeds the cache PRs read, and the build_runner cache keys drop the run_id suffix that guaranteed a miss on every single run. Parakeet GPU tests have failed nightly since 08-03 burning exactly 45 min each with zero logs: the pod is unschedulable (L4 pool at capacity) and the workflow neither noticed nor captured pod events. Gate on PodScheduled within 5 minutes with the real cause in the error, stream pod logs live so deadline kills keep evidence, and include pod events in diagnostics. The capacity fix itself (NLLB squatting the parakeet pool's second GPU) is a cluster change tracked separately. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(repo-checks): make the line-count ratchet invariant under base drift; pin uv 13 of the last 17 Repo Checks failures were the line-count ratchet, and the nondeterministic ones shared one cause: the check demanded byte-exact absolute counts against a synthetic merge with origin/main, which moves under every open PR (~90 pushes/hr). A correct declaration went stale with no author action, in both directions (drifted endpoints, and previously- mandatory exceptions turning fatally 'unused' when main absorbed an edit). The declaration now approves a growth *allowance* — the delta is a property of the PR alone — and an exception the diff no longer needs warns instead of failing. Growth beyond the allowance, duplicates, malformed lines, and unsupported paths still fail. Also pin setup-uv's resolved version (matching backend/Dockerfile's UV_VERSION) so it stops fetching the astral-sh 'latest' manifest, which hard-failed a run. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci: pin setup-uv's resolved version everywhere it was fetching 'latest' Every unpinned astral-sh/setup-uv call resolves 'latest' by fetching the astral-sh/versions manifest from raw.githubusercontent.com at job start — one observed Repo Checks failure was exactly that fetch dying. Pin the same 0.11.13 the backend Dockerfiles already use, which the action resolves locally with no network call. repo-checks.yml was pinned in the previous commit; this covers the remaining call sites and the release-eligibility action (with its byte-exact contract fixture). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(desktop-swift): batch SwiftPM test invocations with per-suite fallback The debug suite spawned one 'xcrun swift test' process per discovered suite — 674 SwiftPM startups at ~5.1s each, ~28.7 min of the verify job's wall clock, dwarfing actual test execution. Workers now run suites in batches of 25 through one SwiftPM process (repeated --filter flags), with batching kept strictly optimistic: a batch that exits non-zero or times out is discarded and every member re-runs through the untouched per-suite path, so isolation semantics, per-suite budgets, and failure attribution are unchanged for anything red. The serial shared-auth cluster is never batched, and OMI_SWIFT_TEST_SUITE_BATCH_SIZE=1 restores the old behavior bit-for-bit. Also fixes a pre-existing bash-3.2 set -u break (empty build_args expansion) that made OMI_SWIFT_TEST_PREBUILD=0 kill every suite, and restores the case-pattern skip list in the workflow's launcher step that check-launcher-test-skips.py parses. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(backend-unit): fail-open module stubs, checkout tail, and honest runner verdicts The suite's recurring 'flake' was one deterministic trap: a module-scope sys.modules['utils'] stub with __path__ = [] turned every new import in routers/conversations.py into a ModuleNotFoundError for whoever ran CI next — main broke twice on 2026-08-25 and four unrelated PRs inherited the red. The stub package now carries the real package __path__, so unlisted submodules resolve to the real module instead of raising (explicit stubs still stub; verified by reproducing the incident against both versions). Checkout p90 was 267s and max 786s from full-blob clones; blob:none keeps the merge-base history scripts/changed-files needs. timeout 45→30 so a hang stops burning a concurrency slot for 25 minutes past the p90. Also two standing runner bugs found while measuring: pytest's exit 5 on total worker death was read as 'no tests selected' and reported green, and the fast-unit duration guard's verdict was silently discarded under xdist (workers now hand offenders to the controller). A single-session xdist partition of the suite was implemented, measured, and disproved on this tree (upb descriptor segfaults, OOM at 815 files, per-worker collection); it ships opt-in via BACKEND_PYTEST_PARALLEL_SESSION=1 with the measurements in the comments, default behavior unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(desktop): changelog fragment for internal CI runner changes Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(desktop-swift): keep measure-block suites out of batches; three workers The first PR run of the batched suite exposed one pathology: the batch carrying MemoryAtlasPerformanceHarnessTests (245s of XCTest measure blocks, slower still under contention) blew its 17-minute budget and paid the 25-suite isolated fallback on top — 15 of the step's 40 minutes. Measure suites are now derived from the source (like the serial cluster, so a new one cannot silently join the batch pool) and run in their own SwiftPM process with a doubled per-suite budget. With per-invocation SwiftPM startup gone — the reason two workers were the measured ceiling — workers now match the runner's three cores. This PR's CI runs are the macOS-runner evidence; drop back to two if the performance harness starts timing out. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> | 12 天前 | |
docs: take operator pages off docs.omi.me Unlisted Mintlify MDX is still a public URL. Move runbooks, flags, invariants, and agent rules next to owning code, add docs/AGENTS.md as the site allow-list, and correct the live kill-switch contract after the JIT authority page leaves the site. Co-authored-by: Cursor <cursoragent@cursor.com> | 7 天前 | |
ci(public-build): env-var removal, per-environment flags, real invoker identity checks (#12582) * ci(public-build): env-var removal, per-environment flags, real invoker identity checks Plaintext Cloud Run env vars survived merge deploys, restricted ingress was applied to development, and TBD placeholders passed presence checks into gcloud. Co-authored-by: Cursor <cursoragent@cursor.com> * ci(public-build): probe actAs via IAM testIamPermissions REST gcloud has no iam service-accounts test-iam-permissions subcommand, so the previous preflight would fail every prod deploy. Call the IAM REST method with urllib and a print-access-token bearer instead. * ci(public-build): reject remove_runtime_env_vars overlapping preserved secrets Cubic review (PRRT_kwDOLkKqys6eWSS0): a runtime name in both preserve_runtime_secrets and remove_runtime_env_vars loaded cleanly, yet deployment emits --remove-env-vars for a binding the contract claims to preserve via the merge update strategies — the removal would strip the preserved secret binding. Extend the dedicated overlap rejection to preserve_runtime_secrets and its mirrored fallback_runtime_secrets, with a regression test. --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: David Zhang <9387252+Git-on-my-level@users.noreply.github.com> | 5 天前 | |
fix: keep the macOS hourly release train moving Bind each candidate to its checked source, recover squash-orphaned changelog branches, install Bun for release eligibility, and age freshness from the oldest unshipped update. Verification: 77 focused release-control tests; actionlint; Python compile; diff check. Failure-Class: new | 18 天前 | |
fix(infra): address PR review comments across deploy control plane Harden composite-action contract guards, fix WMW smoke rollback and Firestore TTL project routing, thread workflow_root/fence rendering in runtime validation, and clean up deploy script edge cases from review. Co-authored-by: Cursor <cursoragent@cursor.com> | 1 个月前 | |
ci: move paid Actions jobs to standard runners (SCA-196) Co-authored-by: multica-agent <github@multica.ai> | 1 个月前 | |
ci: collapse desktop beta to signed-smoke plus hourly freshness (#11588) * ci: collapse desktop beta to signed-smoke plus hourly freshness Qualification never rolled back recent signed-smoke manifests and starved the planner when a push event was missed. Remove the lane, make source-gate failures diagnosable from one command, and alarm when candidate and live beta diverge. Co-authored-by: Cursor <cursoragent@cursor.com> * docs: retarget INV-BETA-1 guards after deleting qualification tests The auto-beta-candidate script is gone with the qualification lane; keep the locked beta-identity invariant pointing at a guard that still exists. Co-authored-by: Cursor <cursoragent@cursor.com> * ci: make skipped desktop checks legible and let a green tip unblock the train Two failure modes survived the beta-train collapse and are fixed here. Skipped jobs published the wrong check name. GitHub does not evaluate a job's `name:` for a SKIPPED job, so the conditional names on `desktop-swift` and `desktop-swift-release-compile` were published verbatim as the raw expression text. On every commit that did not touch desktop paths the required check `Desktop Swift Build & Tests` was therefore ABSENT rather than skipped, and the planner reported "missing" instead of the truth. Observed on f666ddd4a3, 7a79f08329 and 7d7ed62e5, all of which read green. The conditional existed to keep a merged `pull_request.closed` bookkeeping run from publishing a skipped required check onto the merge SHA; dropping the `closed` event removes that hazard at the source and lets both names be literals. A contract test now rejects any expression in a job name. A green tip did not unblock the train. The planner selects the newest desktop-touching commit and, when its checks are red, could only fall back to an OLDER green SHA. On Aug 14 main's tip was green while the newest desktop-touching commit below it was red on a flaky Swift suite, so the train shipped stale code or wedged. A first-parent commit above the blocked SHA contains everything the blocked SHA contains, so its own exact-SHA checks tested a superset of that tree; when they are genuinely green the train may ship from that newer SHA. Tried before the backward fallback, because it ships newer code. Only a real `ready` gate qualifies, so a skipped or absent check still never counts as success. Also keep the beta rollback precondition expressible: beta manifests carry the `signed-smoke` tier, whose frozen-schema truth is `qualification_passed: False`, so the literal T2/True requirement rejected every current rollback target. Failure-Class: none Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): keep the release lifecycle helper compiling under the admin es5 target `newestSparkleVersion` iterated `String.matchAll()` with `for...of`. The admin package sets "target": "es5", where iterating an IterableIterator is TS2802, so `npm run typecheck` failed and took the Web Checks Build job red. Local pre-push does not typecheck the Next.js admin app, so CI was the first place this could surface. Use `exec` loops instead of widening the package's compile target, which would change output for every file to fix one. Verified with the same commands CI runs: `npm run typecheck` clean and `npm test` 88 passed across 14 files. Failure-Class: none Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> | 23 天前 | |
chore(ci): scope the INV-TASK-2 guard to the proposing surfaces The static writer ban covered one function name and could not express which conversations propose. That split is behaviour, so it belongs in the capture tests; the guard keeps the three facts it can actually prove — no creating outcome, no accept in capture, no accept on the delivery client — plus the manifest rule that a shared-policy source declares no create anchor. | 8 天前 | |
feat: converge universal memory and task authority | 26 天前 | |
Admin: real cost data and an honesty sweep for the Omi TV dashboard (#12198) * Infra costs: replace the April estimate table with billed data computeInfraCosts now has a billing mode as the authoritative path: - GCP: daily net spend (cost + credits) from the BigQuery billing export, _PARTITIONTIME-filtered, series ending at D-2 because the export lags ~11h and back-fills — per the gcp-cost-efficiency cost-monitoring contract. New lib/services/gcp-billing.ts (@google-cloud/bigquery dep, creds via GCP_BILLING_SA_JSON or ADC). - Anthropic/OpenAI: org cost APIs (lib/services/provider-costs.ts) behind ADMIN_ANTHROPIC_COST_API_KEY / ADMIN_OPENAI_COST_API_KEY — names chosen because the deploy contract strips ANTHROPIC_API_KEY/OPENAI_API_KEY. A missing leg is partial coverage, never silently $0. - Desktop/mobile split by the omi-cost-analysis usage-weighted shares (LLM pool by memory-event mix, rest by core-event mix), stored as dated config (ADMIN_PLATFORM_COST_SHARES_JSON), replacing the hand-guessed per-service weights. Firestore llm_usage is attribution signal only — its cost_usd rows are provider spend already inside the invoices. - Payload gains costSource/windowEnd/coverage/shares; the legacy hardcoded-table path survives only as a labeled 'estimated' fallback. - profitability: cost and cost-per-user series end at the billing window's D-2 edge instead of padding recent days with the per-user assumption; averages use the same trimmed window; costSource='real' only for billing-backed numbers. Verified: vitest 19 files / 135 tests pass incl. new lib/__tests__/infra-costs-billing.test.ts; tsc --noEmit clean; live fetchGcpBilling(14) against the real export returns windowEnd=2026-08-22 with values penny-identical to the raw bq query. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Wire billing-mode cost secrets into the admin deploy contract GCP_BILLING_SA_JSON -> WEB_ADMIN_GCP_BILLING_SA_JSON (new dedicated SA omi-admin-billing-ro@based-hardware, bigquery.jobUser + dataset READER on gcp_billing_export only; key stored in Secret Manager, runtime SA granted secretAccessor). The two provider cost-API secrets exist as placeholders until org-admin keys are minted — the code treats a rejected key as an unavailable leg (partial), never $0. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Split Anthropic spend by chat mix, not the extraction-token mix The 2026-08-24 methodology audit found the tracked-token base is ~99.9% conversation extraction with essentially no chat in it, while the Anthropic bill in the same window was essentially all chat (floating bar, desktop-backend leak) — splitting Anthropic by the memory-token mix pushed $2.2-3.6k of desktop's own Claude spend onto mobile. Anthropic now splits by a dedicated chat-event share pair (default 0.5464/0.4536, PostHog chat mix); Vertex/Gemini and OpenAI stay on the memory-mix split, which the audit confirmed sound for extraction models. Tests updated; 4/4 pass, tsc clean. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Thread app_platform from clients into gateway accounting The gateway ledger had no platform dimension, so desktop vs mobile chat_agent spend could only be share-split. llm_gateway_headers() now takes an optional platform, emitted as X-Omi-App-Platform only when it normalizes into {desktop, mobile, web} — client-supplied junk never reaches an outbound header and never 403s a request (accounting metadata, not auth). desktop_chat passes its X-App-Platform through; desktop_proactivity hardcodes desktop. The gateway parses the header in ServiceCaller (non-rejecting validator), carries it via AccountingContext into the ledger event as app_platform (null when unknown). 480 backend tests incl. new header/accounting coverage pass; pyright and black clean. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Admin: measure LLM platform split from the gateway ledger; leak metric New lib/services/gateway-ledger.ts reads the llm_gateway_attempts Firestore ledger (~200k docs/day) exclusively through server-side SUM aggregations — never document reads — with per-day immutable caching (volatile last 3 days recomputed). Features roll up to platform classes (desktop-only, shared-chat, shared-extraction, unknown). computeInfraCosts billing mode now derives each day's measured desktop share of extraction-LLM spend from the ledger composition instead of applying the static share directly to the whole pool (desktop-only lanes like desktop_proactive_* finally land on desktop), and emits summary.directPath = provider invoice minus gateway-billed spend — a standing detector for direct-path leaks like the Sonnet 4.6 incident. Missing ledger days keep static shares; a missing invoice leg omits its leak key rather than reading as $0. SUM aggregations need composite indexes including the summed field (verified live — index-merge only serves plain queries), so the six (date, [feature|provider], [payer], estimated_cost_micro_usd) indexes are registered in firestore_index_registry.py and the regenerated manifest; they are already building in prod. Verified: web/admin 142 tests + tsc clean; backend index tests 43 pass after manifest regeneration. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Admin: rebuild the onboarding funnel from the app's real step list STEP_DEFINITIONS had drifted from OnboardingView.swift for months: it still counted a removed Notifications step and a Research step dead since Apr 2026 (both permanent ~0 rows), while HowDidYouHear, DataSources and Exports were missing entirely and Goal lost its skip variant. The list and funnel logic now live in lib/onboarding-funnel.ts, and a test parses OnboardingView.swift directly so the next Swift-side rename fails CI instead of drifting. The grouped query's LIMIT now matches the wrapper's 50k served ceiling and the response carries a truncated flag when it hits the cap. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Admin: make the Stripe metrics honest about what they measure - mrr-trends prices every month at today's prices (a price change rewrites history); the payload now carries pricingBasis: current_prices and the route documents the limitation instead of posing as a historical series. - trialing counts return null on fetch failure instead of a fabricated 0. - monthlyAmount warns loudly on unknown intervals instead of silently adding $0, and excludes non-USD prices from USD totals (surfaced as nonUsdSkipped) instead of blending currencies. - app-subscriptions now counts past_due like every other MRR route, and gains partial/502 degradation matching its siblings. Cache keys bumped v2->v3 so old-shape payloads cannot be served. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Admin: attribute the unknown gateway spend; stop inventing per-user costs The FEATURE_CLASS map only listed model_config feature names, but the gateway's usage-context feature header overrides those on the wire — so the coarse usage-context names (chat, persona, conversation_processing, ...) were the bulk of the unknown bucket. Twelve real features added with verified classifications (workstream_association and chat_structured are desktop-only lanes). Precompute no longer bakes the invented desktop_cost=1.2/mobile_cost=0.3 params into the profitability cache; it now writes the same key the GET route's default parse looks up. In billing mode a zero-active day reports costPerUser null instead of the April $0.20 assumption, and summary averages skip unmeasurable days. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Admin: surface data age and every silent degradation in the stats layer Every payload-cache consumer now stamps freshAt on both cache-hit and fresh-compute paths, so a broken precompute cron no longer serves stale data indistinguishable from live; the cron itself now reports and records ok/failed per metric (precompute-status:v1). macos-versions derives its date label at serve time instead of baking 'today' into the cached payload. Notifications counts explicit true/false/unset instead of counting docs missing the field as enabled. Fabricated values become honest nulls or flags: message-ratings ratio null on no-data days, truncation flags on the capped notification fallback and macos-versions query, voiceSource names which event source reliability metrics used, activation erroredUsers is no longer dropped. crash-rate date keys use UTC to match PostHog bucketing. The classic dashboard stops sending the invented desktop_cost/mobile_cost params and renders unmeasurable cost/user days as gaps, not $0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Guard: scope SCA-118 to inference traffic, not org-admin reporting The admin dashboard's cost legs read spend from the providers' organization namespaces (cost_report / organization/costs), which carry no model traffic. The gateway-only guard now exempts exactly those namespaces via lookahead — inference URLs still fail — with fixtures for both sides. provider-costs.ts stops naming the forbidden inference env vars in a comment. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: r <r@r> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> | 13 天前 | |
fix(ci): follow bounded apt-get into composite actions, and cap it at the caller (#12270) * fix(ci): follow bounded apt-get into composite actions, and cap it at the caller Rebased onto main. The desktop Swift lanes that held this approved PR were red on main itself; #12277 fixed both causes, so this rebase is what turns them green. The checks-manifest edit is re-applied to main's text rather than overwriting it, so the legacy-memory-surface-ratchet entry main added since the merge base survives. One prose fix in the reason field: #12194 is still open, so it proposes the composite action rather than having moved it. * fix(ci): restore the executable bit on check_workflow_apt_network_bounds.py The rebase onto main brought this file across at mode 100644. It is 100755 on main today, and this PR modifies it, so merging as it stood would have changed the file's mode on main -- a tree change no line of the diff describes. Mode only. The blob is byte-identical to the previous head: the tree is built from that tree with one entry, reusing the same blob sha. | 6 天前 | |
fix(preflight): decode Git content as UTF-8 on Windows (#10595) * fix(preflight): decode deployment Git output as UTF-8 * fix(preflight): decode deferred-work Git output as UTF-8 * fix(preflight): decode line-ratchet Git output as UTF-8 The product line-count ratchet read baseline paths and JSON blobs with the Windows host code page. Decode all three Git content reads explicitly as UTF-8 and lock both sharded and legacy paths with a subprocess contract test. --------- Co-authored-by: 491034170 <142008960+491034170@users.noreply.github.com> | 1 个月前 | |
fix: make no-changelog-needed survive the merge boundary The PR label greened PRs and then reddened main because push runs cannot see labels. Require an in-repo kind:none fragment for internal production desktop edits so both lanes agree. Failure-Class: FC-changelog-exemption-lost-at-merge Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: multica-agent <github@multica.ai> | 20 天前 | |
fix(ci): land leftover desktop-beta review fixes from #11588 (#11605) * fix(ci): redact recovery logs and keep qualification reintroduction tested Codemagic log tails could persist query tokens and GitHub PATs in durable issues, the freshness monitor lacked Actions/checks read, and the qualification-trigger mutation test was dropped with the old runner. Co-authored-by: Cursor <cursoragent@cursor.com> * style: drop extra blank line that failed diff-hygiene Co-authored-by: Cursor <cursoragent@cursor.com> * fix(ci): redact OAuth fragments without wiping triage query params Generic query redaction leaked `#code=` values and blanked harmless build/arch/mode fields. Scope assignment redaction to credential-like names, treat `#` as a delimiter, and drop the comment-only qualification mutation that locked in a substring false positive. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> | 23 天前 | |
harden(desktop): rehearse failed Codemagic releases (#11264) * harden(desktop): rehearse failed Codemagic releases * harden(desktop): sanitize recovery diagnostics portably * harden(desktop): redact recovery environment values | 30 天前 | |
fix: keep the macOS hourly release train moving Bind each candidate to its checked source, recover squash-orphaned changelog branches, install Bun for release eligibility, and age freshness from the oldest unshipped update. Verification: 77 focused release-control tests; actionlint; Python compile; diff check. Failure-Class: new | 18 天前 | |
fix(desktop-backend): admit the gateway route in the candidate probe #12337 routed company-paid desktop Gemini traffic through the LLM gateway; since then the desktop proxy stamps X-Omi-Provider: llm_gateway (backend/utils/llm/desktop_gemini_gateway.py), a value the Auto Deploy candidate probe's REAL_GEMINI_PROVIDER_ROUTES did not admit, so "Prove candidate chat compatibility" fail-closed on every promotion (run 33219770852) and dev desktop-backend stayed on the stale revision. Admit llm_gateway: the probed surface (gemini-2.5-flash generateContent) maps to the gateway's desktop-vertex-flash lane, whose primary provider is VertexGeminiProvider with no fallbacks (config_loader.py), so the hop is real Gemini-on-Vertex. Stubs, unknown, and empty routes stay rejected fail-closed; the offline-stub rejection test is unchanged and a new test pins the post-gateway route as admitted end to end through _gemini_request. Verification: python3 .github/scripts/test_desktop_backend_candidate_probe.py 19 tests OK (incl. test_gemini_probe_admits_post_gateway_llm_gateway_route). Failure-Class: FC-client-model-outside-proxy-allowlist Co-authored-by: multica-agent <github@multica.ai> | 9 天前 | |
fix(desktop): make activation and updater health honest Failure-Class: none | 27 天前 | |
ci: collapse desktop beta to signed-smoke plus hourly freshness (#11588) * ci: collapse desktop beta to signed-smoke plus hourly freshness Qualification never rolled back recent signed-smoke manifests and starved the planner when a push event was missed. Remove the lane, make source-gate failures diagnosable from one command, and alarm when candidate and live beta diverge. Co-authored-by: Cursor <cursoragent@cursor.com> * docs: retarget INV-BETA-1 guards after deleting qualification tests The auto-beta-candidate script is gone with the qualification lane; keep the locked beta-identity invariant pointing at a guard that still exists. Co-authored-by: Cursor <cursoragent@cursor.com> * ci: make skipped desktop checks legible and let a green tip unblock the train Two failure modes survived the beta-train collapse and are fixed here. Skipped jobs published the wrong check name. GitHub does not evaluate a job's `name:` for a SKIPPED job, so the conditional names on `desktop-swift` and `desktop-swift-release-compile` were published verbatim as the raw expression text. On every commit that did not touch desktop paths the required check `Desktop Swift Build & Tests` was therefore ABSENT rather than skipped, and the planner reported "missing" instead of the truth. Observed on f666ddd4a3, 7a79f08329 and 7d7ed62e5, all of which read green. The conditional existed to keep a merged `pull_request.closed` bookkeeping run from publishing a skipped required check onto the merge SHA; dropping the `closed` event removes that hazard at the source and lets both names be literals. A contract test now rejects any expression in a job name. A green tip did not unblock the train. The planner selects the newest desktop-touching commit and, when its checks are red, could only fall back to an OLDER green SHA. On Aug 14 main's tip was green while the newest desktop-touching commit below it was red on a flaky Swift suite, so the train shipped stale code or wedged. A first-parent commit above the blocked SHA contains everything the blocked SHA contains, so its own exact-SHA checks tested a superset of that tree; when they are genuinely green the train may ship from that newer SHA. Tried before the backward fallback, because it ships newer code. Only a real `ready` gate qualifies, so a skipped or absent check still never counts as success. Also keep the beta rollback precondition expressible: beta manifests carry the `signed-smoke` tier, whose frozen-schema truth is `qualification_passed: False`, so the literal T2/True requirement rejected every current rollback target. Failure-Class: none Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): keep the release lifecycle helper compiling under the admin es5 target `newestSparkleVersion` iterated `String.matchAll()` with `for...of`. The admin package sets "target": "es5", where iterating an IterableIterator is TS2802, so `npm run typecheck` failed and took the Web Checks Build job red. Local pre-push does not typecheck the Next.js admin app, so CI was the first place this could surface. Use `exec` loops instead of widening the package's compile target, which would change output for every file to fix one. Verified with the same commands CI runs: `npm run typecheck` clean and `npm test` 88 passed across 14 files. Failure-Class: none Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> | 23 天前 | |
ci: collapse desktop beta to signed-smoke plus hourly freshness (#11588) * ci: collapse desktop beta to signed-smoke plus hourly freshness Qualification never rolled back recent signed-smoke manifests and starved the planner when a push event was missed. Remove the lane, make source-gate failures diagnosable from one command, and alarm when candidate and live beta diverge. Co-authored-by: Cursor <cursoragent@cursor.com> * docs: retarget INV-BETA-1 guards after deleting qualification tests The auto-beta-candidate script is gone with the qualification lane; keep the locked beta-identity invariant pointing at a guard that still exists. Co-authored-by: Cursor <cursoragent@cursor.com> * ci: make skipped desktop checks legible and let a green tip unblock the train Two failure modes survived the beta-train collapse and are fixed here. Skipped jobs published the wrong check name. GitHub does not evaluate a job's `name:` for a SKIPPED job, so the conditional names on `desktop-swift` and `desktop-swift-release-compile` were published verbatim as the raw expression text. On every commit that did not touch desktop paths the required check `Desktop Swift Build & Tests` was therefore ABSENT rather than skipped, and the planner reported "missing" instead of the truth. Observed on f666ddd4a3, 7a79f08329 and 7d7ed62e5, all of which read green. The conditional existed to keep a merged `pull_request.closed` bookkeeping run from publishing a skipped required check onto the merge SHA; dropping the `closed` event removes that hazard at the source and lets both names be literals. A contract test now rejects any expression in a job name. A green tip did not unblock the train. The planner selects the newest desktop-touching commit and, when its checks are red, could only fall back to an OLDER green SHA. On Aug 14 main's tip was green while the newest desktop-touching commit below it was red on a flaky Swift suite, so the train shipped stale code or wedged. A first-parent commit above the blocked SHA contains everything the blocked SHA contains, so its own exact-SHA checks tested a superset of that tree; when they are genuinely green the train may ship from that newer SHA. Tried before the backward fallback, because it ships newer code. Only a real `ready` gate qualifies, so a skipped or absent check still never counts as success. Also keep the beta rollback precondition expressible: beta manifests carry the `signed-smoke` tier, whose frozen-schema truth is `qualification_passed: False`, so the literal T2/True requirement rejected every current rollback target. Failure-Class: none Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): keep the release lifecycle helper compiling under the admin es5 target `newestSparkleVersion` iterated `String.matchAll()` with `for...of`. The admin package sets "target": "es5", where iterating an IterableIterator is TS2802, so `npm run typecheck` failed and took the Web Checks Build job red. Local pre-push does not typecheck the Next.js admin app, so CI was the first place this could surface. Use `exec` loops instead of widening the package's compile target, which would change output for every file to fix one. Verified with the same commands CI runs: `npm run typecheck` clean and `npm test` 88 passed across 14 files. Failure-Class: none Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> | 23 天前 | |
refactor(release): unify macOS manifest contract | 1 个月前 | |
ci(mobile): make the batch baseline per workflow and only from settled builds Three defects in the scheduled batch, all found in review: - A failed or cancelled build still became the baseline, so its commit read as already built and the batch never retried it. A broken merge would sit unbuilt until some later app commit happened along. Only finished, success and skipped now advance the baseline; skipped counts because Codemagic looked and decided there was nothing to build. - A baseline that is not an ancestor of HEAD — a rewound or rewritten main — produced an empty range and skipped a batch that was genuinely pending. Any non-ancestor is now treated as pending. - One baseline (iOS) gated both platforms, so a manual iOS-only build from the Codemagic UI would suppress the Android build at that commit. Baselines and dispatch are now per workflow; only the platform with pending work is sent. Verified: 23 tests, up from 18. New coverage for failed/cancelled/in-progress baselines, skipped as a baseline, case-insensitive status, a real orphan commit built with git commit-tree for the non-ancestor path, and per-workflow pending counts in the summary. | 1 个月前 | |
fix(ci): preserve reconciler Cloud Run job arguments Quote the complete --args token for deploy-cloudrun and use one verified 100%-traffic revision during desktop backend rollback and recovery. Failure-Class: FC-agent-vm-bootstrap-contract | 1 个月前 | |
fix(desktop): gate backend candidates before traffic Failure-Class: FC-nonrecoverable-promotion Invariant: INV-DATA-1 | 1 个月前 | |
feat(ci): schedule the failure-class retirement the lifecycle already defines `scripts/failure-class report` computes closure eligibility with a quiet period and requires maintainer confirmation, but it is non-mutating and has no event source: without `--events-file` it returns zero events and a `no_event_source` warning. Nothing scheduled ever supplied that feed, so 26 classes sit at `status: open` with zero retirements since the registry was created. Adds a weekly Repo Hygiene workflow (Mon 09:00 UTC, clear of the release cadence) that builds the merged-PR feed, runs the report, applies the two-field retirement edit, and opens one PR on a reused branch. It never merges and never decides; the reviewable diff is the output, and merging is the confirmation. Three things this deliberately does NOT do: - gate anything: no manifest lane, no pre-push, no release coupling - open an issue for a retirement: a PR has a review queue and its own CI, an advisory issue joins a 400-issue backlog - author the PR with GITHUB_TOKEN, which would arrive with zero checks run; the Omi Bot app token is minted so the diff is verified like any human PR Recurrence of an already-dormant class exits 1 and skips the PR. That is judgment work (AGENTS.md requires a reusable guard surface, not a field flip), and a red weekly run is the notification without new machinery. The feed guard is the load-bearing part. A truncated feed makes every class report "no classified instance", so the job would run green forever while retiring nothing -- the same false-green defect as a check that executes no tests. The run refuses a feed that hit the fetch limit, or an empty one. Verification: python3 .github/scripts/test_failure_class_retirement.py -> Ran 15 tests, OK python3 .github/scripts/test_run_checks.py -> Ran 24 tests, OK (manifest valid) live dry run (real gh, 14d): 748 events spanning 2026-07-11 -> 2026-07-25, 292 carrying a Failure-Class line, 0 eligible, exit 0, no files modified time-travel apply (--now 2026-08-20, real registry copied to a temp root): 13 of 26 classes retired, 13 left open, diff is exactly - "status": "open" + "status": "dormant", + "dormant_since": "2026-08-20T00:00:00Z" per class (39 changed lines across 13 files, no formatting churn) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> | 1 个月前 | |
harden(backend): centralize Firestore read boundaries | 1 个月前 | |
fix(ci): make pre-push portable on Windows Failure-Class: none Route pre-push and preflight shell contracts through Git Bash / native tools on Windows so MSYS/Unicode paths don't break manifest runner, release guards, SwiftLint, firmware checks, and dev-harness scripts. Add _bash_command() and _native_path_from_bash() helpers with Git-for-Windows discovery. Decode Git output as UTF-8. Skip POSIX-only checks on Windows. | 1 个月前 | |
ci: add weekly guardrail baseline health pulse Record baseline counts via existing check counters, append JSONL history, and alert when a nonzero count has not decreased for 30 days (#9454). Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> | 1 个月前 | |
Refactor memory composition error flow | 1 个月前 | |
fix(ci): rename listen helm gateway mode choice off -> disabled GitHub Actions workflow_dispatch returns HTTP 422 for choice value literal "off", which blocked the two-step prod listen cutover step 1. Use disabled in the workflow UI/API and map it to env value off. Failure-Class: none | 30 天前 | |
Desktop core E2E harness, tiered testing, and release blessing gate (#9168) * Add desktop core E2E harness, flow lint, and T0 CI gate. Introduces the tiered desktop-core-harness dispatcher, YAML flow lint against registered bridge actions, CORE_E2E.md, and a Linux self-check job so desktop PRs get a fast static confidence gate. Verification: ./desktop/macos/scripts/desktop-core-harness.sh --self-check Co-authored-by: Cursor <cursoragent@cursor.com> * Add bridge seams for hermetic capture and data snapshots. Adds capture_test_transcript plus conversation/memories/tasks snapshot actions and last_assistant_text in main_chat_snapshot for deterministic harness assertions. Verification: desktop-flow-lint.py (action names registered in bridge source) Co-authored-by: Cursor <cursoragent@cursor.com> * Gate desktop prod promotion on blessed release metadata. Adds bless-release.sh for agent-local T2 blessing, requires blessedSha metadata in check-desktop-release-promotion.py with typed override escape hatch, and extends the promotion workflow/policy enforcement. Verification: python3 .github/scripts/test_check_desktop_release_promotion.py Co-authored-by: Cursor <cursoragent@cursor.com> * Show desktop release blessing status in admin UI. Parses bless keys from GitHub release KEY_VALUE metadata, fixes parseVersion for two-part macOS tags, and adds a Blessed column plus deployable banner on the releases dashboard. Verification: parseVersion regex accepts v11.0+11000-macos Co-authored-by: Cursor <cursoragent@cursor.com> * Fix floating-bar E2E flow assertion to match bridge detail. open_ask_omi returns focused timing fields, not askOmiOpen in action detail; state.expect still covers askOmiOpen on step 3. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix capture E2E seam honesty and compile errors from harness work. Wire capture_test_transcript through TranscriptionStorage with real session IDs, assert conversation_count_increased in the flow, strengthen chat marker checks, and document hermetic vs Omi Dev auth bootstrap in CORE_E2E.md. Verified: swift build, desktop-core-harness --self-check. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix bless-release.sh fatal bugs and hermetic T2 launch path. The manifest pass-check used `raise (X if cond else None)`, which raised TypeError on success; background desktop-run-local and wait for the worktree automation bridge; use make desktop-run-local (alice) instead of omi-auth-seed; discover --port via dev-instance.sh; consolidate KEY_VALUE edits into release-keyvalue.py; add cleanup trap. Verified: bash -n scripts/bless-release.sh; release-keyvalue.py self-test (pass manifest exit 0, fail exit 1, old pattern TypeError on pass). Co-authored-by: Cursor <cursoragent@cursor.com> * Fix desktop-core-harness CI gate, T2 hermeticity, and tier metadata. Split T0 CI to desktop-only static checks (--skip-backend-contracts) so the job no longer duplicates backend pytest; enforce offline dev-stack via config digest + health probes; derive flow lists from tier metadata; record provider_mode in manifests. Verified: bash -n; --self-check (8 pytest contracts); --self-check --skip-backend-contracts; workflow YAML parse; tier-1 metadata = harness-smoke + navigation; probe exits 0 on healthy offline stack and 2 on non-offline digest. Co-authored-by: Cursor <cursoragent@cursor.com> * Strengthen desktop E2E bridge snapshots and flow assertions. memories_snapshot and tasks_snapshot now await real API/store loads and return falsifiable load_state/count-valid fields; flows assert those instead of vacuous ok/refreshed. floating-bar-functional waits for chat idle and asserts the stub marker echo via main_chat_snapshot (shared ChatProvider). Document automationStopCaptureTestSession sync with stopTranscription(). Verified: xcrun swift build -c debug --package-path Desktop (pass); python3 scripts/desktop-flow-lint.py (pass, 22 flows). Live omi-harness runs blocked by concurrent ./run.sh lock contention; manual binary swap breaks Sparkle rpaths so full bundle rebuild needed for live flow runs. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix local harness auth stability and macOS bash 3.2 harness dispatch. Local profile no longer crashes on Auth.auth() or signs out when pi-mono requests a token refresh against the Auth emulator; skip RealtimeHub warm-up and Omi Dev auth seed for desktop-run-local. Replace mapfile (bash 4+) in desktop-core-harness.sh for macOS default bash. Verified: omi-core-e2e boots signed-in as alice on bridge :47956; tier-2 harness runs 9/13 flows pass (capture-lifecycle, chat*, floating-bar still fail). Co-authored-by: Cursor <cursoragent@cursor.com> * Harden bless cleanup, stack ownership probe, and manifest checks. Terminate omi-bless-* app and run.sh process tree on bless cleanup instead of only killing the launch subshell; require harness sentinel + alive owned PIDs before T2 accepts a dev stack; require provider_mode=offline in check-manifest and fail update-blessed when KEY_VALUE block is missing. Verified: bash -n bless-release.sh desktop-core-harness.sh; release-keyvalue.py self-test; probe fixtures (stale PIDs, foreign/absent sentinel → exit 1); cleanup process-tree simulation OK. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix LLM stub to echo markers from the latest user turn only. Dedupe repeated markers in one message and add a regression test so hermetic chat assertions see a single stub echo per turn. Verified: cargo test in Backend-Rust (301 passed). Co-authored-by: Cursor <cursoragent@cursor.com> * Isolate tier-2 chat flows with reset_main_chat harness action. Clear kernel main_chat surface state and UI messages before each chat flow so stub marker assertions are not polluted by prior turns in the same session. Verified: chat-hermetic + chat flows pass on bridge port 47956. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix floating-bar tier-2 assertions for Ask Omi state and chat snapshot. Report askOmiOpen for any open AI conversation (including response view), reset floating-bar kernel/chat on open_ask_omi reset, and use floating-bar snapshot/wait actions so assertions target the correct provider. Verified: floating-bar-functional flow pass on bridge port 47956. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix hermetic capture lifecycle finalization for sub-second sessions. Await segment persistence in the automation inject seam and guarantee finishedAt is at least one second after startedAt so from-segments uploads pass backend validation. Verified: capture-lifecycle flow pass on bridge port 47956. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix dev-up Typesense restart when container dies but manifest PID lingers. Restart unhealthy harness services before dev-up skips them, verify the Typesense Docker container in probe_dev_stack, and retry stack probes after dev-up so tier-2 dispatch passes reliably. Verified: PROVIDER_MODE=offline make dev-up; ./scripts/desktop-core-harness.sh --tier 2 --bundle omi-core-e2e --port 47956 --keep-stack (manifest passed=true, provider_mode=offline); release-keyvalue.py self-test; desktop-core-harness --self-check. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix desktop changelog fragments for pre-push validation. Co-authored-by: Cursor <cursoragent@cursor.com> * Address PR review: harden bless gate, harness safety, and LLM stub fidelity. Gate harness-only chat/capture actions on non-prod bundles, reject bless manifests without tier-2 offline evidence, fix promotion override semantics, and tighten admin blessed parsing. Co-authored-by: Cursor <cursoragent@cursor.com> * Address cubic re-review: harness errors, Gemini markers, capture teardown. - Parse Gemini stub markers from latest user content only; assert in test - Propagate harness chat-reset deletion failures to automation bridge - Clear capture-test state after finalize errors so retries can recover - Share release metadata parsing; restore override_unblessed blessedSha warning - Fix flow-lint typed-flow do-key detection; add promotion gate to lint CI Co-authored-by: Cursor <cursoragent@cursor.com> * Fix CI: OmiSupport imports and clippy unwrap in llm stub. Restore DesktopLocalProfile visibility for harness-gated code paths, replace Response::builder unwraps with response_or_500, and add stderr context to unblessed release promotion test assertions. Co-authored-by: Cursor <cursoragent@cursor.com> * Add e2e flow covers for local-profile Swift changes. Map changed AuthService, agent runtime, floating bar, and transcription files to tier-2 harness flows, restore check-e2e-flow-coverage.py, and align lint.yml with main's strict coverage gate. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix desktop Swift CI: split workflow and update auth guard test. Move the macOS Swift build/test job out of lint.yml into desktop-swift-ci.yml so PR checks no longer show two unrelated jobs under "Lint Check". Update AgentRuntimeProcessTests source-inspection assertions for the merged main auth refresh guard (forceRefreshToken + DesktopLocalProfile). Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> | 2 个月前 | |
fix(ci): fail closed when Codemagic desktop preview dies after dispatch (#10145) (#11171) * fix(ci): fail closed when Codemagic desktop preview dies after dispatch (#10145) Dispatch returning a buildId was treated as success while Codemagic could die at startup with no artifact, so macos.omi.me/preview/<slug> silently fell back to stable. Poll the exact build to a terminal status and keep durable evidence. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(ci): validate Codemagic build IDs before observe (#10145) Reject non-alphanumeric build IDs before GITHUB_OUTPUT, pass observe args via env (not ${{ }} in run) to avoid shell injection from provider data, and drop the unused sys import for Ruff. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> | 13 天前 | |
fix(release): restore native Codemagic tag delivery Failure-Class: FC-release-build-config-source-binding Co-authored-by: multica-agent <github@multica.ai> | 1 个月前 | |
SCA-402: stop firestore query-coverage baseline equality from flapping repo gates (#12562) * chore(ci): stop firestore query-coverage counts from flapping the ratchet Keep failing on new unregistered/unsupported shapes, stale debt entries, and registered-count or percentage decreases. Treat live counts above the committed file as a pass with a regenerate notice so two merges no longer race the baseline-equality check. Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: multica-agent <github@multica.ai> * refactor(tests): split BYOK security tests under the 800-line guardrail Move the 1749-line module into concern-grouped files sharing isolation fixtures, without changing assertions. Collect count stays at 105. Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: multica-agent <github@multica.ai> * docs(backend): add memory_ingestion architecture map Document each source file's role and non-goals, then drop the package from the architecture-guardrails grandfather baseline so the warning cannot return. Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: multica-agent <github@multica.ai> --------- Co-authored-by: David Zhang <9387252+Git-on-my-level@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: multica-agent <github@multica.ai> | 5 天前 | |
fix: enforce the macOS hourly candidate clock Failure-Class: FC-release-train-busy-main-starvation | 18 天前 | |
fix(preflight): keep Windows subprocess text UTF-8 (#12346) * fix(preflight): decode gh metadata as UTF-8 Decode captured gh PR JSON with the producer's UTF-8 encoding instead of the Windows host code page. Add a Unicode metadata contract and register the recurring subprocess locale failure class with #10595 and #10844 as evidence. Tests: MetadataTests passed 14/14; the real #10823 body loaded under preferred_encoding=cp936; Ruff, Black, py_compile, and diff checks passed. Failure-Class: new * fix(preflight): propagate UTF-8 through Windows runner * fix(preflight): pin suggestion subprocess UTF-8 * test(preflight): pin nested subprocess UTF-8 * fix(preflight): emit direct Windows output as UTF-8 --------- Co-authored-by: 491034170 <142008960+491034170@users.noreply.github.com> | 6 天前 | |
fix(preflight): keep Windows subprocess text UTF-8 (#12346) * fix(preflight): decode gh metadata as UTF-8 Decode captured gh PR JSON with the producer's UTF-8 encoding instead of the Windows host code page. Add a Unicode metadata contract and register the recurring subprocess locale failure class with #10595 and #10844 as evidence. Tests: MetadataTests passed 14/14; the real #10823 body loaded under preferred_encoding=cp936; Ruff, Black, py_compile, and diff checks passed. Failure-Class: new * fix(preflight): propagate UTF-8 through Windows runner * fix(preflight): pin suggestion subprocess UTF-8 * test(preflight): pin nested subprocess UTF-8 * fix(preflight): emit direct Windows output as UTF-8 --------- Co-authored-by: 491034170 <142008960+491034170@users.noreply.github.com> | 6 天前 | |
fix(deploy): gate public builds on browser canaries Move browser-public build configuration into a reviewed, versioned source and enforce its contract locally and in CI. Deploy each public target as a no-traffic candidate, prove the built client initializes in a browser, then promote only that revision. Verification: public-build fixtures and static contract check; deployment secret-boundary and concurrency checks; actionlint; targeted client lint/type checks (the web/app legacy next lint runner is incompatible with its installed Next version). | 1 个月前 | |
ci(public-build): env-var removal, per-environment flags, real invoker identity checks (#12582) * ci(public-build): env-var removal, per-environment flags, real invoker identity checks Plaintext Cloud Run env vars survived merge deploys, restricted ingress was applied to development, and TBD placeholders passed presence checks into gcloud. Co-authored-by: Cursor <cursoragent@cursor.com> * ci(public-build): probe actAs via IAM testIamPermissions REST gcloud has no iam service-accounts test-iam-permissions subcommand, so the previous preflight would fail every prod deploy. Call the IAM REST method with urllib and a print-access-token bearer instead. * ci(public-build): reject remove_runtime_env_vars overlapping preserved secrets Cubic review (PRRT_kwDOLkKqys6eWSS0): a runtime name in both preserve_runtime_secrets and remove_runtime_env_vars loaded cleanly, yet deployment emits --remove-env-vars for a binding the contract claims to preserve via the merge update strategies — the removal would strip the preserved secret binding. Extend the dedicated overlap rejection to preserve_runtime_secrets and its mirrored fallback_runtime_secrets, with a regression test. --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: David Zhang <9387252+Git-on-my-level@users.noreply.github.com> | 5 天前 | |
fix(ci): distinguish Windows exit 259 from liveness (#12347) Co-authored-by: 491034170 <142008960+491034170@users.noreply.github.com> | 6 天前 | |
fix(windows): gate releases on live update-feed route Windows builds since #10610 fail closed without /v2/desktop/update-feed/windows. Probe production before tagging so a missing backend deploy cannot mint another broken auto-update cohort. Failure-Class: new Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: multica-agent <github@multica.ai> | 1 个月前 | |
fix(deploy): accept load-balancer-only public-build candidates through their public URL (#12560) The frontend contract restricts Cloud Run ingress to the load balancer so the shared-chat rate-limit subject can trust X-Forwarded-For. That makes the tagged run.app candidate URL answer 404 to CI, so the browser smoke that gates promotion has failed on every prod frontend deploy since the flag landed, and h.omi.me has been serving the May build behind three months of merges. Resolve the acceptance route from the live ingress annotation: open ingress keeps the pre-promotion smoke of the no-traffic candidate; restricted ingress requires a declared public URL in the contract, is promoted first, smoked at that URL, and rolled back to the previously serving revision on failure. A restricted service without a declared URL is refused naming the ingress policy instead of "canary did not become ready". Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> | 6 天前 | |
fix: enforce the macOS hourly candidate clock Failure-Class: FC-release-train-busy-main-starvation | 18 天前 | |
fix: keep the macOS hourly release train moving Bind each candidate to its checked source, recover squash-orphaned changelog branches, install Bun for release eligibility, and age freshness from the oldest unshipped update. Verification: 77 focused release-control tests; actionlint; Python compile; diff check. Failure-Class: new | 18 天前 | |
fix(ci): retire superseded Windows release sync PRs (#10727) (#10960) * feat(ci): retire superseded Windows release sync PRs (#10727) Each Windows release opens a release/windows-v* sync PR to stamp desktop/windows/package.json back onto main. The release tag is authoritative, so older open sync PRs are pure review noise once a newer release has a PR; nine had accumulated. Add a testable Python helper that lists open PRs whose same-repo head matches the release/windows-v* prefix, excludes the current release PR, and closes the rest as superseded with a pointer to the newest. The selection predicate is unit-tested (current PR retained, unrelated heads never selected) and wired into the checks-manifest so the contract runs in CI. Cleanup stays best-effort and non-fatal: publishing and tags remain authoritative. Co-authored-by: CommandCodeBot <noreply@commandcode.ai> * ci(desktop-windows): close superseded version-sync PRs after release After the current release's sync PR exists, invoke the retire helper so older release/windows-v* PRs targeting main are closed as superseded. Failure is non-fatal: the release tag is already published and remains the source of truth. Co-authored-by: CommandCodeBot <noreply@commandcode.ai> * fix(ci): never close fork PRs and page past the list limit cubic review follow-up on #10960: the search matched any open PR whose head starts with release/windows-v*, which could include fork-origin contributor PRs this release job must not touch. Request isCrossRepository in the gh query, default the selection to same-repo only, and add a fork fixture to the contract test. Also pass an explicit --limit so cleanup does not silently stop at the CLI default (30) after a long outage or backlog growth. Co-authored-by: CommandCodeBot <noreply@commandcode.ai> * fix(ci): make --self-test run without release args The script advertised `--self-test` as a hermetic check, but argparse required --current-pr/--version even in that mode, so the documented invocation failed before reaching the test. Make those args optional and validate them only for the real cleanup path. Co-authored-by: CommandCodeBot <noreply@commandcode.ai> * fix(ci): list Windows sync PRs without head: search gh pr list --search head:release/windows-v returns zero same-repo results, so retirement became a no-op. List open main PRs and filter by headRefName prefix locally, with a regression test on the query args. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(ci): paginate Windows sync PR retirement listing gh pr list --limit 100 truncated when main has 100+ open PRs, so older superseded release/windows-v* sync PRs could be missed. Fetch all open PRs via gh api --paginate --slurp and keep local prefix/fork filters. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: CommandCodeBot <noreply@commandcode.ai> Co-authored-by: Cursor <cursoragent@cursor.com> | 10 天前 | |
fix: address remaining Codex P2 review comments (5 items) run-pre-push.sh: - P2: Add backend/*.py pathspec to catch root-level Python files (backend/**/*.py alone misses backend/main.py etc.) - P2: Strip leading 'backend/' from find output to prevent double-prefix in test target paths - P2: Guard --timeout flag behind pytest-timeout plugin detection, gracefully degrade if plugin not installed pre-push.example (same fixes synced): - Same pathspec, prefix-strip, and timeout-guard fixes run-lint.sh: - P2: Propagate fallback lint failures — replace || true with error tracking and exit non-zero if any tool fails .pre-commit-config.yaml: - P2: Add exclude patterns to check-json for known-invalid JSON (Flutter app JSON, Firebase config, IDE files, lock files) | 2 个月前 | |
make bug-class prevention part of CI Use one rename/delete/type-aware changed-path primitive, fail selective backend testing toward full coverage, run the complete desktop agent suite, dedicate import-isolation gates, and encode failure-class closure in agent and PR guidance.\n\nVerified: workflow/path contract tests (15); actionlint on all changed workflows; shell syntax checks; full desktop/backend component suites; independent final review approved. | 1 个月前 | |
feat(desktop): add clickable source citations (#11519) * feat(desktop): add clickable source citations * fix(desktop): harden citation source navigation Failure-Class: none * refactor(repo): derive line-count ratchet from PR base * fix(ci): preserve citation compile contracts | 24 天前 | |
ci(public-build): assert the candidate commit SHA in the promotion smoke (#12581) Bake GITHUB_SHA into the frontend canary so the public-URL promotion smoke can tell the new revision from the previous one behind the LB. Co-authored-by: Cursor <cursoragent@cursor.com> | 5 天前 | |
fix(admin): tidy the Omi TV hero row and drop the duplicate trials tile (#12210) The TV board is read at a distance, and four hero titles were long enough to truncate at TV width ("Users → 1M goal", "Net WAU growth WoW", "Signups — last 7 days", "Activation rate (macOS)"). Shorten them; the board and panel descriptions already carry the qualifiers the titles were spending characters on. "Trials in pipeline" (revenue route) and "Trialing" (subscriptions route) render the same number side by side — two tiles, one metric. Drop the revenue copy and give the three freed columns back to the health-ratio tiles so the row fills the grid again. Activation's description said "reaching the aha moment on day 1"; the Firestore route it reads measures a conversation within 7 days. Say what is measured, and give the mobile board its own description — that tile is repointed to PostHog telemetry, so it must not inherit the Firestore one. Also clamp the WAU chart at zero and thin the retention-curve x labels, both of which were unreadable on the wall. Co-authored-by: r <r@r> | 12 天前 | |
Repair PostHog telemetry ownership and coverage (#10660) * Repair mobile device lifecycle telemetry * Restore Omi device purchase intent telemetry * Track each permissions interstitial presentation * Report backend account deletion outcomes * Add macOS device pairing telemetry * Restore Windows PostHog delivery * Guard analytics emitter reachability * Define PostHog regression alert contracts * Document load-bearing analytics events * Add macOS pairing telemetry changelog * Cover device vendor mapping in desktop flow * Use typed pairing defaults in device tests * Stub the new telemetry module in account-deletion isolation tests Both files close the utils namespace (__path__ = []), so the added utils.integration_telemetry import in account_deletion.py resolved only against sys.modules and raised ModuleNotFoundError in CI. * Stop a Swift property modifier leaking onto the next analytics method The modifier scan walked back from func to the previous brace, so a 'private var x' declared directly above 'func setX' donated its private and the method was audited as an unreachable helper. Main's integrationConnect test seam hit exactly that shape on merge. Baseline records main's integrationConnect call-site counts and its test-only seam. * fix(telemetry): avoid deleted-UID PostHog identity and pin Windows host Account-deletion completion/failure telemetry now uses a service distinct_id with $process_person_profile=false so wiped Firebase UIDs are never re-identified in PostHog. Windows PostHog host is fixed to the CSP-allowed us.i.posthog.com origin. * fix(ci): black-format deletion telemetry and ratchet line counts Format account_deletion.py for black 26.5.1 and raise product-file line-count baselines for users.py and storage.py with justifications; ratchet OmiApp.swift down to the post-repair line count. * fix(ci): classify analytics test seams and re-baseline against main The reachability tripwire audited `set*TelemetryCaptureForTests` as if it were a production emitter, so every new scoped test seam had to be hand-added to `public_orphans` — main's `setSuggestionAssistantTelemetryCaptureForTests` failed the check for exactly that reason. Installing or forwarding to a test capture is production-unreachable by construction, so `emitters()` now drops `*ForTests`/`*ForTesting` methods before the audit and the five seam entries leave the baseline. The baseline also predated main. Regenerating it records main's own drift: 25ea86a1d2 consolidated the onboarding Google-connect flow through ConnectorImportRunner, dropping the four SBOnboardingModel+Steps emit sites (integrationConnectAttempted 3->1, Succeeded/Failed 2->1 — each still requires its remaining call site), and main's live-suggestion work grew suggestionAssistantGateOutcome to 3 and suggestionAssistantDeliveryOutcome to 2. Only those ten entries moved; nothing else was absorbed. Verified: - python3 .github/scripts/check_analytics_reachability.py -> "analytics reachability static tripwire passed" - python3 .github/scripts/test_check_analytics_reachability.py -> Ran 7 tests, OK (new test_test_only_seams_are_not_audited_as_emitters) * fix(desktop): assert the pinned PostHog host at runtime, not in source text The host-pin regression scraped analytics.ts and asserted the file never contains "VITE_POSTHOG_HOST" — which the comment explaining why the override was dropped also matches, so the check failed on its own explanation. It was a static tripwire either way; it never proved the override was inert. It now stubs VITE_POSTHOG_HOST to an origin the renderer CSP does not allow, re-imports the module, and asserts fetch still goes to us.i.posthog.com. Verified in desktop/windows: - npx vitest run src/renderer/src/lib/analytics.test.ts -> 6 passed - restoring the `import.meta.env.VITE_POSTHOG_HOST ||` fallback in analytics.ts fails it with Received "https://not-in-csp.example.com/i/v0/e/", so the test fails for the reason it claims - npx prettier --check on the file -> clean --------- Co-authored-by: Max Carter 祁明思 <136312656+undivisible@users.noreply.github.com> | 1 个月前 | |
Make automatic development backend deploys converge (#12019) * Make automatic development backend deploys converge Automatic development deploys succeeded 8 times in the 25 runs before this change. The 13 failures had three causes, and this addresses the two that are defects rather than configuration. Admission required the Release Eligibility proof SHA to still equal main's tip. Anything merging while eligibility ran therefore rejected a merged, reviewed commit -- 8 of the 13 failures, and why development sat a day behind main. The property that protects the runtime is that the commit is merged, so require ancestry instead. Production is untouched: it deploys only by explicit dispatch, which already required ancestor-of-main rather than tip-equality. Development also had no automatic Firestore migration path. An automatic index reconciliation resolves its environment to prod, and only a manual dispatch ever targeted development, so a merged manifest addition left development with a schema that no longer matched main and every deploy failed its readiness gate until somebody noticed -- 2 more failures, most recently the hourly_usage (year, month) index from #11979. Composite reconciliation is create-only and development carries no required reviewer, so it now converges on the same merge that queues the production migration. Production's approval gate is unchanged. The manual development lane failed separately, at custom_token_signing: its candidate audio gate authenticates against production Firebase, which a development deploy identity cannot sign a custom token for. The probe can now be told which account to sign as. Left unset the behaviour is identical, so this is inert until FIREBASE_PROBE_SIGNER_SERVICE_ACCOUNT is set and the deploy identity is granted token-creator on it. Not addressed here: 3 failures came from GCP_FIRESTORE_READONLY_CREDENTIALS being intermittently unavailable in the development environment. That credential is a deliberate privilege boundary -- readiness executes admitted source and must not hold deploy credentials -- so it wants a configuration fix, not a code change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Update release-vector contract for the development index lane The static migration contract counted --provision-missing across the whole workflow, which asserted 'only one lane applies indexes'. There are now two, one per environment, so count per job instead and pin the development lane to its own environment, concurrency group, and push-only trigger. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Resolve the newest proven main source instead of the triggering one gpt-5.6-sol's review found the previous approach incomplete in two ways, and both are real. Ancestry alone was not safe. Tip-equality was doing more than proving merge status -- it was also a currentness fence. Accepting any ancestor of main lets a late-scheduled run deploy older code than development already had, because Actions concurrency groups are not FIFO, and lets a run for a commit that has since been reverted redeploy the pre-revert tree. This runtime shares production Firestore, Firebase auth, and Stripe, so that is not benign. Ancestry alone was also not sufficient. The scope job green-no-ops any triggering SHA that main has moved past, before it ever inspects changed paths. So a backend commit still never deploys if an unrelated commit merges before scope runs: the backend commit no-ops for being behind, the unrelated commit no-ops on its own diff. The regression test claimed to cover this but built a later main SHA and never passed it to scope, so scope saw the backend commit as main's tip and the assertion proved nothing. Passing it reproduces the strand. Both follow from deploying the triggering commit. Admission now resolves the newest commit on main carrying a first-attempt successful Release Eligibility proof and reachable from current main, and deploys that. Concurrent runs converge on one target rather than racing, a revert is never undone by a late run for the commit it reverted, and a behind trigger still deploys because the target moves forward instead of the run being skipped. Scope's supersession decision is removed as now-redundant, which also deletes its two GitHub API proofs and their fixture -- the contract gains tripwires against reintroducing it. --trigger-is-ancestor-of-sha keeps the resolved target at least as new as the proof that triggered the run, so a stale listing cannot move development backwards from its own trigger. The proof listing is fetched with curl --fail and no error suppression: an unreadable listing refuses to deploy. The guard checkout assertion is gone rather than re-checked-out. sol was right that it had become true by construction and added no independent evidence, and the re-checkout it needed also made an in-flight run execute a newer guard script than the workflow that invoked it. Readiness now needs actions:read to list proofs. The manual lane's readiness job already had exactly that for exactly this lookup, so the contract now expects it for both rather than treating the automatic lane as more restricted. sol's P0 -- that the new development index lane writes to production -- does not hold: RUNTIME_GCP_PROJECT_ID is based-hardware-dev in the development environment and based-hardware in prod, so the two jobs target different projects and cannot race on the same index. It read the value from runtime_env.yaml's runtime_gcp_project rather than the deployed variable. That inconsistency between the checked-in contract and the deployed value is real and worth its own look, but it is not this lane writing to production. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> | 15 天前 | |
merge: integrate current main into web parity Resolve the manifest and brand-check conflicts while retaining both branches’ active CI coverage. Tighten the hex token boundary so the merged brand test rejects embedded literals. Failure-Class: none | 18 天前 | |
One chat shell: every content block renders everywhere, plus chat ergonomics (#12607) * refactor(desktop): mount one chat shell for every account DesktopHomeView held the app on a "Preparing Omi…" card until a network call decided which of two shells to mount, then rendered either ChatFirstShell or a legacy sidebar + DashboardPage tree. Both were the same product with different chrome, and the legacy branch was the only reason `useLegacyHomeDesign`, `useOldestHomeDesign`, DashboardPage's inline chat, SidebarView, and the widget hub still existed. The shell now mounts immediately for everyone. The server-owned capability still resolves — same request, same analytics event, same ChatProvider projection gate — but alongside the mounted shell rather than in front of it, and it now only decides whether the capability-gated kernel features engage. Capability-off renders the same shell. `navigate help` named a "Help from Founder" page no shell had mounted for a long time: the bridge resolved a title and then timed out. It now resolves to Settings → About, where getting help from a person actually lives. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop): render every content block as an interactable component Six of the journal's block kinds — question card, task card, goal link, capture link, conversation link, memory link — were dropped on the floor by every Chat surface except one. `ContentBlockGroup.group` skipped them unless `richBlockRenderingEnabled`, and `ChatBubble.blockView` returned `EmptyView` for each of them again. A turn whose whole content was a task card therefore read as an empty assistant reply in the task panel and in the notch, and as a card you could tick off in the main window. `ChatFirstRichBlockContext` is now non-optional on `QueryShellHome`, `QueryAnswerThread`, `ChatMessagesView` and `ChatBubble`; the task panel and the floating/notch renderers bind the shell's process-wide owners through `.auxiliary`, so a card tapped in the notch summons the main window and routes the one shell. `ChatFirstRichBlockGroupView` is the single renderer all three hosts share. Capability-off degrades rather than disappears: cards render, task check-off works (it binds `TasksStore`, not the projection), links navigate, and a question card shows its options dimmed and unpressable with an explicit "Answering is unavailable right now" line, instead of a question with no visible answers. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): every proactive card says what it is `showNotification` took an optional `kind:` and `FloatingBarNotification` quietly filled it in from `assistantId`, whose default arm is `.general`. Five producers never passed one — trial messaging, onboarding permission help, both notch moments, and the whole generic proactive path — so their cards journaled a bare `notification:<uuid>` continuity key and came back in the transcript badged "Notification" with a bell, a row that says nothing about what Omi actually noticed. `kind:` is now required and never derived inside the value type. The generic proactive path derives it once at the producer edge, from the same `from(assistantId:)` call the category gate already makes three lines earlier. `.trial` and `.onboarding` are new kinds and are excluded from journaling alongside the integration nudge: billing copy and permission help are not observations. `.functional` carries the system notices that used to ride on `.general`. `.general` survives as decode-only so historical bare keys keep reading back, and its badge arm stays for exactly those rows. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(desktop): pin one shell and six live blocks, and retire the second-shell flows The tests that protected the old shape are the reason it would come back: `ChatFirstRichBlockTests` asserted that a caller without an explicit context got *nothing*, and three flows waited on `shellVariant: legacy`. - `OneChatShellRichBlockTests` builds one turn carrying prose plus all six interactable kinds and asserts the grouping keeps every one of them, in transcript order, on the same entry point the notch and task panel call. It then drives each link's typed navigation target through the real navigation owner, and pins capability-off to "options dimmed", never "options gone" — via a new `ChatFirstQuestionCardOptionsPolicy` that separates *answered* and *retired* (hide) from *capability-off* (disable). - `ProactiveNotificationKindTests` walks every assistant id a producer ships and proves none of them derives `.general`, so no producer can mint a bare `notification:<uuid>` key; historical bare keys still decode. - `check-single-chat-shell.py` (+ manifest entries, with a self-test) is the tripwire for the vocabulary that made a second shell expressible. - `home-stage.yaml` and `dashboard.yaml` described `DashboardPage` and are deleted with it; `chat-first-capability-isolation.yaml` is repurposed to the assertion that now matters — capability-off mounts the same shell. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): the notch and the main window share one navigation owner `ChatFirstRichBlockContext.auxiliary` binds `ChatFirstShellNavigation.shared` so a card tapped in the notch or the task panel routes the shell. The root was still creating its own instance, so those taps would have moved a navigation object nothing rendered — the card would appear to do nothing. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(desktop): delete the intelligence store nothing renders any more `DashboardIntelligenceStore` fetched recommendations, projected them, and kept a feedback outbox for a section that only `DashboardPage` mounted. With the page gone its 750 lines had no renderer and no caller — the grep is exact: the class was referenced only by its own file and its own tests. Its 1,117-line test file went with it, because a test for a store nothing mounts is coverage of nothing. What stays is the part other surfaces still use: `TaskNavigationRequestStore`, the exact-record handoff `QueryShellHome` and the chat-first task card give the Tasks page instead of a tab index. The file is named for it now. Three source-reading tests still pointed at `DashboardPage.swift` and failed on the missing file rather than on anything real; they move onto `QueryAnswerThread` / `QueryShellHome`, which is where Home's error card and its colour tokens actually live. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): an explicit settings section still wins over the help default `navigate help` pre-selects About because that is where getting help from a person lives. A caller that also names a section meant that section. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): let a settled chat answer be selected, copied and seen as cut off Three things a reader could not do with a reply on screen. Select it. `OmiMarkdown` disabled native text selection outright, so a date or a name in an answer could only be retyped. The reason was real — one AppKit selection overlay per `Text`, on a row that rewrites its body every streaming flush, is a non-converging layout loop (400 segments, 2 s hangs) — but it is a reason about *streaming* rows. Selection is now opt-in through `\.chatTextSelectable`, and `ChatTextSelectionPolicy` grants it to settled rows only, on every surface that shows chat prose: the transcript, the notch, the expanded floating bar, and onboarding. The three `.textSelection(.enabled)` calls outside `OmiMarkdown` in the floating surfaces were dead — the inner `.disabled` won — and are replaced rather than left as decoration. Copy it without hunting. The copy button lived only in the hover-revealed strip, and a user turn had no copy affordance at all. Every row now has a "Copy Message" context menu over the same pasteboard write, and the copy button takes ⌘C while its row's strip holds keyboard focus — not window-wide, which would take the shortcut from selected prose and the composer. See that it stopped mid-sentence. A voice barge-in persists the partial answer with a terminal failed status, and the only failure affordance was a stamp for a row with no text — so "…arrive on Saturday," rendered exactly like a finished reply. `ChatTurnFailurePresentation` decides between that stamp and a quiet trailing "Interrupted" mark, and the empty-row case is unchanged. Also: the hover strip is now `accessibilityHidden` when it is invisible (opacity and hit-testing hid it from the eye and the mouse but not VoiceOver), each row carries a You/Omi label, and a row whose whole content is a rich block reserves no metadata band — a memory card stamps its own time and has nothing to copy or rate. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): stop charging the transcript twice for the metadata band Two consecutive one-line answers sat roughly 100 device pixels apart, and a memory card floated in symmetric dead space. Measured on the real views: 44 pt between two settled replies, and a 48 pt card row inflated to 69 pt. Two causes, both double-charges. The hover strip is 28 pt of real reserved height under every settled reply. The stack then added a full 16 pt inter-exchange gap on top of it, so the separation the band already provides was paid for twice. `ChatTranscriptLayout.spacing` now asks whether the row above reserves a band and takes a hairline when it does; every other rung of the ladder is unchanged, and a reply still binds to its question more tightly than to the next exchange. `ChatOmiMarkPlacement.rowHeight` reserved 32 pt on every assistant row for a mark that only needs it when an empty streaming reply has no height of its own. On a settled row the reservation did nothing but centre short content in a box taller than itself — which is what put equal dead space above and below the memory card. It now applies while streaming, top-aligned. Measured after: 32 pt between two replies, 48 pt for the card row, and a five-row transcript 296 pt tall instead of 341. Also collapses adjacent repeats. Dedup only ran on messages over 200 characters, so three push-to-talk tries at the same ~90-character question stuttered down the transcript untouched — each press mints a distinct `voice:<uuid>` turn, so those are three legitimate journal rows and journal identity is not the place to fix it. `adjacentDuplicateIDs` collapses a short answer repeated in the row immediately below it within ten minutes, and folds a failed barge-in fragment into the answer it is a strict prefix of. It stays behind the existing expandable "Duplicate message" chip, so nothing is hidden outright, and non-adjacent, distant, or cross-sender repeats are left alone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(app): decode chat content blocks into typed models Mobile stored `content_blocks` as raw maps and, since #12015, hid any message whose blocks were only desktop chat-first chrome (goal/task/ question) because there was no renderer for them. Both halves are now wrong: the components are coming, so the schema needs a typed projection and the hide filter has to go. Add `ChatContentBlock`, a sealed model mirroring the canonical schema in `desktop/macos/agent/src/runtime/types.ts` and the Swift codec's required-field rules, decoding both the camelCase (desktop/agent) and snake_case (chat-first spec) dialects. Malformed blocks are dropped; unknown types become `UnknownContentBlock` so the message keeps its synthesized fallback text instead of losing content. The raw list stays authoritative on the wire — `toJson` is unchanged. Delete `hideFromMobileChat` / `visibleOnMobile` and their four call sites in MessageProvider, replacing the "is this body only the fallback dump?" test with `textIsStructuredFallback`, which the renderer uses to decide whether components replace the body or sit beside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(app): add l10n keys for chat content-block components Ten new keys for the block eyebrows, destination actions, unavailable state, and the conversation link's recommended-steps header, translated into all 48 non-template locales. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(app): render chat content blocks as interactable components Every block type that macOS renders as a control now has a mobile component, driven from the same wire schema (#12598): - taskCard: a live checkbox wired to the single tasks mutation path, ActionItemsProvider.updateActionItemState. The tasks API is list-only, so the card resolves against the loaded list and mirrors the macOS loading / unavailable states rather than inventing a fetch-by-id. - goalLink: mobile has no goal detail route, so the card opens a bottom sheet with the goal's title and progress resolved from GoalsProvider, and shows the unavailable state when the id is not in the list. - captureLink / conversationLink: push ConversationDetailPage through the citation preamble already shipped in chat (grouped-map hit, then fetch by id). conversationLink also lists its recommended action items as plain rows; mobile creates tasks from the tasks surface, so the block mutates nothing. - memoryLink: opens the existing memory sheet for the resolved memory. - questionCard: options send their preparedAnswer down the normal chat send path, so the runtime stays authoritative for what an answer means. A deferral option is not special — it sends its own prepared answer. Once selectedOptionId is set only the chosen option remains, disabled, so no stale chip ever looks tappable. text/thinking/toolCall/discoveryCard/citation/agentSpawn/agentCompletion and unknown types render nothing extra — the body (or its synthesized fallback) already carries them — but they never hide the message. Where the body is only that fallback, the components replace it instead of repeating it. Every interactive element carries Key('chat-block-<type>-<id>...'). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * docs: record the presentation-cohort-drops-journaled-content failure class Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): give the reader a selectable copy instead of a selectable transcript The selection change this branch shipped is reverted. `OmiMarkdown` disables native text selection again at both sites, explicitly, and the `\.chatTextSelectable` environment key, `ChatTextSelectionPolicy` and `OmiChatTextSelectability` are gone along with the host wiring in `AIResponseView`, `FloatingControlBarView` and `OnboardingChatView` — those hosts now carry no `textSelection` modifier around `OmiMarkdown` at all, since the ones that were there before this branch were dead code under the inner `.disabled`. The settled-row gate was not enough. PR #10834 made the same argument and reopened FC-selection-overlay-layout-loop in Omi Beta 0.12.146: every sampled main-thread stack sat in `SelectionOverlay`, `setFont`, intrinsic-size invalidation and AttributeGraph, and memory grew without bound. A settled row is still rebuilt by transcript loading, scrolling, window resize and parent-state updates, which is all that loop needs. `.github/scripts/check_chat_selection_boundary.py` rejects the escape hatch and names the remedy: the existing copy actions, or a separate non-live reading surface. So this adds the reading surface. "Select Text…" sits on the row's context menu next to "Copy Message" and as an ibeam button in the hover strip, on assistant and user rows alike. It opens `ChatSelectableTextPopover`: one `NSTextView` over one message's copyable text — `isEditable` false, `isSelectable` true, ⌘A and ⌘C native, Escape closes, sized to content with a 360 pt cap and internal scrolling. It is outside the transcript's layout and does not mount until the reader asks for it, so it cannot take part in the loading, scrolling and resize passes that made in-place selection unsafe. No SwiftUI `textSelection` anywhere in it — AppKit selection is what an `NSTextView` already is. The strip now carries four controls plus the timestamp (thumbs, thumbs, copy, select, info-when-present). At 24 pt each that still leaves the timestamp its own room, so both affordances stay rather than context-menu only. Everything else on this branch is unchanged: right-click Copy, focus-gated ⌘C, the metadata-band policy, the transcript rhythm, the interrupted-turn marker, adjacent duplicate collapse, and the accessibility fixes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(desktop): drop the deleted shell files from the static guards Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(desktop): repoint ptt-lifecycle covers at ChatToolExecutor RealtimeConversationToolProjection.swift was folded into ChatToolExecutor in 758cd3f1fb and the flow kept the stale path, which fails desktop-flow-lint. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): a reply the reader watched arrive stays whole when it settles `ChatBubbleTruncation` clamps any body over 500 characters, and only once `isStreaming` goes false. So an answer rendered in full while it streamed — with the transcript following it down — collapsed to its own first paragraph the instant it finished. A forty-item list became three items and a "Show more", and the document shrank by thousands of points under a reader pinned to the live edge. Watching that happen reads as "the chat stopped scrolling": everything you just followed is taken back at the end. Truncation is for restored history, where a long transcript should not be mostly one old reply. An answer that just settled here is the opposite case, so it keeps its full body. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(desktop): the assistant reads the same task list the Tasks page shows Asking by voice what was on the list answered "you don't have any tasks overdue or due today" to someone looking at thirty of them. `get_tasks` is backed by `TasksStore.loadDashboardTasks`, which narrowed the list twice in ways the Tasks page never has: - a seven-day recency window, as a lower bound on overdue due dates and as a `createdAfter` cutoff on undated rows. The page buckets on `dueAt < startOfTomorrow` alone and ages nothing out, so a backlog older than a week was invisible to the assistant and only to the assistant. - a source filter that dropped every AI-capture row. That gate was written when a capture could still land in `action_items` unreviewed. INV-TASK-2 has since made capture suggestion-only — `TaskCaptureModePolicy.usesLegacyStaging` is false for every mode, and a capture stays a Candidate until an explicit gesture accepts it — so the gate no longer separated reviewed from unreviewed. It hid the user's own backlog, and anything they created by voice, since `create_action_item` comes back stamped `conversation`. On the reporting account those two took thirty visible tasks to zero. The buckets now carry what the page carries: 82 overdue + 3 due today against the page's "Today: 85". The per-bucket cap moves 50 → 500, because the count is spoken and 50 would have understated it. `isPendingSuggestion` stays for proactive nudges, the one consumer still asking whether a capture pipeline wrote a row. `DashboardTaskLanePolicy` had no other caller and goes. The voice tool descriptions said "overdue + due today", which is where the spoken wording came from; they now describe the whole open list. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(desktop): the transcript's own follow-scroll is not the reader taking it Found while tracking down the streaming-scroll report. `UserScrollDetector` promotes an open mouse press to reader ownership as soon as the clip view moves, because a scrollbar-track click repositions the viewport without ever emitting a drag. That test cannot tell the app's own follow-scroll from the reader's, and while an answer streams the transcript re-reaches the live edge every `ChatScrollFollowThrottle.interval` — so a press still open when one lands ends follow mode for the rest of the answer. Now that every content block is something you can click, a press inside a streaming transcript is ordinary. Two lifecycle holes alongside it: a release delivered to another window — the "Select Text…" popover and context menus present in their own — was dropped by the same-window guard, leaving the press candidate open for the life of the scroll view with every later follow-scroll able to promote it; and a second press registered its bounds observer without removing the first. The transcript now records when it moves its own viewport, and movement inside that window re-baselines the press instead of promoting it. A drag and a scrollbar-track click still take ownership, which the existing harness cases pin. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(desktop): the tasks flow checks the lanes the assistant answers from The Tasks page and `tasks_snapshot` read the same rows now, so the flow can say so: S6 compares the page's Today count against overdue_count + today_count. That is the comparison the bug failed — thirty tasks on the page, zero in the lanes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(desktop): the components the agent renders reach the transcript `render_chat_blocks` had never once been offered to the model on this Mac, and on the turns where it did run its cards were deleted three seconds later. Four separate gates, one confusion between them: *which surface is this*. One shell means main Chat and the floating bar project the same conversation, so a session the bar registered carries `surface_kind = floating_chat` while main-Chat runs execute on it. Every chat-first gate admits `main_chat` only, and three of them read the session's registration instead of the run's: the adapter metadata that sets `OMI_CHAT_FIRST_UI`, the tool-capability broker that admits the call, and the surface Swift re-validates before executing. The fourth was separate — run admission built its own context snapshot and dropped the capability on the way, because the capability map lived on `KernelSessions` where run admission could not reach it. Any one of the four was enough on its own; advertised tools go 41 -> 46 with all four fixed. Then the cards died anyway. `monotonicAcceptContentBlocks` protected exactly two block kinds across terminalization, and terminalization applies the projection *Swift* assembled from the adapter stream — text and tool calls, never a block the agent appended mid-turn, because that append is a journal mutation the surface never saw. So the replace erased every task card, goal link and memory link the turn had rendered, moments after the tool reported `ok`. The protected set is now the kinds the kernel authors and the surface never does. Confirmed live: the tool executes, and three `taskCard` blocks now survive on the finished turn. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(desktop): ticking a task card shows it done, not gone Checking a box replaced the card with "Task is no longer available" — the one message that means the row is not the user's any more. Nothing had gone away. Ticking moves the task between the store's two arrays and the toggle awaits SQLite first, so `liveTask` is briefly nil. That flips the card's `hydrationKey`, SwiftUI cancels the in-flight hydration, and `TasksStore.isCurrent` folds `!Task.isCancelled` into its lease check — so `resolveCanonicalTask` then returns nil *by construction*, not because the task is missing. The old code published that nil, cleared the retained task and set `hydrationFinished`: exactly the pair that draws the unavailable placeholder. A superseded hydration now speaks for nothing, and a store that cannot vouch for a row is no longer read as the row being retired. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(chat): components replace the writing instead of doubling it Two ways the same answer was being said twice. The tool's own instructions were the first. They told the model to render components "whenever you retrieve, create, or summarize" those entities and "do not leave them as a Markdown table/list" — so a summary of yesterday stacked three conversation cards above the prose that already said it. That is backwards: reading an entity to answer a question makes it a *source*, and sources belong in citations. A component is for the entity that IS the answer — the one the user asked to see or act on, or the one this turn created or changed — and when components are rendered they are the list, so the prose above them is one lead-in sentence at most. The second was ours. A turn that carried only components still printed the blocks' own degradation text underneath them, so three cards sat above the three lines they were made from. `ChatStructuredFallbackText` mirrors the producer case for case, and a body that is only that projection is recognised as the cards talking to themselves rather than as answer text. Mobile drew six of the nine kinds the desktop transcript draws and silently skipped the rest. Discovery cards and the two agent-run blocks have components now, and a parity test fails if the desktop grows a tenth without one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(desktop): selection lives in the words, not in a box beside them "Select Text…" opened a popover that re-printed the message in raw Markdown next to the row the reader was already looking at, and only assistant rows had it — a user turn has no hover strip, so their own words were never selectable by any means. SwiftUI's selection stays barred: PR #10834 put it back on settled rows and reopened FC-selection-overlay-layout-loop in Omi Beta 0.12.146, every sampled main-thread stack in `SelectionOverlay` and `setFont` while memory climbed. But that bar is on `SelectionOverlay`, not on selecting. An `NSTextView` *is* one selection: one view owns it, a rebuild replaces a string, and nothing per-`Text` is mounted for a parent to thrash. `ChatSelectableTextPopover` said so in its own header — it just kept that view outside the transcript. So the words are the surface now. Chat prose renders through one text view per block, parsed from exactly what the SwiftUI renderer parses, and the reader drags across an answer in place with ⌘C copying what they highlighted. Citation markers become link ranges rather than buttons, which is what made the surrounding line selectable at all — a chip in a flow layout forced the prose to be chopped into per-segment views — and they still open their source and still preview it on hover. Two things had to be got right. Height is measured beside the live view rather than inside it, because a container left holding a measurement width wrapped the answer to a width the transcript never granted; and the measurement is synchronous, because publishing a height a frame late broke the transcript's follow-scroll. The gesture harness reads pixels, and `cacheDisplay` stopped seeing prose once it was drawn from a layer, so the probe composites the layer tree and can measure the row again. The boundary check now guards this file too: the AppKit path may never quietly acquire the SwiftUI one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(chat): a rendered card is the citation, and it survives the next update Two leaks left over from making components reachable. The first: terminalization was not the only replace. The streaming projection pushes the surface's own block list several times a turn, and `updateJournalTurn` took the replacement literally — so cards that survived the terminal commit died to the very next update instead, which is why they appeared on one turn and not the next. Both paths now apply a projection over the journal rather than in place of it: an id the projection carries is the projection's to define, which is how a question card's options still get retired, and an id it omits survives only when the kernel wrote it. The second is what the reader saw. When the model renders components and skips inline markers, we append a compact `Sources: [1][2][3]` rail so provenance stays discoverable. That was a sensible garnish when components were rare; now that a component turn is one lead-in line and three cards, the rail is a row of bare markers printed under the very things they point at. Sources the turn already draws are dropped from it, and anything with no component of its own keeps its marker. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(desktop): every citation marker the transcript writes is selectable and live The AppKit prose path matched `[\d{1,3}]` of its own invention. Ordinals run to four digits and the model also writes the kind beside them — `[5004]`, `[memory 5023]` — so real markers rendered as dead text in the middle of an answer. It uses `ChatCitationMarkup.numericMarkerPattern` now, which is the transcript's own definition of a marker rather than a second copy of it. Also pins what the harness proved by hand: the mounted transcript hosts a selectable text view on rows at more than one inset, so it is both senders' words that can be dragged across, not just Omi's. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(desktop): the cohesive chat flow covers the selection renderer that replaced the popover Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(ci): the selection boundary fixtures cover the AppKit surface too The remedy is an NSTextView owning its own selection; a SwiftUI rewrite of that file would put SelectionOverlay back in the transcript under a name no pattern check can see, so the fixtures pin that rule alongside the existing ones. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(agent): the component cap is not a licence to write the list out instead Asking to see your tasks and getting seventeen of them numbered in prose is the shape components exist to replace; the cap bounds how many are drawn, not whether any are. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(agent): tasks the user asked to see are rendered, not written out The render-vs-cite rule lived only on `render_chat_blocks`, where the model reads it after it has already decided what to write. Asking "what are my open tasks right now?" still returned seventeen of a hundred as a numbered list — prose that names the tasks but cannot tick one off, which is the shape components exist to replace. The rule now sits on `get_action_items` too, at the moment the tasks arrive: if the user asked to see, review, pick from or work through them, render the few that matter and say how many more there are; if a task is only evidence for a question answered in prose, cite it and render nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(agent): a rendered task list is counted, not named The `get_action_items` guidance said to render the cards and "say how many more there are". The model read *say* as *list*: it rendered three task cards and bulleted the same three titles above them, printing every task twice — once as words that cannot be ticked off and once as the card. Say the count, never the names, and say why: naming them is what doubles the answer. Also corrects the OmiMarkdown header comment, which still described the deleted selection popover. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(desktop): stop the Removed lane tombstoning live tasks Ticking a task on a chat task card replaced it with "Task is no longer available" — the reader's own completion, erased. The tombstone was local and it was fabricated. `fetchDeletedPage` asks for retired rows with a `deleted=true` query item that `GET /v1/action-items` has never had: FastAPI drops the unknown item, and that handler's stream skips soft-deleted documents outright, so the "deleted lane" answered with the owner's live first page. Every row came back stamped `.retired()` and was synced into SQLite, so each visit to Removed tombstoned another hundred live tasks. Completing one read the tombstone back — 190 ms after the local write, before the backend was called at all. This Mac was carrying 100 of them, every one with no `deletedBy` and a canonical status still `active`. Three changes, because each is load-bearing on its own: - The lane keeps the rows the response itself reports retired and drops the rest. Removed showing fewer rows is a gap; manufacturing retirement is data loss. - A migration clears the tombstones already written, and only those: a deletion the owner performed records `deletedBy`, a server-side retirement arrives as canonical status, and a row with neither witness was retired by nothing but the stamp. - The card treats the reader's tick as authoritative. A retirement discovered afterwards does not get to undo a completion the app accepted, whatever put the retirement there. Until the backend grows a real `deleted` filter, Removed can only show deletions made on this Mac. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(desktop): clearing Chat takes the daily summary with it The summary card is chrome above the thread rather than a turn — INV-CHAT-1 keeps transcript authorship in the kernel, so nothing on the Swift side may write a journal row for it. The consequence was that Clear could not reach it: the day's summary sat alone in a chat the reader had just emptied, which reads as a clear that did not work. Clearing now records which summary was on screen and withdraws the card on the same frame. Recording the id rather than a flag is what lets tomorrow's summary come back on its own, and the key is owner-scoped like the announcement's, so clearing on one account cannot blank another reader's day on a shared Mac. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(desktop): stop one test post killing the whole test host `swift test` has been exiting 1 on this machine, before this branch and after it. Not a failing test — the runner died with SIGTRAP partway through, taking every suite alphabetically after it with no message and no crash report. The two suites involved each pass alone. `JITProactivityDeliveryTests` posts `.runtimeOwnerDidChange` from an async test body, so it lands on a cooperative thread. `NotificationCenter` delivers synchronously on whichever thread posted, and the observers are `@MainActor` types whose sink closures carry an isolation check on entry. Once `InterjectWiringTests` had mounted `FloatingControlBarManager`, that check failed `dispatch_assert_queue` and trapped the process. Nothing inside the closure could have guarded it: the check runs before the first statement. The fix is the post, not the observers. `performEffectiveOwnerTransition` — the only production poster — already posts inside `await MainActor.run`, so no shipping code does what the test did. And the observers' synchronous delivery is load-bearing, not incidental: it is what lets a surface fence itself *during* an owner transition rather than a runloop later (see `IntegrationNudgeCoordinator`). Hopping delivery downstream instead was tried first and was wrong twice over — it broke `FloatingOwnerProjectionTests`, which correctly caught the previous account's pills surviving the switch. So: the test posts on the main thread as production does, and the contract is written down where the notification is declared. The suite now runs to completion for the first time: 6695 tests, no crash. Two `JITProactivityRuntimeTests` failures remain, pre-existing and unrelated — they assert an outcome that only holds when the local mirror database is absent, which is true only when they run alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(app): restore the ephemeral iOS package Flutter regenerated Running `flutter gen-l10n` and `flutter test` on this Mac rewrote `app/ios/Flutter/ephemeral/Packages/FlutterGeneratedPluginSwiftPackage/Package.swift`, emptying its dependency list because the iOS packages were never resolved here. That is build residue from the merge, not a change anyone made. Restored to origin/main's content. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * style(app): dart format the files this branch touches `message.dart` picked up its formatting drift from the merge resolution; the generated wire file is regenerated output. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Revert "style(app): dart format the generated wire file" `subscription_usage_wire.g.dart` is generated by `backend/scripts/generate_dart_models.py`, so reformatting it makes it stale against its own generator. Restored to origin/main's content; the `message.dart` formatting in the previous commit stands. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(app): regenerate subscription_usage wire models The committed file was stale against `backend/scripts/generate_dart_models.py --group subscription_usage` — the preflight contract check rejects it on origin/main's copy too, so this is drift the merge surfaced rather than anything this branch changed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(desktop): a follow-up's citation borrows the ordinal it points back at Ordinals are assigned per attempt, so a turn that retrieved nothing has no ordinals of its own — and when the model answers "pick one conversation from that day" without a tool call, the `[1]` it writes is the first conversation of the previous answer. Left unbound, it drew as plain text beside a title the reader could not open. Only ordinals the turn cannot resolve itself are borrowed, each from the nearest earlier assistant turn that persisted it, so a turn's own provenance always outranks the past. The binding runs at finalize and again over every journal projection, so restored history opens the same source the reader could open live. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop): fold a reply after two screens of text, not five lines Five hundred characters was five lines, so every real answer collapsed and the reader clicked "Show more" under nearly everything they asked. The budget is now measured in viewports of rendered text: a reply may fill two screens before the transcript offers to fold it, and one that long starts folded. Lines are estimated per source line at the column's width, so a bulleted list is measured as the lines it takes, not the characters it holds, and the cut lands on a line boundary with any open fence closed. The transcript publishes its viewport to the rows through the environment, republished only at a coarse step so a resize drag does not re-evaluate every row per frame. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop): stream as a flow, and glide the live edge instead of snapping The wire delivers an answer in bursts — a provider chunk, a paragraph the moment a tool returns — and a flush that dumped everything it had made the transcript lurch by a sentence and then sit still. Each 35 ms flush now reveals a paced slice: a small backlog drains over about five flushes, a large one is let through fast enough that the reader is never more than two lines behind, and the tail of every burst tapers. Boundary flushes (a tool starting, the turn settling) still land everything at once, so block order is unchanged. The follow scroll during a stream eases to the live edge rather than jumping a row at a time; restores and sends stay instant, and Reduce Motion disables the glide. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(agent): native components are the default answer for what Omi draws "Pick one conversation from that day" came back as a bold title with a citation number where a conversation card belonged. The render policy now says it plainly: default to a component whenever the user asks for a task, goal, memory, conversation or capture, and reserve prose for a request to read rather than open or act — a summary, a recap, an analysis, a comparison, a count, or a list too long to render. The conversation and memory retrieval tools carry the same rule with their own block shapes, including follow-ups that narrow an earlier result. Regenerated tool surfaces and fixture. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(desktop): the hermetic chat flow covers the streaming buffer The paced reveal lives in ChatStreamingBuffer.swift, which no e2e flow declared; the hermetic chat flow streams an answer through it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(desktop): the offline readiness fixture asks for a port nothing holds The harness verifies a launched bundle's /health whenever --port is bound. On a Mac with any dev bundle up on the default 47777 the offline-only readiness cases reached that check and sourced app-config.sh, which the fixture never provides, so the pre-push launcher gate failed on a test that never left the box. The fixture now probes for a closed port and passes it explicitly. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): a conversation citation opens the conversation it names The agent cites whatever its conversation tools retrieved, and most of this user's library is desktop and phone recordings. But a citation of kind .conversation routed the capture focus, and the capture archive resolves that focus through a strictly source-scoped fetch (GET v1/conversations/{id}?source=omi). For a desktop recording that request is a 404, so the hub landed on the Conversations list with nothing opened — the report's "Open Conversation just brings me to the conversation page". Citations of omi-device captures kept working, which is why only some markers failed. Citation routing now fetches the unscoped record first and lets the record's own provenance pick the route: omi captures keep the capture focus so the transcript moment still plays; every other conversation opens as the exact fetched record through open(conversation:), which presents it even when the paginated list does not contain it. A failed or mismatched fetch navigates nowhere instead of stranding the reader. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): launch no longer prints the daily summary before the chat Opening the app admitted the daily summary card while the initial history was still loading, so launch read as a summary page that then yanked away the moment the transcript landed at the live edge. The card is chrome (INV-CHAT-2), and chrome must not outrun the thread: admission now defers for the whole initial load, the loading-complete observer admits it before the live-edge restore measures geometry, and an empty thread still meets its card. The decision lives in one pure, tested place — ChatDailySummaryAdmission. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * perf(desktop): returning to Chat no longer re-parses and re-measures every prose block Chat is torn down and remounted on every route change (the shell keys the destination by route), and each mount rebuilt the whole visible transcript window from scratch: one Foundation Markdown parse plus one NSAttributedString build per prose block, and one throwaway TextKit stack per sizeThatFits query. ChatProseRenderCache memoizes both, keyed on the exact render inputs (markdown, style, fontSize, font scale in thousandths, sorted citation ordinals), so a hit is identical to a recompute. Bounded LRU (192 entries) absorbs a streaming row minting a key per 35 ms flush; per-entry measured heights are keyed on the proposed width and dropped wholesale at eight widths, which is what a window resize wants anyway. Measured by instrumenting ChatSelectableProse.attributedString and ChatSelectableProseText.height on the mounted transcript (ChatTranscriptGestureHarnessTests.Harness, 120-message journal, compact 50-row window, debug build), cache disabled vs enabled in the same binary: - before: every mount, including every return to Chat, paid 50 parses (4-8 ms) and 600 height measures (54-62 ms of TextKit layout) - after: the cold mount pays 50 parses and 50 measures once (4-17 ms, 4-14 ms); every remount pays 0 parses and 0 measures New ChatProseRenderCacheTests pin hit identity, key sensitivity to every render input, nil-produce never being cached, the LRU bound, and the per-width height memo and its drop bound. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): the follow glide runs on the run loop, so it lands in the app and in the tests The streaming follow rode a SwiftUI withAnimation transaction, which advances on the display cycle. In the mounted-transcript guard host that cycle never runs, so the follow silently did nothing there and the three live-edge drift guards failed at full stream growth: 462-484 pt of drift against a 120 pt limit, bisected to the commit that introduced the glide. A follow that only works where nothing can observe it is exactly what INV-CHAT-2's guard tests exist to catch. ChatFollowGlide moves the clip view on a 60 Hz run-loop timer with the same ease-out and the same 0.16 s duration (now one shared constant). The run loop is pumped by both the app and the harness, so the glide the reader feels is the glide the tests measure. A newer follow retargets the clock in flight, and reader input or teardown cancels it through cancelAllPendingScrolls; the snap path stays the fallback for pre-resolution frames and for Reduce Motion. Measured on the mounted transcript (60-message window, 40x35 ms stream flushes): worst drift 462-484 pt before, 48.0 pt after; all 20 ChatTranscriptGestureHarnessTests pass again. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): the streaming flush beat keeps its period when a flush overruns The paced flush re-armed from the moment the last flush finished, so the reveal's period was interval + render work, and jitter in that work shifted every later beat back. Measured with a flush workload of 0.75x the interval driving ChatStreamingBuffer directly: completion-anchored beats ran at 151-160 ms against an 80 ms interval; anchored beats hold 80 ms while the work still fits inside one. Beats now anchor to the beat before them, clamped to now so an overrun beat fires at once and resynchronizes instead of firing late. A drained backlog ends the chain, so the next delta starts a fresh beat rather than inheriting a deadline that already passed. Instrumented per-flush costs on the mounted transcript, for the record: the markdown re-parse grows 0.8 to 4.1 ms and the TextKit height measure 1.4 to 4.9 ms over a 0.25-10 KB answer (the O(n) per-flush terms; a stable-prefix incremental parse was assessed and declined - the parse is only about a quarter of the O(n) work, the rest is TextKit measure and the live view's own relayout, which incremental parsing does not touch). The storage isEqual compare is 0.01-0.08 ms and the transcript re-evaluates one bubble body per flush - both cleared as stutter candidates. The follow glide retargets from its mid-glide position with no step: worst sample-to-sample jump 11 pt over 210 samples of a 30-flush stream. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): the settle frame folds the tool trace away instead of teleporting When isStreaming flips, the row drops its tool-call groups and pre-tool commentary and swaps in the terminal answer, citations and metadata band in one layout change - measured on the mounted transcript with a three-tool trace, a 62 pt height change landing in a single frame, which reads as a jump. The fold is now animated: an .animation scoped on message.isStreaming wraps the row content, so that one frame eases out over the motion table's settle duration while every per-token streaming frame, where isStreaming did not change, stays exactly as instantaneous as before. Reduce Motion folds instantly. Final state is unchanged: the settle probe measures identical streaming and settled heights before and after (2620 / 2558), the group-drop assertions in ChatFirstRichBlockTests, OneChatShellRichBlockTests and ChatBubbleLayoutRegressionTests pass unchanged, and the mounted-transcript harness - which streams forty-flush answers through a row carrying this modifier - stays green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(desktop): the beat-period test declares its real-time wait The flush-beat test sleeps for real by design - the subject is the period the scheduler holds under genuine elapsed work - so it carries the wall-clock-wait escape annotation the desktop test-quality gate requires, keeping the counted baseline where it was. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(desktop): prune merge survivors of views this branch deletes Taking main's side of ViewExporter kept exports for the old Dashboard page and its score gauge, and main's automation registry picked up a Home-knows action family for the legacy hub - all three views and the feature behind them are deleted on this branch, so the entries and the action family go too. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(desktop): retire the knows-list flow and pin the merged test debt The knows-list rotation flow covers the HomeKnows feature this branch deletes with the legacy Home hub, and the merged suite's real source-inspection debt is 53 files / 143 sites - lower than the inherited ceiling. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): restore the memory-review fixture the knows-list prune removed The prune of the retired knows-list action family took `enum MemoryReviewFixture` down with it, but the surviving `seed_memory_review_fixture` / `memory_review_snapshot` / `memory_review_vote` actions and MemoryReviewCardTests still reference it — the branch did not compile. The enum moves to its own automation file, content unchanged from 98ddcffab9^. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): finish the merge-fallout repairs the tree still needed Two more of the same class as the fixture restore, on tracked files the merge left referencing deleted things: - `AnimatedGIFView.swift` went out with the retired onboarding wizard views, but main's own PermissionsPage and BrowserExtensionSetup still render it, and the merge kept those references. Restored verbatim from origin/main. - The ViewExporter prune meant to drop the full-dashboard export with the deleted Dashboard page but left a bare `AnyView` expression behind; the dead entry is now actually gone. The untracked Dashboard WIP files belong to a concurrent lane and are deliberately untouched. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop): decode the recap sections the dedicated page renders DailySummaryRecord gains unresolved_questions, decisions_made, and knowledge_nuggets, conversation_ids on topic highlights, and source_conversation_id on action items — the wire contract mobile's daily_summary.dart already mirrors and the backend has served since the sections shipped. Everything is decodeIfPresent, so an older backend just leaves the sections absent. Highlight and ActionItem get explicit inits with defaulted new fields so existing call sites are untouched. Tests pin the new keys verbatim and pin the older-backend case: absent sections decode to nil rather than dropping the summary. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop): the daily recap becomes a page, and Chat shows a slim pill Recaps are now clickable into a dedicated surface. DailyRecapPage presents the whole record — date and headline, overview, stat chips, highlights, tasks with their completed state, unresolved questions, decisions, the memories-learned review rows, and learnings — and every row carrying a conversation id deep-links through the hub-owned conversation detail. Rows carry no badges or buttons anywhere else: the pill in Chat and the pill on an Activity day show only title and summary, and clicking either opens this page as a sheet. A sheet, not a ChatFirstRoute: route values persist across launches and are gated to primary destinations, while a recap is a transient read over the surface that produced it. The Chat pill replaces the pinned full card. The old bar expanded in place over the transcript, and its measured inset tracked a high-water mark that could never shrink — one tall expand padded the thread forever, and a shrink left the bar lapping over the newest message. The pill never expands (the page is the expansion), so the inset now tracks the live height both ways, bounded to a couple of lines. The full-card view is retired; its pure section projections move to ChatDailySummaryPresentation next to the other recap rules. Review rows gained a daily_summary_detail source so votes from the page are counted as page votes, matching mobile's telemetry. INV-CHAT-2 admission is untouched. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop): the Activity day folds its recap into a doorway pill The day header was already the recap's toggle — thin when the day is folded, the recap inside when it is open. What it contained changes: the stored recap renders as a pill of title and summary only. The highlight chips, the Ask-about-this-day button, and the Regenerate button leave the list entirely — the full record, its badges, and its actions are the dedicated recap page the pill opens in a sheet, reusing the hub's own onOpenConversation so a recap's conversation links can never open a second detail owner beside it. Generate stays for a day with no recap: that is the way the recap comes to exist, not recap chrome. memory-review.yaml follows the rows to the page: the pill renders no rows, so the flow opens it before reading and voting, and the section now reports the daily_summary_detail source. chat-first-cohesive covers the new pill, page, and fixture files and drives pill-to-sheet. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(desktop): follow SpineDayRecapRow to its doorway signature The row no longer takes now/calendar — the follow-up question moved to the dedicated recap page with the rest of the recap actions — so the two host-view tests construct it with content and dateKey only. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): the restoring phase always resolves, even when its attempt is superseded Three dev launches hung in appState=launching / isSignedIn=false / isRestoringAuth=true for their whole session (10:31-11:05), then an identical launch came up clean. The hung logs end at the auth listener's "skipping REST validation while launch restore is in flight" with nothing after, and the phase never leaves .restoring. The deeper defect: the restoring phase's only guaranteed escape - the 5s watchdog in configure() - was gated on isSessionAttemptCurrent(attempt). Any newer session attempt defuses it, and every fenced exit in the restore flow (validateRestoredSessionNow's entry/post-refresh guards, refreshIdToken's post-network guard, the saveAuthState/commitRestoredSession commits) then returns silently, leaving .restoring stuck with no remaining resolver and no log line. The hung trail is exactly that shape; the begin itself is invisible because the flow's OMI AUTH NSLog lines are not captured in the dev log, and the closed set of beginSessionAttempt call sites includes the restore flow's own invalidation branches, which fire on a transient empty credential read - the same launch whose Firebase cached user read also came back nil. The watchdog now resolves on the phase alone: while the app still reports restoring and no user-driven sign-in owns the UI (isLoading), it lands in the same recoverable state the ungated case always did. It arms before the restore's first await, so the restore's own awaits cannot delay the arming, and it logs - visibly in the dev log - whether the launch attempt was superseded, so the next occurrence names its race. Tests: AuthRestoreWatchdogTests drives the seam with an injectable timeout. testRestoringPhaseResolvesEvenWhenTheLaunchAttemptWasSuperseded fails on the attempt-gated watchdog (phase stuck, verified) and passes on the fix; testTheWatchdogLeavesARestoringPhaseAloneWhileASignInOwnsTheUI pins the isLoading deferral; a source-wiring pin asserts the watchdog arms before the restore await and resolves on phase alone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): re-follow the live edge when the recap pill's inset moves The live-tracked inset is document space above the reader, so when the pill's measured height changes after admission — the overview arriving a beat after the headline is enough — the document shifts under a stationary viewport and the newest row ends up buried under the pill. Caught in the running bundle: the pill rendered thin, but the first transcript row slid under it. A reader following the live edge gets admission's own answer, a re-follow; a reader scrolled away is left exactly where they are. Verified visually: chat now shows the thin pill with the live edge readable beneath it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): close the activation-actions extension the knows-list prune truncated The knows-list retirement left DesktopAutomationActivationActions.swift one brace short of its extension, so a clean checkout of the branch does not parse. Closing it restores the build without touching the pruned flow. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop): the daily recap becomes a full page on the Chat-first shell The sheet is dead. `ChatFirstRoute` gains `.dailyRecap(DailyRecapRouteRef)` — a non-primary destination in the `.more` family, hosted by the shell on the glass page lane like the other full pages. The route carries identity only (record id + date), so a relaunch onto the persisted route re-fetches through the new `getDailySummary(id:)` read; a record that is gone degrades to an honest unavailable state. `ChatFirstShellNavigation.openDailyRecap` captures the opening surface's route as a transient origin, and the page's back chevron — and Escape — return there; a tab select supersedes it and an owner change clears it. The page keeps the punch-list layout: back chevron + Daily-recap eyebrow, single-line action pills (Ask primary, Regenerate secondary — `fixedSize` so no label can ever wrap), dayEmoji + full weekday date eyebrow + headline + lede in a 720 pt column, one row of equal-width single-line stat chips that scrolls rather than folds a label, and the deep-linked sections. The eyebrow's wider date format lives beside the pill's in `ChatDailySummaryPresentation` so the arithmetic stays in one place. Chat and Activity open the same route; their flows wait on `visibleChatFirstRoute: daily-recap`. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop): the chat recap becomes an in-history day boundary The pinned overlay bar, its measured-height inset, the high-water and live-tracking preference plumbing, and INV-CHAT-2's admission gating are all deleted from ChatMessagesView. A row that lives in the thread's own history cannot shift the live edge, so none of the banner constraints apply any more — the transcript simply renders `ChatDailyRecapRow` as part of its row data, anchored above the first message on or after the recap's day. When that boundary is outside the loaded window (no message reaches the day, or the day's start is at the top of the window with more history above), the thread renders nothing rather than a marker the history cannot back up. A cleared thread keeps its recap withdrawn through the coordinator's existing clear contract. The row is a quiet, centered history marker — day label with emoji, a one-line headline, two lines of overview, whole row clickable — and it opens the typed recap route. The old pinned pill file is gone; its identifier and covers follow the rename. The admission unit tests become placement tests. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop): the Activity day card grows its recap as one continuous body The open day's recap no longer floats as a second rounded surface under the day header. The header becomes the card's top half — rounded at the top, square at the bottom, no gap — and the recap renders as the card's body on the same material, square where they meet and rounded where the card ends, with the hairline only on the outer edges so no line runs across the seam. The header's content resolution is shared with the slot, so the attached shape and the slot's content can never disagree; folding the day keeps the thin header-only card, and the generate affordance for a recap-less day stays its own small surface. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(desktop-e2e): identity seam and drive path for the recap page `daily_summary_snapshot` reports the record's opaque id (identity, not content), and a new `open_daily_recap_page` action opens the dedicated page through the same `ChatFirstShellNavigation.openDailyRecap` the recap rows call — so flows and harnesses can reach the page without a click. Launch always lands on chat (`openMainAppChat`), so the persisted recap route never survives a relaunch on screen; the route's persistence stays harmless exactly as intended. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): stat chips hold their full label on one line 106 pt folded "4h watching" into "4h… watc…" — the exact mid-word wrap the page layout exists to forbid. The equal-width slot widens to fit the longest label ("8h 5m listening", "13 conversations"); a narrower lane scrolls the row rather than folding a chip. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * style(desktop): the in-history recap row sits below the messages Review feedback: centered bordered prose read as another answer. The row is now left-aligned one type step under the bubbles, borderless on a softer fill, with the chevron trailing - a day marker in the thread, not a competing card. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * perf(desktop): a live stream pins the following viewport to the live edge The streaming follow still jittered: every 80 ms the throttle fired a 0.16 s ease-out toward the bottom as measured when the follow fired, but the paced flush grows the document every 35 ms - so each target was stale on arrival, and the viewport's velocity cycled between decelerate, stall, accelerate for the whole answer. Mid-stream re-wraps that briefly shrank the document even drove the glide upward. While the last message is streaming and the reader is following, a 60 Hz ChatLiveEdgePinner now moves the clip view to the current maximum scroll every tick - the bottom is read and taken in the same tick, so there is no target to go stale. Armed from the streaming follow path; disarmed on stream settle, reader input, mode change, conversation switch, and teardown through cancelAllPendingScrolls; every move marks the programmatic-scroll signal so the user-scroll detector keeps classifying tracking as the app's own. The glide remains for discrete jumps (button, restore settling). Measured on the mounted transcript: worst between-flush drift 0.0 pt (the replaced glide's was 48 pt); Reduce Motion applies - pinning is positional tracking, not animation. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(desktop-e2e): the activity day header answers to spine-day-recap-header A stable handle for flows (and a reviewer) to drive a day's fold without reaching into the spine's private collapse state. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): recap stat chips hug their labels The fixed 152 pt slot was the defect: short labels left dead space inside their capsule while the row's tail clipped past the column edge. Chips now size to their content with one uniform gap, and the row scrolls only when the window is genuinely too narrow. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): the activity day card stops repeating its recap's emoji and arrow The header already carries the recap's emoji and is the fold control, so the recap body showed both a second time - two emojis stacked, and the fold chevron sitting directly over the row's own arrow. The body is now headline + summary only, and its whole surface still opens the page. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): all six recap stat chips fit the column at min window width * feat(desktop): the chat recap row shows the day's stats, compactly The chat row was the only recap surface without the day's numbers. A strip of micro chips - same data as the Activity day card, bubble-row weight, content-hugging so nothing clips - sits under the overview. Badges and actions stay page-only; this is the number line, nothing more. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): the tasks list stops nesting a lazy stack inside a lazy item TaskCategorySection is itself an item of tasksListView's LazyVStack, and it laid its rows out in a nested LazyVStack. A lazy container nested inside a lazy item makes the section's measured height non-convergent — the inner stack reports estimates while the outer measures it and real heights once placed — so every layout pass mutated lazy-item phases and re-signaled prefetch, each scheduling another transaction. The flush never drained: a permanent beachball the moment the list was scrolled with enough data (sampled: 2627/2627 main-thread samples inside GraphHost.flushTransactions, dominated by LazyStack measureEstimates and LazyLayoutViewCache updateItemPhase). The row stack is now eager, so a section's measured height is deterministic and the flush drains. The outer stack stays the virtualizer. A source inspection guard pins the no-nesting contract. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * perf(desktop): the switch into Chat stops paying a per-node preference reduce and a fixed settle ladder Navigating to Chat felt delayed. Two measured causes: 1. The composer height travelled through QueryComposerHeightKey, written at the composer and read at the surface root, so SwiftUI ran a reduce over one combiner pair per node of everything mounted between them — the whole transcript — on every switch. Sampled: 40% of the mount transaction's observer work. The height now travels through onGeometryChange into the same guarded state write; the preference key is gone. 2. The initial restore re-pinned the live edge on a fixed [0.05, 0.2, 0.5, 1.0] ladder, so every switch waited out the full ladder even when the document had stopped reflowing after the first pass — 1.16s of pure waiting at the p50. ChatInitialRestoreSettle tightens the ladder and settles early on a height-stability check (two consecutive passes measuring the same laid-out document); the last pass still completes unconditionally. Also adds the opt-in OMI_SWITCH_PERF switch-span telemetry (route select, destination teardown, transcript mount, first laid-out document, settled restore, main-thread stall watchdog) that measured all of this, and the chat-first flow cover for it. Measured: restoreSettled p50 1156ms -> 241ms against memories as the source page. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * perf(desktop): the bubble identity check stops re-deriving the answer text ChatBubbleIdentity.equal compared lhs.copyableText == rhs.copyableText, but copyableText derives purely from (contentBlocks, text, isStreaming) — every one of which the same conjunction already compares. SwiftUI runs this equality for every bubble on every transcript body pass, and each evaluation re-derived and whitespace-normalized the full answer text. Sampled mid-switch: the copyableText chain was the largest app-code cost of the mount. Removing the redundant term is behavior-preserving and dropped the navigate-to-chat settle from ~1.3s to ~240ms at the p50; a regression test pins that block-only edits stay visible through the blocks term. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * perf(desktop): bubble identity compares content blocks field-by-field, not through JSON ChatBubbleIdentity.equal's blocks term went through ChatContentCodec.comparisonData, which JSON-encodes the persistence dictionary with sorted keys — two full serializations per bubble per transcript pass, for every unchanged row, ~28 times a second while an answer streams. ChatContentBlock now conforms to Equatable with hand-written field-by-field comparison (questionCard options deep-compare as NSDictionary; ToolCallInput synthesizes). The comparison is deliberately one notch stricter than the encoder was — in-flight toolCall statuses (.running/.slow/.stalled) no longer collapse together, which can only force an extra re-render mid-turn, never stale UI: the isStreaming guard already re-renders through those states. Also gives the task-chat panel the compact transcript window: it was the one ChatMessagesView host still on the 500-row default that measured 910 ms and 607 native views against 114 ms and 84 compact. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): goals cards stop showing impossible progress and echoing their own titles UI audit findings. A goal created without an explicit target keeps the creation default (1.0) while progress updates push currentValue past it, and the row rendered that raw: "Progress: 20 of 1". When the current value has overshot that unconfigured default target, the target carries no information, so the summary reports the value alone. desiredOutcome falling back to the title echoed every card's title under its title; it is hidden when it matches. Both the focused-goal card and the list rows capped at 680pt while the page header spanned the panel — the cards now fill the lane like the rest of the page. The Rewind capture control said "Capture Off" next to a switch knob in the on position, because the knob shows the setting and the label showed health, and the two can disagree. "Off" now only ever describes the setting; a dead-but- enabled capture reads "Capture Stopped". Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * perf(desktop): the task-chat panel gets the compact transcript window Missed from d27ca379ed: TaskChatPanel is the file carrying the transcriptWindowPolicy one-liner; without it the panel mounts the 500-row default (910 ms / 607 views vs 114 ms / 84 compact for a 400-message thread). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(desktop): the content-block equality lives in the codec file ChatProvider.swift is under a hard convergence limit and has one job; the ChatContentBlock: Equatable conformance belongs beside the block codec. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(desktop): regenerate tool surfaces at the merged manifest Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(desktop): the tasks list flattens to a single lazy level TaskCategorySection was one lazy item holding an eager stack of all its rows. Eager-in-lazy meant the outer list could not estimate a section without laying out every row, so all sections materialized: leaving the page tore down every task row in the profile (~330 trees, ~600 ms of main-thread layout per navigate away), the reason navigate-to-Chat from Tasks cost ~2x the same switch from Memories. The list is now one LazyVStack whose items interleave a category header with one item per task row (TasksListItem). Rows virtualize individually, nothing nests inside a lazy item — which also retires the lazy-in-lazy estimate livelock class the eager stack had traded the freeze for. Row ids keep keyboard navigation and per-row state; the header (icon, count, Today menu, top drop zone, accessibility identifiers) survives as TaskCategorySectionHeader; per-row drag-drop and inline-create composition is byte-for-byte the old one. Collapse was never wired at the only call site; the dead parameters stay compiling with a note that a real implementation must filter the items where they are built. Measured on a ~330-task profile: freeze stress clean (zero flushTransactions / propagate_dirty samples), the teardown storm gone (TaskRow ≈ 3-5 of ~2370 samples mid-switch), tasks→chat switch parity with Memories. Interactions exercised: snapshot, create, toggle, reorder (persisted), capture. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(desktop): release compile and the Dart analyzer ratchet TaskCardView's body type-checked past the release optimizer's expression budget ("unable to type-check this expression in reasonable time" at ChatFirstContentBlockViews.swift:198); the branches are now extracted into renderedCard / unavailableBlock / unavailableDescription so each sub-expression type-checks on its own. Drops the one new unused import the Dart analyzer ratchet flagged in the content-block parity test. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(desktop): the beat-anchoring contract holds on a simulated clock The beat-period test drove a real Thread.sleep and then asserted the sleep stayed inside the beat budget; on a loaded CI runner the sleep itself overshot (periods measured 125-194 ms against a 120 ms ceiling) and the test failed on machines whose streaming was fine. The contract lives in the scheduling math, so nextBeatDeadline is extracted as a pure function and the test simulates the beat chain against it: work inside the interval holds the grid exactly, an overrunning beat resynchronizes to now instead of compounding. Deterministic, and it pins more than the wall-clock version could — the old assertion tolerated any drift under a ceiling, this one requires the grid to hold exactly. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> | 2 天前 | |
fix: dispatch every macOS candidate directly to Codemagic Failure-Class: FC-release-train-busy-main-starvation | 18 天前 | |
fix(ci): exclude gitignored files from the dead-code ratchet (#12756) * fix(ci): exclude gitignored files from the dead-code ratchet check_dead_code.py enumerated each area with a raw filesystem walk, so any gitignored file under app/lib, backend/ or a desktop src tree was treated as production source. Those files are absent from CI's checkout and from every diff, so the gate failed only on developer machines and read as broken. Reproduced on this repo: app/lib/firebase_options_dev.dart is gitignored at .gitignore:176 and generated during Flutter setup. With it present, --area flutter reported it as newly dead and failed; removing the file turned the same check green. Filter each scan through git ls-files --others --ignored --exclude-standard, mirroring _git_ignored_paths in backend/scripts/generate_plan_catalog.py, which was added for this same failure in #12476. --directory collapses an ignored directory to a single entry, so membership is tested through ancestors. When git cannot answer the set is empty and the previous pure-filesystem behaviour stands, which keeps the checker hermetic. Verdicts are unchanged for agent, windows and backend on this checkout; only flutter moves, from a false failure to ok. * fix(ci): fail closed on an ignored entry and harden the git filter Review follow-ups on the gitignore filter. scan_flutter seeds reachability from app/lib/main.dart. Excluding an ignored entry from the scan set while still walking from it would strand every other file and report the whole area dead, so an ignored entry now errors the way a missing one does, matching what scan_ts already did when no entry remained. git ls-files output is decoded with errors="surrogateescape", and UnicodeError joins the fallback path. A filename whose bytes are invalid for the locale previously raised out of the checker instead of degrading to the filesystem scan the fallback exists to provide. Tests: the shared scan_ts path now has its own ignored-file case, since the agent and windows areas went through it untested; the flutter and backend fixtures stage their files so the "still caught when tracked" assertions describe what they actually exercise; and the ignored-entry guard has a regression test that fails without it. | 2 天前 | |
fix(preflight): decode Git content as UTF-8 on Windows (#10595) * fix(preflight): decode deployment Git output as UTF-8 * fix(preflight): decode deferred-work Git output as UTF-8 * fix(preflight): decode line-ratchet Git output as UTF-8 The product line-count ratchet read baseline paths and JSON blobs with the Windows host code page. Decode all three Git content reads explicitly as UTF-8 and lock both sharded and legacy paths with a subprocess contract test. --------- Co-authored-by: 491034170 <142008960+491034170@users.noreply.github.com> | 1 个月前 | |
feat(desktop-release): isolate drift self-test from hook GIT_DIR Pre-push exports GIT_DIR for the outer checkout, so the fixture repo inherited that namespace and failed to commit. Strip GIT_* and disable hooks/gpgsign in the self-test. Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: multica-agent <github@multica.ai> | 4 天前 | |
fix(deploy): select desktop-backend's runtime image by container name (#12156) status.imageDigest is the legacy single-container scalar in the Cloud Run v1 Revision API and comes back empty on any multi-container revision. The Managed Prometheus sidecar (#11998) made every desktop-backend candidate two-container, so the field went empty and image lineage verification has failed closed on every dev/prod candidate since. Extract the runtime image from the named desktop-backend-1 container instead, asserting exactly one match so the guard still fails closed on an ambiguous topology. Co-authored-by: r <r@r> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> | 14 天前 | |
chore(ci): make traffic-regression gate robust to hook git env and python 3.14 - argparse help: escape the literal % (python 3.14 applies %-formatting to help strings and rejects a bare '% t') - _git_is_ancestor + fixtures: strip GIT_DIR/GIT_WORK_TREE/GIT_INDEX_FILE/ GIT_OBJECT_DIRECTORY/GIT_NAMESPACE so the probes and scratch repos target the current directory's repository even when run from a pre-push hook Co-authored-by: multica-agent <github@multica.ai> | 20 天前 | |
fix: keep the macOS hourly release train moving Bind each candidate to its checked source, recover squash-orphaned changelog branches, install Bun for release eligibility, and age freshness from the oldest unshipped update. Verification: 77 focused release-control tests; actionlint; Python compile; diff check. Failure-Class: new | 18 天前 | |
feat: add gateway eligibility-proof break-glass hatch AGENTS.md requires every gated surface to have a break-glass hatch, but the standalone LLM gateway production admission added in this PR had no operational bypass when the Release Eligibility proof system is unavailable (flaky/red check poisoning the SHA, API outage, etc.). That made the new gate a hard deploy blocker for gateway-only releases during proof-system outages. Add the same break-glass shape the backend workflow already uses: skip_eligibility_proof + break_glass_confirm=deploy-without-proof + break_glass_reason, with a dedicated record_break_glass job that opens a release-gate-failure tracking issue. The hatch may skip the eligibility proof but never the merged-main ancestry check -- unreviewed code still cannot reach production by any path. This completes the PR's stated goal of aligning the gateway's trusted source-admission contract with the backend's. Also pins the hatch properties in the direct-production-admission contract checker with a new mutation test so the hatch cannot be removed without a failing check. Failure-Class: none Verification: - python3 .github/scripts/check-direct-backend-production-admission.py — passed - python3 .github/scripts/check_backend_deploy_source_admission.py — passed - python3 .github/scripts/test_check_direct_backend_production_admission.py — 5 passed - python3 .github/scripts/test_check_backend_deploy_source_admission.py — 28 passed - actionlint -ignore 'SC2086|SC2016' .github/workflows/gcp_llm_gateway.yml — passed - make preflight — 15/15 selected checks passed | 1 个月前 | |
docs: correct stale references surfaced in review - apps-marketplace e2e: covers now lists category_section.dart (the flow's S1-S7 browsing renders CategorySection/SectionAppItemCard; AppListItem is only rendered in search-filter results, which the flow never enters) - goals-tracking e2e: GoalsWidget sits after the processing-conversations widget in conversations_page.dart (TodayTasksWidget renders in home_content.dart); with zero goals the widget hides and the home-screen Add Goal entry point is ActionItemsPage._buildGoalsRow - goals_widget.dart: same correction in the hide-when-empty comment - 00-INDEX.md: focus/stats.ts and AutoCreatedTasksStep.tsx entries updated to match the deletion note — both actionable suggestions are now marked resolved-by-deletion instead of presented as live work - plugins/README.md: document monolith (plugins/requirements.txt -> ./omi-plugin-sdk) vs per-service (own requirements.txt -> ../omi-plugin-sdk) SDK install paths separately | 3 天前 | |
fix(ci): count a landed change once in the failure-class guard ratchet (#12119) The guard ratchet counts first-parent integration changes declaring a class, then adds the `--pr-body-file` declaration so the declaring PR fails first. Both sources can describe the same change: on a main push HEAD *is* that integration change, and `git log --first-parent` already counted it. So one change is counted twice and the total depends on where the check runs. A class with one prior declaration reads as 2 on the PR (green) and as 3 on the push of that same PR (red at the default threshold), which is exactly the after-merge red the main-push metadata path exists to prevent. Locally the count is inflated the same way, so pre-push does not predict its own PR. Subtract what HEAD already contributed: a declaration the body carries counts only when HEAD's own message did not carry it. A `pull_request` run checks out the merge ref, which declares nothing, so the pending change still counts. Failure-Class: new | 14 天前 | |
fix(release): harden direct beta evidence admission Bind both stable and Beta signed-smoke results concurrently, align release metadata and tag contracts, and close the remaining review guard and test gaps without restoring qualification or adding a sequential release wait. Failure-Class: none | 27 天前 | |
fix(repo): reject Git identities that can never be attributed (#12244) * fix(repo): reject Git identities that can never be attributed PR #12239 was authored entirely by `r <r@r>` from a stale clone-local `user.email`. This guard ran on every one of those commits — 66 over five days — and printed "OK: Git author identity is not a test fixture" each time. It was enumerating the previous incident rather than the invariant. After #11525 minted `Ratchet Test <ratchet-test@example.invalid>`, the check learned that exact display name, that exact local-part, and three reserved TLDs. `r@r` is none of those, so it passed: `email_tld("r@r")` returns `"r"`, because the domain has no dot to split on. The rule is now the property that makes an identity wrong. An address whose domain cannot resolve can never be delivered or attributed to anyone, whatever it is called. That single predicate covers both incidents and the family they belong to, and it needs no allowlist to maintain. `failures_for` already applied its predicate to clone-local config as well as to the commit range, so one change covers every lane at once: - pre-commit (`--pending`) rejects the override before a commit is minted - the local lane of `git-author-identity` rejects it again before push - the CI lane never sees a developer's `.git/config`, so the commit range is the only unbypassable place this can be caught — and it now is Real addresses, including `users.noreply.github.com`, are untouched; a name-only identity carries an empty email and is not judged on it. The report wording said "test fixture", which is exactly what let a non-fixture bad identity read as a pass. It now says unattributable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(failure-classes): add FC-guard-encodes-incident-literals The class this PR's fix belongs to, and the reason it took a second incident to notice: a guard written after an incident encoded that incident's literal values instead of the property that made it wrong, so the next instance passed while the check reported success. evidence_prs is empty because this PR is the first instance; the adding commit is the evidence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> | 12 天前 | |
Enforce lifecycle headers for rollout scaffolding | 1 个月前 | |
fix(desktop): repair Xcode 16.4 XCTest concurrency Remove only direct XCTestCase lifecycle calls that transfer @MainActor test instances across the Xcode 16.4 non-Sendable boundary, while retaining each suite's own setup and teardown.\n\nAdd a fixture-tested manifest guard before the full Swift build and save SwiftPM state after failed test attempts so retrying diagnostics can reuse validated incremental work.\n\nVerified: desktop/macos/scripts/swift-test-suites.sh (300 suites, 4 workers); actionlint .github/workflows/desktop-swift-ci.yml; targeted guard and workflow contract tests. | 1 个月前 | |
fix(monitoring): accept the Cloud Run metrics egress filter at Cloud Monitoring (#12099) * fix(monitoring): accept the Cloud Run metrics egress filter at Cloud Monitoring The dedicated Stackdriver exporter added in #11998 has never imported a single series. Cloud Monitoring rejects a filter that mixes AND with OR across resource.labels restrictions, so every descriptor query returned HTTP 400 while the exporter stayed Available, its Prometheus target stayed up, and Grafana showed empty panels that read as no traffic. Express the namespace disjunction as one_of(...), which the filter grammar defines for this case. Keep the namespace scope: dropping it would import every omi_ series from every Cloud Run service in the project. Add omi-cloud-run-metrics-egress-query-rejected, alerting on the exporter's own upstream error rather than on its liveness. Extend the exporter contract test to cover the dev values file, which was unasserted and is what the automatic post-merge rollout installs. Correct the runbook's verification step, which queried the mangled metric name that a healthy deployment no longer produces. * fix(monitoring): key the Cloud Run observer exemption on the monitored resource The production-data-plane-routing guard exempts the Stackdriver egress values files from its retired-GKE-desktop-backend rule only when they contain the literal resource.labels.namespace="desktop-backend". That pins the exemption to one spelling of a filter rather than to what makes the file a Cloud Run observer, so rewriting the disjunction as one_of(...) to satisfy Cloud Monitoring's grammar made a read-only metrics reader look like retired GKE ownership. Key the exemption on resource.labels.cluster="__run__" instead. That is Cloud Run's reserved pseudo-cluster, so together with the prometheus.googleapis.com/ prefix it identifies the monitored resource directly and survives any future edit to the namespace set. --------- Co-authored-by: r <r@r> | 15 天前 | |
ci: collapse desktop beta to signed-smoke plus hourly freshness (#11588) * ci: collapse desktop beta to signed-smoke plus hourly freshness Qualification never rolled back recent signed-smoke manifests and starved the planner when a push event was missed. Remove the lane, make source-gate failures diagnosable from one command, and alarm when candidate and live beta diverge. Co-authored-by: Cursor <cursoragent@cursor.com> * docs: retarget INV-BETA-1 guards after deleting qualification tests The auto-beta-candidate script is gone with the qualification lane; keep the locked beta-identity invariant pointing at a guard that still exists. Co-authored-by: Cursor <cursoragent@cursor.com> * ci: make skipped desktop checks legible and let a green tip unblock the train Two failure modes survived the beta-train collapse and are fixed here. Skipped jobs published the wrong check name. GitHub does not evaluate a job's `name:` for a SKIPPED job, so the conditional names on `desktop-swift` and `desktop-swift-release-compile` were published verbatim as the raw expression text. On every commit that did not touch desktop paths the required check `Desktop Swift Build & Tests` was therefore ABSENT rather than skipped, and the planner reported "missing" instead of the truth. Observed on f666ddd4a3, 7a79f08329 and 7d7ed62e5, all of which read green. The conditional existed to keep a merged `pull_request.closed` bookkeeping run from publishing a skipped required check onto the merge SHA; dropping the `closed` event removes that hazard at the source and lets both names be literals. A contract test now rejects any expression in a job name. A green tip did not unblock the train. The planner selects the newest desktop-touching commit and, when its checks are red, could only fall back to an OLDER green SHA. On Aug 14 main's tip was green while the newest desktop-touching commit below it was red on a flaky Swift suite, so the train shipped stale code or wedged. A first-parent commit above the blocked SHA contains everything the blocked SHA contains, so its own exact-SHA checks tested a superset of that tree; when they are genuinely green the train may ship from that newer SHA. Tried before the backward fallback, because it ships newer code. Only a real `ready` gate qualifies, so a skipped or absent check still never counts as success. Also keep the beta rollback precondition expressible: beta manifests carry the `signed-smoke` tier, whose frozen-schema truth is `qualification_passed: False`, so the literal T2/True requirement rejected every current rollback target. Failure-Class: none Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): keep the release lifecycle helper compiling under the admin es5 target `newestSparkleVersion` iterated `String.matchAll()` with `for...of`. The admin package sets "target": "es5", where iterating an IterableIterator is TS2802, so `npm run typecheck` failed and took the Web Checks Build job red. Local pre-push does not typecheck the Next.js admin app, so CI was the first place this could surface. Use `exec` loops instead of widening the package's compile target, which would change output for every file to fix one. Verified with the same commands CI runs: `npm run typecheck` clean and `npm test` 88 passed across 14 files. Failure-Class: none Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> | 23 天前 | |
build(devops): add offline OpenTofu foundation guard Establish an intentionally empty OpenTofu foundation module for #9842.\n\nThe module has a GCS backend shape, separate environment placeholders, and a locked Google provider. Its guard allows only durable identity, additive IAM, secret metadata, and state-bucket resource families; it rejects release resources, data sources, and secret values. The validation workflow is credentials-free and can only validate the source plus plan a backend-free temporary copy.\n\nVerification:\n- python3 .github/scripts/test_check_opentofu_foundation.py\n- python3 .github/scripts/check_opentofu_foundation.py --plan-json /tmp/omi-opentofu.wkAqDX/foundation-clean-plan.json\n- actionlint -shellcheck '' .github/workflows/opentofu-foundation-validate.yml\n- checksum-verified OpenTofu 1.12.4 init -backend=false, validate, and offline -refresh=false empty plan with Google credentials unset | 1 个月前 | |
Document memory architecture and enforce package maps | 1 个月前 | |
ci: regression guardrails — PR scope advisor, test-discovery ratchet, sync annotation ratchet Three mechanical regression guardrails derived from an internal audit of merged PRs: 1. PR scope advisor (advisory-only, never blocks): warns at 1,500 and 3,000 changed production-source lines so reviewers calibrate skepticism to diff size. 2. Unit-test discovery ratchet: fails when any test file is not discovered by a verified runner. Allowlists only shrink. 3. Sync parameter annotation ratchet: pyright executionEnvironments enforces reportMissingParameterType=error for utils/sync/. Co-authored-by: skanderkaroui <skanderkaroui@users.noreply.github.com> | 1 个月前 | |
ci: make CI fast and reliable — every job under 20 minutes, recurring flake removed (#12247) * ci(desktop-windows): cache electron & electron-builder downloads The Linux package job failed on main (EAI_AGAIN github.com) and on PRs (read ECONNRESET) because electron-builder re-downloads its packaging tools on every run. Cache the tool and Electron archives so steady-state builds never depend on those endpoints, on both the Linux and Windows packaging jobs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(hermetic-e2e): stop minting red checks from cancelled runs; harden sync stack readiness The merge gate ran under always() and failed hard on 'cancelled' upstream results whenever cancel-in-progress superseded a run — 3 of the last 4 re-run-to-green cycles on this workflow were exactly that. Gate now goes neutral on cancellation and defers to the superseding run. The sync Cloud Tasks stack's one genuine flake: scenarios run back-to-back, a dying child of the previous stack can still answer a TCP probe, and the authoritative health check had the tightest budget (20s) while the weak port probe had the generous one. close() now drains the stack's ports and health gets 60s. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(desktop-swift): collapse macOS-runner demand that was queueing every run Desktop Swift CI's tail (p90 150 min, max 218 min) was queue time: the repo's only macos-15 consumers demanded ~4.9 concurrent runners against a cap of ~5. Cut the demand instead of the deadline: - Run the release-mode UserNotifications regression inside the Release Compile job, next to the release build it consumes, instead of paying a second from-scratch release build (~31 min) on the verify runner. The planner still requires the Release Compile check by exact name, and the Build & Tests aggregate now hard-fails on the release job's result, so release tagging and PR merges both still gate on it (#11373/#11374 protection preserved). - Cache SwiftPM dependencies and the release Desktop/.build in the Release Compile job, saved from main pushes so every ref can restore it, and drop the forced rm -rf that guaranteed a ~20-min cold build. - Run launcher script tests 3-way parallel with per-test logs (9.3 min serial p50), and widen the build-lock recovery budgets that saturated runners pushed past 4s. - blob:none the verify checkout; full history stays for diffing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(mobile,web,parakeet): kill checkout/cache tails and 45-min unschedulable burns Mobile and Web Checks' p90-to-max tail was almost entirely actions/checkout pulling all blobs of a ~1.25GB repo with fetch-depth: 0 (a 53-minute checkout-only Web run happened); history-needing jobs now use blob:none and worktree-only jobs go shallow, with timeout caps so no job can idle unbounded again. The Android compile smoke paid a ~700s cold Gradle build every run because the default cache policy never wrote from main; main now seeds the cache PRs read, and the build_runner cache keys drop the run_id suffix that guaranteed a miss on every single run. Parakeet GPU tests have failed nightly since 08-03 burning exactly 45 min each with zero logs: the pod is unschedulable (L4 pool at capacity) and the workflow neither noticed nor captured pod events. Gate on PodScheduled within 5 minutes with the real cause in the error, stream pod logs live so deadline kills keep evidence, and include pod events in diagnostics. The capacity fix itself (NLLB squatting the parakeet pool's second GPU) is a cluster change tracked separately. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(repo-checks): make the line-count ratchet invariant under base drift; pin uv 13 of the last 17 Repo Checks failures were the line-count ratchet, and the nondeterministic ones shared one cause: the check demanded byte-exact absolute counts against a synthetic merge with origin/main, which moves under every open PR (~90 pushes/hr). A correct declaration went stale with no author action, in both directions (drifted endpoints, and previously- mandatory exceptions turning fatally 'unused' when main absorbed an edit). The declaration now approves a growth *allowance* — the delta is a property of the PR alone — and an exception the diff no longer needs warns instead of failing. Growth beyond the allowance, duplicates, malformed lines, and unsupported paths still fail. Also pin setup-uv's resolved version (matching backend/Dockerfile's UV_VERSION) so it stops fetching the astral-sh 'latest' manifest, which hard-failed a run. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci: pin setup-uv's resolved version everywhere it was fetching 'latest' Every unpinned astral-sh/setup-uv call resolves 'latest' by fetching the astral-sh/versions manifest from raw.githubusercontent.com at job start — one observed Repo Checks failure was exactly that fetch dying. Pin the same 0.11.13 the backend Dockerfiles already use, which the action resolves locally with no network call. repo-checks.yml was pinned in the previous commit; this covers the remaining call sites and the release-eligibility action (with its byte-exact contract fixture). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(desktop-swift): batch SwiftPM test invocations with per-suite fallback The debug suite spawned one 'xcrun swift test' process per discovered suite — 674 SwiftPM startups at ~5.1s each, ~28.7 min of the verify job's wall clock, dwarfing actual test execution. Workers now run suites in batches of 25 through one SwiftPM process (repeated --filter flags), with batching kept strictly optimistic: a batch that exits non-zero or times out is discarded and every member re-runs through the untouched per-suite path, so isolation semantics, per-suite budgets, and failure attribution are unchanged for anything red. The serial shared-auth cluster is never batched, and OMI_SWIFT_TEST_SUITE_BATCH_SIZE=1 restores the old behavior bit-for-bit. Also fixes a pre-existing bash-3.2 set -u break (empty build_args expansion) that made OMI_SWIFT_TEST_PREBUILD=0 kill every suite, and restores the case-pattern skip list in the workflow's launcher step that check-launcher-test-skips.py parses. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(backend-unit): fail-open module stubs, checkout tail, and honest runner verdicts The suite's recurring 'flake' was one deterministic trap: a module-scope sys.modules['utils'] stub with __path__ = [] turned every new import in routers/conversations.py into a ModuleNotFoundError for whoever ran CI next — main broke twice on 2026-08-25 and four unrelated PRs inherited the red. The stub package now carries the real package __path__, so unlisted submodules resolve to the real module instead of raising (explicit stubs still stub; verified by reproducing the incident against both versions). Checkout p90 was 267s and max 786s from full-blob clones; blob:none keeps the merge-base history scripts/changed-files needs. timeout 45→30 so a hang stops burning a concurrency slot for 25 minutes past the p90. Also two standing runner bugs found while measuring: pytest's exit 5 on total worker death was read as 'no tests selected' and reported green, and the fast-unit duration guard's verdict was silently discarded under xdist (workers now hand offenders to the controller). A single-session xdist partition of the suite was implemented, measured, and disproved on this tree (upb descriptor segfaults, OOM at 815 files, per-worker collection); it ships opt-in via BACKEND_PYTEST_PARALLEL_SESSION=1 with the measurements in the comments, default behavior unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(desktop): changelog fragment for internal CI runner changes Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(desktop-swift): keep measure-block suites out of batches; three workers The first PR run of the batched suite exposed one pathology: the batch carrying MemoryAtlasPerformanceHarnessTests (245s of XCTest measure blocks, slower still under contention) blew its 17-minute budget and paid the 25-suite isolated fallback on top — 15 of the step's 40 minutes. Measure suites are now derived from the source (like the serial cluster, so a new one cannot silently join the batch pool) and run in their own SwiftPM process with a doubled per-suite budget. With per-invocation SwiftPM startup gone — the reason two workers were the measured ceiling — workers now match the runner's three cores. This PR's CI runs are the macOS-runner evidence; drop back to two if the performance harness starts timing out. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> | 12 天前 | |
docs: take operator pages off docs.omi.me Unlisted Mintlify MDX is still a public URL. Move runbooks, flags, invariants, and agent rules next to owning code, add docs/AGENTS.md as the site allow-list, and correct the live kill-switch contract after the JIT authority page leaves the site. Co-authored-by: Cursor <cursoragent@cursor.com> | 7 天前 | |
ci(public-build): env-var removal, per-environment flags, real invoker identity checks (#12582) * ci(public-build): env-var removal, per-environment flags, real invoker identity checks Plaintext Cloud Run env vars survived merge deploys, restricted ingress was applied to development, and TBD placeholders passed presence checks into gcloud. Co-authored-by: Cursor <cursoragent@cursor.com> * ci(public-build): probe actAs via IAM testIamPermissions REST gcloud has no iam service-accounts test-iam-permissions subcommand, so the previous preflight would fail every prod deploy. Call the IAM REST method with urllib and a print-access-token bearer instead. * ci(public-build): reject remove_runtime_env_vars overlapping preserved secrets Cubic review (PRRT_kwDOLkKqys6eWSS0): a runtime name in both preserve_runtime_secrets and remove_runtime_env_vars loaded cleanly, yet deployment emits --remove-env-vars for a binding the contract claims to preserve via the merge update strategies — the removal would strip the preserved secret binding. Extend the dedicated overlap rejection to preserve_runtime_secrets and its mirrored fallback_runtime_secrets, with a regression test. --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: David Zhang <9387252+Git-on-my-level@users.noreply.github.com> | 5 天前 | |
ci: make CI fast and reliable — every job under 20 minutes, recurring flake removed (#12247) * ci(desktop-windows): cache electron & electron-builder downloads The Linux package job failed on main (EAI_AGAIN github.com) and on PRs (read ECONNRESET) because electron-builder re-downloads its packaging tools on every run. Cache the tool and Electron archives so steady-state builds never depend on those endpoints, on both the Linux and Windows packaging jobs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(hermetic-e2e): stop minting red checks from cancelled runs; harden sync stack readiness The merge gate ran under always() and failed hard on 'cancelled' upstream results whenever cancel-in-progress superseded a run — 3 of the last 4 re-run-to-green cycles on this workflow were exactly that. Gate now goes neutral on cancellation and defers to the superseding run. The sync Cloud Tasks stack's one genuine flake: scenarios run back-to-back, a dying child of the previous stack can still answer a TCP probe, and the authoritative health check had the tightest budget (20s) while the weak port probe had the generous one. close() now drains the stack's ports and health gets 60s. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(desktop-swift): collapse macOS-runner demand that was queueing every run Desktop Swift CI's tail (p90 150 min, max 218 min) was queue time: the repo's only macos-15 consumers demanded ~4.9 concurrent runners against a cap of ~5. Cut the demand instead of the deadline: - Run the release-mode UserNotifications regression inside the Release Compile job, next to the release build it consumes, instead of paying a second from-scratch release build (~31 min) on the verify runner. The planner still requires the Release Compile check by exact name, and the Build & Tests aggregate now hard-fails on the release job's result, so release tagging and PR merges both still gate on it (#11373/#11374 protection preserved). - Cache SwiftPM dependencies and the release Desktop/.build in the Release Compile job, saved from main pushes so every ref can restore it, and drop the forced rm -rf that guaranteed a ~20-min cold build. - Run launcher script tests 3-way parallel with per-test logs (9.3 min serial p50), and widen the build-lock recovery budgets that saturated runners pushed past 4s. - blob:none the verify checkout; full history stays for diffing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(mobile,web,parakeet): kill checkout/cache tails and 45-min unschedulable burns Mobile and Web Checks' p90-to-max tail was almost entirely actions/checkout pulling all blobs of a ~1.25GB repo with fetch-depth: 0 (a 53-minute checkout-only Web run happened); history-needing jobs now use blob:none and worktree-only jobs go shallow, with timeout caps so no job can idle unbounded again. The Android compile smoke paid a ~700s cold Gradle build every run because the default cache policy never wrote from main; main now seeds the cache PRs read, and the build_runner cache keys drop the run_id suffix that guaranteed a miss on every single run. Parakeet GPU tests have failed nightly since 08-03 burning exactly 45 min each with zero logs: the pod is unschedulable (L4 pool at capacity) and the workflow neither noticed nor captured pod events. Gate on PodScheduled within 5 minutes with the real cause in the error, stream pod logs live so deadline kills keep evidence, and include pod events in diagnostics. The capacity fix itself (NLLB squatting the parakeet pool's second GPU) is a cluster change tracked separately. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(repo-checks): make the line-count ratchet invariant under base drift; pin uv 13 of the last 17 Repo Checks failures were the line-count ratchet, and the nondeterministic ones shared one cause: the check demanded byte-exact absolute counts against a synthetic merge with origin/main, which moves under every open PR (~90 pushes/hr). A correct declaration went stale with no author action, in both directions (drifted endpoints, and previously- mandatory exceptions turning fatally 'unused' when main absorbed an edit). The declaration now approves a growth *allowance* — the delta is a property of the PR alone — and an exception the diff no longer needs warns instead of failing. Growth beyond the allowance, duplicates, malformed lines, and unsupported paths still fail. Also pin setup-uv's resolved version (matching backend/Dockerfile's UV_VERSION) so it stops fetching the astral-sh 'latest' manifest, which hard-failed a run. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci: pin setup-uv's resolved version everywhere it was fetching 'latest' Every unpinned astral-sh/setup-uv call resolves 'latest' by fetching the astral-sh/versions manifest from raw.githubusercontent.com at job start — one observed Repo Checks failure was exactly that fetch dying. Pin the same 0.11.13 the backend Dockerfiles already use, which the action resolves locally with no network call. repo-checks.yml was pinned in the previous commit; this covers the remaining call sites and the release-eligibility action (with its byte-exact contract fixture). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(desktop-swift): batch SwiftPM test invocations with per-suite fallback The debug suite spawned one 'xcrun swift test' process per discovered suite — 674 SwiftPM startups at ~5.1s each, ~28.7 min of the verify job's wall clock, dwarfing actual test execution. Workers now run suites in batches of 25 through one SwiftPM process (repeated --filter flags), with batching kept strictly optimistic: a batch that exits non-zero or times out is discarded and every member re-runs through the untouched per-suite path, so isolation semantics, per-suite budgets, and failure attribution are unchanged for anything red. The serial shared-auth cluster is never batched, and OMI_SWIFT_TEST_SUITE_BATCH_SIZE=1 restores the old behavior bit-for-bit. Also fixes a pre-existing bash-3.2 set -u break (empty build_args expansion) that made OMI_SWIFT_TEST_PREBUILD=0 kill every suite, and restores the case-pattern skip list in the workflow's launcher step that check-launcher-test-skips.py parses. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(backend-unit): fail-open module stubs, checkout tail, and honest runner verdicts The suite's recurring 'flake' was one deterministic trap: a module-scope sys.modules['utils'] stub with __path__ = [] turned every new import in routers/conversations.py into a ModuleNotFoundError for whoever ran CI next — main broke twice on 2026-08-25 and four unrelated PRs inherited the red. The stub package now carries the real package __path__, so unlisted submodules resolve to the real module instead of raising (explicit stubs still stub; verified by reproducing the incident against both versions). Checkout p90 was 267s and max 786s from full-blob clones; blob:none keeps the merge-base history scripts/changed-files needs. timeout 45→30 so a hang stops burning a concurrency slot for 25 minutes past the p90. Also two standing runner bugs found while measuring: pytest's exit 5 on total worker death was read as 'no tests selected' and reported green, and the fast-unit duration guard's verdict was silently discarded under xdist (workers now hand offenders to the controller). A single-session xdist partition of the suite was implemented, measured, and disproved on this tree (upb descriptor segfaults, OOM at 815 files, per-worker collection); it ships opt-in via BACKEND_PYTEST_PARALLEL_SESSION=1 with the measurements in the comments, default behavior unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(desktop): changelog fragment for internal CI runner changes Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(desktop-swift): keep measure-block suites out of batches; three workers The first PR run of the batched suite exposed one pathology: the batch carrying MemoryAtlasPerformanceHarnessTests (245s of XCTest measure blocks, slower still under contention) blew its 17-minute budget and paid the 25-suite isolated fallback on top — 15 of the step's 40 minutes. Measure suites are now derived from the source (like the serial cluster, so a new one cannot silently join the batch pool) and run in their own SwiftPM process with a doubled per-suite budget. With per-invocation SwiftPM startup gone — the reason two workers were the measured ceiling — workers now match the runner's three cores. This PR's CI runs are the macOS-runner evidence; drop back to two if the performance harness starts timing out. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> | 12 天前 | |
fix: dispatch every macOS candidate directly to Codemagic Failure-Class: FC-release-train-busy-main-starvation | 18 天前 | |
fix(infra): harden admitted GCP deploy control plane Resolve the main merge and repair deploy control-source staging, rollback, admission, and runtime configuration guards. Failure-Class: FC-workflow-control-source-identity | 1 个月前 | |
ci: move paid Actions jobs to standard runners (SCA-196) Co-authored-by: multica-agent <github@multica.ai> | 1 个月前 | |
test(ci): drop the guard cases for the retired writer scan 8 tests, all passing: the outcome, accept and client-protocol facts the guard still enforces. | 8 天前 | |
test: update version filename fixture | 26 天前 | |
Admin: real cost data and an honesty sweep for the Omi TV dashboard (#12198) * Infra costs: replace the April estimate table with billed data computeInfraCosts now has a billing mode as the authoritative path: - GCP: daily net spend (cost + credits) from the BigQuery billing export, _PARTITIONTIME-filtered, series ending at D-2 because the export lags ~11h and back-fills — per the gcp-cost-efficiency cost-monitoring contract. New lib/services/gcp-billing.ts (@google-cloud/bigquery dep, creds via GCP_BILLING_SA_JSON or ADC). - Anthropic/OpenAI: org cost APIs (lib/services/provider-costs.ts) behind ADMIN_ANTHROPIC_COST_API_KEY / ADMIN_OPENAI_COST_API_KEY — names chosen because the deploy contract strips ANTHROPIC_API_KEY/OPENAI_API_KEY. A missing leg is partial coverage, never silently $0. - Desktop/mobile split by the omi-cost-analysis usage-weighted shares (LLM pool by memory-event mix, rest by core-event mix), stored as dated config (ADMIN_PLATFORM_COST_SHARES_JSON), replacing the hand-guessed per-service weights. Firestore llm_usage is attribution signal only — its cost_usd rows are provider spend already inside the invoices. - Payload gains costSource/windowEnd/coverage/shares; the legacy hardcoded-table path survives only as a labeled 'estimated' fallback. - profitability: cost and cost-per-user series end at the billing window's D-2 edge instead of padding recent days with the per-user assumption; averages use the same trimmed window; costSource='real' only for billing-backed numbers. Verified: vitest 19 files / 135 tests pass incl. new lib/__tests__/infra-costs-billing.test.ts; tsc --noEmit clean; live fetchGcpBilling(14) against the real export returns windowEnd=2026-08-22 with values penny-identical to the raw bq query. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Wire billing-mode cost secrets into the admin deploy contract GCP_BILLING_SA_JSON -> WEB_ADMIN_GCP_BILLING_SA_JSON (new dedicated SA omi-admin-billing-ro@based-hardware, bigquery.jobUser + dataset READER on gcp_billing_export only; key stored in Secret Manager, runtime SA granted secretAccessor). The two provider cost-API secrets exist as placeholders until org-admin keys are minted — the code treats a rejected key as an unavailable leg (partial), never $0. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Split Anthropic spend by chat mix, not the extraction-token mix The 2026-08-24 methodology audit found the tracked-token base is ~99.9% conversation extraction with essentially no chat in it, while the Anthropic bill in the same window was essentially all chat (floating bar, desktop-backend leak) — splitting Anthropic by the memory-token mix pushed $2.2-3.6k of desktop's own Claude spend onto mobile. Anthropic now splits by a dedicated chat-event share pair (default 0.5464/0.4536, PostHog chat mix); Vertex/Gemini and OpenAI stay on the memory-mix split, which the audit confirmed sound for extraction models. Tests updated; 4/4 pass, tsc clean. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Thread app_platform from clients into gateway accounting The gateway ledger had no platform dimension, so desktop vs mobile chat_agent spend could only be share-split. llm_gateway_headers() now takes an optional platform, emitted as X-Omi-App-Platform only when it normalizes into {desktop, mobile, web} — client-supplied junk never reaches an outbound header and never 403s a request (accounting metadata, not auth). desktop_chat passes its X-App-Platform through; desktop_proactivity hardcodes desktop. The gateway parses the header in ServiceCaller (non-rejecting validator), carries it via AccountingContext into the ledger event as app_platform (null when unknown). 480 backend tests incl. new header/accounting coverage pass; pyright and black clean. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Admin: measure LLM platform split from the gateway ledger; leak metric New lib/services/gateway-ledger.ts reads the llm_gateway_attempts Firestore ledger (~200k docs/day) exclusively through server-side SUM aggregations — never document reads — with per-day immutable caching (volatile last 3 days recomputed). Features roll up to platform classes (desktop-only, shared-chat, shared-extraction, unknown). computeInfraCosts billing mode now derives each day's measured desktop share of extraction-LLM spend from the ledger composition instead of applying the static share directly to the whole pool (desktop-only lanes like desktop_proactive_* finally land on desktop), and emits summary.directPath = provider invoice minus gateway-billed spend — a standing detector for direct-path leaks like the Sonnet 4.6 incident. Missing ledger days keep static shares; a missing invoice leg omits its leak key rather than reading as $0. SUM aggregations need composite indexes including the summed field (verified live — index-merge only serves plain queries), so the six (date, [feature|provider], [payer], estimated_cost_micro_usd) indexes are registered in firestore_index_registry.py and the regenerated manifest; they are already building in prod. Verified: web/admin 142 tests + tsc clean; backend index tests 43 pass after manifest regeneration. Failure-Class: none Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Admin: rebuild the onboarding funnel from the app's real step list STEP_DEFINITIONS had drifted from OnboardingView.swift for months: it still counted a removed Notifications step and a Research step dead since Apr 2026 (both permanent ~0 rows), while HowDidYouHear, DataSources and Exports were missing entirely and Goal lost its skip variant. The list and funnel logic now live in lib/onboarding-funnel.ts, and a test parses OnboardingView.swift directly so the next Swift-side rename fails CI instead of drifting. The grouped query's LIMIT now matches the wrapper's 50k served ceiling and the response carries a truncated flag when it hits the cap. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Admin: make the Stripe metrics honest about what they measure - mrr-trends prices every month at today's prices (a price change rewrites history); the payload now carries pricingBasis: current_prices and the route documents the limitation instead of posing as a historical series. - trialing counts return null on fetch failure instead of a fabricated 0. - monthlyAmount warns loudly on unknown intervals instead of silently adding $0, and excludes non-USD prices from USD totals (surfaced as nonUsdSkipped) instead of blending currencies. - app-subscriptions now counts past_due like every other MRR route, and gains partial/502 degradation matching its siblings. Cache keys bumped v2->v3 so old-shape payloads cannot be served. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Admin: attribute the unknown gateway spend; stop inventing per-user costs The FEATURE_CLASS map only listed model_config feature names, but the gateway's usage-context feature header overrides those on the wire — so the coarse usage-context names (chat, persona, conversation_processing, ...) were the bulk of the unknown bucket. Twelve real features added with verified classifications (workstream_association and chat_structured are desktop-only lanes). Precompute no longer bakes the invented desktop_cost=1.2/mobile_cost=0.3 params into the profitability cache; it now writes the same key the GET route's default parse looks up. In billing mode a zero-active day reports costPerUser null instead of the April $0.20 assumption, and summary averages skip unmeasurable days. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Admin: surface data age and every silent degradation in the stats layer Every payload-cache consumer now stamps freshAt on both cache-hit and fresh-compute paths, so a broken precompute cron no longer serves stale data indistinguishable from live; the cron itself now reports and records ok/failed per metric (precompute-status:v1). macos-versions derives its date label at serve time instead of baking 'today' into the cached payload. Notifications counts explicit true/false/unset instead of counting docs missing the field as enabled. Fabricated values become honest nulls or flags: message-ratings ratio null on no-data days, truncation flags on the capped notification fallback and macos-versions query, voiceSource names which event source reliability metrics used, activation erroredUsers is no longer dropped. crash-rate date keys use UTC to match PostHog bucketing. The classic dashboard stops sending the invented desktop_cost/mobile_cost params and renders unmeasurable cost/user days as gaps, not $0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Guard: scope SCA-118 to inference traffic, not org-admin reporting The admin dashboard's cost legs read spend from the providers' organization namespaces (cost_report / organization/costs), which carry no model traffic. The gateway-only guard now exempts exactly those namespaces via lookahead — inference URLs still fail — with fixtures for both sides. provider-costs.ts stops naming the forbidden inference env vars in a comment. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: r <r@r> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> | 13 天前 | |
fix(ci): follow bounded apt-get into composite actions, and cap it at the caller (#12270) * fix(ci): follow bounded apt-get into composite actions, and cap it at the caller Rebased onto main. The desktop Swift lanes that held this approved PR were red on main itself; #12277 fixed both causes, so this rebase is what turns them green. The checks-manifest edit is re-applied to main's text rather than overwriting it, so the legacy-memory-surface-ratchet entry main added since the merge base survives. One prose fix in the reason field: #12194 is still open, so it proposes the composite action rather than having moved it. * fix(ci): restore the executable bit on check_workflow_apt_network_bounds.py The rebase onto main brought this file across at mode 100644. It is 100755 on main today, and this PR modifies it, so merging as it stood would have changed the file's mode on main -- a tree change no line of the diff describes. Mode only. The blob is byte-identical to the previous head: the tree is built from that tree with one entry, reusing the same blob sha. | 6 天前 | |
fix(ci): isolate candidate probe deadlines on Windows (#12344) Keep orthogonal candidate-probe fixtures from entering the POSIX-only signal deadline boundary. Apply an autospecced suite-wide seam while retaining the real deadline contract in a dedicated platform-gated test class. Tests: candidate probe suite passed on Windows with 17 passed and 1 POSIX-only skip; Ruff, Black, py_compile, and diff checks passed. Failure-Class: new Co-authored-by: 491034170 <142008960+491034170@users.noreply.github.com> | 6 天前 | |
fix: make no-changelog-needed survive the merge boundary The PR label greened PRs and then reddened main because push runs cannot see labels. Require an in-repo kind:none fragment for internal production desktop edits so both lanes agree. Failure-Class: FC-changelog-exemption-lost-at-merge Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: multica-agent <github@multica.ai> | 20 天前 | |
fix(ci): land leftover desktop-beta review fixes from #11588 (#11605) * fix(ci): redact recovery logs and keep qualification reintroduction tested Codemagic log tails could persist query tokens and GitHub PATs in durable issues, the freshness monitor lacked Actions/checks read, and the qualification-trigger mutation test was dropped with the old runner. Co-authored-by: Cursor <cursoragent@cursor.com> * style: drop extra blank line that failed diff-hygiene Co-authored-by: Cursor <cursoragent@cursor.com> * fix(ci): redact OAuth fragments without wiping triage query params Generic query redaction leaked `#code=` values and blanked harmless build/arch/mode fields. Scope assignment redaction to credential-like names, treat `#` as a delimiter, and drop the comment-only qualification mutation that locked in a substring false positive. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> | 23 天前 | |
ci: consolidate impact routing | 1 个月前 | |
fix(ci): keep unconfigured proactive monitor neutral (#11483) * fix(ci): keep unconfigured proactive monitor neutral * docs(ci): align proactive monitor evidence contract | 25 天前 | |
fix: address cubic review feedback — validation, test, timeout, and messaging - desktop_release_manifest: reject non-boolean qualification_passed to prevent 0 == False set-membership bypass (numeric 0 admitted as 'passed: False') - test_desktop_release_manifest: re-indent assertRaisesRegex back inside the subTest loop so both cases run; add regression test for numeric truthiness - codemagic.yaml: add per-request connect/total timeouts (--connect-timeout 10 --max-time 30) to the Beta promotion curl loop so hung requests retry promptly - beta_breakglass_evidence: parameterize _smoke_evidence error messages by caller label so normal signed-Beta failures don't read 'emergency target' - update codemagic workflow contract fixture digests for the codemagic.yaml edit Addresses cubic threads: #3756205182 (0==False), #3756205195 (dedented test), #3756205198 (no curl timeouts), #3756205221 (misleading error messages). Failure-Class: none | 27 天前 | |
fix(release): harden direct beta evidence admission Bind both stable and Beta signed-smoke results concurrently, align release metadata and tag contracts, and close the remaining review guard and test gaps without restoring qualification or adding a sequential release wait. Failure-Class: none | 27 天前 | |
fix: keep the macOS hourly release train moving Bind each candidate to its checked source, recover squash-orphaned changelog branches, install Bun for release eligibility, and age freshness from the oldest unshipped update. Verification: 77 focused release-control tests; actionlint; Python compile; diff check. Failure-Class: new | 18 天前 | |
fix(ci): keep desktop main health current (#12285) | 10 天前 | |
ci: harden desktop follow-up verification Keep merged-PR close events from cancelling exact-SHA macOS evidence and retry transient Electron postinstall downloads in the Linux package smoke. Add workflow contract coverage for both PR #11447 failures. | 25 天前 | |
test(ci): make mobile cadence fixture hermetic | 1 个月前 | |
fix(ci): preserve reconciler Cloud Run job arguments Quote the complete --args token for deploy-cloudrun and use one verified 100%-traffic revision during desktop backend rollback and recovery. Failure-Class: FC-agent-vm-bootstrap-contract | 1 个月前 | |
feat(ci): schedule the failure-class retirement the lifecycle already defines `scripts/failure-class report` computes closure eligibility with a quiet period and requires maintainer confirmation, but it is non-mutating and has no event source: without `--events-file` it returns zero events and a `no_event_source` warning. Nothing scheduled ever supplied that feed, so 26 classes sit at `status: open` with zero retirements since the registry was created. Adds a weekly Repo Hygiene workflow (Mon 09:00 UTC, clear of the release cadence) that builds the merged-PR feed, runs the report, applies the two-field retirement edit, and opens one PR on a reused branch. It never merges and never decides; the reviewable diff is the output, and merging is the confirmation. Three things this deliberately does NOT do: - gate anything: no manifest lane, no pre-push, no release coupling - open an issue for a retirement: a PR has a review queue and its own CI, an advisory issue joins a 400-issue backlog - author the PR with GITHUB_TOKEN, which would arrive with zero checks run; the Omi Bot app token is minted so the diff is verified like any human PR Recurrence of an already-dormant class exits 1 and skips the PR. That is judgment work (AGENTS.md requires a reusable guard surface, not a field flip), and a red weekly run is the notification without new machinery. The feed guard is the load-bearing part. A truncated feed makes every class report "no classified instance", so the job would run green forever while retiring nothing -- the same false-green defect as a check that executes no tests. The run refuses a feed that hit the fetch limit, or an empty one. Verification: python3 .github/scripts/test_failure_class_retirement.py -> Ran 15 tests, OK python3 .github/scripts/test_run_checks.py -> Ran 24 tests, OK (manifest valid) live dry run (real gh, 14d): 748 events spanning 2026-07-11 -> 2026-07-25, 292 carrying a Failure-Class line, 0 eligible, exit 0, no files modified time-travel apply (--now 2026-08-20, real registry copied to a temp root): 13 of 26 classes retired, 13 left open, diff is exactly - "status": "open" + "status": "dormant", + "dormant_since": "2026-08-20T00:00:00Z" per class (39 changed lines across 13 files, no formatting churn) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> | 1 个月前 | |
ci: add weekly guardrail baseline health pulse Record baseline counts via existing check counters, append JSONL history, and alert when a nonzero count has not decreased for 30 days (#9454). Failure-Class: none Co-authored-by: Cursor <cursoragent@cursor.com> | 1 个月前 | |
fix(ci): fail closed when Codemagic desktop preview dies after dispatch (#10145) (#11171) * fix(ci): fail closed when Codemagic desktop preview dies after dispatch (#10145) Dispatch returning a buildId was treated as success while Codemagic could die at startup with no artifact, so macos.omi.me/preview/<slug> silently fell back to stable. Poll the exact build to a terminal status and keep durable evidence. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(ci): validate Codemagic build IDs before observe (#10145) Reject non-alphanumeric build IDs before GITHUB_OUTPUT, pass observe args via env (not ${{ }} in run) to avoid shell injection from provider data, and drop the unused sys import for Ruff. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> | 13 天前 | |
fix(release): restore native Codemagic tag delivery Failure-Class: FC-release-build-config-source-binding Co-authored-by: multica-agent <github@multica.ai> | 1 个月前 | |
fix: enforce the macOS hourly candidate clock Failure-Class: FC-release-train-busy-main-starvation | 18 天前 | |
fix(ci): distinguish Windows exit 259 from liveness (#12347) Co-authored-by: 491034170 <142008960+491034170@users.noreply.github.com> | 6 天前 | |
fix(ci): keep desktop main health current (#12285) | 10 天前 | |
fix(ci): distinguish Windows exit 259 from liveness (#12347) Co-authored-by: 491034170 <142008960+491034170@users.noreply.github.com> | 6 天前 | |
fix(ci): probe stable channel and add break-glass probe bypass Address trusted Codex review feedback on #10883: P1: Probe both beta AND stable channels before tagging. The backend's default channel is 'stable' and it deliberately does not fall stable through to beta. A beta-only probe could pass while stable 404s, shipping a broken default cohort (fresh installs, un-toggled settings). P2: Add workflow_dispatch input 'bypass_update_feed_probe_reason' so a transient probe failure (broken probe, GitHub runner cannot reach api.omi.me) does not indefinitely block Windows hotfixes. Requires a non-empty reason tracked in the run log for audit. Tests: python3 .github/scripts/test_probe_windows_update_feed.py (10 passed, +2) | 1 个月前 | |
fix: enforce the macOS hourly candidate clock Failure-Class: FC-release-train-busy-main-starvation | 18 天前 | |
fix: keep the macOS hourly release train moving Bind each candidate to its checked source, recover squash-orphaned changelog branches, install Bun for release eligibility, and age freshness from the oldest unshipped update. Verification: 77 focused release-control tests; actionlint; Python compile; diff check. Failure-Class: new | 18 天前 | |
fix(ci): retire superseded Windows release sync PRs (#10727) (#10960) * feat(ci): retire superseded Windows release sync PRs (#10727) Each Windows release opens a release/windows-v* sync PR to stamp desktop/windows/package.json back onto main. The release tag is authoritative, so older open sync PRs are pure review noise once a newer release has a PR; nine had accumulated. Add a testable Python helper that lists open PRs whose same-repo head matches the release/windows-v* prefix, excludes the current release PR, and closes the rest as superseded with a pointer to the newest. The selection predicate is unit-tested (current PR retained, unrelated heads never selected) and wired into the checks-manifest so the contract runs in CI. Cleanup stays best-effort and non-fatal: publishing and tags remain authoritative. Co-authored-by: CommandCodeBot <noreply@commandcode.ai> * ci(desktop-windows): close superseded version-sync PRs after release After the current release's sync PR exists, invoke the retire helper so older release/windows-v* PRs targeting main are closed as superseded. Failure is non-fatal: the release tag is already published and remains the source of truth. Co-authored-by: CommandCodeBot <noreply@commandcode.ai> * fix(ci): never close fork PRs and page past the list limit cubic review follow-up on #10960: the search matched any open PR whose head starts with release/windows-v*, which could include fork-origin contributor PRs this release job must not touch. Request isCrossRepository in the gh query, default the selection to same-repo only, and add a fork fixture to the contract test. Also pass an explicit --limit so cleanup does not silently stop at the CLI default (30) after a long outage or backlog growth. Co-authored-by: CommandCodeBot <noreply@commandcode.ai> * fix(ci): make --self-test run without release args The script advertised `--self-test` as a hermetic check, but argparse required --current-pr/--version even in that mode, so the documented invocation failed before reaching the test. Make those args optional and validate them only for the real cleanup path. Co-authored-by: CommandCodeBot <noreply@commandcode.ai> * fix(ci): list Windows sync PRs without head: search gh pr list --search head:release/windows-v returns zero same-repo results, so retirement became a no-op. List open main PRs and filter by headRefName prefix locally, with a regression test on the query args. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(ci): paginate Windows sync PR retirement listing gh pr list --limit 100 truncated when main has 100+ open PRs, so older superseded release/windows-v* sync PRs could be missed. Fetch all open PRs via gh api --paginate --slurp and keep local prefix/fork filters. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: CommandCodeBot <noreply@commandcode.ai> Co-authored-by: Cursor <cursoragent@cursor.com> | 10 天前 | |
fix(ci): load the merged PR body for main-push Hygiene citations #10965 passed the squash commit message as the PR body. This repo squashes with the commit list, not the PR description, so INV-* citations that made PR Hygiene green vanish on main (#11835). Append the live (#NNNN) PR body when resolving. Failure-Class: none | 18 天前 | |
fix(release): verify pre-fix Stable candidates Failure-Class: none | 15 天前 | |
test(desktop): guard SwiftLint baseline and suppressions Ticket 07 of #9843 strict Swift enforcement program. Add a portable Python guard (check-swiftlint-baseline.py) that: 1. Compares baseline records semantically against the merge-base version. Candidate entries must be a subset of base entries — additions fail, removals pass. File paths are normalized from absolute file:// URLs to relative paths for cross-machine comparison. Bootstrap mode allows the initial baseline-introducing commit. 2. Enforces swiftlint:disable policy: every suppression must name a rule, have a local reason (-- ...), and not be a blanket disable. Net new suppressions are rejected by comparing against the base count. Adversarial fixtures (9 tests) cover: addition (rejected), removal (allowed), reorder (allowed), blanket disable (rejected), disable without reason (rejected), multi-rule without reason (rejected), valid reasoned suppression (accepted), absolute path normalization, same-file-different-line identity. The guard is portable (pure Python). The macOS lint producer (SwiftLint binary) remains platform-routed via the manifest. Verification: - test_swiftlint_baseline_guard.py: 9/9 pass - check-swiftlint-baseline.py --base origin/main --bootstrap: OK - make preflight: all checks pass (e2e carry-over from Ticket 04) | 1 个月前 | |
fix(ci): make pre-push portable on Windows Failure-Class: none Route pre-push and preflight shell contracts through Git Bash / native tools on Windows so MSYS/Unicode paths don't break manifest runner, release guards, SwiftLint, firmware checks, and dev-harness scripts. Add _bash_command() and _native_path_from_bash() helpers with Git-for-Windows discovery. Decode Git output as UTF-8. Skip POSIX-only checks on Windows. | 1 个月前 | |
Make automatic development backend deploys converge (#12019) * Make automatic development backend deploys converge Automatic development deploys succeeded 8 times in the 25 runs before this change. The 13 failures had three causes, and this addresses the two that are defects rather than configuration. Admission required the Release Eligibility proof SHA to still equal main's tip. Anything merging while eligibility ran therefore rejected a merged, reviewed commit -- 8 of the 13 failures, and why development sat a day behind main. The property that protects the runtime is that the commit is merged, so require ancestry instead. Production is untouched: it deploys only by explicit dispatch, which already required ancestor-of-main rather than tip-equality. Development also had no automatic Firestore migration path. An automatic index reconciliation resolves its environment to prod, and only a manual dispatch ever targeted development, so a merged manifest addition left development with a schema that no longer matched main and every deploy failed its readiness gate until somebody noticed -- 2 more failures, most recently the hourly_usage (year, month) index from #11979. Composite reconciliation is create-only and development carries no required reviewer, so it now converges on the same merge that queues the production migration. Production's approval gate is unchanged. The manual development lane failed separately, at custom_token_signing: its candidate audio gate authenticates against production Firebase, which a development deploy identity cannot sign a custom token for. The probe can now be told which account to sign as. Left unset the behaviour is identical, so this is inert until FIREBASE_PROBE_SIGNER_SERVICE_ACCOUNT is set and the deploy identity is granted token-creator on it. Not addressed here: 3 failures came from GCP_FIRESTORE_READONLY_CREDENTIALS being intermittently unavailable in the development environment. That credential is a deliberate privilege boundary -- readiness executes admitted source and must not hold deploy credentials -- so it wants a configuration fix, not a code change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Update release-vector contract for the development index lane The static migration contract counted --provision-missing across the whole workflow, which asserted 'only one lane applies indexes'. There are now two, one per environment, so count per job instead and pin the development lane to its own environment, concurrency group, and push-only trigger. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Resolve the newest proven main source instead of the triggering one gpt-5.6-sol's review found the previous approach incomplete in two ways, and both are real. Ancestry alone was not safe. Tip-equality was doing more than proving merge status -- it was also a currentness fence. Accepting any ancestor of main lets a late-scheduled run deploy older code than development already had, because Actions concurrency groups are not FIFO, and lets a run for a commit that has since been reverted redeploy the pre-revert tree. This runtime shares production Firestore, Firebase auth, and Stripe, so that is not benign. Ancestry alone was also not sufficient. The scope job green-no-ops any triggering SHA that main has moved past, before it ever inspects changed paths. So a backend commit still never deploys if an unrelated commit merges before scope runs: the backend commit no-ops for being behind, the unrelated commit no-ops on its own diff. The regression test claimed to cover this but built a later main SHA and never passed it to scope, so scope saw the backend commit as main's tip and the assertion proved nothing. Passing it reproduces the strand. Both follow from deploying the triggering commit. Admission now resolves the newest commit on main carrying a first-attempt successful Release Eligibility proof and reachable from current main, and deploys that. Concurrent runs converge on one target rather than racing, a revert is never undone by a late run for the commit it reverted, and a behind trigger still deploys because the target moves forward instead of the run being skipped. Scope's supersession decision is removed as now-redundant, which also deletes its two GitHub API proofs and their fixture -- the contract gains tripwires against reintroducing it. --trigger-is-ancestor-of-sha keeps the resolved target at least as new as the proof that triggered the run, so a stale listing cannot move development backwards from its own trigger. The proof listing is fetched with curl --fail and no error suppression: an unreadable listing refuses to deploy. The guard checkout assertion is gone rather than re-checked-out. sol was right that it had become true by construction and added no independent evidence, and the re-checkout it needed also made an in-flight run execute a newer guard script than the workflow that invoked it. Readiness now needs actions:read to list proofs. The manual lane's readiness job already had exactly that for exactly this lookup, so the contract now expects it for both rather than treating the automatic lane as more restricted. sol's P0 -- that the new development index lane writes to production -- does not hold: RUNTIME_GCP_PROJECT_ID is based-hardware-dev in the development environment and based-hardware in prod, so the two jobs target different projects and cannot race on the same index. It read the value from runtime_env.yaml's runtime_gcp_project rather than the deployed variable. That inconsistency between the checked-in contract and the deployed value is real and worth its own look, but it is not this lane writing to production. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> | 15 天前 | |
fix(infra): harden admitted GCP deploy control plane Resolve the main merge and repair deploy control-source staging, rollback, admission, and runtime configuration guards. Failure-Class: FC-workflow-control-source-identity | 1 个月前 | |
Make automatic development backend deploys converge (#12019) * Make automatic development backend deploys converge Automatic development deploys succeeded 8 times in the 25 runs before this change. The 13 failures had three causes, and this addresses the two that are defects rather than configuration. Admission required the Release Eligibility proof SHA to still equal main's tip. Anything merging while eligibility ran therefore rejected a merged, reviewed commit -- 8 of the 13 failures, and why development sat a day behind main. The property that protects the runtime is that the commit is merged, so require ancestry instead. Production is untouched: it deploys only by explicit dispatch, which already required ancestor-of-main rather than tip-equality. Development also had no automatic Firestore migration path. An automatic index reconciliation resolves its environment to prod, and only a manual dispatch ever targeted development, so a merged manifest addition left development with a schema that no longer matched main and every deploy failed its readiness gate until somebody noticed -- 2 more failures, most recently the hourly_usage (year, month) index from #11979. Composite reconciliation is create-only and development carries no required reviewer, so it now converges on the same merge that queues the production migration. Production's approval gate is unchanged. The manual development lane failed separately, at custom_token_signing: its candidate audio gate authenticates against production Firebase, which a development deploy identity cannot sign a custom token for. The probe can now be told which account to sign as. Left unset the behaviour is identical, so this is inert until FIREBASE_PROBE_SIGNER_SERVICE_ACCOUNT is set and the deploy identity is granted token-creator on it. Not addressed here: 3 failures came from GCP_FIRESTORE_READONLY_CREDENTIALS being intermittently unavailable in the development environment. That credential is a deliberate privilege boundary -- readiness executes admitted source and must not hold deploy credentials -- so it wants a configuration fix, not a code change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Update release-vector contract for the development index lane The static migration contract counted --provision-missing across the whole workflow, which asserted 'only one lane applies indexes'. There are now two, one per environment, so count per job instead and pin the development lane to its own environment, concurrency group, and push-only trigger. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Resolve the newest proven main source instead of the triggering one gpt-5.6-sol's review found the previous approach incomplete in two ways, and both are real. Ancestry alone was not safe. Tip-equality was doing more than proving merge status -- it was also a currentness fence. Accepting any ancestor of main lets a late-scheduled run deploy older code than development already had, because Actions concurrency groups are not FIFO, and lets a run for a commit that has since been reverted redeploy the pre-revert tree. This runtime shares production Firestore, Firebase auth, and Stripe, so that is not benign. Ancestry alone was also not sufficient. The scope job green-no-ops any triggering SHA that main has moved past, before it ever inspects changed paths. So a backend commit still never deploys if an unrelated commit merges before scope runs: the backend commit no-ops for being behind, the unrelated commit no-ops on its own diff. The regression test claimed to cover this but built a later main SHA and never passed it to scope, so scope saw the backend commit as main's tip and the assertion proved nothing. Passing it reproduces the strand. Both follow from deploying the triggering commit. Admission now resolves the newest commit on main carrying a first-attempt successful Release Eligibility proof and reachable from current main, and deploys that. Concurrent runs converge on one target rather than racing, a revert is never undone by a late run for the commit it reverted, and a behind trigger still deploys because the target moves forward instead of the run being skipped. Scope's supersession decision is removed as now-redundant, which also deletes its two GitHub API proofs and their fixture -- the contract gains tripwires against reintroducing it. --trigger-is-ancestor-of-sha keeps the resolved target at least as new as the proof that triggered the run, so a stale listing cannot move development backwards from its own trigger. The proof listing is fetched with curl --fail and no error suppression: an unreadable listing refuses to deploy. The guard checkout assertion is gone rather than re-checked-out. sol was right that it had become true by construction and added no independent evidence, and the re-checkout it needed also made an in-flight run execute a newer guard script than the workflow that invoked it. Readiness now needs actions:read to list proofs. The manual lane's readiness job already had exactly that for exactly this lookup, so the contract now expects it for both rather than treating the automatic lane as more restricted. sol's P0 -- that the new development index lane writes to production -- does not hold: RUNTIME_GCP_PROJECT_ID is based-hardware-dev in the development environment and based-hardware in prod, so the two jobs target different projects and cannot race on the same index. It read the value from runtime_env.yaml's runtime_gcp_project rather than the deployed variable. That inconsistency between the checked-in contract and the deployed value is real and worth its own look, but it is not this lane writing to production. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> | 15 天前 | |
fix: scope gateway proof rerun enforcement Failure-Class: none | 1 个月前 | |
fix(desktop): reject attestation lineage child Failure-Class: FC-runtime-image-lineage Co-authored-by: multica-agent <github@multica.ai> | 1 个月前 | |
ci: collapse desktop beta to signed-smoke plus hourly freshness (#11588) * ci: collapse desktop beta to signed-smoke plus hourly freshness Qualification never rolled back recent signed-smoke manifests and starved the planner when a push event was missed. Remove the lane, make source-gate failures diagnosable from one command, and alarm when candidate and live beta diverge. Co-authored-by: Cursor <cursoragent@cursor.com> * docs: retarget INV-BETA-1 guards after deleting qualification tests The auto-beta-candidate script is gone with the qualification lane; keep the locked beta-identity invariant pointing at a guard that still exists. Co-authored-by: Cursor <cursoragent@cursor.com> * ci: make skipped desktop checks legible and let a green tip unblock the train Two failure modes survived the beta-train collapse and are fixed here. Skipped jobs published the wrong check name. GitHub does not evaluate a job's `name:` for a SKIPPED job, so the conditional names on `desktop-swift` and `desktop-swift-release-compile` were published verbatim as the raw expression text. On every commit that did not touch desktop paths the required check `Desktop Swift Build & Tests` was therefore ABSENT rather than skipped, and the planner reported "missing" instead of the truth. Observed on f666ddd4a3, 7a79f08329 and 7d7ed62e5, all of which read green. The conditional existed to keep a merged `pull_request.closed` bookkeeping run from publishing a skipped required check onto the merge SHA; dropping the `closed` event removes that hazard at the source and lets both names be literals. A contract test now rejects any expression in a job name. A green tip did not unblock the train. The planner selects the newest desktop-touching commit and, when its checks are red, could only fall back to an OLDER green SHA. On Aug 14 main's tip was green while the newest desktop-touching commit below it was red on a flaky Swift suite, so the train shipped stale code or wedged. A first-parent commit above the blocked SHA contains everything the blocked SHA contains, so its own exact-SHA checks tested a superset of that tree; when they are genuinely green the train may ship from that newer SHA. Tried before the backward fallback, because it ships newer code. Only a real `ready` gate qualifies, so a skipped or absent check still never counts as success. Also keep the beta rollback precondition expressible: beta manifests carry the `signed-smoke` tier, whose frozen-schema truth is `qualification_passed: False`, so the literal T2/True requirement rejected every current rollback target. Failure-Class: none Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): keep the release lifecycle helper compiling under the admin es5 target `newestSparkleVersion` iterated `String.matchAll()` with `for...of`. The admin package sets "target": "es5", where iterating an IterableIterator is TS2802, so `npm run typecheck` failed and took the Web Checks Build job red. Local pre-push does not typecheck the Next.js admin app, so CI was the first place this could surface. Use `exec` loops instead of widening the package's compile target, which would change output for every file to fix one. Verified with the same commands CI runs: `npm run typecheck` clean and `npm test` 88 passed across 14 files. Failure-Class: none Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> | 23 天前 | |
fix(release): harden retained stable promotion Failure-Class: FC-split-mutation-authority | 1 个月前 | |
fix(infra): harden admitted GCP deploy control plane Resolve the main merge and repair deploy control-source staging, rollback, admission, and runtime configuration guards. Failure-Class: FC-workflow-control-source-identity | 1 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 3 天前 | ||
| 3 天前 | ||
| 4 天前 | ||
| 1 个月前 | ||
| 1 天前 | ||
| 11 天前 | ||
| 18 天前 | ||
| 20 天前 | ||
| 4 天前 | ||
| 13 天前 | ||
| 1 个月前 | ||
| 15 天前 | ||
| 7 天前 | ||
| 22 天前 | ||
| 22 天前 | ||
| 1 个月前 | ||
| 7 天前 | ||
| 4 天前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 15 天前 | ||
| 7 天前 | ||
| 2 天前 | ||
| 2 天前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 4 天前 | ||
| 20 天前 | ||
| 1 个月前 | ||
| 3 天前 | ||
| 14 天前 | ||
| 1 个月前 | ||
| 12 天前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 18 天前 | ||
| 12 天前 | ||
| 7 天前 | ||
| 5 天前 | ||
| 18 天前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 23 天前 | ||
| 8 天前 | ||
| 26 天前 | ||
| 13 天前 | ||
| 6 天前 | ||
| 1 个月前 | ||
| 20 天前 | ||
| 23 天前 | ||
| 30 天前 | ||
| 18 天前 | ||
| 9 天前 | ||
| 27 天前 | ||
| 23 天前 | ||
| 23 天前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 30 天前 | ||
| 2 个月前 | ||
| 13 天前 | ||
| 1 个月前 | ||
| 5 天前 | ||
| 18 天前 | ||
| 6 天前 | ||
| 6 天前 | ||
| 1 个月前 | ||
| 5 天前 | ||
| 6 天前 | ||
| 1 个月前 | ||
| 6 天前 | ||
| 18 天前 | ||
| 18 天前 | ||
| 10 天前 | ||
| 2 个月前 | ||
| 1 个月前 | ||
| 24 天前 | ||
| 5 天前 | ||
| 12 天前 | ||
| 1 个月前 | ||
| 15 天前 | ||
| 18 天前 | ||
| 2 天前 | ||
| 18 天前 | ||
| 2 天前 | ||
| 1 个月前 | ||
| 4 天前 | ||
| 14 天前 | ||
| 20 天前 | ||
| 18 天前 | ||
| 1 个月前 | ||
| 3 天前 | ||
| 14 天前 | ||
| 27 天前 | ||
| 12 天前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 15 天前 | ||
| 23 天前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 12 天前 | ||
| 7 天前 | ||
| 5 天前 | ||
| 12 天前 | ||
| 18 天前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 8 天前 | ||
| 26 天前 | ||
| 13 天前 | ||
| 6 天前 | ||
| 6 天前 | ||
| 20 天前 | ||
| 23 天前 | ||
| 1 个月前 | ||
| 25 天前 | ||
| 27 天前 | ||
| 27 天前 | ||
| 18 天前 | ||
| 10 天前 | ||
| 25 天前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 13 天前 | ||
| 1 个月前 | ||
| 18 天前 | ||
| 6 天前 | ||
| 10 天前 | ||
| 6 天前 | ||
| 1 个月前 | ||
| 18 天前 | ||
| 18 天前 | ||
| 10 天前 | ||
| 18 天前 | ||
| 15 天前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 15 天前 | ||
| 1 个月前 | ||
| 15 天前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 23 天前 | ||
| 1 个月前 | ||
| 1 个月前 |