| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
perf(persist): fetch a narrow internal auth context on the auth hot path (#6800) - Persist authenticates every request (half of the [current](https://us3.datadoghq.com/notebook/269063?cell_id=qeu4s4b9&from_ts=1781602210008&to_ts=1784194210008&refresh_mode=sliding&tpl_var_env=%2A&tpl_var_metric=%2A) traffic in production — the single largest query family on the main DB, ~45% of all client-observed query time) by resolving the full account context: a 6-relation query serializing five `row_to_json` blobs plus two secret decryptions. Its consumers only read `environment.{id,name}`, `account.id` and two plan fields. This PR adds `accountService.getPersistAuthContext()` — a 4-relation, 6-scalar-column lookup — and switches persist to it. - New `PersistAuthContext` type (`Pick`-based, no default values): downstream parameter types narrow to the fields they actually read, so the compiler guarantees no consumer touches a field the query doesn't fetch. Adding a consumer that needs more is a type change + query extension, never a silent gap. - Self-hosted keys provisioned via `NANGO_SECRET_KEY_*` env vars may not exist in `api_secrets`, so a miss falls back to the full lookup and narrows the result. On cloud the full lookup filters on the same predicate and can never match a miss, so unknown keys fail fast (avoids paying a second pbkdf2 + heavy query per 401). ## Production numbers `EXPLAIN (ANALYZE, BUFFERS)` on the production writer, same key, warm cache: | | current query | new query | Δ | |---|---|---|---| | Execution time | 0.175 ms | 0.076 ms | **−57%** | | Execution buffers | 20 shared hits | 13 | −35% | | Planning time | 2.118 ms | 0.555 ms | −74% | | Planning buffers | 103 | 49 | −52% | The dropped `default_secret` self-join re-fetched the row already matched by the `WHERE` clause; the `pending_secret` join was unread. The accounts lookup becomes an Index Only Scan (zero heap fetches) since only `id` is selected. Planning matters at this rate too: pg re-plans one-shot statements per execution, and the 4-relation tree plans in ~¼ the time. At ~3,400 req/s the execution savings alone are ~0.34 CPU-seconds/s on the writer, plus the planning-side reduction. A side benefit: persist no longer materializes decrypted secret keys, `hmac_key`, or customer OTLP headers per request — the resolved context is three ints and two short strings. ## Rollout: percentage feature flag The middleware branches between two fully isolated auth paths, selected per request by the `persist-light-auth-context` flag (default `off`): the legacy full account context lookup, verbatim and under its original trace span, and the light `PersistAuthContext` one. No account is known before the lookup, so the Unleash gradual-rollout strategy should use random stickiness (per-request bucketing). Each path keeps its own span name, so the rollout ratio and per-path health are directly visible in APM. Once the ramp completes, deleting the flag and the legacy branch is a middleware-only change. ## Notes for reviewers - The persist auth span renames to `persist.middleware.auth.getPersistAuthContext` — dashboards/monitors referencing `trace.persist.middleware.auth.getAccountContextByApiKey.*` need updating after deploy. - The server's script-auth path (`access.middleware`, ~15% of this query's volume) intentionally stays on the full lookup; migrating it is a future, independent step (its locals type is shared with the customer-key path). - PR #6779 (context cache) will be reworked on top of this once merged: the cache moves to `getPersistAuthContext`, caching the narrow context instead of the full one. ## Test plan - [x] Exact-shape integration test (`toStrictEqual`) — fails if the query ever fetches more or fewer fields than the type declares - [x] Unknown key returns null; env-var keys resolve through the fallback - [x] Full monorepo build passes with the narrowed types (the compiler-level proof that no consumer reads an unfetched field) - [x] Persist integration suite (real middleware) and records unit suite green - [x] Legacy-path test: flag off authenticates through the legacy lookup end-to-end - [ ] Create `persist-light-auth-context` in Unleash (gradual rollout, random stickiness); deploy with flag off — no behavior change - [ ] Ramp to a small percentage: persist 401s stay ≈0, no failed `getPersistAuthContext` spans, latency ≤ legacy - [ ] Ramp to 100%: `api_secrets` join rate on `postgres-nango` drops accordingly - [ ] Update DD dashboards referencing the renamed auth span 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated description by cubic. --> <a href="https://cubic.dev/pr/NangoHQ/nango/pull/6800?utm_source=github" target="_blank" rel="noopener noreferrer" data-no-image-dialog="true"><picture><source media="(prefers-color-scheme: dark)" srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source media="(prefers-color-scheme: light)" srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img alt="Review in cubic" src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a> <!-- End of auto-generated description by cubic. --> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> | 2 个月前 | |
refactor(records): share one helper for composite model names (#7415) ## Problem Records are keyed by a composite model name — the bare model for the `base` variant, `Model::variant` otherwise. Three places built that string by hand, each spelled a little differently, and nothing made them agree. A format change that misses one doesn't fail loudly. The lookup just misses, and the sync reports zero records. Raised by @kaposke reviewing #7369. ## Solution - Add `recordModelName` to `packages/records`, which owns the format, with unit tests for the base / variant / missing cases. - Move both server callers onto it: the syncs endpoint and the records endpoint. Two copies stay, deliberately: - `manager.service.ts` — `shared` has no runtime dependency on `records`, and `RecordsServiceInterface` exists precisely to avoid one. Adding that dependency to collapse a single line inverts the boundary, and [NAN-6820](https://linear.app/nango/issue/NAN-6820) reworks that path anyway. - `syncTargetId` in `controllers/sync/helpers.ts` — byte-identical, but it builds *audit target ids*. Folding them together would mean a change to the audit format silently breaking record lookups. Fixes [NAN-6913](https://linear.app/nango/issue/NAN-6913) ## Testing 3 unit tests on the helper, plus the existing suites for both migrated endpoints — 23 integration tests across `getSyncs` and `getRecords`. No behaviour change: the two implementations already agreed, one just also handled `null`, which the shared version keeps. | 20 天前 | |
fix(server): Clamp proxy retry to maximum duration (#7264) <!-- Describe the problem and your solution --> While investigating a spike in server latency at the top of each hour, we discovered long running proxy operations for customer Centralize that was causing requests to stack up at the top of each hour. The root cause seems to be our proxy handler respecting long retries (1 hour+) retrying operations long after the client request will have timed out. There should be a maximum value for retry-after beyond which proxy operations fast fail and return retry information to the caller. <!-- Issue ticket number and link (if applicable) --> [NAN-6759: fix(server): Add maximum retry-after in proxy calls](https://linear.app/nango/issue/NAN-6759/fixserver-add-maximum-retry-after-in-proxy-calls) <!-- Testing instructions (skip if just adding/editing providers) --> To test this I had an agent stand up a mock server temporarily. A quick summary is: 1. Edit `providers.yaml` for a provider like `private-api-bearer` or `unauthenticated` to support retry headers. There might be a provider you could use for this already, but I didn't see it 🤷 : ``` retry: after: - 'retry-after' ``` 2. Standup a mock endpoint that always returns a long retry header based on what you are using from 1 3. Create a new integration and connection for provider from step 1 4. Allow Nango to send traffic to your mock endpoint by setting the following in your `.env`: ``` NANGO_OUTBOUND_URL_POLICY={"blockPrivateIps":false} ``` 5. Proxy a request through Nango connection. It should fail fast with this change rather than hanging for the duration of the retry header you provided from the mock endpoint. <!-- This is an auto-generated description by cubic. --> <a href="https://cubic.dev/pr/NangoHQ/nango/pull/7264?utm_source=github" target="_blank" rel="noopener noreferrer" data-no-image-dialog="true"><picture><source media="(prefers-color-scheme: dark)" srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source media="(prefers-color-scheme: light)" srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img alt="Review in cubic" src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a> <!-- End of auto-generated description by cubic. --> --------- Co-authored-by: cubic-dev-ai[bot] <191113872+cubic-dev-ai[bot]@users.noreply.github.com> | 30 天前 | |
feat: add @nangohq/kms package (#6471) add @nangohq/kms package: - resolve data encryption key from NANGO_ENCRYPTION_KEY and load/unwrap wrapped key in the background to verify equality - refactor encryption logic in the rest of the codebase to use the `kms` package. No change in behavior so far. Encryption key is still loaded from NANGO_ENCRYPTION_KEY. Once I have verified that infra is all setup correctly I'll make a PR to load the wrapped key instead. <!-- Issue ticket number and link (if applicable) --> <!-- Testing instructions (skip if just adding/editing providers) --> | 3 个月前 | |
fix(records): split records_seen entries to limit size of ids array (#6485) When records_seen entries contains more than 100+ ids, their size goes beyond pg TOAST threshold (~2kb). We observe that pg can wait a lot of IO when "untoasting" records_seen in deleteOutdatedRecords query. This commit is an attempt at limiting the size of each records_seen entry so it doesn't go beyond the TOAST threshold. In production, it would only affect ~10-15% of the rows. <!-- Describe the problem and your solution --> <!-- Issue ticket number and link (if applicable) --> <!-- Testing instructions (skip if just adding/editing providers) --> <!-- This is an auto-generated description by cubic. --> <a href="https://cubic.dev/pr/NangoHQ/nango/pull/6485?utm_source=github" target="_blank" rel="noopener noreferrer" data-no-image-dialog="true"><picture><source media="(prefers-color-scheme: dark)" srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source media="(prefers-color-scheme: light)" srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img alt="Review in cubic" src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a> <!-- End of auto-generated description by cubic. --> | 3 个月前 | |
feat(webapp): Records list page (#5862) ## Summary <img width="1089" height="650" alt="Screenshot 2026-04-14 at 6 10 55 PM" src="https://github.com/user-attachments/assets/2a8023cd-48ce-4d69-aff8-227093510a47" /> <img width="1089" height="682" alt="Screenshot 2026-04-14 at 6 11 08 PM" src="https://github.com/user-attachments/assets/10efc874-c015-471d-bfb8-73bf9f8a8e3a" /> <img width="1151" height="580" alt="Screenshot 2026-04-14 at 6 11 15 PM" src="https://github.com/user-attachments/assets/40c41f8a-6ebe-4dcc-bf82-48b334dc9309" /> - Adds a new Records experience to connection pages, including a model list, variant-aware drilldown, paginated records table, and lazy-loaded payload viewer. - Adds the private connection records API, shared types, and webapp hooks needed to list record models/variants and fetch records for a connection. - Lays groundwork for the upcoming connection page redesigns by introducing the shared connection sidebar layout and moving the current connection pages onto that structure. - Hardens records pagination/cursor handling and adds coverage for invalid cursor handling, pagination, and model/variant listing. ## Issue - `NAN-5080` ## Testing - Open a connection with records and verify the new Records tab: - lists models and variants correctly - paginates through records correctly - opens the payload dialog only when requested - Verify the shared connection page layout still renders correctly for the existing connection tabs. - `npm run test:unit -- packages/records/lib/cursor.unit.test.ts` - `npm run test:integration -- packages/server/lib/controllers/v1/connections/connectionId/records/getModels.integration.test.ts packages/server/lib/controllers/v1/connections/connectionId/records/getRecords.integration.test.ts` - `npx tsc -p packages/records/tsconfig.json --noEmit` - `npx tsc -p packages/server/tsconfig.json --noEmit` - `npx tsc -p packages/webapp/tsconfig.json --noEmit` --------- Co-authored-by: Matej Vobornik <vobornik.matej@gmail.com> | 3 个月前 | |
feat(webapp): Records list page (#5862) ## Summary <img width="1089" height="650" alt="Screenshot 2026-04-14 at 6 10 55 PM" src="https://github.com/user-attachments/assets/2a8023cd-48ce-4d69-aff8-227093510a47" /> <img width="1089" height="682" alt="Screenshot 2026-04-14 at 6 11 08 PM" src="https://github.com/user-attachments/assets/10efc874-c015-471d-bfb8-73bf9f8a8e3a" /> <img width="1151" height="580" alt="Screenshot 2026-04-14 at 6 11 15 PM" src="https://github.com/user-attachments/assets/40c41f8a-6ebe-4dcc-bf82-48b334dc9309" /> - Adds a new Records experience to connection pages, including a model list, variant-aware drilldown, paginated records table, and lazy-loaded payload viewer. - Adds the private connection records API, shared types, and webapp hooks needed to list record models/variants and fetch records for a connection. - Lays groundwork for the upcoming connection page redesigns by introducing the shared connection sidebar layout and moving the current connection pages onto that structure. - Hardens records pagination/cursor handling and adds coverage for invalid cursor handling, pagination, and model/variant listing. ## Issue - `NAN-5080` ## Testing - Open a connection with records and verify the new Records tab: - lists models and variants correctly - paginates through records correctly - opens the payload dialog only when requested - Verify the shared connection page layout still renders correctly for the existing connection tabs. - `npm run test:unit -- packages/records/lib/cursor.unit.test.ts` - `npm run test:integration -- packages/server/lib/controllers/v1/connections/connectionId/records/getModels.integration.test.ts packages/server/lib/controllers/v1/connections/connectionId/records/getRecords.integration.test.ts` - `npx tsc -p packages/records/tsconfig.json --noEmit` - `npx tsc -p packages/server/tsconfig.json --noEmit` - `npx tsc -p packages/webapp/tsconfig.json --noEmit` --------- Co-authored-by: Matej Vobornik <vobornik.matej@gmail.com> | 3 个月前 | |
feat: add @nangohq/kms package (#6471) add @nangohq/kms package: - resolve data encryption key from NANGO_ENCRYPTION_KEY and load/unwrap wrapped key in the background to verify equality - refactor encryption logic in the rest of the codebase to use the `kms` package. No change in behavior so far. Encryption key is still loaded from NANGO_ENCRYPTION_KEY. Once I have verified that infra is all setup correctly I'll make a PR to load the wrapped key instead. <!-- Issue ticket number and link (if applicable) --> <!-- Testing instructions (skip if just adding/editing providers) --> | 3 个月前 | |
refactor(records): share one helper for composite model names (#7415) ## Problem Records are keyed by a composite model name — the bare model for the `base` variant, `Model::variant` otherwise. Three places built that string by hand, each spelled a little differently, and nothing made them agree. A format change that misses one doesn't fail loudly. The lookup just misses, and the sync reports zero records. Raised by @kaposke reviewing #7369. ## Solution - Add `recordModelName` to `packages/records`, which owns the format, with unit tests for the base / variant / missing cases. - Move both server callers onto it: the syncs endpoint and the records endpoint. Two copies stay, deliberately: - `manager.service.ts` — `shared` has no runtime dependency on `records`, and `RecordsServiceInterface` exists precisely to avoid one. Adding that dependency to collapse a single line inverts the boundary, and [NAN-6820](https://linear.app/nango/issue/NAN-6820) reworks that path anyway. - `syncTargetId` in `controllers/sync/helpers.ts` — byte-identical, but it builds *audit target ids*. Folding them together would mean a change to the audit format silently breaking record lookups. Fixes [NAN-6913](https://linear.app/nango/issue/NAN-6913) ## Testing 3 unit tests on the helper, plus the existing suites for both migrated endpoints — 23 integration tests across `getSyncs` and `getRecords`. No behaviour change: the two implementations already agreed, one just also handled `null`, which the shared version keeps. | 20 天前 | |
feat(syncs): paginate and virtualize the connection Syncs tab (NAN-6819) (#7369) ## Problem The Syncs tab fetched every sync for a connection and re-polled the whole list every 5s. [NAN-6818](https://linear.app/nango/issue/NAN-6818) stopped it 500ing past 1000 syncs, but at Levity's 1581 syncs it still took 2.65s to paint with 1.85s of blocked main thread, 38,642 DOM nodes, and 1.2MB every 5 seconds. ## Solution - `GET /api/v1/sync` becomes paginated: `page`/`limit` with a `total`, plus exact `name`/`variant` filters for point lookups. Changed in place rather than added alongside — it's `webAuth`, never was in the typed registry, and ships with its only client, so there was no contract to preserve and nothing left dead behind. - Deletes the implementation it replaces: `getSyncsByParams`, `addRecordCount`, and `sync.service.getSyncs` with its duplicate latest-job query. - Filtering and paging run over sync ids alone, so the latest-job lateral runs `limit` times rather than `offset + limit` times, and record counts cover only the page's models. - **Behaviour change:** a failed orchestrator schedule lookup degrades a row to `schedule_status: null` instead of throwing. That's what let the endpoint drop its self-hosted empty-list short-circuit, and what makes it testable. - The tab moves to infinite scroll with virtualized rows, scrolling with the dashboard's own container so it adds no second scrollbar. **Deploy note:** the response goes from a bare array to `{data, pagination}`, so a tab loaded before the deploy breaks until refresh. Fixes [NAN-6819](https://linear.app/nango/issue/NAN-6819) ## Testing 14 integration tests, including a correct `total` on an out-of-range page, ordering with no gaps across pages, variant-scoped record counts, and a 200 when the orchestrator is unreachable. Verified in the browser against a 1601-sync connection: 223ms to first rows, zero long tasks, 573 DOM nodes, page 0 at 22.8KB. ## Follow-ups - Polling still refetches every loaded page, so cost grows with scroll depth — deliberate, `maxPages` is the lever if it bites. - No sticky header; making one work means changing the shared `Table` component. - `getCountsByModel` still folds in O(n²) — [NAN-6821](https://linear.app/nango/issue/NAN-6821). | 1 个月前 | |
feat(records): bound getRecords response by byte budget (NAN-5409) (#6187) ## Summary Bound peak per-request memory in `GET /records` by a configurable byte budget, mitigating the INC-107 OOM pattern. Implements **Option B** from the NAN-5409 [design doc](https://linear.app/nango/document/bounding-peak-memory-per-request-for-get-records-02af52a107df), riding on the `size_bytes` column from #6164. **Rollout is staged through dry-run** — we observe the impact before enforcing. ## Two env knobs - `RECORDS_MAX_RESPONSE_SIZE_BYTES` (default `0` = disabled). When > 0, `getRecords` walks results in cursor order, accumulates `records.size_bytes`, and identifies the point where the next record would exceed the budget. Always keeps ≥ 1 record so pagination can advance even when a single record is larger than the budget. `next_cursor` is emitted from the last kept record; strict `(updated_at, id) > cursor` resumes exactly at the first un-fetched row. - `RECORDS_MAX_RESPONSE_SIZE_DRY_RUN` (default `true`). In dry-run the truncation point is identified but the page is returned in full. A counter (`nango.records.budgetDryRunTruncate`) is emitted with `{ accountId, service: 'server' | 'persist' }` so we can see which accounts would be affected via which path. ## Rollout plan (intentional, please don't enforce yet) 1. **Ship in dry-run** across dev/staging/prod (`RECORDS_MAX_RESPONSE_SIZE_DRY_RUN=true`, budget 50 MiB). Companion PR: NangoHQ/nango-environments#70. 2. **Observe** the `nango.records.budgetDryRunTruncate` counter for ~1 week. Decide: - Is the current 50 MiB threshold the right number? Tighten or relax based on the distribution metric (`nango.server.getRecords.responseSizeBytes`) and the dry-run counter. - Do we enforce on **both** server and persist, or only one? The runner SDK already honors `next_cursor` on both `listRecords` and `getRecordsByIds`, so enforcing the persist path is safe — but the impact data tells us if it's worth it. 3. **Flip `DRY_RUN` to false** to enforce. **Remove the dry-run flag entirely** as a follow-up — it's transitional infra, not a long-term feature. ## Trade-offs (in the code as comments) `NULL size_bytes` is counted as 0. Two transient populations are affected: - Legacy `records.json` rows (pre-data-table migration) - `records_data` rows written before #6164 Both drain as customers re-upsert (every write path now populates `size_bytes`). Worst case during the transition: one un-backfilled row overshoots the budget by its own size — bounded by data shape, not unbounded. Chose this over a fallback IN-query to keep the hot path on a single round-trip. `size_bytes` is `pg_column_size(records_data.data)` — compressed on-disk size, smaller than wire bytes by 1.5–3×. The budget is therefore looser than its name suggests; the operator chooses a value comfortable with the worst-case multiplier. ## Test plan - [x] `npm run test:integration -- packages/records/lib/models/records.integration.test.ts -t "getRecords"` — 10/10 pass (5 existing + 5 new, including the dry-run-does-not-truncate test). - [x] TypeScript clean across utils, records, types, server, persist. - [ ] CI green. - [ ] Companion env-config PR merged **after** this lands in a deployable image — otherwise the env is set on a server that doesn't read it (harmless no-op, but pointless). ## Observability Adds a counter and a span tag: - Counter `nango.records.budgetDryRunTruncate` (dimensions `accountId`, `service`) — only emitted in dry-run, group by `accountId` to find affected customers, group by `service` to see whether the impact comes from the public endpoint or the runner path. - Span tag `nango.records.budgetTruncated` on `nango.records.getRecords` — only set when enforcement actually truncated (dry-run off). 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> | 4 个月前 |