| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
Implement cgroups conditions memory handling (#2906) * Implement cgroups conditions memory handling Signed-off-by: stafot <stafot@gmail.com> * fix: resolve container cgroup path via /proc/self/cgroup The initial implementation read memory.max and memory.current from fixed paths under /sys/fs/cgroup/, which only works when the container runtime sets up a cgroup namespace (cgroupns) — remounting the pod's own cgroup as the root inside the container. Without cgroupns the container sees the host cgroup hierarchy, so the pod's memory files live under a nested path such as /sys/fs/cgroup/kubepods.slice/.../scope/memory.max rather than at the root. The guard would silently fall back to node-level sysinfo, never firing for the container. Fix: read /proc/self/cgroup at runtime to resolve the process's actual cgroup path (0::<rel_path> for v2, memory:<rel_path> for v1), then construct the full file paths as cgroup_root + rel_path + /memory.max etc. This works correctly both with and without cgroupns. get_cgroup_memory signature updated to accept injectable proc_self_cgroup and cgroup_root paths for testing. Test fixtures restructured to simulate a real cgroup filesystem layout under test/resources/cgroupfs/. Signed-off-by: stafot <stafot@gmail.com> * Revert unneeded changes Signed-off-by: stafot <stafot@gmail.com> * Reorg tests Signed-off-by: stafot <stafot@gmail.com> * Fix test sandbox fixtures Signed-off-by: stafot <stafot@gmail.com> --------- Signed-off-by: stafot <stafot@gmail.com> | 3 个月前 | |
String sort: allow missing values to be sent to either end. | 4 年前 | |
fix(analytics): use meta fields in query id generation (#2766) * test: add test for unique id generation with meta fields - verify `popular_queries` generates distinct ids per `meta_fields` value - cover combinations of `filter_by` and `analytics_tag` * fix(analytics): include meta fields in popular queries id generation | 7 个月前 | |
Parameterize configs used by remote embedding API. | 8 个月前 | |
add prefix auth_key in access logs (#2915) * add prefix auth_key in access logs * Log resolved API key prefixes for scoped and multi-search requests * add empty query check --------- Co-authored-by: Kishore Nallan <kishorenc@gmail.com> | 4 个月前 | |
add: personalization model APIs (#2001) * add: basic recommendations endpoint * add: POST `/recommendations/models` recommendations endpoint * add: get specific model endpoint * add: DELETE `/recommendations/models/:id` and GET all models recommendations endpoint * add: ONNX model upload endpoint * add: save the file as a .tar.gz instead of ONNX * add: verify_tar_gz function * add: support for model upload * add: PUT `/recommendations/model/:id` to update the model, name, collection * refactor: recommendations -> personalization * add: personalization tests * add: personalization_model_manager_tests * fix: review comments * fix: review comments | 1 年前 | |
Properly memset after malloc. | 5 年前 | |
Highlighting should include all search fields, and not just the best matched field. | 8 年前 | |
Address asan test warnings. | 1 年前 | |
Tighten bounds during typo exploration. | 8 个月前 | |
add prefix auth_key in access logs (#2915) * add prefix auth_key in access logs * Log resolved API key prefixes for scoped and multi-search requests * add empty query check --------- Co-authored-by: Kishore Nallan <kishorenc@gmail.com> | 4 个月前 | |
Read configuration from a configuration file. | 7 年前 | |
Fix alias replay ordering (#3055) * Fix alias replay ordering * Test alias replay after failed swap * Fix alias mutation ordering and preserve pending dependencies | 11 天前 | |
Support indexing of bool fields. | 8 年前 | |
feat: add existence filter for optional fields (#2777) * feat: add existence filter for optional fields * feat: add optional_index field to collection tests * refactor: use only missing index for existence check and update related tests * docs: clarify comment on _exists logic in filter_result_iterator * feat: add lazy evaluation for existence filter * clarify logic for applying not equals in filter initialization * fix bugs and rename to `track_missing_values`exist * Add tests for missing filter on dynamic fields * Add track_missing_values support for the fallback field * use id_list_t::iterator * Refactor test by removing redundant validity checks * Refactor filter_result_iterator by removing unused methods and simplifying logic * Fix validity assignment in filter_result_iterator reset method * remove unncessary indent --------- Co-authored-by: Kishore Nallan <kishorenc@gmail.com> | 5 个月前 | |
update: curation refactor (#2574) * add: v30 overrides * update: UTs * fix: all UTs * add: migration support and integration tests * add: migration support and integration tests * fix: CI * rename: curation * rename: curation * rename: curation * fix: CI * add: curation_sets API tests * fix: unintentional override changes | 11 个月前 | |
Preserve reference facet identity for pinned hits (#3084) * test: reproduce pinned reference facet identity loss * fix: preserve reference facet identity for pinned hits * test: cover pinned reference facet edge cases * test: cover pinned reference facets through API * test: cover reference-only pinned facets over HTTP | 2 天前 | |
restrict schema alter types to safe transitions (#2892) | 4 个月前 | |
perf: compute a lopsided `&&` from its selective side alone (#3049) * perf: compute a lopsided `&&` from its selective side alone `filter_by=A && B` costs what its wider side costs, whatever the conjunction matches. `compute_iterators()` materialises both subtrees and intersects the two results, so a filter that narrows to a hundred documents still pays for the leaf that matches two million. The estimate the node makes at construction time is right -- `min(left, right)` really is the size of the result -- but a small result says nothing about the cost of computing it, and here that cost is bounded by the larger side. An `&&` yields at most as many ids as its narrower side. When the other side is at least `AND_PROBE_RATIO` times wider, materialise the narrow side alone and ask the wide side about each of the ids it yields, through the `is_valid(id)` that `and_filter_iterators()` already uses to advance a lazy subtree. The ids come out in the same order, so the node's result, its validity, its `approx_filter_ids_length` and everything downstream of them are unchanged. It is a different plan for the same query. The plan applies where the wide side is still lazy when its parent `&&` is initialised: any string leaf above `string_filter_ids_threshold`, in every configuration, and an integer or float leaf when `enable_lazy_filter` is on. A range-index, geo or `id:*` leaf has materialised itself inside its own `init` before the `&&` node exists; probing it is still correct, and each probe is then a binary search, but it cannot give back a cost that was paid one level down. It is skipped, and both sides materialised as before, when either subtree filters on a referenced collection -- intersecting drops an id whose reference results have nothing in common, and probing has no such step -- when the node is an object filter root, and when the wide side is an operator rather than a leaf. Timeouts keep working the way they did: both children get a copy of the budget before anything is computed, the counted check inside `is_valid` cuts the probe loop short, and the forced check after it covers a loop too short to reach the clock. A -1 from `is_valid` is not read as a timeout -- in a lopsided `&&` it almost always means the wide side ran out of ids, which is an ordinary complete result -- so `validity` is what distinguishes the two. Measured on release builds, 2M documents, one selective and one broad string value, single request: the conjunction goes from 92 ms to 1 ms while its selective leaf alone costs 1 ms and its broad leaf 137 ms, from either operand order, returning the same 95 documents. A numeric wide side under the default configuration is unchanged at 37 ms, which is the scope limit above. Tests ----- `compute_iterators()` deletes both subtrees before it returns, so nothing left on the node says which plan computed it. `computed_by_probe` does, and `_get_computed_by_probe()` exposes it the way the other `_get_*` helpers do, so a test fails if a later change quietly goes back to materialising both sides. `AndProbeSelectiveSide` builds a collection whose fields differ in selectivity by a known factor and asserts both plans return the same ids in the same order for the reported shape and either operand order, for a wide side that runs out of ids first, for a `!=` wide side, for three leaves, for an operator on either side, and for a side that matches nothing or everything. `AndProbeNumericSide` does the same for a numeric wide side, lazy and eager, the eager case documenting the scope limit. `AndProbeRatioBoundary` pins the gate at 39 and 40 ids against a narrow side of 5. `AndProbeTimeout` covers a probe loop cut short by the counted check in `is_valid`, one too short to reach the clock and left to the forced check, and the two cases without a budget that must not report a timeout -- including the one where the wide side runs out of ids, which returns the same -1 a timeout does. `AndProbeKeepsReferences` joins a collection and asserts the node keeps intersecting and its reference results survive. End to end, `SelectiveAndDoesNotMaterializeWideSide` sizes its fixture around the two thresholds a test build inverts: 25 documents on the narrow side, over the 20 that make `Index::search` compute the iterator, against 500 on the wide one. Benchmark --------- The corpus the benchmark already downloads has the shape a conjunction needs to be expensive. `release_group_types:[Album,Single,Compilation]` matches 960,372 of the million songs and `primary_artist_name:Nirvana` matches 431, so the `&&` of the two returns 431 documents and, before this change, cost what the 960,372 cost: 31 ms before, 1 ms after, the same 431 documents. The scenario was chosen by measuring the alternatives. Against the same narrow side the gain tracks the width of the wide side: `release_decade:2000s` (631,584) gives 13 -> 2 ms, `release_group_types:Album` (913,296) gives 20 -> 1 ms, and the three release types above give 31 -> 1 ms. Going narrower than a few hundred documents on the selective side drives the result below the millisecond the API reports, which would leave the harness computing a percentage change against zero, so the narrow side stops there. `filter_complex` already covers this shape by accident -- `&&` and `||` are left-associative, so it parses as `(genres:Rock && primary_artist_name:Queen) || primary_artist_name:Led Zeppelin`, and that inner conjunction is 960 against 214,468. It went from 8 ms to 2 ms on the same corpus. But it measures the plan only incidentally, mixed in with an `||` and a second artist, so a regression there would be hard to read. The new scenario isolates it. The threshold matches `filter_simple`, the closest scenario in absolute cost: the `milliseconds` ceilings are p95 under 50 and 100 virtual users, and the k6 stack needs Docker, InfluxDB and Grafana, so this was not measured under load. The `percentage` check is what makes the scenario a regression guard regardless. `BenchmarkConfigSchema` derives its keys from `searchScenarios` and refines that every one of them has a threshold, so the scenario and its threshold cannot be separated. * perf: gate the `&&` probing plan on a 32x ratio and the wide side's kind What decides between probing a wide side and materializing it is the ratio between the two sides rather than their absolute sizes. Measured over 10,000,000 documents, probing loses by 3x when the wide side is 10x the narrow one -- equally so whether the narrow side holds 10,000 ids or 100,000 -- breaks even around 32x, and wins by 9.7x at 900x. Two kinds of wide side can never repay a probe. One that has already materialized its ids has no materialization left to skip, and intersecting it is a linear merge where probing would be a binary search per id. A lazy numeric range holds one id list iterator for every value it matches and walks all of them on each probe, so a single probe costs as much as a pass over the range: at 10,000,000 documents that measured 134s against 7.6s to intersect, and returned a truncated result once the search ran out of budget. Both keep the intersecting plan. The gate decides on `approx_filter_ids_length`, an estimate a string leaf overshoots. It is re-checked against the narrow side's exact count once that side is computed and before any seek into the wide one, where falling back costs nothing. The benchmark gains scenarios for the band the ratio decides, at 17x and 32x, and for a numeric wide side both materialized and lazy. * chore: re-run CI | 12 天前 | |
fix: bound group-by second-pass aggregation | 2 个月前 | |
When infix search does not find highlight, use normal search. | 2 年前 | |
Preserve reference facet identity for pinned hits (#3084) * test: reproduce pinned reference facet identity loss * fix: preserve reference facet identity for pinned hits * test: cover pinned reference facet edge cases * test: cover pinned reference facets through API * test: cover reference-only pinned facets over HTTP | 2 天前 | |
Add index for resolving facet string value. | 1 年前 | |
Skip remote model endpoint validation on collection load (#2988) (#3058) * fix: skip remote model endpoint validation on collection load (#2988) A restart while a remote embedding endpoint was unreachable dropped the auto-embedding field from the persisted schema, because init_collection treated any validate_and_init_model failure as fatal for the field. When the field's dimensions are already persisted (num_dim > 0), the load-time endpoint ping guarantees nothing and the embedder it would register is identical to one registered without it. Register remote embedders directly and let embedding requests surface endpoint problems; a malformed persisted config still fails construction and drops the field as before. Approach and patch from Kishore's review on Slack. * fix: reject query embeddings whose size does not match the field dims A remote model that starts returning a different dimensionality (e.g. the model behind a fixed URL was swapped) would otherwise feed mismatched vectors into the vector index. Fail the search request with a clear 400 instead, in all three places query embeddings are consumed. | 13 天前 | |
Index document having nested `optional: false` and `index: false` field. (#2603) | 11 个月前 | |
Should allow analytics popularity field to be int64. | 1 年前 | |
dynamic faceting based on occurrence ratio (#2822) * add support for dynamic faceting based on occurrence ratio * separate logic of filtering out of populating facets | 6 个月前 | |
restrict schema alter types to safe transitions (#2892) | 4 个月前 | |
Add a `sum` mode to `_eval` so multiple matching filter scores add up (#3011) * feat: add a `sum` mode to `_eval` so matching filter scores add up `_eval([(a):1, (b):2, (c):4])` returns the score of the first matching expression and stops, so a sort_by written as a list of additive boosts does not add. Working around it means enumerating every combination, which needs 2^n - 1 expressions. `mode: sum` adds the scores of every expression that matches, so a linear model over n boolean signals needs n expressions. _eval([(a):1, (b):2, (c):4], mode: sum):desc _eval((brand:nike), mode: sum):desc The default stays first_match. Summing by default would silently re-rank every released query and break the if/else-if idiom, where a document that is both in stock and backordered should score as in stock rather than as the sum of both. The array form is delimited by `]`, so its parameters are unambiguous. The single expression form has no such delimiter and a filter value may hold a top level comma, as `title:Hello, World` does, so parameters are recognised there only when the expression is wrapped in its own parentheses. `_eval(brand:nike, mode: sum)` therefore stays a filter, exactly as it parses today. Scores are combined with a saturating add clamped to INT64_MIN + 1, since descending sorts negate the score and -INT64_MIN is undefined. Fixes #2014 * Fix ASAN warning. * Added per-referenced-document scoring with DESC/max and ASC/min reduction, deduplication, zero-score handling, and saturation. * Accumulate _eval sum-mode weights in __int128 and clamp once per document. * Parse nested commas in reference include_fields sorting. * Preserve INT64_MIN eval sort scores. * Consume the response of import api to avoid search in parallel to index operation. --------- Co-authored-by: Mutharasu Archunan <mutharasu39@gmail.co> Co-authored-by: Harpreet Sangar <happy_san@protonmail.com> | 1 个月前 | |
fix(validator): stop dropping store:false non-optional docs on reload | 1 个月前 | |
test: cover infix matches with a curated hit and a filter_by that matches nothing (#3051) * test: cover infix matches with a curated hit and a filter_by that matches nothing * fix: pin doc IDs in tests, assert more in GroupedHitsExcludeInfixMatchesWhenFilterMatchesNothing | 16 天前 | |
Fix synonym matching for long ART prefixes (#3054) * Fix synonym matching for long ART prefixes * Fix zero-typo synonym prefix matching | 15 天前 | |
Fix broken reference on restart if alias collection name is used. (#2919) * Fix broken reference on restart if alias collection name is used. Improve referenced_ins persistence. * Persist `$REFERENCED_INS` after `_populate_referenced_ins()` call. Populate `referenced_field_name` correctly. * Fix test. | 4 个月前 | |
test: cover partial missing group allowlist | 2 个月前 | |
feat(conversation): enhance SSE handling and fix async write callbacks (#2837) | 6 个月前 | |
add prefix auth_key in access logs (#2915) * add prefix auth_key in access logs * Log resolved API key prefixes for scoped and multi-search requests * add empty query check --------- Co-authored-by: Kishore Nallan <kishorenc@gmail.com> | 4 个月前 | |
update: curation refactor (#2574) * add: v30 overrides * update: UTs * fix: all UTs * add: migration support and integration tests * add: migration support and integration tests * fix: CI * rename: curation * rename: curation * rename: curation * fix: CI * add: curation_sets API tests * fix: unintentional override changes | 11 个月前 | |
Fix test asan warnings. | 3 年前 | |
Refactor fuzzy search state transition. Handle extra chars in the middle of a query. | 3 年前 | |
Fix facet count iterator after node reinsertion | 5 个月前 | |
perf: compute a lopsided `&&` from its selective side alone (#3049) * perf: compute a lopsided `&&` from its selective side alone `filter_by=A && B` costs what its wider side costs, whatever the conjunction matches. `compute_iterators()` materialises both subtrees and intersects the two results, so a filter that narrows to a hundred documents still pays for the leaf that matches two million. The estimate the node makes at construction time is right -- `min(left, right)` really is the size of the result -- but a small result says nothing about the cost of computing it, and here that cost is bounded by the larger side. An `&&` yields at most as many ids as its narrower side. When the other side is at least `AND_PROBE_RATIO` times wider, materialise the narrow side alone and ask the wide side about each of the ids it yields, through the `is_valid(id)` that `and_filter_iterators()` already uses to advance a lazy subtree. The ids come out in the same order, so the node's result, its validity, its `approx_filter_ids_length` and everything downstream of them are unchanged. It is a different plan for the same query. The plan applies where the wide side is still lazy when its parent `&&` is initialised: any string leaf above `string_filter_ids_threshold`, in every configuration, and an integer or float leaf when `enable_lazy_filter` is on. A range-index, geo or `id:*` leaf has materialised itself inside its own `init` before the `&&` node exists; probing it is still correct, and each probe is then a binary search, but it cannot give back a cost that was paid one level down. It is skipped, and both sides materialised as before, when either subtree filters on a referenced collection -- intersecting drops an id whose reference results have nothing in common, and probing has no such step -- when the node is an object filter root, and when the wide side is an operator rather than a leaf. Timeouts keep working the way they did: both children get a copy of the budget before anything is computed, the counted check inside `is_valid` cuts the probe loop short, and the forced check after it covers a loop too short to reach the clock. A -1 from `is_valid` is not read as a timeout -- in a lopsided `&&` it almost always means the wide side ran out of ids, which is an ordinary complete result -- so `validity` is what distinguishes the two. Measured on release builds, 2M documents, one selective and one broad string value, single request: the conjunction goes from 92 ms to 1 ms while its selective leaf alone costs 1 ms and its broad leaf 137 ms, from either operand order, returning the same 95 documents. A numeric wide side under the default configuration is unchanged at 37 ms, which is the scope limit above. Tests ----- `compute_iterators()` deletes both subtrees before it returns, so nothing left on the node says which plan computed it. `computed_by_probe` does, and `_get_computed_by_probe()` exposes it the way the other `_get_*` helpers do, so a test fails if a later change quietly goes back to materialising both sides. `AndProbeSelectiveSide` builds a collection whose fields differ in selectivity by a known factor and asserts both plans return the same ids in the same order for the reported shape and either operand order, for a wide side that runs out of ids first, for a `!=` wide side, for three leaves, for an operator on either side, and for a side that matches nothing or everything. `AndProbeNumericSide` does the same for a numeric wide side, lazy and eager, the eager case documenting the scope limit. `AndProbeRatioBoundary` pins the gate at 39 and 40 ids against a narrow side of 5. `AndProbeTimeout` covers a probe loop cut short by the counted check in `is_valid`, one too short to reach the clock and left to the forced check, and the two cases without a budget that must not report a timeout -- including the one where the wide side runs out of ids, which returns the same -1 a timeout does. `AndProbeKeepsReferences` joins a collection and asserts the node keeps intersecting and its reference results survive. End to end, `SelectiveAndDoesNotMaterializeWideSide` sizes its fixture around the two thresholds a test build inverts: 25 documents on the narrow side, over the 20 that make `Index::search` compute the iterator, against 500 on the wide one. Benchmark --------- The corpus the benchmark already downloads has the shape a conjunction needs to be expensive. `release_group_types:[Album,Single,Compilation]` matches 960,372 of the million songs and `primary_artist_name:Nirvana` matches 431, so the `&&` of the two returns 431 documents and, before this change, cost what the 960,372 cost: 31 ms before, 1 ms after, the same 431 documents. The scenario was chosen by measuring the alternatives. Against the same narrow side the gain tracks the width of the wide side: `release_decade:2000s` (631,584) gives 13 -> 2 ms, `release_group_types:Album` (913,296) gives 20 -> 1 ms, and the three release types above give 31 -> 1 ms. Going narrower than a few hundred documents on the selective side drives the result below the millisecond the API reports, which would leave the harness computing a percentage change against zero, so the narrow side stops there. `filter_complex` already covers this shape by accident -- `&&` and `||` are left-associative, so it parses as `(genres:Rock && primary_artist_name:Queen) || primary_artist_name:Led Zeppelin`, and that inner conjunction is 960 against 214,468. It went from 8 ms to 2 ms on the same corpus. But it measures the plan only incidentally, mixed in with an `||` and a second artist, so a regression there would be hard to read. The new scenario isolates it. The threshold matches `filter_simple`, the closest scenario in absolute cost: the `milliseconds` ceilings are p95 under 50 and 100 virtual users, and the k6 stack needs Docker, InfluxDB and Grafana, so this was not measured under load. The `percentage` check is what makes the scenario a regression guard regardless. `BenchmarkConfigSchema` derives its keys from `searchScenarios` and refines that every one of them has a threshold, so the scenario and its threshold cannot be separated. * perf: gate the `&&` probing plan on a 32x ratio and the wide side's kind What decides between probing a wide side and materializing it is the ratio between the two sides rather than their absolute sizes. Measured over 10,000,000 documents, probing loses by 3x when the wide side is 10x the narrow one -- equally so whether the narrow side holds 10,000 ids or 100,000 -- breaks even around 32x, and wins by 9.7x at 900x. Two kinds of wide side can never repay a probe. One that has already materialized its ids has no materialization left to skip, and intersecting it is a linear merge where probing would be a binary search per id. A lazy numeric range holds one id list iterator for every value it matches and walks all of them on each probe, so a single probe costs as much as a pass over the range: at 10,000,000 documents that measured 134s against 7.6s to intersect, and returned a truncated result once the search ran out of budget. Both keep the intersecting plan. The gate decides on `approx_filter_ids_length`, an estimate a string leaf overshoots. It is re-checked against the narrow side's exact count once that side is computed and before any seek into the wide one, where falling back costs nothing. The benchmark gains scenarios for the band the ratio decides, at 17x and 32x, and for a numeric wide side both materialized and lazy. * chore: re-run CI | 12 天前 | |
Allow large float values for `default_sorting_field`. Fixes https://github.com/typesense/typesense/issues/94 | 6 年前 | |
Return error when `sort: false` is set for geopoint/geopoint[] field. (#1872) * Return error when `sort: false` is set for geopoint/geopoint[] field. * Fix failing tests. | 2 年前 | |
fix(geopolygon): make candidate selection selective for polygon queries (#3061) | 8 天前 | |
More tests for grouping. | 6 年前 | |
perf: pool curl handles with shared dns and tls sessions on curl 8.22 (#3057) * perf: bump curl to 8.22 and lease pooled handles with shared dns and tls sessions * test: fix HTTP client pool test portability * build: disable auto-detected brotli and zstd in bundled curl * fix(http_client): notify stream waiters when curl handle init fails * fix(http_client): count unpooled curl handles in the shutdown drain * fix(http_client): refuse pooled curl leases once the pool is down * fix(http_client): log when the curl shutdown drain times out * fix(server): delete http server before curl global cleanup * fix(http_client): wait for in-flight transfers on curl pool shutdown * fix(test): block SIGPIPE in HTTP test server connection threads - Block SIGPIPE only in test server connection threads; see https://man7.org/linux/man-pages/man3/pthread_sigmask.3.html. - Prevent SSL_shutdown alert writes to closed sockets from terminating isolated tests; see https://docs.openssl.org/3.0/man3/SSL_shutdown/ and https://man7.org/linux/man-pages/man2/write.2.html. - Verified the isolated test and all 10 pool tests with uncached Bazel runs in typesense-dev. | 10 天前 | |
Add sampling for value based faceting. | 2 年前 | |
Fixed a bug in bulk indexOf forarray search. | 9 年前 | |
Fixed an edge case in trie fuzzy search. | 5 年前 | |
Fix edge case in updating empty array strings. | 3 年前 | |
Fix an edge case in match score calculation. | 5 年前 | |
Fix failing test. | 8 年前 | |
Cmake compatible bazel build. | 3 年前 | |
Fix group by curation. | 1 年前 | |
feat(nl-search): add configurable schema prompt facet parameters (#3046) * feat(nl-search): add configurable schema prompt facet parameters * fix(nl-search): route preset params through auth_manager parsing * fix(nl-search): reject non-scalar preset params and missing collection | 10 天前 | |
feat(nl): add azure openai support for natural language search (#2478) | 1 年前 | |
num_tree iterator (#1538) * Add `num_tree_t::iterator_t`. * Add `num_tree_t::iterator_t` tests. * Add `bool_iterator` in `filter_result_iterator_t`. * Fix `filter_result_iterator_t::compute_iterators`. | 2 年前 | |
Float array field should accept integer values. | 6 年前 | |
Return range_index property when enabled. | 2 年前 | |
Allow fields to be marked as optional in the schema. Downside: optional fields cannot be used for sorting or marked as default sorting field. | 6 年前 | |
Include/Exclude reference fields. | 3 年前 | |
add: personalization model manager disposer for using manager across tests | 1 年前 | |
fix: num_dims (#2529) * fix: num_dims * update: assertions | 1 年前 | |
add: auto-embedding support | 1 年前 | |
refactor: analytics (#2321) * add: doc analytics * add: query analytics * add: capture_search_request (internal query event) * add: analytics manager files renaming * feat: delayed write analytics write queue * add: migrate v29 to v30 * add: mising UTs * fix: app metrics UTs * update: POST /analytics/rules endpoint * add: POST /flush and GET /status for analytics * update: integration with timeout removed * rm: event_type in POST /events * fix: CI * update: no-phase test * update: CI with download artifact * fix: CI * fix: old rule migration bug * update: api_tests .gitignore * fix: PATCH rule with PUT (upsert) rule * fix: add removed tests * update: DELETE rule api response | 1 年前 | |
Fix edge condition in posting list update. | 10 个月前 | |
Fix async-reference replay stalls and harden batched indexer snapshot recovery. (#2997) * Fix reference dependency chains to avoid quadratic draining. * Rename to `add_reference_request`. * Fix legacy reference tail reconstruction ordering. * Fix live snapshot batched indexer replacement * Fixed legacy main-queue dependency compaction * Fixed snapshot load readiness gating * Fixed snapshot restore failure handling * Fixed empty batched indexer snapshot restoration | 2 个月前 | |
Use `nlohmann::json::dump()` to handle JSON escaping (#2780) * Use `nlohmann::json::dump()` to handle JSON escaping * Fix tests to account for `nlohmann::json::dump()` not adding space after `:` like we manually used to before. | 7 个月前 | |
Fixed an edge case in fuzzy search with SKU-like tokens. | 6 年前 | |
Fix test asan warnings. | 3 年前 | |
don't transliterate stopword tokens (#2781) * don't transliterate stopword tokens * add test * rename boolean to do_transliterate and make it default enabled | 7 个月前 | |
More efficient store contains. | 6 年前 | |
Fix reference facet_by parsing issue. (#2468) * Fix reference facet_by parsing issue. * Correctly handle reference facet_by at position other than first. | 1 年前 | |
add: synonym index manager tests (#2516) | 1 年前 | |
Implement cgroups conditions memory handling (#2906) * Implement cgroups conditions memory handling Signed-off-by: stafot <stafot@gmail.com> * fix: resolve container cgroup path via /proc/self/cgroup The initial implementation read memory.max and memory.current from fixed paths under /sys/fs/cgroup/, which only works when the container runtime sets up a cgroup namespace (cgroupns) — remounting the pod's own cgroup as the root inside the container. Without cgroupns the container sees the host cgroup hierarchy, so the pod's memory files live under a nested path such as /sys/fs/cgroup/kubepods.slice/.../scope/memory.max rather than at the root. The guard would silently fall back to node-level sysinfo, never firing for the container. Fix: read /proc/self/cgroup at runtime to resolve the process's actual cgroup path (0::<rel_path> for v2, memory:<rel_path> for v1), then construct the full file paths as cgroup_root + rel_path + /memory.max etc. This works correctly both with and without cgroupns. get_cgroup_memory signature updated to accept injectable proc_self_cgroup and cgroup_root paths for testing. Test fixtures restructured to simulate a real cgroup filesystem layout under test/resources/cgroupfs/. Signed-off-by: stafot <stafot@gmail.com> * Revert unneeded changes Signed-off-by: stafot <stafot@gmail.com> * Reorg tests Signed-off-by: stafot <stafot@gmail.com> * Fix test sandbox fixtures Signed-off-by: stafot <stafot@gmail.com> --------- Signed-off-by: stafot <stafot@gmail.com> | 3 个月前 | |
fix(tokenizer): preserve original utf8 byte length for offsets (#2912) * test(tokenizer): add greek locale byte offset validation * test(collection): add utf8-safe highlighting test for greek transliterations * fix(tokenizer): preserve original utf8 byte length for offsets * fix(tokenizer): correct offset calculation for korean normalization | 4 个月前 | |
Fix union search returning fewer hits than `found` (#3035) * Keep every document in a union search result The union topster keyed its entries by hash_combine(collection_id, seq_id) (or search_index, seq_id). hash_combine is not injective for small integers: (3, 1000) and (4, 932) hash alike, and so do any two documents of neighbouring collections whose sequence ids differ by about 64. The topster took such pairs for the same document and kept only one, so a union returned fewer hits than `found` while reporting no cutoff. Both parts of the key are 32-bit, so pack them into the 64-bit key instead of hashing them. * Use the union key for curated deduplication too The curated skip in do_union computed hash_combine(collection_id, distinct_key) on its own, with the same collisions the topster key had: a curated document could suppress an unrelated raw document of a neighbouring collection. Reuse Union_KV::get_key for both insertion and lookup. | 1 个月前 | |
add ssl wrapper for remote model http calls (#2907) * add ssl wrapper for model http api calls * Fix SSE proxy handling for SSL verification failures --------- Co-authored-by: Kishore Nallan <kishorenc@gmail.com> | 4 个月前 | |
fix: preserve filters in phrase and infix searches | 18 天前 | |
add ssl wrapper for remote model http calls (#2907) * add ssl wrapper for model http api calls * Fix SSE proxy handling for SSL verification failures --------- Co-authored-by: Kishore Nallan <kishorenc@gmail.com> | 4 个月前 | |
Make cors domains a separate parameter. Also fixes --enable-cors flag parsing issue. | 4 年前 | |
Fix vector query format validation error messages (#2063) * fix: correct vector query format validation error message - Fix parser to properly detect missing colon in vector query format - Remove redundant condition check for colon validation - Ensure format `fieldname:([values])` is strictly enforced - Update test cases to match expected error messages * fix(vector-search): make missing colon error more descriptive * chore: update comments to reflect proper vector query formatting | 1 年前 | |
Sorting on popularity metric - WIP. Still has bugs. | 10 年前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 3 个月前 | ||
| 4 年前 | ||
| 7 个月前 | ||
| 8 个月前 | ||
| 4 个月前 | ||
| 1 年前 | ||
| 5 年前 | ||
| 8 年前 | ||
| 1 年前 | ||
| 8 个月前 | ||
| 4 个月前 | ||
| 7 年前 | ||
| 11 天前 | ||
| 8 年前 | ||
| 5 个月前 | ||
| 11 个月前 | ||
| 2 天前 | ||
| 4 个月前 | ||
| 12 天前 | ||
| 2 个月前 | ||
| 2 年前 | ||
| 2 天前 | ||
| 1 年前 | ||
| 13 天前 | ||
| 11 个月前 | ||
| 1 年前 | ||
| 6 个月前 | ||
| 4 个月前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 16 天前 | ||
| 15 天前 | ||
| 4 个月前 | ||
| 2 个月前 | ||
| 6 个月前 | ||
| 4 个月前 | ||
| 11 个月前 | ||
| 3 年前 | ||
| 3 年前 | ||
| 5 个月前 | ||
| 12 天前 | ||
| 6 年前 | ||
| 2 年前 | ||
| 8 天前 | ||
| 6 年前 | ||
| 10 天前 | ||
| 2 年前 | ||
| 9 年前 | ||
| 5 年前 | ||
| 3 年前 | ||
| 5 年前 | ||
| 8 年前 | ||
| 3 年前 | ||
| 1 年前 | ||
| 10 天前 | ||
| 1 年前 | ||
| 2 年前 | ||
| 6 年前 | ||
| 2 年前 | ||
| 6 年前 | ||
| 3 年前 | ||
| 1 年前 | ||
| 1 年前 | ||
| 1 年前 | ||
| 1 年前 | ||
| 10 个月前 | ||
| 2 个月前 | ||
| 7 个月前 | ||
| 6 年前 | ||
| 3 年前 | ||
| 7 个月前 | ||
| 6 年前 | ||
| 1 年前 | ||
| 1 年前 | ||
| 3 个月前 | ||
| 4 个月前 | ||
| 1 个月前 | ||
| 4 个月前 | ||
| 18 天前 | ||
| 4 个月前 | ||
| 4 年前 | ||
| 1 年前 | ||
| 10 年前 |