| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
fix: Python env, doctor accuracy, Lightpanda URL; add warmup --verify - Route every Python operation (pip installs, doctor imports, runtime subprocesses) through getPythonBin(dataDir) — uses SearXNG venv when present, falls back to system python3. Stops the "installed in pyenv, runtime can't find it" class of failures. - Doctor: stop flagging "not running" SearXNG as DEGRADED (it starts on-demand). Detect packages without __version__ (FlashRank). - Lightpanda: switch to lightpanda-io/browser nightly URL (old one 404'd). - BackendStatus.markBootstrapping(): one-shot "still starting up" warning pointing at warmup --all, distinct from the unhealthy fallback warning. - warmup --verify (also runs at end of --all): starts SearXNG, runs a test search, import-checks every Python extra, then shuts SearXNG down. - README: warmup-first Quick Start, expanded Troubleshooting, drop Roadmap. | 5 个月前 | |
feat(server): add BackendStatus one-shot warning surface | 5 个月前 | |
feat: evidence shape + max_tokens_out (closes #6) (#18) * feat(search): add gpt-tokenizer-based token budget util Adds countTokens + truncateByTokens with sentence > paragraph > heading boundary preference. cl100k-base; ~5-15% drift on non-OpenAI tokenizers documented in JSDoc. Foundation for max_tokens_out across tools. * feat(search): track char spans + nearest heading on passages Foundation for evidence shape: passages now carry {text, charStart, charEnd} and a sectionHeading derived from the markdown heading tree. * refactor(search): tighten passage span/text invariant + token boundary tests charEnd now tracks text length when a paragraph exceeds MAX_PASSAGE_LENGTH so markdown.slice(charStart, charEnd) === text. Strengthens truncateByTokens test to verify sentence/paragraph/heading boundary preference distinctly, not just the hard fallback. Names the annotated passage shape (AnnotatedPassage). * feat(search): EvidenceItem type + stable citation_id helper Defines the canonical evidence shape (title/url/section_heading/excerpt/ score/citation_id/source_span). Highlights now optionally carry source_span and section_heading. citation_id is sha1(url#start) — stable across pagination. * feat(search): hard-rename format enum to {answer,stream_answer} Drops 'full'/'context'/'highlights'. Default omitted = evidence shape. Retired values reject with migration error pointing at the new shape. Schema gains max_tokens_out, include_full_markdown, citation_format (applied to all six tool schemas). * chore(search): remove dead applyHighlightsFormat after format rename * feat(search): default returns evidence list; citation_format wired Search now produces output.evidence by default, sized to max_tokens_out. include_full_markdown=false (default) drops markdown_content. Three citation_format styles supported: numbered (default), json, anthropic_tags. * refactor(search): move evidence helpers to src/search/evidence.ts; surface extraction failure * chore(tests): remove tests for retired format values + opt into markdown Deletes search-context.test.ts and search-highlights.test.ts (whole files for format=context/full/highlights, removed in ticket #6 hard-rename). Drops format=context/full assertions from search-answer.test.ts and multi-query.test.ts. Updates rerank-pipeline.test.ts and search-tool.test.ts to opt into include_full_markdown:true where they assert on markdown_content. * feat(tools): apply evidence shape across fetch/find_similar/crawl/research/agent Every tool now emits an evidence list by default. include_full_markdown opt-in restores full bodies. max_tokens_out budget is honoured per tool. * fix(crawl): place evidence on CrawlResultItem per spec Plan calls for per-page evidence on each CrawlResultItem; aggregate max_tokens_out cap walking pages in input order. Move evidence field from CrawlOutput to CrawlResultItem. * chore(evidence): guard marker-only excerpts; tighten test casts Skip evidence items whose truncated excerpt is only the truncation marker (small max_tokens_out edge case). Replace `as any` in newly added crawl/research tests with `as unknown as <Type>` to match project convention. * feat(tools): enforce max_tokens_out across all tools; tokens > chars precedence * test(search): guard citation_id stability across pagination * refactor(budget): share aggregate token budget across search markdown; DRY tool loops Plan T7 says per-tool max_tokens_out is an aggregate cap walked in score order. Search was capping each result's markdown_content at the full budget independently. Extract the running-tally pattern into a shared applyAggregateMarkdownBudget helper, reuse from search evidence + find_similar + crawl + research + agent. * docs(mcp): document evidence default + max_tokens_out + citation_format Drops format=highlights/full/context from advertised behaviour. Routing table now points at evidence-by-default. Per-tool descriptions document include_full_markdown, max_tokens_out, citation_format and stay inside the #8 token budgets. * docs: changelog + api-reference for evidence default + max_tokens_out * fix(instructions): trim WIGOLO_INSTRUCTIONS under 1000 words T9 additions pushed word count to 1014; consolidated three new bullets into one and tightened the routing-table evidence-excerpt cell. | 5 个月前 | |
test(instructions): split assertions between WIGOLO_INSTRUCTIONS and _FULL A7 trimmed WIGOLO_INSTRUCTIONS by moving routing tables, performance budgets, and the Extras section to WIGOLO_INSTRUCTIONS_FULL behind the wigolo://docs/usage resource. The pre-existing v3 instructions tests were still asserting moved content against the short version. Re-target the now-moved assertions (localhost, use_auth, multi-query, routing table intents, full-text search syntax) against _FULL while keeping 'must always be visible per-session' checks on the short version. | 4 个月前 | |
chore: remove internal dev-process references from tracked files Strip development breadcrumbs (codename, slice/phase/wave/flaw and audit-item tags, internal design-doc pointers) from source comments, SQL migration headers, tests, CHANGELOG, and a benchmark baseline. Correct stale diff/watch tool-schema comments that described them as unimplemented stubs. Reword to preserve the real rationale; no code, control-flow, or test-assertion changes. | 3 个月前 | |
chore: remove internal dev-process references from tracked files Strip development breadcrumbs (codename, slice/phase/wave/flaw and audit-item tags, internal design-doc pointers) from source comments, SQL migration headers, tests, CHANGELOG, and a benchmark baseline. Correct stale diff/watch tool-schema comments that described them as unimplemented stubs. Reword to preserve the real rationale; no code, control-flow, or test-assertion changes. | 3 个月前 | |
feat: boot no longer probes ONNX provider; add boot-negative spy (D2) The eager embedding probe is gone from boot as a consequence of the init() split — this commit fixes the stale server.ts comment to describe the lazy model load, and adds a boot-negative test to server-factory.test.ts asserting initSubsystems inits the embedding store (positive control) but never calls ensureProviderReady() at startup. | 2 个月前 | |
refactor(server): extract tool schemas to tool-schemas module for testability | 5 个月前 | |
feat(server): warm engine origins on MCP server start A cold first search dropped engines whose DNS+TLS handshake didn't finish inside the merge deadline (fresh-process query #1 = ~12s/2 engines; #2 = ~3s/3 engines, warmed via OS caches). The long-running MCP server now warms the five primary engine origins on start (fire-and-forget, non-blocking, never throws), so the first tool call gets the full pool. WIGOLO_WARM_ENGINES=0 disables. | 2 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 4 个月前 | ||
| 3 个月前 | ||
| 3 个月前 | ||
| 2 个月前 | ||
| 5 个月前 | ||
| 2 个月前 |