| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
feat(libsy): Algorithms select a `Category` (e.g. "efficient") not a specific model. (#630) Instead of giving the available models to the algorithm in new we pass them alongside the request in run_stream where they go in the Driver. Algorithms work with Category instead of ModelId. Later libsy selects a model from the list for that category. The mapping from Category to ModelId lives in the Driver. See #588 for the originating idea. Assisted-by: Codex:GPT 5.6 Sol high Assisted-by: Claude:Opus 5 medium Reviewed-by: Claude:Opus 5 medium Signed-off-by: Graham King grahamk@nvidia.com | 21 天前 | |
feat: Failure cooldown after exhausted transient failures (#887) If a model is unavailable, we would retry the configured number of times and then fall back to the next model in the list. But we would do this on every request. Now we briefly remember the model is not available. We don't retry the model until after a cooldown period. Signed-off-by: Graham King <grahamk@nvidia.com> | 21 小时前 | |
fix(stage): match tool semantics by bare MCP tool name (#831) Stage `tool_semantics` now match MCP tools by their bare tool name. For example, `mutate = ["send_payment_request"]` now matches: - A Codex Responses call with `"name": "send_payment_request", "namespace": "mcp__billing"`. - A Claude Code call named `mcp__billing__send_payment_request`. The full name `mcp__billing__send_payment_request` still matches both. The full name is checked first, and the bare name only when the full name matches nothing. Changes: 1. `libsy` finds the bare name in one of two ways: - Codex: from the namespace mapping that the Responses decoder stores on the request. - Claude Code: from the `mcp__<server>__<tool>` form. The name is split once, after the server name, so tool names that contain `__` still work. 2. The namespace key and its two read helpers (`tool_namespaces`, `split_qualified_name`) move from `switchyard-translation` into a new `switchyard_protocol::codex_namespaces` module. This lets `libsy` read the mapping without depending on `switchyard-translation`. Fixes: https://linear.app/nvidia/issue/SWITCH-1456 Assisted-by: Claude:Opus 5.5 medium Signed-off-by: Graham King <grahamk@nvidia.com> | 9 天前 | |
feat(protocol): add decision model request and response types (#880) * feat(protocol): add decision model request and response types Signed-off-by: nachiketb <nachiketb@nvidia.com> * refactor(protocol): simplify decision types to public data Signed-off-by: nachiketb <nachiketb@nvidia.com> * fix(protocol): check decision numbers during deserialization Signed-off-by: nachiketb <nachiketb@nvidia.com> * refactor(protocol): remove decision numeric decoding checks Signed-off-by: nachiketb <nachiketb@nvidia.com> --------- Signed-off-by: nachiketb <nachiketb@nvidia.com> | 1 天前 | |
feat(server): forward upstream response headers (#571) Preserve upstream HTTP response headers for buffered and streaming LLM calls. Forward an allowlisted set of tracing, request ID, processing, rate-limit, and upstream headers to downstream clients. Keep cookies, body headers, and Switchyard-owned headers private, and let Switchyard’s own headers take precedence. Signed-off-by: Lars van der Zande <lmvanderzande@gmail.com> Signed-off-by: Graham King <grahamk@nvidia.com> | 17 天前 | |
docs(rust): repair libsy and protocol references (#263) * docs(rust): repair libsy and protocol references Signed-off-by: nachiketb <nachiketb@nvidia.com> * docs(libsy): list all public algorithms Signed-off-by: nachiketb <nachiketb@nvidia.com> * docs(libsy): make quickstart directly runnable Signed-off-by: nachiketb <nachiketb@nvidia.com> * docs(rust): streamline crate landing pages Signed-off-by: nachiketb <nachiketb@nvidia.com> * docs(libsy): focus the crate landing page Signed-off-by: nachiketb <nachiketb@nvidia.com> * docs(protocol): construct the simple request directly Signed-off-by: nachiketb <nachiketb@nvidia.com> * docs(rust): use readmes as crate documentation Signed-off-by: nachiketb <nachiketb@nvidia.com> * docs(libsy): run classifier in quick start Signed-off-by: nachiketb <nachiketb@nvidia.com> * docs(libsy): clarify crate purpose Signed-off-by: nachiketb <nachiketb@nvidia.com> * docs(rust): link crate landing page APIs Signed-off-by: nachiketb <nachiketb@nvidia.com> * docs(rust): complete API contract audit Signed-off-by: nachiketb <nachiketb@nvidia.com> --------- Signed-off-by: nachiketb <nachiketb@nvidia.com> | 1 个月前 | |
feat(protocol): add decision model request and response types (#880) * feat(protocol): add decision model request and response types Signed-off-by: nachiketb <nachiketb@nvidia.com> * refactor(protocol): simplify decision types to public data Signed-off-by: nachiketb <nachiketb@nvidia.com> * fix(protocol): check decision numbers during deserialization Signed-off-by: nachiketb <nachiketb@nvidia.com> * refactor(protocol): remove decision numeric decoding checks Signed-off-by: nachiketb <nachiketb@nvidia.com> --------- Signed-off-by: nachiketb <nachiketb@nvidia.com> | 1 天前 | |
fix(translation): preserve structured-output enforcement (#892) Record schema enforcement in the shared IR. Enable OpenAI strict mode for compatible schemas. Add `is_schema_enforced` to the IR output params so a request's schema enforcement survives translation. Decoders resolve provider defaults: Anthropic always enforces, OpenAI reads the `strict` flag. When enforcement cannot carry over, such as advisory mode on Anthropic or a schema outside OpenAI's strict subset, the codec emits a lossy diagnostic. Fixes: https://github.com/NVIDIA-NeMo/Switchyard/issues/467 Assisted-by: Pi:GPT 6 Astra medium Reviewed-by: Pi:Grok 4.7 high Reviewed-by: Pi:Kimi K3 high Signed-off-by: Graham King <grahamk@nvidia.com> | 22 小时前 | |
fix(routing): send Codex spawned threads to the subagent route (#748) Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com> | 14 天前 | |
feat: Introduce ModelId type (#373) * feat: Introduce ModelId type Identify model ID's (e.g. "openai/gpt-oss-120b") with their own type `ModelId`. That makes it clear what we are passing, and prevent accidentally passing the wrong thing. `ModelId` behaves like a string. It can be printed, deref-ed, compared, etc. Previously that was either `String` or `LlmTarget`. This PR removes `LlmTarget` and `LlmTargetSet`, in favor of `ModelId` and `Vec<ModelId>`. To review start with new file `crates/protocol/src/model_id.rs` which contains the type. Assisted-by: Claude:Opus 5 medium Signed-off-by: Graham King <grahamk@nvidia.com> * Minor doc fix Signed-off-by: Graham King <grahamk@nvidia.com> * Extend ModelId to synthetic model IDs such as "switchyard/random" Signed-off-by: Graham King <grahamk@nvidia.com> --------- Signed-off-by: Graham King <grahamk@nvidia.com> | 1 个月前 | |
fix(translation): retain response IDs arriving after stream start (#771) Some model providers start sending a reply before they supply its response ID or model name. Switchyard previously checked only the first part of the reply, so it could miss those details even when they arrived later. This PR captures the first nonempty ID and model name whenever they arrive, then passes them to monitoring and response collection. It also handles an important detail when converting replies to OpenAI’s Responses format: if Switchyard has already sent an ID to the caller, that ID stays unchanged. The provider’s late ID is still recorded internally. The updated test checks that monitoring receives the late ID, reply text and token counts survive, and the caller sees the same response ID at the beginning and end. Signed-off-by: nachiketb <nachiketb@nvidia.com> | 14 天前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 21 天前 | ||
| 21 小时前 | ||
| 9 天前 | ||
| 1 天前 | ||
| 17 天前 | ||
| 1 个月前 | ||
| 1 天前 | ||
| 22 小时前 | ||
| 14 天前 | ||
| 1 个月前 | ||
| 14 天前 |