| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
feat(providers): add latest Claude, Gemini, Kimi, and Grok models (#993) * feat(providers): add latest model catalog entries Co-authored-by: tejs1 <tejsrelax@gmail.com> * fix(ai): preserve reasoning settings across provider paths * fix(providers): define Grok output limits * fix(kimi): preserve reasoning across tool calls * fix(kimi): scope reasoning replay to K3 * fix(kimi): gate pi reasoning replay by model --------- Co-authored-by: tejs1 <tejsrelax@gmail.com> | 2 个月前 | |
Fix Haiku 4.5 thinking controls (#501) Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 5 个月前 | |
Add Atlas Cloud LLM provider (#948) * Add Atlas Cloud provider * Account for Atlas Cloud non-thinking model metadata * test: verify Atlas Cloud thinking payloads * style: format Atlas Cloud payload test * test: type Atlas Cloud fetch mock --------- Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> | 1 个月前 | |
fix(ai): let LLM calls outlive undici's 300s headers timeout (#1404) * fix(ai): let LLM calls outlive undici's 300s headers timeout A non-streaming completion only receives response headers once the whole completion exists, so undici's default 300 s headers timeout bounds total generation time. Thinking models (GLM-5.x at effort=max, MiniMax M3, …) routinely exceed that on large prompts — the request then dies with "Cannot connect to API: Headers Timeout Error" at exactly 300 s and, with maxRetries 0, the scene never generates (observed on a self-hosted deployment with glm-5.2 scene generation). Route every outbound LLM request through an undici Agent with a 15-minute headers/body timeout, attached at the shared transportFetch seam so the OpenAI-compatible, native OpenAI/Responses, Azure, Anthropic-compatible and Google transports all inherit it (same lazy-import dispatcher pattern the Google proxy transport already uses). Native transports now always install the shared transport instead of only when the server injects a validating fetch. * fix(ai): harden the LLM transport timeout fallback Review follow-ups for the undici headers-timeout fix: - getLlmDispatcher(): drop the cached promise on rejection so one transient undici import/Agent failure can't brick every later call; transportFetch falls back to the base transport (no dispatcher) when the dispatcher can't be built, warning once per failure episode - bedrock: install fetch: transportFetch like every other transport - google proxy: build ProxyAgent with the same headers/body timeout budget (http/https proxies; undici's Socks5ProxyAgent drops these options for socks5://) - transportFetch no longer overwrites a caller-supplied dispatcher - tests: dispatcher failure fallback + retry, bedrock transport, caller-dispatcher preservation, google proxy budget --------- Co-authored-by: XingJia He <hexingjia@mobirit.com> | 19 天前 | |
fix(ai): let LLM calls outlive undici's 300s headers timeout (#1404) * fix(ai): let LLM calls outlive undici's 300s headers timeout A non-streaming completion only receives response headers once the whole completion exists, so undici's default 300 s headers timeout bounds total generation time. Thinking models (GLM-5.x at effort=max, MiniMax M3, …) routinely exceed that on large prompts — the request then dies with "Cannot connect to API: Headers Timeout Error" at exactly 300 s and, with maxRetries 0, the scene never generates (observed on a self-hosted deployment with glm-5.2 scene generation). Route every outbound LLM request through an undici Agent with a 15-minute headers/body timeout, attached at the shared transportFetch seam so the OpenAI-compatible, native OpenAI/Responses, Azure, Anthropic-compatible and Google transports all inherit it (same lazy-import dispatcher pattern the Google proxy transport already uses). Native transports now always install the shared transport instead of only when the server injects a validating fetch. * fix(ai): harden the LLM transport timeout fallback Review follow-ups for the undici headers-timeout fix: - getLlmDispatcher(): drop the cached promise on rejection so one transient undici import/Agent failure can't brick every later call; transportFetch falls back to the base transport (no dispatcher) when the dispatcher can't be built, warning once per failure episode - bedrock: install fetch: transportFetch like every other transport - google proxy: build ProxyAgent with the same headers/body timeout budget (http/https proxies; undici's Socks5ProxyAgent drops these options for socks5://) - transportFetch no longer overwrites a caller-supplied dispatcher - tests: dispatcher failure fallback + retry, bedrock transport, caller-dispatcher preservation, google proxy budget --------- Co-authored-by: XingJia He <hexingjia@mobirit.com> | 19 天前 | |
feat(attribution): send X-APP-URL on TokenDance gateway requests (#1675) Co-authored-by: Percy <percy@PercydeMacBook-Pro.local> | 6 天前 | |
fix(ai): let LLM calls outlive undici's 300s headers timeout (#1404) * fix(ai): let LLM calls outlive undici's 300s headers timeout A non-streaming completion only receives response headers once the whole completion exists, so undici's default 300 s headers timeout bounds total generation time. Thinking models (GLM-5.x at effort=max, MiniMax M3, …) routinely exceed that on large prompts — the request then dies with "Cannot connect to API: Headers Timeout Error" at exactly 300 s and, with maxRetries 0, the scene never generates (observed on a self-hosted deployment with glm-5.2 scene generation). Route every outbound LLM request through an undici Agent with a 15-minute headers/body timeout, attached at the shared transportFetch seam so the OpenAI-compatible, native OpenAI/Responses, Azure, Anthropic-compatible and Google transports all inherit it (same lazy-import dispatcher pattern the Google proxy transport already uses). Native transports now always install the shared transport instead of only when the server injects a validating fetch. * fix(ai): harden the LLM transport timeout fallback Review follow-ups for the undici headers-timeout fix: - getLlmDispatcher(): drop the cached promise on rejection so one transient undici import/Agent failure can't brick every later call; transportFetch falls back to the base transport (no dispatcher) when the dispatcher can't be built, warning once per failure episode - bedrock: install fetch: transportFetch like every other transport - google proxy: build ProxyAgent with the same headers/body timeout budget (http/https proxies; undici's Socks5ProxyAgent drops these options for socks5://) - transportFetch no longer overwrites a caller-supplied dispatcher - tests: dispatcher failure fallback + retry, bedrock transport, caller-dispatcher preservation, google proxy budget --------- Co-authored-by: XingJia He <hexingjia@mobirit.com> | 19 天前 | |
feat(llm): retry once on a fallback model for retryable failures (#1614) * feat: configurable model escalation scheduler ## Motivation Course generation can stall when the model configured for a heavy stage (e.g. `scene-content:interactive` / `scene-content:pbl`) is repeatedly rate-limited, times out, or returns empty output. Today the operator must manually switch models in `MODEL_ROUTES` and re-trigger the page; long-running generations die on a single bad call. This PR makes the fallback automatic and observable: an optional per-stage **escalation policy** that re-runs the failed scene once with an explicitly chosen upgrade model, and logs every decision to a durable ledger viewable from a new Settings section. ## Design - **Disabled by default, zero behavior change.** No `data/model-schedule.json` → the engine behaves exactly as before. The config is hot-read per call (250 ms cache), so saving from the panel takes effect without a server restart. - **Explicit escalation only.** The upgrade model is resolved through the existing `resolveModel` pipeline **outside** `MODEL_ROUTES` (unrouted model string wins), reusing the full provider/model parsing (`openai:...`, `qwen:...`). - **Trigger semantics.** `onTimeout` requires a timeout error; `onRetryableError` accepts any retryable generation error (`isRetryableGenerationError`). Content-safety style rejections are **not** retryable and **never** escalate — the safety boundary is preserved. - **Budget guard.** Optional `budget.dailyEscalationCap` stops escalations for the day once exhausted; per-stage `max` limits retries. - **Ledger.** Every escalation appends one line to `data/schedule-events.jsonl` (stage, scene, base → used, error class, reason); the Settings panel shows the last 20. ## Screenshots Settings > Model Scheduling (config present):  Ledger after an escalation (example entry):  ## Verification - `pnpm lint` — 0 errors - `pnpm check` (prettier) / `pnpm check:i18n-keys` — pass (12 locales) - `pnpm build` — Next.js production build passes - Engine logic script — 8/8 assertions (disabled default, policy parsing, budget guard, ledger I/O) - API A/B flow — GET null without config; PUT template → served immediately (hot reload); DELETE → back to null ## Notes - The config schema is minimal free-form JSON (optional `models` / `budget` / `escalation` / `strictMode`); a stricter TS schema can follow in a later PR. - `strictMode` is reserved (not enforced yet). - The 3-tier preset template lives in the panel's "Load 3-tier preset template" button; no example file is shipped. * feat: configurable model escalation scheduler ## Motivation Course generation can stall when the model configured for a heavy stage (e.g. `scene-content:interactive` / `scene-content:pbl`) is repeatedly rate-limited, times out, or returns empty output. Today the operator must manually switch models in `MODEL_ROUTES` and re-trigger the page; long-running generations die on a single bad call. This PR makes the fallback automatic and observable: an optional per-stage **escalation policy** that re-runs the failed scene once with an explicitly chosen upgrade model, and logs every decision to a durable ledger viewable from a new Settings section. ## Design - **Disabled by default, zero behavior change.** No `data/model-schedule.json` → the engine behaves exactly as before. The config is hot-read per call (250 ms cache), so saving from the panel takes effect without a server restart. - **Explicit escalation only.** The upgrade model is resolved through the existing `resolveModel` pipeline **outside** `MODEL_ROUTES` (unrouted model string wins), reusing the full provider/model parsing (`openai:...`, `qwen:...`). - **Trigger semantics.** `onTimeout` requires a timeout error; `onRetryableError` accepts any retryable generation error (`isRetryableGenerationError`). Content-safety style rejections are **not** retryable and **never** escalate — the safety boundary is preserved. - **Budget guard.** Optional `budget.dailyEscalationCap` stops escalations for the day once exhausted; per-stage `max` limits retries. - **Ledger.** Every escalation appends one line to `data/schedule-events.jsonl` (stage, scene, base → used, error class, reason); the Settings panel shows the last 20. ## Screenshots Settings > Model Scheduling (config present):  Ledger after an escalation (example entry):  ## Verification - `pnpm lint` — 0 errors - `pnpm check` (prettier) / `pnpm check:i18n-keys` — pass (12 locales) - `pnpm build` — Next.js production build passes - Engine logic script — 8/8 assertions (disabled default, policy parsing, budget guard, ledger I/O) - API A/B flow — GET null without config; PUT template → served immediately (hot reload); DELETE → back to null ## Notes - The config schema is minimal free-form JSON (optional `models` / `budget` / `escalation` / `strictMode`); a stricter TS schema can follow in a later PR. - `strictMode` is reserved (not enforced yet). - The 3-tier preset template lives in the panel's "Load 3-tier preset template" button; no example file is shipped. * Delete omai-upload-tree directory * fix: ensure trailing newline in locale files fix: ensure trailing newline in locale files * chore: english comments for upstream review chore: english comments for upstream review * Delete components/settings/model-schedule-settings.tsx * Delete lib/server/model-schedule.ts * Delete app/api/model-schedule/route.ts * Delete assets/model-schedule/ledger-with-entry.png * Delete assets/model-schedule/settings-panel.png * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Update .env.example * Enhance access control comments in .env.example Added additional context and warnings for access control configuration. * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * Add files via upload * run ci * Add files via upload * Add files via upload * Add files via upload * Supplement. env.example Supplement. env.example * Add files via upload * Update .env.example * Update fallback notes in .env.example Clarified fallback behavior in callLLM layer documentation. * Update llm.ts * Update llm-fallback.test.ts * fix(llm): round-4 fallback hardening fix(llm): round-4 fallback hardening — server-managed gate, content-filter refusal, APICallError unwrap stop, outlines fullStream errors * fix(llm): rebase round-4 onto main v1.1.1 fix(llm): rebase round-4 onto main v1.1.1 * fix(llm): round4c-prettier formatting+retry test fix(llm): round4c-prettier formatting+retry test * test(llm): round-4d-align outline-stream test(llm): align outline-stream and title-generator tests with round-4 call signatures --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> | 2 天前 | |
[codex] Add Amazon Bedrock LLM provider support (#538) * Enable AWS-hosted text models without exposing ambient credentials Rebase the Bedrock provider onto current main and close the review-requested security, credential lifecycle, usage attribution, configuration, and metadata gaps. Constraint: Bedrock may use ambient AWS credentials only when explicitly enabled by the server operator Rejected: Trust client-supplied provider types | permits built-in keyless IDs to reach Bedrock credentials Confidence: high Scope-risk: moderate Directive: Keep provider ID/type validation at both request resolution and model construction boundaries Tested: targeted Bedrock/resolver/config/usage tests; TypeScript; ESLint; i18n alignment; production build Not-tested: Full suite has 27 failures in unchanged quiz/runtime and runtime/chat-storage tests * Give Bedrock a recognizable provider identity Use the official AWS architecture service icon so Bedrock no longer falls back to the generic provider cube in settings and model selection. Constraint: Preserve the AWS-provided artwork without redesigning the service mark Confidence: high Scope-risk: narrow Tested: SVG XML validation; provider unit test; TypeScript; ESLint; production build --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> | 1 个月前 | |
feat: add Claude Opus 4.8 and MiniMax M3 (#635) Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 3 个月前 | |
feat: support GPT-5.6 model family (#907) * feat: support GPT-5.6 model family * test: cover GPT-5.6 SDK validation * fix: canonicalize GPT-5.6 Sol alias * fix: complete GPT-5.6 alias lookups --------- Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 2 个月前 | |
feat(token-plan): add TokenDance one-key preset for every modality (#1525) * feat(token-plan): add TokenDance one-key preset for every modality TokenDance is a model gateway: chat and images are OpenAI-compatible at /gateway/v1, and the same key authenticates vendor-protocol routes on the same host (Ark, MiniMax, Bocha). The preset reuses the existing adapters with those route prefixes as base URLs, so one key lights up LLM, image, video, TTS and web search from Settings -> Token Plan. - providers: add a built-in `tokendance` OpenAI-compatible provider (TOKENDANCE_* env prefix, logo, provider name in all locales) - token-plan: add the TokenDance preset (Seedream image, MiniMax H3 video, MiniMax speech TTS, Bocha web search) - seedream: use a base URL that already ends in a version segment verbatim, so gateway routes like `/ark/v3` do not get `/api/v3` appended - minimax-video: route H3-family models through the v2 task API (content array submit, task-envelope poll); connectivity checks for H3 probe auth on the v2 query route instead of submitting a billable task - README: add a one-key quick example and replace the Gemini-specific model recommendation with a provider-agnostic setup recommendation Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qrrq9CPwb718mpouz8Y2KL * fix(token-plan): accept preset web-search base URLs and report H3 dimensions per ratio - web-search: the client base URL allowlist also accepts the exact base URL a built-in token plan preset writes for that provider, derived from TOKEN_PLAN_PRESETS. Applying a plan whose web-search route is not an official vendor host previously stored a URL that the route rejected with 400. Any other client URL is still rejected. - minimax-video: report H3 v2 clip dimensions for 16:9, 9:16, 4:3 and 1:1 instead of assuming landscape for every non-portrait ratio. - tests: pin the allowlist for every preset, the 1:1 H3 dimensions, and clear TOKENDANCE_* in the provider-config env isolation list. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qrrq9CPwb718mpouz8Y2KL --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> | 14 天前 | |
fix: keep long Grok relay generations alive and inline image bytes (#1364) Two independent Grok failures seen when the provider is reached through a relay (a custom base URL) instead of api.x.ai directly: - lib/ai/providers.ts: a long non-streaming chat completion was cut off by the relay with a 504 at its idle timeout (~5 min), because nothing is sent upstream until the model has the whole answer. Adding 'grok' to the existing streaming-compat path (OPENAI_COMPAT_USE_STREAMING_CHAT=true) keeps bytes flowing across the idle window; the SSE is buffered back into a normal JSON response for the caller. The path is for relays only. usesCustomOpenAIBaseUrl recognises OpenAI's origin alone, so Grok's own api.x.ai also read as "custom" and was forced onto the compat transport; the provider's native endpoint is now excluded. - lib/media/adapters/grok-image-adapter.ts: response_format 'url' returns a link on the relay's CDN host (imgen.x.ai), which may be unreachable from the server's network. The generation then failed at the follow-up fetch through /api/proxy-media even though the image had been produced successfully. 'b64_json' inlines the bytes and removes that second hop. Inline bytes declare no media type, so the adapter reports one on ImageGenerationResult and returns a typed data URL, which is the shape openrouter-image-adapter already uses. Consumers take the type from there: agent image persistence records it, the client's stored media row keeps the type its data URL states, and classroom-media-generation names the file with the matching extension. A JPEG is no longer stored, served or named as a PNG. The process-wide undici timeout that previously accompanied these changes is dropped: upstream #1404 now gives LLM calls their own undici headers/body timeouts, which covers the same failure without raising the defaults globally. Co-authored-by: ciclou1 <ciclou1@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> | 8 天前 | |
fix: pass reasoning_content back to DeepSeek in thinking mode (multi-turn) (#1486) * fix: pass reasoning_content back to DeepSeek in thinking mode (multi-turn) DeepSeek rejects multi-turn thinking-mode requests whose assistant messages lack reasoning_content: 400 invalid_request_error: "The `reasoning_content` in the thinking mode must be passed back to the API." Two gaps caused it: 1. Round-trip loss. Responses fold reasoning_content into an inline <think> block (then extractReasoningMiddleware splits it into reasoning parts), but on the next turn the OpenAI chat adapter drops reasoning parts. Kimi already had a preservation mechanism (marker encoding + restore); extend it to DeepSeek, in both compatFetch and the middleware assembly. 2. Missing field. A response that skipped reasoning (direct tool call) produces an assistant message with no reasoning part at all. When thinking is enabled, DeepSeek still requires the field, so inject an empty reasoning_content for such messages. Also: send the thinking toggle alone (no reasoning_effort) when the route sets no explicit effort. Tool-using transports (maic-agent-driver) cannot combine reasoning_effort with function tools; DeepSeek applies its own default effort. Verified: vitest (reasoning-sse, thinking-config, openai-sdk-integration), Docker production build, and end-to-end multi-turn agent sessions with thinking enabled (DeepSeek v4-pro + v4-flash) completing successfully. * fix: close the DeepSeek reasoning round-trip on the agent-driver path Review follow-up on the DeepSeek reasoning_content fix. - One shared predicate, preservesReasoning(providerId, modelId), derived from the resolved request adapter (deepseek + atlascloud deepseek models) plus Kimi K3, now gates both the request-side wiring and the agent driver's includeReasoning. The driver previously gated on kimi-k3 alone, so every thinking block was dropped before the AI SDK saw the messages and the middleware had nothing to encode. - Driving the gate through the predicate covers atlascloud deepseek models, which share the request adapter but were previously excluded by a literal provider-id check. - Scope the effort-less thinking toggle to requests that carry tools: a tool-carrying request can never set reasoning_effort, while non-tool calls keep the historical default effort. - Inject the catalog thinking default for preserved providers so the wire always states the mode explicitly, and gate both restore and backfill on that resolved mode (a disabled turn carries neither markers nor field). - Tests: DeepSeek counterparts of the Kimi K3 recorded-shape test covering the round-trip across tool continuations, the empty backfill, the reasoning_effort scoping, a non-preserving provider, the non-streaming path, and the predicate itself. Verified they fail when the predicate is disabled. * fix: key the reasoning round-trip on the resolved thinking mode; cover the driver wire Round-2 review follow-up. - A disabled turn no longer ships the internal marker. The request-side step now keys on the resolved ThinkingConfig instead of the serialized body, so a disabled turn strips the sentinel instead of sending it verbatim; this also covers Kimi K3, whose adapter emits reasoning_effort and no thinking object. - Backfill an empty reasoning_content only for the DeepSeek request adapter. - Add the missing driver-wire test: createCallLlmStreamFn with a DeepSeek turn carries the prior thinking block as non-empty reasoning_content, and a disabled turn strips the marker. Reverting the driver gate to the old Kimi-K3-only check fails the first test; disabling the strip fails the second. - Formatting: prettier clean for the touched files. * fix: scope the effort-less thinking toggle to tool-carrying requests Review follow-up. `options.hasTools || config.effort === undefined` also dropped reasoning_effort for NON-tool calls that enable thinking without an explicit effort (e.g. {enabled:true}), where the historical wire carried reasoning_effort:'high'. Only a tool-carrying request must omit the effort (the transport rejects function tools combined with it); every other request keeps the explicit value or the default. Updated the scoping test to assert reasoning_effort:'high' for the non-tool no-effort case, and verified the assertion fails when the old condition is restored. * fix: derive the round-trip decision from the wire, not the requested mode Round-3 review follow-up. Two remaining issues: 1. Kimi K3 disabled strips reasoning_content while the request still reasons. getThinkingMode({enabled:false}) is 'disabled', but the openai-adapter body builder emits reasoning_effort:'low' (no 'none' level exists), so thinking was never actually off on the wire. getCompatThinkingBodyParams now also reports disablesThinking, computed per adapter from the params it really emits (true only where the wire carries an explicit off switch); the round-trip gate keys on that instead of the requested config. Same fix covers DeepSeek {effort:'none'}, which sends thinking:{type:'disabled'} but previously went through restore. 2. Empty reasoning leaks the private marker into content. extractKimiReasoning now reports a found flag, restore/strip act on it (marker removed in both cases; field emitted only by restore), and the preservation middleware no longer encodes empty reasoning parts, so the degenerate `:0:` marker is never produced. Restore/strip share one internal helper. Tests: Kimi K3 {enabled:false} recorded body (field kept + reasoning_effort low present, no thinking object); DeepSeek {effort:'none'} (disabled object, no field, no marker); driver-wire empty-thinking cases (enabled: field '' + clean, disabled: no field + clean); Kimi K3 no-backfill pin; unit cases for the empty marker. Each pinned by mutation: old gate fails T1+T2, serialized absent-strip gate fails T1 plus the pre-existing Kimi test, found-revert fails the empty-marker tests, widened backfill fails T4. Full suites green, tsc and prettier clean. --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> | 3 天前 | |
fix: pass reasoning_content back to DeepSeek in thinking mode (multi-turn) (#1486) * fix: pass reasoning_content back to DeepSeek in thinking mode (multi-turn) DeepSeek rejects multi-turn thinking-mode requests whose assistant messages lack reasoning_content: 400 invalid_request_error: "The `reasoning_content` in the thinking mode must be passed back to the API." Two gaps caused it: 1. Round-trip loss. Responses fold reasoning_content into an inline <think> block (then extractReasoningMiddleware splits it into reasoning parts), but on the next turn the OpenAI chat adapter drops reasoning parts. Kimi already had a preservation mechanism (marker encoding + restore); extend it to DeepSeek, in both compatFetch and the middleware assembly. 2. Missing field. A response that skipped reasoning (direct tool call) produces an assistant message with no reasoning part at all. When thinking is enabled, DeepSeek still requires the field, so inject an empty reasoning_content for such messages. Also: send the thinking toggle alone (no reasoning_effort) when the route sets no explicit effort. Tool-using transports (maic-agent-driver) cannot combine reasoning_effort with function tools; DeepSeek applies its own default effort. Verified: vitest (reasoning-sse, thinking-config, openai-sdk-integration), Docker production build, and end-to-end multi-turn agent sessions with thinking enabled (DeepSeek v4-pro + v4-flash) completing successfully. * fix: close the DeepSeek reasoning round-trip on the agent-driver path Review follow-up on the DeepSeek reasoning_content fix. - One shared predicate, preservesReasoning(providerId, modelId), derived from the resolved request adapter (deepseek + atlascloud deepseek models) plus Kimi K3, now gates both the request-side wiring and the agent driver's includeReasoning. The driver previously gated on kimi-k3 alone, so every thinking block was dropped before the AI SDK saw the messages and the middleware had nothing to encode. - Driving the gate through the predicate covers atlascloud deepseek models, which share the request adapter but were previously excluded by a literal provider-id check. - Scope the effort-less thinking toggle to requests that carry tools: a tool-carrying request can never set reasoning_effort, while non-tool calls keep the historical default effort. - Inject the catalog thinking default for preserved providers so the wire always states the mode explicitly, and gate both restore and backfill on that resolved mode (a disabled turn carries neither markers nor field). - Tests: DeepSeek counterparts of the Kimi K3 recorded-shape test covering the round-trip across tool continuations, the empty backfill, the reasoning_effort scoping, a non-preserving provider, the non-streaming path, and the predicate itself. Verified they fail when the predicate is disabled. * fix: key the reasoning round-trip on the resolved thinking mode; cover the driver wire Round-2 review follow-up. - A disabled turn no longer ships the internal marker. The request-side step now keys on the resolved ThinkingConfig instead of the serialized body, so a disabled turn strips the sentinel instead of sending it verbatim; this also covers Kimi K3, whose adapter emits reasoning_effort and no thinking object. - Backfill an empty reasoning_content only for the DeepSeek request adapter. - Add the missing driver-wire test: createCallLlmStreamFn with a DeepSeek turn carries the prior thinking block as non-empty reasoning_content, and a disabled turn strips the marker. Reverting the driver gate to the old Kimi-K3-only check fails the first test; disabling the strip fails the second. - Formatting: prettier clean for the touched files. * fix: scope the effort-less thinking toggle to tool-carrying requests Review follow-up. `options.hasTools || config.effort === undefined` also dropped reasoning_effort for NON-tool calls that enable thinking without an explicit effort (e.g. {enabled:true}), where the historical wire carried reasoning_effort:'high'. Only a tool-carrying request must omit the effort (the transport rejects function tools combined with it); every other request keeps the explicit value or the default. Updated the scoping test to assert reasoning_effort:'high' for the non-tool no-effort case, and verified the assertion fails when the old condition is restored. * fix: derive the round-trip decision from the wire, not the requested mode Round-3 review follow-up. Two remaining issues: 1. Kimi K3 disabled strips reasoning_content while the request still reasons. getThinkingMode({enabled:false}) is 'disabled', but the openai-adapter body builder emits reasoning_effort:'low' (no 'none' level exists), so thinking was never actually off on the wire. getCompatThinkingBodyParams now also reports disablesThinking, computed per adapter from the params it really emits (true only where the wire carries an explicit off switch); the round-trip gate keys on that instead of the requested config. Same fix covers DeepSeek {effort:'none'}, which sends thinking:{type:'disabled'} but previously went through restore. 2. Empty reasoning leaks the private marker into content. extractKimiReasoning now reports a found flag, restore/strip act on it (marker removed in both cases; field emitted only by restore), and the preservation middleware no longer encodes empty reasoning parts, so the degenerate `:0:` marker is never produced. Restore/strip share one internal helper. Tests: Kimi K3 {enabled:false} recorded body (field kept + reasoning_effort low present, no thinking object); DeepSeek {effort:'none'} (disabled object, no field, no marker); driver-wire empty-thinking cases (enabled: field '' + clean, disabled: no field + clean); Kimi K3 no-backfill pin; unit cases for the empty marker. Each pinned by mutation: old gate fails T1+T2, serialized absent-strip gate fails T1 plus the pre-existing Kimi test, found-revert fails the empty-marker tests, widened backfill fails T4. Full suites green, tsc and prettier clean. --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> | 3 天前 | |
feat(providers): add GLM-5.3 and GLM-5.3-Flash support (#1401) Zhipu released GLM-5.3 (2026-08-19, text flagship) and GLM-5.3-Flash (2026-08-26, native multimodal, 320B MoE / 18B active). Both serve a 1M-token context with 128K max output and always-on thinking controlled by reasoning_effort low/high/max — the API rejects thinking {"type":"disabled"} with code 1210 ("该模型始终思考,不支持关闭思考"). - catalog: glm-5.3 (text) and glm-5.3-flash (vision) entries, 1M/128K - metadata: glm53Effort capability — effort-only, not toggleable, default effort max (same forced-thinking pattern as Fable 5) - adapter: for non-toggleable GLM models, degrade a "disabled" thinking request to the lightest effort instead of sending the API-rejected {type:'disabled'}, mirroring the guard in the Anthropic adapter; this also un-breaks POST /api/verify-model, which hardcodes thinking off Co-authored-by: XingJia He <hexingjia@mobirit.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> | 19 天前 | |
fix(ai): share thinking context across bundle evaluations (#1378) Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> | 19 天前 | |
feat(ai): add Xiaomi MiMo V2.6 models [AI-assisted] (#1655) * feat(ai): add Xiaomi MiMo V2.6 models * docs(ai): preserve Xiaomi provider references --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> | 7 天前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 2 个月前 | ||
| 5 个月前 | ||
| 1 个月前 | ||
| 19 天前 | ||
| 19 天前 | ||
| 6 天前 | ||
| 19 天前 | ||
| 2 天前 | ||
| 1 个月前 | ||
| 3 个月前 | ||
| 2 个月前 | ||
| 14 天前 | ||
| 8 天前 | ||
| 3 天前 | ||
| 3 天前 | ||
| 19 天前 | ||
| 19 天前 | ||
| 7 天前 |