| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
fix(media): refuse redirects on adapter generation and poll calls (#1636) #930 made the connectivity probes in the media adapters pass `redirect: 'manual'`. The generation and poll calls in the same 14 files were left following redirects. Those requests carry the provider credential and go to a base URL that comes from provider settings a caller can supply, so a 3xx would replay the credential at a host the caller chose — and the redirect target can be an address the outbound guard already refused. Every such call now passes `redirect: 'manual'` and rejects a 3xx through a shared `assertNotRedirected` helper, which reports it as "<provider>: Redirects are not allowed (HTTP <status>)" instead of letting the generic failure path describe it as a provider error. - 26 call sites across the image adapters (seedream, openai, qwen, grok, lemonade, minimax, nano-banana), the video adapters (seedance, kling, grok, happyhorse, minimax, veo) and ComfyUI's submit and image fetch. - ComfyUI's `pollHistory` keeps its contract of handing the caller a retryable failure rather than aborting the generation: it logs the refusal and returns null. - ComfyUI's same-origin workflow load is deliberately untouched — it reads the app's own public/ asset, carries no credential and is not provider-influenced. - The two adapters added since #930 (OpenRouter image and video) already did this. tests/media/adapter-redirects.test.ts covers one case per adapter family. Each serves a 302 and asserts that the call rejects with the redirect message and that every request carrying an init object asked fetch not to follow redirects; each case fails if its adapter stops passing `redirect: 'manual'`. The HappyHorse test asserted the exact request init, so it now includes the new option. AI-assisted commit Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> | 5 天前 | |
feat(persistence): turn on the server-owned asset lifecycle and release assets on course deletion (#1007 amendment, part 2) (#1473) App wiring for the @openmaic/storage 0.31.0 lifecycle: reference tracking and document references are paired unconditionally and declared at startup, course deletion withdraws references inside the tombstone transaction, ASSET_PENDING_TTL_MS configures the pending window, and the dead client-side reclamation code is removed. | 12 天前 | |
feat(dsl): standardize the asset manifest and converge the export paths (#1007 part 3) (#1117) * feat(dsl): standardize the document asset manifest Add asset-manifest.ts to @openmaic/dsl: the canonical AssetManifestEntry shape (ref, kind, and byteSize/mimeType/duration/voice/prompt metadata where available) plus enumerateAssetManifest, the pure document-to-manifest enumeration. An entry's ref is the reference exactly as the document holds it -- the manifest is the id-based reference enumeration with metadata, not a content hash and not a resolution result. The traversal walks the stage whiteboard, each scene's canvas/whiteboards/speech actions, and the stage video-manifest keys in document order, with logical-owner reference counts that match the accounting duplication-safe replacement uses. This settles the media-ref + asset-manifest schema question (#779 open question 4) on the side the asset-pool RFC already implied: the schema is a function of the id semantics decided there. The type lives in the dsl rather than a new @openmaic/exporter package because the enumeration is pure over document types the dsl already owns (Stage/Scene/Slide/Action), so a separate package would add a published artifact and release-workflow surface without adding a capability; the storage contract comment now points at the module. Refs #1007 * refactor(export): drive the classroom ZIP from the asset manifest collectMediaFiles used to scan the whole mediaFiles table for the stage, so any row the document no longer references -- an orphan left by an edit or a superseded regeneration -- rode along into the archive. Both ZIP collectors now take their reference sets from the standardized asset manifest (buildStageAssetManifest wraps the dsl enumeration with the compatibility rows' metadata): only referenced assets are archived, and a referenced asset whose bytes exist only in the pool is still collected via a synthesized record. Byte resolution is unchanged: pool first through resolveStoredBytes / resolveAudioBlob, with the compatibility row kept as the legacy byte fallback and as the metadata source. mediaIndex is now a serialized view of the manifest, and the missing-audio report derives from the manifest's audio entries instead of a second action walk. The audioRef mapping and the legacy audioUrl fetch path (collectLegacyAudioForExport) are untouched. Refs #1007 * refactor(video-export): take the timeline's reference sets from the manifest createVideoTimelineDeps scanned the whole mediaFiles table for the stage and derived its audio id set from its own action walk -- a third, independent answer to "which media does this course use?". Both record loads now key off the standardized asset manifest: media rows are read per manifest ref by compound key instead of by table scan, and the audio id set is the manifest's audio entries. Orphan rows were never reachable through the scene-scoped elementId-to-mediaRef bridge; now they are not even read. The bridge itself is untouched: element ids recur across scenes, so the elementId-to-mediaRef mapping stays scoped per scene, and the legacy audioUrl fallback keeps its own action walk because a URL is not a manifest ref. AssetPlan remains the video IR's view of the same references. Refs #1007 * refactor(export): resolve PPTX media through the shared resolver only Each PPTX element branch carried its own resolution chain: a task-state renderable-URL lookup first, then -- gated on the legacy placeholder predicate -- a stored-bytes override, with the poster block repeating the pattern. One helper now owns resolution for backgrounds, images, video / audio sources, and posters: opaque refs (allocated ids and legacy placeholders alike, no placeholder-pattern gate) resolve pool-first through resolveStoredBytes and embed as data URLs, concrete addresses resolve through the media state machine and keep the caller's fetch path. exportMediaResolution and the resolveStoredMediaBlob wrapper fold into the helper; resolvePptxMediaBinding stays as the state-machine entry the resolution-surface test matrix drives. Refs #1007 * refactor(export): retire the export-side Dexie byte fallbacks Export call sites no longer read bytes off compatibility rows directly. The ZIP collectors and the video timeline's audio load resolve bytes only through the shared resolvers (resolveStoredBytes / resolveAudioBlob), which answer pool-first and keep the compatibility row as their internal legacy fallback level; the row reads that remain at the call sites supply metadata (format/duration/voice/mime/size/prompt) only. The rows themselves stay for legacy and regeneration readers -- what goes is the export paths' own fallback logic. One observable tightening: a failed media row (error set, empty placeholder blob) no longer ships a 0-byte file into the classroom ZIP, and an evicted row no longer ships its empty local blob; referenced-but- byteless assets are simply absent from the archive, as they already were when no row existed. Refs #1007 * test(media): cover the enriched stage asset manifest builder Pins the join between the pure dsl enumeration and the compatibility rows: metadata attaches by ref, rows no document reference names never appear, and a referenced asset with no row keeps a metadata-free entry. Refs #1007 * fix(video-export): widen the deps stage input for the manifest enumeration enumerateAssetManifest reads the stage's whiteboard and videoManifest, so createVideoTimelineDeps declares them on its input instead of the bare id; callers pass only the id today and the optional fields stay absent. Also applies the repo prettier formatting to the files this branch touched. Refs #1007 * fix(dsl): enumerate slide audio elements in the asset manifest Slide audio elements carry their own src, and the manifest skipped them, so a manifest-driven collector could never archive their bytes. The audio slot maps to kind 'audio' alongside narration ids. Refs #1007 * fix(media): harden ref-keyed lookups against prototype-named asset refs AssetRef is an unconstrained string alias, so a media reference can legitimately be "__proto__", "constructor", or any other Object.prototype member. Plain objects keyed by such refs silently drop assignments or answer lookups with the prototype object, which rewrite paths then accept as a mapped id. Convert the remaining ref-keyed lookup tables introduced by the export convergence to prototype-safe structures: the classroom import media/poster alias maps and the legacy-conversion video-manifest reconstruction now use Map / null-prototype containers with explicit membership checks, and every consumed value is validated as a string before it is written into a src / mediaRef / audioId slot. The shared media-task lookup receives the same treatment: one centralized own-property-checked lookupMediaTask now serves the stored-bytes resolver, the PPTX embeddable-src path, the video collection path, and the element/background task resolution, so a prototype-named placeholderRef can no longer hide a re-keyed task from the fallback chain. Adversarial tests drive "__proto__" and "constructor" refs through the import round trip, the PPTX fallback path, the legacy conversion commit path, and the media-task fallback end to end, including a buildPptxBlob regression with a task re-keyed to an allocated id while retaining a prototype-named placeholderRef. * fix(export): use safe archive asset paths * refactor(dsl): centralize slide media slot roles * refactor(export): derive consumer refs from manifest * fix(export): sanitize classroom archive extensions * fix(video-export): preserve narration speech order * fix(export): enforce kind-coherent archive media * fix(export): define media coherence boundary * fix(export): carry task-owned poster binding for PPTX export A video element with no explicit poster falls back to its media task's generated poster URL, but resolveVideoMediaForElement left posterTask undefined for that case, so the PPTX manifest guard saw a foreign URL with no task-ownership exemption and dropped the video element instead of using the established runtime poster fallback. Carry the poster task binding whenever the task poster is the effective poster: the task-owned URL then satisfies the guard's objectUrl exemption end to end. A concrete explicit element poster still stays element-owned and never borrows the binding, and the guard's foreign-ref rejection is preserved (and exported as a directly testable predicate). Coverage: an element with no poster plus a task-provided poster embeds the task poster as the PPTX cover (red at the pre-fix head, green now), and a genuinely unrelated URL with no task ownership is still rejected by the guard. * fix(export): preserve legacy narration source refs in the media index The explicit sourceRef contract was partial: primary audio and generated media entries carried it, but legacy URL narration serialized no source ref. The legacy URL itself is the natural source ref — it is known at fetch time — so wire it through the collected blob into the mediaIndex entry. Import already registers serialized sourceRefs as aliases, so the URL now round-trips as an explicit mapping instead of being reconstructed only from the action's audioRef. Poster siblings are deliberately NOT given their own mediaIndex entry: a sibling poster (media/asset-<n>.poster.<ext>) is a legacy byte copy written from the video record and is not an independently referenced document asset — when the poster is a real document asset it already has its own indexed entry with a sourceRef, and import reconstructs the sibling by path derivation from its parent video entry, reusing the poster's own indexed allocation when one exists. The PR description is narrowed to match; corrected paragraph: "Archive names never interpolate refs — sequential safe paths (media/asset-<n>.<ext>, audio/audio-<n>.<ext>) with the original ref preserved through an explicit sourceRef mapping on every independently indexed media entry: generated media assets, poster assets, primary narration, and legacy URL narration (the legacy URL itself is the entry's sourceRef). Extensions are allowlisted per kind. The one exception is the legacy sibling poster byte copy (media/asset-<n>.poster.<ext>, written next to its video when the video record still carries the pre-pool poster bytes): it is not an independently referenced document asset, so it has no mediaIndex entry or sourceRef of its own — its identity is derivable from its parent video entry (same index), and import reconstructs it by sibling-path derivation from that video entry, reusing the poster's own indexed allocation when one exists." --------- Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 1 个月前 | |
feat(media): write generated media through the asset pool under server-backed persistence (#1392) * feat(media): store generated media in the asset pool when persistence is server-backed With server-backed persistence the document is durable and shared, but generated media stayed in the producing browser: the document kept its gen_img_* / gen_vid_* placeholder and narration kept a browser-derived audio id. Every new browser that opened such a course re-ran generation for every slide, and it never converged, because the address of the generated bytes was never written back into the document. Under server-backed persistence only, the classic generation chain now stores bytes in the asset pool first and writes the id the pool allocated into the document. - The client bootstrap configures the asset seam alongside the document and runtime seams: an HttpAssetStore over the persistence endpoint carrying the same credentials the document store carries, marked server-backed. The seam preflight now covers all three, so a failure still cannot half-configure persistence. - Image, video and TTS generation commit in one fixed order: provider, pool, document, local cache, task. A reference reaches the document only after put returned an id, so a document can never name bytes that were not stored. A failure before the write-back leaves the placeholder with the provider called exactly once; the retry happens on the next owner load. - The write-back is a per-slot rewrite through mutateDocument, which re-reads the current document under the per-stage lock, so it cannot clobber a newer scene. The open course is refreshed with the same rewrite without being marked dirty. - "Has this already been generated?" is answered by the document (the slide exists and no longer holds the placeholder) instead of by this browser's task table. - The classroom's resume effect fails closed on ownership: only a resolved owner starts generation, so a viewer opening a shared course spends nothing. - The local media and audio tables become a per-tab cache. A failed cache write costs a re-download, never the media. Browser-only mode is unchanged: every new call site sits behind the server-backed gate, the local tables stay authoritative there, and placeholders stay in the document. Rendering and export needed no changes. HttpAssetStore.resolve mints an object URL exactly as the browser store does, and the export byte resolver was already pool-first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(classroom): make the generation owner gate a three-outcome rule and apply it everywhere The gate refused everything but a resolved owner, which read a sidecar that answered "no ownership fact exists for this course" as a reason to block. That is the answer a deployment without the sidecar's server-side prerequisites gives for every course, and the answer a course with no ownership record gives: in both, there is nobody the operator's budget needs protecting from, and refusing strands the course's own author behind a question that can never be answered. Ownership is now four states over the sidecar's three outcomes. A definite answer splits into owner and not-owner. An absent record is its own answer, ownerless, and generation proceeds — the behaviour such a deployment had before the gate existed. Only the absence of an answer, a transport failure or a load that has not asked yet, stays unresolved and fails closed: "we could not ask" must never be read as "nobody owns this". One mapper turns a sidecar result into that state, and one predicate decides on it. The workbench classroom pane runs the same resume effect and had no ownership input at all, so a viewer opening a shared course there could still spend the budget. It now asks the sidecar once per course, in parallel with its load and feeding only the generation gate, so its read-only and edit behaviour is unchanged. The shared progressive-load policy carries the gate for it, with both new inputs required rather than defaulted so a future caller cannot omit them into an open budget. Its stale comment claiming ownership could not be expressed here is corrected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the write-back survive autosave, arrive before the scene does, and never leak Independent reviews of the write-back found three ways a durable document could still end up naming a placeholder, and two ways the gate that protects the operator's budget could be walked around. An autosave round captures the store synchronously and writes that capture, so a round already in flight when a rewrite landed wrote the placeholder straight back over the allocated id, and nothing marked the store dirty again to correct it. The rewrite now marks the units it changed, which leaves a corrective flush queued behind the stale one; re-saving a scene that already holds the id is idempotent, losing the id is not. Media is generated from outlines in parallel with scene content and usually finishes first, so the slide that will carry the placeholder does not exist yet and the write-back has nothing to rewrite. That was the ordinary path, not a tail case, and its result was discarded: the task was marked done, the scene was added afterwards with its placeholder intact, and a second pass in the same run could call the provider again. The allocation is now held under the placeholder — which also answers the skip test, so nothing pays twice — and applied when that scene is committed, before its first save. One complete pass now leaves no placeholder behind. A failed commit used to abandon what it had already allocated. A poster upload that failed threw away a stored video and sent the retry to submit the most expensive job in the system again; a rejected write-back left registry rows that name bytes nothing references, which the byte collector cannot reclaim because it only collects blobs no row names. A poster failure now costs the poster, and a write-back that reached nothing reclaims what it allocated. A partial write is left alone, because the document already names it. The ownership gate is fail-closed again. Treating the sidecar's 404 as permission was wrong: the client cannot tell "this course has no owner" from "this deployment told me nothing", so a visitor who opened a shared course could bill the operator. The root cause was the sidecar itself, which gated on the agent runtime although every persisted course has an owner regardless — the persistence route resolves one for every request. It now gates on server persistence, so the configuration that made 404 the universal answer has real ownership facts to report, and the gate can refuse everything but a named owner. Retry affordances answered to no gate at all. A viewer of a shared course with one failed image was shown a Retry button that called the provider. Both retry entry points and every surface that draws them now read one shared permission, so what is offered and what is allowed are the same value. Also: narration regeneration no longer pretends it can replace bytes behind a live id — the exclusivity proof that would allow it is refused by construction once references leave the browser, so it forks to a fresh id and says so; the "already generated" test lets a finished deck answer from the document alone, since scene order stops identifying an outline once slides are inserted or deleted; stored assets record a specific media type rather than a generic transfer type; the pane no longer asks the sidecar in browser-only mode; and the funnel's docstring now states what the per-stage lock actually guarantees, which is same-browser serialization and not a cross-browser compare-and-swap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): park allocations in the deciding turn, never reclaim on an ambiguous write A delta review of the write-back found the first-pass fix still had a window, and the reclamation it added could delete media the document already names. The allocation was parked after an awaited local cache write. A scene committed in that window reconciled against a registry that did not hold it yet, so the document kept the placeholder — and the entry recorded a moment later then answered the skip test as "already handled", so nothing could correct it. Parking now happens inside the write-back, in the same synchronous turn as the decision that nothing could take the reference; no await separates the live check from the park. Allocations parked by an earlier pass are handed to their slides at the start of the next one, before anything decides what still needs generating, so a held allocation whose scene has since arrived becomes a rewrite rather than an answer. Reclaiming on a rejected write was unsound: a rejection does not prove the server did not apply the write, so deleting the asset could break the scene that now names it. The funnel decides instead, and says so: it reclaims only when no store write was ever issued and nothing took the reference. Anything else is placed if its slide exists and parked if it does not, so the next pass reuses the bytes instead of paying for them again. When a write fails after part of it landed, the live store is brought up to the document before the error is rethrown — otherwise the next ordinary flush would overwrite the half that did land, with the ids deliberately not reclaimed. Parked allocations are now cleared with the course. Classic placeholders are reused across runs, so one surviving an interrupted run would be handed to a different slide of the next deck: the previous picture, on a slide whose provider was never asked. Both classroom surfaces clear the arriving course, the deletion cascade clears the deleted one, and clearing the database clears them all. Two more ways generation could start without asking the gate are closed. An overlapping pass — an outline retry re-enters generation with every outline while the first is still working — re-requested elements whose provider call was already in flight; a task that is not done is an answered request, not an unanswered one. And narration regeneration in the timeline editor called the TTS provider and allocated a pool asset with no ownership check at all; it now reads the same permission, which withholds both the per-line and whole-timeline controls and refuses the call. Finally, a pane opened during the stage-link availability gap recorded the sidecar's 404 for a course that was moments from existing and never asked again, leaving the real owner locked out of generation until it remounted. Ownership is re-fetched once the document becomes available; the gate stays closed until an answer arrives, so asking again can only open it for someone entitled to it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the asset routes reachable, and stale snapshots harmless A full-branch audit found that the deployment this project documents could not store a single generated asset, and that several routes into durable storage could still write a placeholder over a reference that had already landed. The persistence route sent asset requests through the development authenticator, which refuses outright in a production build that has not explicitly opted into it — and the documented server-persistence recipe produces exactly that build. Every store and every read answered 401, so images and video failed on every slide while re-billing the provider on each retry, and a narration failure stopped the deck at its first slide. Assets live in one shared partition by design, so there was never anything per-caller for that authenticator to decide: the route now resolves the asset principal itself, alongside the owner it already resolves for documents. Runtime sessions are genuinely per-learner and keep the development authenticator until real session verification replaces it. And narration that cannot be stored no longer fails its scene: the line stays unvoiced and retryable, which is what an image that cannot be stored does to its slide. Placeholders could also come back from behind. A queued autosave's snapshot, an editor-history entry replayed by an undo, the departing save a course switch flushes — each captures content at its own moment, and any of those moments can predate a write-back. Point fixes at each producer would leave the next producer to rediscover the bug, so the check lives at the write boundary every producer passes through, and the allocation record it consults now outlives the parked queue: a placeholder whose rewrite landed long ago is exactly the case it catches. Two ways generation could be lost or repeated are closed. A pass now claims the elements it will reach and releases them however it ends, so an overlapping pass stands down while an aborted one strands nothing — previously its tasks stayed `pending` and every later pass skipped them with no retry control to recover them. And the media abort controller is aborted before being replaced, so a superseded pass stops calling providers instead of running on for a course the user has left. The remaining two are narrower. The workbench pane asks for ownership only after a document load succeeds, and after every later one, mirroring the page route: the load is what creates the ownership row the first time a course is opened, so asking beforehand asked about a course that did not exist yet and locked its author out for the mount. And the ownership gate on the timeline editor now withholds narration regeneration alone; listening back to existing narration and seeing whether a line has any spend nothing and stay available. Known limitation, unchanged and now stated plainly in the comments that used to point at it as a solution: nothing reclaims an unreferenced pool asset. The registry sweep is written but not wired up, and the byte collector only reclaims blobs no registry row names, so every narration regeneration and every abandoned allocation leaves storage behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): gate asset mutations, and make claims and allocation records survive a handoff Opening the asset routes opened all of them. Reads and allocations are meant to be as open as document reads and creates already are, but no authorization hook was supplied, so the handler's default admitted PUT and DELETE too — and those scope by principal key alone, which is one shared constant. Any caller who learned an id, and a document read hands out every id its slides name, could overwrite or destroy another author's media. Mutations now require the deployment's credential, which in a production build without the development-auth opt-in means they are refused outright; reads and allocations stay open. The route comment says what the posture is and what it is not: the deployment-level fence is the access code, and no per-principal quota is configured. The client's own reclaim is best effort to match — losing an argument about deleting an asset must not cost a task its retry, and the bytes are left for server-side reclamation. The pass claim could not survive the handoff it was written for. A retry aborts the live media pass and starts its replacement in the same synchronous block, long before the aborted pass's cleanup runs, so the replacement saw every element still claimed, collected nothing, and returned — leaving each unreached element at pending with nobody coming back for it and no retry control to recover it, which is the exact failure the claim was introduced to prevent. A claim now carries its pass's signal and is retired the moment that signal aborts, and a pass releases only claims it still owns, so a late unwind cannot take its replacement's work. Claims are also acquired at the single point every request passes through, so a single-task retry participates too — previously a retry awaiting its provider was invisible to a pass starting alongside it and both called it. The allocation record could outlive the bytes it named. It was written before the write-back attempted anything and survived the reclaim that followed a failure, so when the slide finally arrived the write boundary stamped a deleted id into the document — and the placeholder it replaced was gone, which reads as already generated and stops anything from retrying. The record is now written only where the allocation is retained, and forgotten wherever a reclaim removes the bytes, including the narration rollback path. The tests follow. The route test drives the real storage handler against an in-memory registry instead of a stub, so it can see what the resolved principal is then allowed to do; the handoff test performs a real abort mid-pass rather than starting from an already-aborted signal; and the guards that could only assert file layout now assert the property they care about, or have been replaced by behaviour. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make media passes serial per course instead of tracking element ownership Three rounds of per-element claims each produced a new way to lose an element. Whole-pass reservations swallowed a Retry for an element the same pass had already failed, leaving it pending with the affordance gone. Retiring a claim by its signal freed an element whose commit was still uploading, so the replacement pass paid for it twice. A claim held for a failed element stranded its retry. The bookkeeping is the defect: every refinement of "who owns this element right now" answered the question at a moment when the answer was already stale. Passes for one course are now serial. A replacement aborts its predecessor, as before, and then waits for it to settle before collecting. That removes the question entirely: a commit already under way finishes — its bytes stored and its reference written, so the new pass sees a resolved slide and skips it — and an element the aborted pass never reached is still a placeholder and gets collected like any other. The claim set, the reservations, the signal retirement and the identity-checked release are all gone. The task table is consulted for one thing only: an element that is generating right now is a single-element retry running alongside the pass, and taking it too would pay twice. Pending is deliberately not a skip reason — it means a pass once intended to reach an element, which an abandoned pass leaves behind with nobody acting on it, and reading that as answered is what stranded elements before. A retry runs concurrently with a pass, because a pass never revisits an element it has processed, and it re-reads the task after its own await and refuses before touching it: marking first and refusing afterwards destroyed the failed state that draws the affordance. Browser-only mode is back to exactly what it was. The abort is now conditional, the waiting does not apply, and the original status-based skip is restored verbatim. Two baseline lines remain changed in each of the two files, and both are behind a server-backed fork whose else-branch is the original. Two smaller things. The allocation record becomes visible when a write goes on the wire rather than when the round trip ends, and the write boundary reconciles under the document lock rather than before it — a save queued during a write-back was otherwise captured with the placeholder and, for a course the user had left, had no corrective flush to follow. And the comments that said a refused reclaim leaves its bytes for server-side reclamation were wrong: nothing collects them, because the registry entry still names its blob and the sweep that would remove it is not wired up. They now say the bytes leak. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a deferred pass re-earn its right to run, and bound the commit it waits on Serializing passes moved their body out of the block that launched them, and three things followed from that. A pass now wakes when its predecessor settles, which can be after the user has left the course. It enqueued before it looked at its signal, into a task table keyed by element id alone — and placeholder ids are not unique across courses, which is why the classroom clears that table on arrival. So a departing course's pass seeded the arriving course's table with tasks carrying the wrong stage id, and a Retry routes by that id: the reference went into the wrong document. A pass now re-validates after the wait, before touching anything shared. The same lateness broke the skip test. `documentSkipIndex` answers only while the live store is on the pass's stage, and returning nothing put the collection loop on the browser-only rule — a silent demotion from "the document is the authority" to "this browser's task table is", on exactly the path where that table has just been cleared. Every element the predecessor had committed was collected again, paid for again, and its second write-back found no placeholder to rewrite, so its bytes were parked where nothing will ever reference them. In server-backed mode an unreadable document now means the pass stands down. And waiting was unbounded. A commit is uncancellable: the asset client takes no signal, and a document write cannot be half-undone. One stalled upload therefore froze the course's media generation for the session — the replacement never collected, the element sat on a skeleton that draws no Retry, and only a reload recovered. The pass's signal is now threaded into the media proxy fetch, and the commit is bounded by a deadline. The deadline is on the wait, not the work: the commit carries on, and if it lands late the document simply ends up correct, while the element becomes retryable and the queue moves on. The tests that were meant to pin the previous round were not sensitive to it. Two asserted end states where the mechanism only changes ordering, and one of them rigged the document read so the assertion held whether or not the pass had waited; a third covered half of what it claimed. They now observe the ordering directly — nothing is issued while another pass for the course is working; in browser-only mode a second pass reaches its provider immediately — and the reconciliation under the document lock has a test that fails when it moves back outside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * revert(media): drop the commit deadline and the abortable download The deadline bought less than it cost. Abandoning a commit after two minutes makes the element retryable while the real commit is still running, so a Retry starts a second commit for the same placeholder against the first: two provider calls, two allocations, and whichever lands second stamps its result over the other's task by element id. The allocation record is keyed by placeholder, so the loser's cleanup erases the winner's record, and the write boundary then puts the raw placeholder back into the document. That is the overlap serial passes were built to remove, reopened through the one door serialization never covered. So a stalled commit holds the course's media queue until it settles or the page is reloaded, and that is written down rather than papered over. The wait is unbounded on purpose: every ceiling on it turns out to be a way of running two commits for one element. Threading the pass signal into the download was also a mistake, in the other direction. The provider call that produced the URL has already been billed, so cancelling the download throws away work that is paid for — and the shared proxy cache records a cancelled request as a transient failure against that URL, which after three of them blocks it for every consumer in the session. Browser-only mode never asked for this: it had no way to observe an abort there, which is exactly why the bytes were kept. The signal is gone from the download again, and `fetchAsBlob` is byte-for-byte what it was before this branch. The regression guard for the stranded-element rule is restored alongside the timing test that was meant to supersede it. It catches a different rule — a task left pending being read as answered — and nothing else does: making the pass skip pending leaves every other suite green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): stop asking the pool for refs it never issued, bound it, and adopt cached bytes Four things a deployment found once this was running for real. A reference this application mints itself — a generation placeholder, a derived narration key — was never in the pool, because the pool allocates every id it holds. Asking anyway used to be an IndexedDB miss; once the pool is server-backed it is a request that answers 404, one per element per load, forever on a course that still holds placeholders. Every lease and probe now checks first. The check is a negative test on shapes this application owns, not an id validator: the pool's id domain stays unconstrained, and anything that is not one of ours is still asked about. The asset store can bound how much one principal holds, and enforces it inside the write transaction, but nothing ever passed the number. It does now, with a default rather than an opt-in: allocation is reachable by any caller a deployment admits, and with one shared principal an unbounded store is unbounded database growth with no operator-visible brake. Refusing asset mutations to unauthenticated callers was not enough, because every authenticated caller resolves to that same shared principal — so authentication decided nothing, and any signed-in visitor could delete any id they learned. Since this branch began storing media the registry is the only copy a course has. Replacing and deleting are now refused to everyone, and the browser no longer tries: an entry nothing references waits for server-side reclamation instead. What a browser must still do is forget its own record of an allocation that reached nothing, or a later save would stamp an id the document has no reason to trust. And a course generated before any of this holds placeholders in its document with its bytes only in the author's browser. Those bytes are paid for, so the author's next load converts them — stored to the pool and written back through the ordinary commit path, with no provider call — instead of buying them again. A row that records only a hosted URL is treated as absent: that URL is the provider's address, not something a document may hold. One renderer expectation moved with this. An untracked placeholder used to paint as pending on first render because asking the pool left a lease in flight; it settled to disabled a moment later either way, and now says so from the start. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): surface a full store as a refusal and convert legacy narration A quota refusal reached the browser as HTTP 500 with a generic message, which reads as a transient failure: the element kept a Retry that would pay a provider again and be refused again. The store raises the contract's own error and the handler maps it to 507, but the store answering a request is not always built by the same bundle as the handler -- the persistence provider is reached from the route bundle and from instrumentation, which is why its state lives on a Symbol.for global -- and `instanceof` is false across that boundary while the declared code is still right. Classify on the code as well as the class, and make the code a permanent, persisted refusal in the browser: recorded locally so it survives a reload, shown as "storage is full", and refused by the retry entry point so a stale button cannot buy a second generation. Every other storage failure stays retryable. Convert what a pre-server-backed course still holds. Generated media is adopted under either key this application has used for it -- the placeholder, and the allocated id of a course converted once and later rolled back -- instead of only the first. Narration is converted by a load-time pass over the open course's speech actions, since nothing re-enters generation for an action that already has an id: bytes to the pool, id written back through a funnel that mirrors the media one, owner-only and server-backed-only. A line whose bytes are in no browser is left alone rather than re-synthesized. Also: the pool guard is now a positive `ast_` test rather than an enumeration of the shapes we mint (imports never reach the pool, so this is safe in both modes); the slide ref collection is an exported pure function so its four lease sites are covered behaviourally; ASSET_QUOTA_BYTES treats every spelling of zero as opting out and refuses a malformed value at startup instead of falling back; the abort signal is re-checked after the cache read, before an uncancellable commit; and the unused `removeAsset` and pool `replace` surfaces are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * chore(storage): release 0.29.1 The asset HTTP handler now recognises a store refusal by the contract code it declares as well as by its class, so a quota refusal raised in another module realm answers 507 instead of 500. Same contract, stricter recognition, no API change: a patch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a full store recoverable and adoption course-safe Narration adoption read the local audio row by its derived key alone. That key carries no stage id and the table is keyed by id alone, so two courses can mint the same one -- a PPTX import numbers its scenes and actions deterministically, which gives every imported deck's first slide `tts_s1_speech-scene-p1`. Locally a collision only means one course plays another's clip in one browser; adopting it wrote that clip into the shared document permanently, for every device and every visitor. A row that names a course is now adopted only into that course, and a row from before that column existed only when the text it recorded is the text of the action being converted. A full asset store was made permanent last round, which was wrong three times over: it overwrote the refused bytes with an empty blob -- on the conversion path that row is a course's only copy of its own media -- it kept sending the rest of the deck to a provider against a ceiling it already knew was reached, and it left no way back once an operator raised that ceiling. A full store is neither the content's fault nor the configuration's, so it is now its own case: the bytes are kept, the pass stops at the first refusal, and the element shows the reason together with a Retry that re-attempts the upload from those bytes. Nothing retries automatically, so no one is re-billed. The narration write-back now reaches the write boundary every producer of a durable write passes through, not only the dirty mark: adoption never deletes the derived row, so a snapshot that reverts the rewrite is adopted again on the next load and allocates a fresh asset every time. Adoption is also mounted by both classroom surfaces rather than one, takes the course's abort signal, and re-validates that this browser still has the course open before each write. ASSET_QUOTA_BYTES is validated from instrumentation, where the README and the docstring already claimed it was: its only other consumer is lazy and memoised, so a malformed ceiling let the process boot and then failed every persistence request, documents and runtime included. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): remember a full store per course, and never lose retained bytes A stopped pass left the elements it never reached as placeholders with no persisted record -- deliberately, since nothing was attempted for them. But that left the next load with no reason not to try: it called a provider for the next placeholder and was refused at exactly the same point, once per reload, indefinitely. A full store is not a property of any slide. It belongs to the deployment and changes for reasons the document knows nothing about, so it is now remembered once per course in the browser's device KV. A pass that finds the marker stands down before spending anything and leaves every placeholder its "storage is full" state and its Retry; the first upload that succeeds clears it and the next pass runs normally. Narration adoption latched per course so it runs once per load, and the latch outlived the abort that leaving a course performs. On a surface that stays mounted across switches -- the workbench pane is one component for every course it shows -- owner course A, visitor course B, then back to A skipped exactly the clips the abort had cut off, and nothing else converts them. The latch is released with the abort now, and a course adopts one run at a time so a re-entry cannot hand a clip a second allocation while the previous run's uncancellable tail is still settling. A quota-blocked element retried into a network error or a 500 lost the bytes that were kept for it: the retry deleted the row before attempting the upload and wrote no replacement for an error carrying no structured code, so the next retry went back to a provider for media this browser had a moment earlier. The row now survives until an upload succeeds, the failure handler keeps whatever bytes the attempt was given, and the retry asks the question a pass asks -- does this browser already hold bytes for this element -- rather than reading an error code that a second failure has already overwritten. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): adopt real legacy narration, queue re-entries, report attempt outcomes Narration adoption admitted a stage-less row only when the text it recorded matched the action being converted. Both of those columns were added to the local audio table by the very change that moved narration onto allocated ids, so a row still carrying a derived key has neither: the rule refused every real pre-allocation course and passed only on fixtures built from post-allocation rows. What the row cannot say, the key can. A derived key names two clips only when two courses share a scene order and an action id, and an action id repeats only when something other than the generator minted it -- an import numbers them by slide position. So a key built from a generated action id is adopted on that basis, a key an import could have reproduced still needs matching text, and a row that names another course is refused however unique its key looks. Handing a re-entering caller the adoption run already in flight undid the latch release it was paired with: that run is bound to the signal the departure just aborted, so it stops at its next clip while the caller -- which has the course open and a live signal -- is told the work is done, and an effect replayed as mount, cleanup, mount adopts nothing at all. A later caller now waits for the uncancellable tail and scans again, which costs a lookup on a course that has nothing left and finishes the clips the abort cut off on one that does. One attempt at an element now reports both facts its callers need instead of a bare boolean: whether the store refused it for room, and whether bytes actually reached the store. Leaving a course clears the task table, so a retry that landed afterwards read "no failed task" as success and deleted the row holding the only copy of the media. Nothing is inferred from that table any more. Reading the localStorage property can throw where storage is denied by policy, typeof included, so the availability check moved inside the guard: this metadata is best-effort, and a rejection here strands a generation pass that has already enqueued its tasks. A retry is never blocked by the per-course "store is full" marker, but a retry that is refused again re-sets it, and adoption now reads and writes the same marker rather than issuing one refused upload per clip on every load. The two canvas element renderers and both thumbnail renderers show the reason beside the Retry, so a full store does not look like an ordinary failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): probe a full store instead of standing down, and pair notices with a Retry Narration adoption was given both halves of the per-course "the store is full" marker last round: it stood down when the marker was set, and it set the marker when its own upload was refused for room. Those halves are only safe together if something can lift the marker, and for adoption nothing could. It has no affordance of its own, it stood down before reaching its own clear, the media pass returns before its marker gate when there is nothing to generate -- so a narration-only deck, or one whose slides are already satisfied, painted no storage-full element and offered no Retry -- and narration generated rather than adopted allocates directly rather than through the media commit. The course's cached narration was then lost for good, where before it converted on the first load after the ceiling was raised. The gate is a probe now. A marked course attempts exactly one clip per load: refused, it stops and the marker stands, which costs what standing down cost; stored, it lifts the marker and finishes the course. Adoption spends no provider money, so the whole cost of probing a store that is still full is one refused upload. Generated narration lifts the marker too. The three surfaces that gained a failure notice last round drew it for any failure with a reason, including the one refusal that is reachable without server-backed persistence, so a browser-only deck painted something it had not painted before. The notice is drawn beside a Retry and nowhere else, which is what it was added for and what leaves browser-only output unchanged. Both are now asserted through the render harness the surface matrix already had. A caller arriving while a rescan is queued shares it rather than appending another. One rescan converts whatever the run in flight left and every later one would find an allocated id on every action, so a chain bought nothing and turned a single stalled upload into a course that never adopts again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): treat a refusal for room as a fact about one clip, not the deck The asset store checks each write against the headroom it has left, so a store that refuses a long opening clip can still hold every short clip behind it. Narration adoption assumed the opposite: it broke the deck at the first refusal and then re-attempted that same first clip on every later load, because the document names it first. A deck whose longest clip exceeds current headroom therefore never converted the clips that would have fit, with no affordance to recover it -- the state the probe was introduced to remove, reached through a narrower door. An unmarked load now attempts every clip, skipping the ones that do not fit, and remembers the condition only if the load ends with clips it still could not store. A marked load spends its single upload on the smallest clip left rather than the first one named: that is the clip that answers the question the marker asks, because if the smallest does not fit nothing does. The media pass keeps stopping at its first refusal, and for a reason adoption does not share -- every element it attempts costs a provider call. A rescan several callers share took the newest caller's signal, and the newest caller is not necessarily the one still there: a surface that opened a course and closed it again would stop work a surface still showing that course was waiting for, and that surface is latched, so it would never ask again. The shared run now takes a signal that is aborted only once every caller has left. The comment claiming the shared rescan contains a stalled upload was wrong -- the rescan is chained off the run in flight, so a stalled upload leaves every caller pending exactly as a chain would. It claims the bounded queue it actually provides, and the stall is recorded as a limitation. The failed-state containers took their stacking classes unconditionally, so markup differed in browser-only mode even though nothing moved on screen. Those classes are applied only when there is a notice to stack, and the tests assert the exact class attribute rather than a substring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): stop narration adoption writing the media pass's store-full marker The marker means "do not call a provider for this course". A path is entitled to write it only if its own refusal cost a provider call, and narration adoption's refusals cost nothing: it uploads bytes this browser already holds. The store also checks each write against the headroom it has left, so a clip that does not fit says nothing about whether a slide's image would. Adoption was writing it anyway, and one over-long narration clip was therefore enough to stand a course's entire image pass down on every later load -- on a store that had just accepted adoption's other clips. The author could still recover each element by hand, every load, for ever. Three rounds of narrowing this seam produced a finding each time, so it is removed rather than narrowed again. Gone: the marker read, the single-clip probe, the smallest-clip selection, and the up-front read of every row into an array -- which also retires a sampled-then-stale flag and the retention of a whole deck's blobs for the length of a run, and returns the loop to streaming one row at a time. Adoption's rule is now that every load attempts every clip it holds, once; any failure skips that clip and the load continues. The noise the coupling was meant to avoid does not arise, because after the first load the clips still outstanding are exactly the ones that did not fit -- normally none, or one. A successful write still clears the marker, and that is a different kind of statement: a write that went through is a fact this run established, where a refusal is an inference about what some other write would cost. For a course whose media needs nothing, adoption and generated narration are also the only paths that can establish it. The failure module still documented the deck-wide premise this contradicts. It now says what is true: the check is per write, and the media pass stops the deck as a judgement about cost rather than about certainty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bound a full store's cost from the store's own arithmetic Removing the store-full marker from narration adoption removed its bound too, and the code then asserted the bound was unnecessary. It is, on a store with room for most of a deck. On the store the whole mechanism exists for -- the ceiling reached, nothing fitting -- the outstanding set after every load is the entire deck, so a thirty-clip course posted thirty full blobs on every load, indefinitely. Each of those is not a cheap refusal: the bytes are uploaded, the server hashes the whole payload, and only then takes a per-principal lock and sums every entry that principal owns before saying no. The bound needs no flag, no key and nothing carried between loads. The store asks whether `used + addedBytes` exceeds the ceiling, and `used` only grows while a run is uploading, so a clip refused for want of room implies every clip at least that large is refused for the rest of that run. The run keeps the smallest size it has been refused and skips anything no smaller without uploading it; a smaller clip is still attempted, because it may fit. A deck the store refuses entirely now costs one upload per successive size minimum instead of one per clip, and a deck it has room for costs nothing extra, because nothing is refused. Only a refusal for room lowers the bar: a dropped connection says nothing about how much room there is. The deck-wide certainty premise the failure module retracted last round still stood verbatim at the site that implements the stand-down. Both copies now say the same thing: the check is per write, and the pass stops the deck as a judgement about cost rather than about certainty. The comment on adoption's marker clear now names its price. Narration of a few hundred bytes fits in headroom an image does not, so a proven write can let the next pass buy one more image that is refused again -- bounded at one, and the price of the alternative being a course whose media never generates again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(media): state the adoption bound exactly, and stop three comments describing the old rule The comment introducing the in-load bound gave its cost as "at most a handful, and the first load pays the most". Neither clause is a property of the rule. A clip is skipped only when something no larger was already refused, so a fully-refused deck costs one upload per successive size minimum in document order: one when the clips grow, about ln N for an arbitrary order, and one per clip when they only shrink -- a long opener followed by terser lines is exactly that shape. And no load is cheaper than the first, because the bound resets per run and a refused clip stays outstanding. The comment now says that, and points at what would make it exactly one for any ordering: the store returning its remaining headroom in the refusal's existing details channel, which the server leaves empty today. Two other comments still described the previous rule -- "attempts every clip it holds, every load" -- one of them twenty lines above the paragraph that introduces the bound, in the same block. Both now say what the code does. The bound's soundness is worth stating where a maintainer will look for it: quota is charged at full length with no discount for a duplicate, the sum it is checked against joins entries to blobs so the collector cannot lower it, the check takes a per-principal lock before summing, and replace and delete are refused to every browser. Nothing a run can do makes room appear inside it. One test installed a row implementation and replaced it wholesale a few lines later, so the first was dead and the survivor dropped the text the first clip's import-shaped key needs for the ownership rule -- it passed on the coincidence that the fixture's default text is the action's. Merged into one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): keep a Retry from re-buying parked media, and state the store seam once Five findings from an inline review. The standalone classroom route asked the ownership sidecar once per load and recorded only stage ownership when that ask failed. Every non-answer fails closed, so one transient 5xx left the genuine author with no resume, no Retry affordance and no legacy narration converted for the rest of the load, with nothing to change it short of a reload. The failure now records the fail-closed answer explicitly -- an answer an earlier load established must not outlive the failure that replaced it -- and an unresolved answer is asked for again, a few times over a few seconds. A real answer, however unwelcome, is final. A Retry could pay a provider for media the pool already held. When the bytes are stored and only the write-back fails in a way that keeps the allocation, it is parked and no local row exists, because that row is written only after a successful write-back. Retry now reads the parked queue exactly as the pass does and re-attempts the write-back: it re-keys the task done when the document takes it, leaves the entry parked when the slide still does not exist, and stays failed and retryable when the document refuses again. Object URLs a parked allocation owns are revoked when the entry is dropped. The commit path leaves them alone while the entry is parked, because it is then the only thing holding bytes this tab can render, so a course switch or a stage deletion was pinning the whole blob for the life of the tab. An entry a slide has already taken is left alone: the task table is displaying those URLs. The fallback lookup for cached bytes is a stage-scoped scan, and the keyed lookup misses for every row the commit path writes, so a pass was materializing and sorting the course's whole media table once per element. One scan per pass now, built on the first miss. It is sound and not merely cheaper: an element asks only for its own placeholder, and every row a pass writes carries the placeholder of the element that wrote it. "The store accepted a write, so it is not out of room" was enforced at three call sites under slightly different conditions, which made it a convention the next pool write path could silently break. It is stated once, in putAsset, for the course whose bytes it just stored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 16 天前 | |
release: OpenMAIC 1.0.0 — the agent workbench (#1228) * feat(storage): add an agent-session store with PG backend and layered contracts (#1163) * feat(storage): add agent-session store with PG backend and layered contracts * test(storage): avoid BigInt literals for pre-ES2020 root typecheck * fix(storage): close agent-session store review findings * docs(storage): align hook ordering and contention-probe claims with the code * ci: run on the agent-workbench integration branch * chore(storage): bump to 0.5.0 for the agent-session store * fix(storage): carry replay compaction across page boundaries * feat(agent): add the driver model contract and stage route dialect (#1165) * feat(agent): add the driver model contract and stage route dialect * fix(agent): validate route context windows and clarify dialect precedence * feat(agent): adapt the agent-session store and runtime foundations (#1167) * feat(agent): adapt the agent-session store and runtime foundations * feat(agent): resolve request owner identity via an anonymous cookie * docs(agent): document the opt-in compaction default and harden edge cases * feat(agent): add the background session runner (#1169) * feat(agent): add the background session runner * feat(agent): wire the runner into startup behind feature flags * fix(agent): stop clean interruptions from consuming the attempt budget * fix(storage): charge the attempt budget for abandoned leases but not clean parks * docs(storage): document the attempt-charging contract and decouple its tests * feat(agent): add agent session and owner event streams (#1170) * feat(agent): add agent session and owner event streams * fix(agent): close the session-existence oracle and document the owner seam * feat(agent): add agent session lifecycle routes (#1171) * feat(agent): add agent session lifecycle routes * fix(agent): validate session-create input and preserve the owner cookie on errors * refactor(storage): drop the unused active-stage API from the agent-session contract (#1174) * refactor(storage): drop the unused active-stage API from the agent-session contract Tools address stages explicitly on every call, so the store keeps no mutable session-level stage pointer. Removes resolveActiveStage and setActiveStage from the store interface, their PG implementations, the active_stage_changed lifecycle event, the session_active_stage owner event variant, and the contract tests pinning them. The active_stage_id column and the DDL check constraint stay untouched for schema compatibility. * chore(storage): bump @openmaic/storage to 0.7.0 for the contract removal * docs: document the agent runtime configuration surface (#1176) * fix(agent): repair orphaned and late tool results across interruption boundaries (#1180) * fix(agent): repair orphaned and late tool results across interruption boundaries A crash, shutdown, or provider failure can leave the durable transcript with tool calls that have no result, or with results ordered illegally for the provider. Three failure modes were fixed: - Orphaned tool calls: a run that died between an assistant tool-call frame and its result left a dangling call in the entry tree. Resume no longer synthesizes and persists receipts for it: interrupted results are a read-time provider view owned by a shared read-boundary repair, which returns the original array for a healthy transcript and never mutates the tree. - Late parallel results: a parallel tool can finish while pi unwinds an aborted assistant frame, leaving result(A), assistant(aborted), result(B) in durable order. Strict providers reject non-contiguous results, so the read-boundary repair moves existing results next to their owning assistant frame (in call order), omits incomplete unwind frames, and synthesizes receipts only for genuinely missing calls. - Interrupted calls at the write boundary: a call still in flight when the run winds down (shutdown, lease loss, cancellation, provider failure) had no receipt at all. The runner now tracks in-flight calls from their assistant frames and, before the terminal flush, appends an interrupted-result receipt for each still-orphaned call through the same attempt-fenced write chain, so a lease-stealing zombie never writes and the next claim sees a provider-safe transcript. * test(agent): pin the runner wiring for interruption-boundary tool repair * feat(agent): add neutral tool foundation libraries (#1184) * feat(agent): register a web_search tool on the session runner (#1185) * feat(storage): add a per-session URL trust gate (#1186) * feat(agent): add the skills system (#1189) * feat(agent): add the skills system (builtin directories and durable user skills) * fix(storage): serialize the user-skill quota check-and-insert per owner Two concurrent creates at the 50-skill boundary both counted 49 rows and both inserted (READ COMMITTED, no lock), overshooting the quota contract. The create transaction now takes a per-owner pg_advisory_xact_lock first, and the same-name idempotency check runs before the count check so an at-least-once retry of the create that committed as the owner's 50th row still returns its durable receipt instead of a quota error. The 23505 backstop is retained for writes that do not take the lock. * fix(agent): share unstorable-character validation and align skill lookup * feat(agent): add session materials and a fetch_url tool behind the URL trust gate (#1190) * feat(agent): add session materials and a fetch_url tool behind the URL trust gate * fix(agent): harden session material fetching * feat(storage): add an ownership scope to stage documents (#1191) * feat(agent): add material read and search tools (#1192) * feat(agent): add stage read and patch tools (#1194) * feat(agent): add page generation and deck editing tools (#1198) * test(storage): keep the PG contract suite order-independent (#1200) * fix(agent): revoke deleted-session URL authority and reject private ISATAP endpoints (#1199) * fix(storage): revoke deleted session URL authority * fix(ssrf): reject private ISATAP endpoints in strict fetches * chore(storage): bump to 0.11.1 for the session-URL authority fix * feat(agent): add roster and voice registration tools (#1201) * feat(agent): add folder organisation tools (#1202) * feat(api): add stage and material HTTP routes (#1203) * feat(workbench): add the client data layer (#1204) * feat(workbench): add the client data layer * docs(workbench): write the ported comments in English * chore(edit): remove the in-editor agent panel (#1210) * chore(edit): remove the in-editor agent panel * style: apply prettier formatting * fix(agent): report the runtime as unusable without a database (#1207) * fix(agent): report the runtime as unusable without a database * style: apply prettier formatting * feat(agent): add image, video and pptx import tools (#1211) * feat(workbench): add the agent chat surface (#1205) * feat(workbench): add the agent chat surface * docs(workbench): write the ported comments in English * fix(workbench): label the folder and rename tools on the timeline * fix(workbench): label the roster and voice tools on the timeline The reconciliation test iterates every tool the runner registers and requires a display label of its own. The roster and voice-clone tools (list_voices, set_roster, clip_audio, register_voice) reached the integration base with the roster/voice-registration tools but never gained presentation rows, so they fell through to the default branch and rendered their wire names. Port their rows from the reference implementation (labels and i18n keys verbatim) and extend the reconciliation allowlist with ROSTER_TOOL_NAMES and VOICE_CLONE_TOOL_NAMES, so a future tool cannot enter the product without a label. * feat(agent): add the material extraction lifecycle (#1212) * feat(storage): add material extraction lifecycle * feat(agent): execute queued material extraction * style: apply prettier formatting * style: satisfy prefer-const in the extraction runner * test: give material fixtures the extraction lifecycle fields The media-tools slice and the extraction lifecycle slice were each green in isolation but never compiled together: the lifecycle made derivedFrom and extraction required on AgentSessionMaterial while the media-tool fixtures predate them. * chore: remove stray task notes * fix(workbench): label the extraction lifecycle tools on the timeline * feat(workbench): add the workspace shell (#1206) * feat(workbench): add the workspace shell * docs(workbench): write the ported comments in English * i18n(workbench): align workspace keys across locales * fix(workbench): adopt the landed data layer and label the extraction tools - replace the sibling-slice seam stubs with the real data-layer modules - drop ambient declarations now shadowed by landed files - port timeline labels for the extraction lifecycle tools from the reference - align the new i18n keys across all locales * ci: retrigger * feat(api): folder routes, stage-meta viewer surfaces, and the material upload contract (#1215) * fix(storage): restore capability-based stage access * fix(api): bind document access to request owner * fix(agent): restore three-state stage access on the tool layer Port probeStageAccess and the three-state StageAccess (owned / foreign / missing / tombstoned) and gate every stageId-bearing stage tool on an owned probe, mirroring the reference per tool: - move_to_folder, rename_stage, read_stage_outline refuse a non-owned stage with the single not-yours message before touching the store. - The course/DSL toolset and the roster toolset are wrapped by withOwnerStageAuthorization: read_stage, patch_stage, grep_stage and every writer refuse a foreign stage with the same message and refusal shape. - Scene preview keeps its own probe and its own refusal text, and is registered beside the course toolset (never double-gated). - The runner injects one probe factory at the three call sites. Tests: the dsl cross-owner test premise (a foreign stage is readable by id) encoded an invented capability-read policy that the reference does not have at the tool layer; it now asserts foreign read/patch/grep are all refused while the owner still reads. Curriculum cross-owner assertions were already the reference's and now pass with the probes in place. * docs: correct per-file test counts in the fidelity report * test: fix type errors in stage-access fidelity test * test: adapt media-tool and gate suites to the owner-scoped store seam * feat(api): add owner-scoped course-folder HTTP routes Port the reference implementation's /api/folders family (list, create, rename, delete with ungroup/remove modes, and folder membership) onto the owner-bound document store, replacing its provider-based auth with the existing withRequestOwnerId / owner-scoped store seams. The storage package's folder store grows the pieces the routes need: DocumentFolder.order (schema column + max+1 assignment + ordering), renameFolder, deleteFolder(mode) with captured member ids, and setStageFolder(stageId, folderId | null) with idempotent un-filing. FolderNameError moves into folder-name-validation.ts (stage-storage re-exports it, keeping import sites intact). Every route gates on the configured agent runtime (plain 404 when off or unconfigured), keeps the reference's machine codes and envelopes, and is covered by gate tests plus a behavior suite. * feat(api): add stage-meta viewer surfaces for the classroom Port the reference implementation's viewer-facing stage state — can-edit / collected / published / generation-complete — on top of the stage-access base (stage_meta + tombstones). stage_meta gains published_at and generation_complete columns plus a stage_bookmarks table; the reference's deployment-specific origin/claimed_at columns are stripped. New gated routes: GET /api/stage-meta/[stageId] (per-viewer facts, 404 for absent/tombstoned, never returns the owner id), GET /api/stages/[id]/status, POST generation-complete / publish / unpublish (owner-only), POST /api/bookmarks. The resolver lives in lib/server/stage-access.ts. Wiring: a fetchStageMeta client with the reference's three-outcome contract, stage-store isOwner/isBookmarked/readOnly fields (upstream single-user defaults, no-op until the sidecar answers) plus setViewerAccess, the classroom apply path computing readOnly = !(isOwner || isBookmarked), the Stage editability gate, and a sidecar probe after each classroom load. A sidecar 'absent' answer keeps the editable default here because the classroom also serves local-only courses; server writes stay owner-enforced. * feat(api): port the reference material upload contract Rewrite POST /api/materials to the reference implementation's upload shape so the workbench uploader (uploadWorkbenchMaterial, which posts no session id and expects a flat 201 view) works unchanged: owner-scoped upload with mime normalization/validation (415), per-class size caps checked on the declared content-length and the streamed body (413), empty body (400), quota (429), sha256 reserve->store->finalize lifecycle with abandon on failure, flat { materialId, originalName, bytes, mime, extraction } 201, and an x-request-id echo. Adds the owner-scoped material library (owner_material table + quota + 24h lazy sweep, bytes in the host's asset registry as the neutral replacement for the reference's object-storage byte path) and the material cap configuration. The session-scoped GET list is left as-is; the reference's owner-material extraction worker is not ported (the branch's session-material extraction lifecycle already covers extraction). Gate tests now cover all 23 persistence routes across the three runtime env states; the materials behavior suite pins the new contract. * feat(media): add an optional local ffmpeg media extractor (#1213) Adds a local ffmpeg/ffprobe pipeline as a second media extraction provider behind the extractor registry, ported faithfully from the reference implementation: duration probing, keyframe-safe chunking, per-chunk ASR with timeout and deadline budgets, and timestamped transcript assembly. - Availability probing feeds the registry's candidate selection: the provider simply is not a candidate when ffmpeg/ffprobe are absent. - With neither ffmpeg nor a cloud provider configured, extraction fails with an actionable message naming both enablement paths. - Media materials route through the same extraction lifecycle and lease fence as documents; no parallel queue. - Tests inject the executable resolver so the missing-ffmpeg path is the default-tested one; the real pipeline test is skip-if-unavailable. - @openmaic/storage 0.13.0 -> 0.14.0 (media routing in the material lifecycle surface). * feat(storage): per-scene monotonic revisions via database triggers (#1214) * feat(storage): per-scene monotonic revisions via database triggers Restore the reference implementation's freshness granularity: a per-scene monotonic revision maintained by database triggers, so every writer (HTTP routes, agent tools, jobs, manual SQL) bumps it without application cooperation. - Companion revision tables + trigger functions in the storage package's idempotent schema bootstrap, with the lock-order invariant, pg_notify wakeup and the suppression switch for batch writers. - ensureDocumentSchema gained a dollar-quote-aware statement splitter. - The freshness and manifest routes serve per-scene revisions. - Mutation-verified: dropping the triggers turns the revision tests red. - @openmaic/storage 0.13.0 -> 0.14.0. * fix: forward the freshness manifest through the owner-bound store * feat(workbench): add the Pro entry points and preserve the mode-transition semantics (#1208) * feat(workbench): add the Pro entry points * feat(workbench): preserve Pro mode transition semantics * fix(workbench): drop ambient declarations shadowed by landed slices * fix(workbench): drop ambient declarations shadowed by the landed shell * feat: port workspace shell sibling modules Port the 16 leaf modules the Pro workspace shell imports but that were only ambient-declared, replacing the compile-time bridge with real implementations adapted from the sibling-slice reference: pure workbench helpers (session title, rail tab, course-chat bootstrap, created-course tabs, course-tabs memory, workspace navigation, pane navigation, pro-edit sizing, existing-course minting, first-message session), the neutral brand context and course-rename server API, the server-action session delete, the home discovery hook, the classroom pane host with its load-policy leaf, the theme toggle and floating-layer owner, plus the floating-layer-owner wiring the dialog/dropdown/tooltip portals stamp. Also add the workbench-shell locale copy for all 12 locales, port the reference tests for the ported modules, and drop types/workbench-sibling-slices.d.ts now that every declaration has a real implementation. * docs: keep ported comments in English and deployment-neutral * docs: announce 1.0.0 and refresh the feature overview (#1216) * docs: announce 1.0.0 and refresh the feature overview * docs: finalize 1.0.0 README after feature merge * fix(agent): control-plane routes answer 404, not 500, without a database The agent control-plane routes gated only on the runtime flag, so an enabled-but-unconfigured deployment (flag on, DATABASE_URL empty) answered 500 from a store that cannot connect. Gate them on the configured check instead, matching the stage/material routes: the whole surface is cleanly absent until both the flag and the database are present. The status probe keeps reporting both bits. * test: mock both runtime gate exports in the control-plane route suites * fix(agent): abort in-flight TTS on cancel and bound each provider request with a timeout (#1217) The generate_tts / scene-tts path checked the runner's AbortSignal between actions but never created the provider HTTP requests with it, so a session cancel left a hung synthesis fetch in flight until a restart repaired the tool result. Thread the signal end-to-end: TTSModelConfig carries an optional signal, generateTTS combines it with a per-request timeout (TTS_REQUEST_TIMEOUT_MS, default 30s, ported from the reference runtime's TTS bounds) via AbortSignal.any, and every provider fetch (openai, azure, glm, qwen incl. voice-clone + audio download, voxcpm, minimax, doubao, elevenlabs, lemonade) is created with that signal. A timeout now fails the tool call with TTSRequestTimeoutError (a clear retryable error) instead of wedging the session; a caller cancel propagates as the interruption so the runner settles the session as cancelled without a restart. Tests: hung-provider simulation rejects at the timeout with the retryable error; abort mid-flight aborts the captured request signal and surfaces the interrupted shape; removing the signal wiring makes the abort tests fail (red), restoring them turns green. * fix(workbench): PG-mode home listing via owner stages; keep the interrupted terminal course card (#1218) Finding 1: with server persistence on, listStages resolved to the generic GET /api/persistence/documents listing, which the capability model deliberately answers 403 FORBIDDEN_DOCUMENTS for (reads by id, listings owner-only). The home/workspace library now lists through the owner-scoped GET /api/stages surface (same anonymous-owner cookie the workbench uses) when server persistence is enabled; the server-side 403 is untouched. Finding 2: a run interrupted (session_interrupted) and repaired (session_resumed) that ends cancelled before agent_end stranded its pending classroom sightings, so the timeline's terminal card lost the course the answer produced. session_end (cancelled) now flushes the pending sightings into the same course card set agent_end paints, before the stopped caption. * chore(workbench): remove the bookmark concept and the saved-courses drawer (#1219) * chore(classroom): remove the bookmark ('collected') concept entirely The stage-meta viewer port introduced a bookmark surface (stage_bookmarks table, POST /api/bookmarks, the isBookmarked sidecar field, and a readOnly rule that let a saved course stay editable). The product has no such concept, so remove it as a closure: - delete the /api/bookmarks route and the stage_bookmarks table plus its query helpers from the persistence bootstrap - drop isBookmarked from GET /api/stage-meta/[stageId] - simplify the classroom read-only rule to readOnly = !isOwner across the sidecar client, ownership signal, classroom load, stage store and the classroom page - keep publish/unpublish, generation-complete, isOwner and isPublic exactly as they were - update the gate and stage-meta route suites and the README mentions The workspace rail's Bookmark glyphs and comments describe the upstream saved-courses (favorites) section, which is driven by isOwner and renders no collect affordance; they are kept as unrelated homonyms. * chore(workbench): remove the saved-courses drawer UI The first pass removed the bookmark data model but kept the rail's "Saved courses" drawer, judging it a separate surface driven by `isOwner === false`. The home/workspace listing is owner-scoped, so that flag can never occur: `allSaved` is permanently empty and the drawer (plus the collapsed-rail Bookmark mini-button) is a dead affordance. Remove it: the SavedDrawer component and its mount, the savedOpen / savedSection state, the allSaved / matchedSaved derivations, the 'saved' variant of the course-list renderers, the mini Bookmark glyph, the drawer-only CSS, and the drawer's i18n keys from all 12 locales. The courses tab is now exactly one folders tree. The authored/favorites split in workspace-tree.ts goes with it; the tree module no longer reads `isOwner`. The discovery course type keeps the field — the shell still reads it for read-only gating. Upstream has no collect concept; the drawer could only ever render empty here. The reference implementation HAS this drawer (its favorites come from its account system), so this removal is a deliberate upstream product decision, not a fidelity bug. * fix(workbench): restore the attach entry, add the rail settings entry, pin all three entry points (#1221) * fix(workbench): restore the composer attach entry by gating it on the live runtime The AttachButton's rollout probe read a `materialsEnabled` field that this branch's /api/agent/runtime never answers (the materials routes gate on the runtime itself, like the stages), so the gate could never pass and the attach button never rendered — the Pro launch and chat composers showed only the @-mention and enhance glyphs. Substitute the field with the runtime's `enabled` value, which IS the upload action's precondition: POST /api/materials answers 404 whenever it is false, so the render condition now equals the action precondition (no dead button). The button's label (`proMode.attach`) is a user-visible string that becomes visible again; port the reference implementation's own translations verbatim into the 11 locales that still carried the Chinese copy. * feat(workbench): add the settings entry to the rail's bottom-left cluster The reference's rail foot carries a cluster of utilities (its saved-courses drawer, the language switcher, the display toggle). This branch removed the drawer — it could only ever render empty here — and the product decision is to fill that freed spot with the settings entry. Add a settings trigger to the foot cluster (expanded rail, beside the language and display toggles, and on the collapsed strip) and mount the model/provider SettingsDialog in the rail, wired to the trigger. It is the same dialog the classic home opens from its header pill; the workspace had no settings entry of its own, so nothing is duplicated within a surface. * test(workbench): pin the restored upload, attach, and settings entry points Covers the three restored entry points: - the courses-tab upload control: rendered beside the course name filter, wired to the discovery hook's ZIP import trigger, disabled while an import runs, and gated by the same condition as its action (the courses tab); - the composer attach control: an actual render of AttachButton under both probe answers (visible when the runtime says the upload path is live, hidden otherwise), its mounts in the launch and chat composers, the branch's runtime-field substitution in the probe, and the reference's own `proMode.attach` copy in all 12 locales; - the settings entry: the trigger in the rail's foot cluster (expanded and collapsed), beside the language and display toggles, opening the SettingsDialog the rail mounts. * chore(config): the Pro workbench flag implies the MAIC Editor gate (#1223) A workbench build without the editor toggle has no way to edit a course: enabling NEXT_PUBLIC_PRO_WORKBENCH_ENABLED while forgetting NEXT_PUBLIC_MAIC_EDITOR_ENABLED produced exactly that split-brain bundle. The workbench IS Pro mode, so its flag now implies the editor gate; the standalone flag remains for deployments that want the classroom editor without the workbench. Documents both flags in .env.example. * fix(agent): wake SSE tails and the runner on durable deltas (streaming fidelity) (#1222) The Pro workbench chat did not stream: the session/owner SSE routes polled the durable event log on a 5s/30s clock with no wakeup, so message_update deltas (written at 150ms cadence) reached the browser in poll-sized blocks and the thinking strip only mounted after the whole reasoning text had accumulated. Port the reference's LISTEN/NOTIFY delta path: - storage: add in-transaction wake hooks (onSessionEventAppended, onOwnerEventAppended, onCancelRequested) so a host queues pg_notify in the same transaction as the durable append; align readEventsAfterForReplay to rank the bounded page so the first delta after the cursor is always kept (the live tail can never starve). Bump @openmaic/storage to 0.18.0. - app: port the process-wide event-notify bus (dedicated LISTEN client, self-check probe, reconnect backoff; notify through the storage transaction surface), wire the store hooks, subscribe both SSE routes before the initial read with the reference's initializing gate, and give the runner one {kind:'session'} subscription whose wake runs the cancel check and the message drain. Polls stay as the lossy-NOTIFY backstop. - lifecycle: start/stop the bus from instrumentation. Tests: storage hook + compaction contract; route wakeup latency; runner wakeup wiring with a fake agent; bus unit tests; PG contracts proving a real append wakes the routes and a live SSE route forwards a message_update on the wakeup, and that a rolled-back append never wakes. Also fix the pre-existing park-attempt-budget PG test TRUNCATE (missing CASCADE against newer FK tables). * fix(storage): asset writes self-deadlocked against pooled PostgreSQL (#1225) * fix(storage): refuse the non-transactional byte-write deadlock configuration A byte store whose plain write() runs on its own pooled connection cannot be invoked from inside a registry write transaction: after the transaction has claimed the blob-row lock, that write blocks on the lock the transaction just took while the transaction waits on the write - a self-deadlock PostgreSQL cannot detect (one side is idle in transaction). There is no lock-safe ordering for such a writer: bytes must be written after the row claim (writing before it lets the collector delete the bytes while the upsert waits), and any second-connection write after the claim is the deadlock. The configuration is therefore detected and refused: - AssetByteStore gains writesOutsideRegistryDatabase?: true, declaring that the layer's plain byte operations cannot contend for the registry's row locks. - PgAssetStore refuses put()/replace() up front (and defends coordinatedWrite) when the byte store has no writeWith and does not declare the flag, throwing a clear configuration error before any row is claimed. - The collector mirrors the guard on its delete path (deleteWith or a declared out-of-registry layer, else a configuration error). - The object store declares the flag (its out-of-transaction write remains legitimate); the in-registry PostgreSQL byte column provides writeWith / deleteWith instead. - Write transactions (put/replace/remove) set SET LOCAL lock_timeout = 30s so any future lock-contention variant fails loudly instead of hanging. Bumps @openmaic/storage to 0.18.0. * fix(persistence): forward the transactional byte methods through the lazy asset byte-store wrapper The no-bucket case of lazyAssetByteStore returned a bare { write, read, delete } and dropped writeWith/readWith even though the underlying PgAssetByteStore has them. The registry's hasTransactionalWriter duck check then failed and put() fell back to the byte store's own pooled connection, which blocks forever on the blob-row lock the registry transaction just took when the bytes live in the same PostgreSQL - the production self-deadlock. The no-bucket layer is statically PgAssetByteStore, so its transaction-pinned methods are forwarded eagerly (typed against the real signatures via PgForwardedByteStore). The bucket case keeps its lazy-probing semantics: no transactional writer exists there, the signed-URL method stays absent or lazy exactly as documented, and the wrapper now declares writesOutsideRegistryDatabase so the registry may run the plain write inside its transaction. New tests pin the wrapper's transactional capability red-to-green and assert put()/resolve() route byte traffic through the transaction-pinned queryable. * fix(home): cap the generate-prep ingest drain at 3s so Generate never waits the full server budget The classic home flow's Generate click drained in-flight ingests for the full 15s server budget. Cap the wait at GENERATE_DRAIN_CAP_MS (3000ms, documented as a UX bound) and reuse the existing timeout fallback: sources that miss the cap proceed on the legacy byte path and each late-resolving id is released. * chore(storage): bump to 0.19.0 over the concurrently landed 0.18.0 * fix(agent): bound every tool call with a timeout; never resurrect a cancelled session (#1226) * fix(agent): bound every tool call with a global timeout and settle it on cancel A tool await that neither resolves nor rejects wedges the session forever: the lease keeps heartbeating and the driver never reaches its next cancel checkpoint. Race every tool execution (in buildAgent) against a hard budget (OPENMAIC_AGENT_TOOL_TIMEOUT_MS, default 10 min, per-tool overrides for known long runners) and against the caller's AbortSignal, so even a signal-ignoring await cannot keep a cancelled session running. On timeout the call rejects with AgentToolTimeoutError; the agent loop turns the rejection into a structured error tool-result the agent can retry or proceed from, and the abort signal is delivered to the tool's in-flight work through a derived controller. Zombie-tool updates after settlement are dropped. * fix(storage): never re-lease a cancel-requested session; settle it as cancelled on claim The claim scan treated a session with cancel_requested_at set as a normal claim candidate: after a restart it re-leased the same session for attempt N+1 and resumed generating despite the pending cancel. claimNextSession now settles such candidates as cancelled under the claim lock (status cancelled, attempt reset, lease and cancel request cleared, terminal session_end event and owner projection) instead of leasing them, then keeps scanning. Bump @openmaic/storage to 0.18.0. * docs: takeaway-style 1.0.0 announcement with bilingual guide links The 1.0.0 head is now a short takeaway block — badge links to the official user guides (English and Chinese), five one-line highlights, and pointers into Features and the workbench setup section — instead of six dense paragraphs. The detailed provider-neutrality and freshness notes move into the Features workbench section, phrased database- neutrally (the announcement no longer names a specific database). Release date corrected to August 27. * fix(workbench): restore editor chrome, mode transition, streaming, materials, mentions, folders (#1229) * fix(workbench): wire workspace folder routes * fix(editor): restore reference workbench chrome * fix(workbench): persist composer materials and course refs * fix(workbench): preserve live reasoning frames * fix(persistence): back off failed streaming saves * chore(workbench): retire stale slice seams * test(editor): cover element pin layer * chore(storage): bump to 0.21.0 for the user-message ref/material fields * chore(editor): translate ported code comments to English * fix(agent): fence durable tool writes and consume cancel requests atomically (#1230) * fix(agent): enforce provider force-off in agent tools and scrub vendor identity from tool results (#1231) * fix(materials): serialize per-owner quota reservations and make crashed uploads reclaimable (#1232) * fix(editor): resolve dock-bar i18n keys, remove dock height drag, wire element referencing (#1233) * fix(workbench): send the opening session message exactly once with refs intact (#1234) * feat(editor): port timeline TTS preview single-flight and voice-all state latching (#1235) * fix(media): restore the reference classic media chain (#1236) * fix(import): adapt imported PPTX canvas size so decks render without overflow (#1237) * fix(editor): complete element referencing — renderer DOM contract and GenUI picking aligned with the reference (#1238) * test(providers): reconcile the provider-config vendor-token debt count after the main merge The integration line's AK/SK fallback for the managed document provider adds occurrences that main's allowlist snapshot predates. Same mixed-composition debt category the group already documents; no new vendor behavior. * test(providers): reconcile vendor-token debt counts with the integration line The main-merge brought main's neutrality-guard snapshot next to integration features it predates (media-extractor fallback chain, local voice-profile deletion semantics, the enabled-TTS helper). Same debt categories the guard already documents; counts updated to the guard's own tally and two grouped entries added. No new vendor behavior. * fix(agent): carry reasoning through the completions dialect so the thinking strip renders (#1239) * feat(skills): add Feynman and spiral curriculum methods (#1240) * feat(agent): port missing reference tools and skills (parity audit) (#1241) * feat(media): retire asset-registry wiring; media and materials follow the reference byte model (#1242) * fix(classroom): center adapted canvases in the stage and send back navigation home during generation (#1243) * feat(settings): skill management with real list, download, delete, and upload (#1244) * feat(settings): skill management section with real list, detail, and zip download * feat(skills): owner skill delete and upload across storage, API, and settings * fixup! feat(settings): skill management section with real list, detail, and zip download chore: neutralize a reference note in the settings header comment * fix(media): persist origin-independent classroom-media references from the agent runtime (#1245) * feat(editor): float the insert toolbar in the outer frame with collapse (#1246) The insert strip was bounded to the slide card, so it could only ever sit on top of slide content: the card's overflow clipped it and it could not be parked in the padding beside the slide. Move it into the studio frame the element picker's panel already roams (CanvasOverlayPortal + the frame selector), so both canvas overlays share one bounding container and their handles behave the same. While picking, the strip rises over the picker and goes inert, which is the z-order CANVAS_OVERLAY_Z already documents. Add a fold beside the grip: the chevron collapses the strip to that grip row and back, with the buttons unmounted rather than hidden. The fold is session-local state owned by EditShell, next to the drag offset, so a surface swap keeps it; nothing is persisted. Expanding a strip parked at the bottom edge re-clamps through the same bounds rule the keyboard move uses. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * fix(workbench): align the chat timeline's left edge with the composer (#1247) * fix(agent): fence session claims while an ask_user question is outstanding (#1248) * fix(agent): settle-time rescue tracks real delivery instead of a count offset (#1249) * fix(persistence): migrate owner_material to oss_key and drop legacy asset_id (#1250) * docs(readme): surface the 1.0.0 user guide badges at the top (#1253) * fix(workbench): show newly created folders in the sidebar without reload (#1254) * docs(readme): add the release version prefix and drop the opt-in framing * fix(workbench): single-source the chat gutter so timeline and composer share a left edge (#1255) The transcript and the composer each established their own column: their own `px-*` gutter and their own `mx-auto w-full max-w-*` centering wrapper. Equal padding values were never enough, because the two columns are centered inside different containing blocks — the transcript's is a scroll container, whose content box is narrower than the composer footer's by the scrollbar's width: transcript text left = pad + (pane - 2*pad - scrollbar - measure) / 2 composer box left = pad + (pane - 2*pad - measure) / 2 The padding cancels out of the difference and what remains is `-scrollbar/2` at every padding value, so the transcript sat half a scrollbar to the left of the composer and tuning the two paddings against each other could not move it. The column is now established once, by the nearest common ancestor of both (`chatColumn`), and the scroll viewport and the composer footer are siblings inside it that add no horizontal inset of their own. The cap carries the gutter on top of the 760px reading measure, so the text column keeps its width. The handed-over question row drops the padding that indented it past the agent's prose; framed rows keep their own inner padding, which is what a card's border sitting on the column edge means. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * fix(workbench): lock pane-embedded classroom to edit mode (#1256) The workspace right pane painted the full learning chrome — speed control, play button, learner avatars, mic bar — for a course the agent had just created, then flipped to edit once the first scene landed. resolveStageChromeMode treated playback as the DEFAULT branch for a hosted classroom, so every shortfall fell into it: a course whose tab opens at stage_link time has no scenes yet, so currentSceneId is null and isHostedSceneEditable is false. A folded pane parked the playback root behind the fold and cross-faded it out over the pane on unfold, and a failed editor chunk dropped into playback permanently. Lock it at the pane instead of defaulting per entry path: - WorkbenchPanelProvider — the single element that mounts a classroom into the workspace — publishes editPinned (visible && !playback). Every entry path passes through it, so none of them decides. - The hosted resolution can no longer degrade to playback. Start Learning (workbenchLearning, new input, split out from pane visibility) is the one door; everything else resolves between the neutral loading shell and edit. - Stage's chrome dispatch is exhaustive on chromeMode, so the playback root is no longer the else-branch of a condition about the current scene. No flicker: chromeMode is resolved during render, and preloadEditor now answers synchronously (isEditorPreloaded) so a remount with the chunk already registered paints edit on the first frame. A failed import is no longer cached forever, so the lock cannot strand the pane. Standalone classrooms keep their stored mode unchanged. --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> | 1 个月前 | |
feat(storage): wire the asset backend into the app (#1007) (#1089) * fix(persistence): stop corrupting binary response bodies in the route adapter The embedded persistence route bridges a Node `RequestListener` to the Fetch API by hand. Its response object is cast with `as unknown as ServerResponse`, so the compiler checks none of that surface, and two parts of it were wrong in ways only a byte-carrying handler would hit. `end` decoded any `Uint8Array` chunk with `Buffer.from(chunk).toString()`, which is UTF-8. `ServerResponse.end` accepts a `Uint8Array`, and those bytes are not necessarily valid UTF-8, so every unpaired byte became U+FFFD -- silent corruption with no error anywhere. The adapter now buffers chunks as bytes and builds the response from them. `write` was absent entirely, making any chunked handler a runtime TypeError rather than a compile error. It is part of the surface this object claims to implement, so it is now implemented. Neither reaches a caller today: the document and runtime handlers only ever end with JSON strings, which took the string branch. Both are traps laid for the first handler that carries bytes. Both cases are pinned by tests that fail before this change -- the binary one on a byte comparison, the chunked one with a 500 from the TypeError. The existing adapter round-trip test, whose comment already called this the most bug-prone code in the route, covered only the string path. * feat(storage): wire the asset backend into the app (#1007) The server asset registry shipped with #1007 but was reachable from nothing: the persistence route never passed an `assetStore`, and no file under `app/`, `lib/` or `components/` imported the asset server or client modules. This connects it, additively -- no DSL change, no version bump, and no behavioural change when server persistence is off. The route now builds a `PgAssetStore` on the same pool and transaction as the document store, with `PgAssetByteStore` by default and `S3AssetByteStore` when `ASSET_S3_BUCKET` opts a deployment into it. The AWS SDK stays an optional peer, reached only through a dynamic import on that branch. The offline collector is deliberately not scheduled here, and the route says so where a reader would look for it, because leaving it unscheduled grows storage without bound and scheduling it is the deployment's decision. On the browser side the asset pool gains a single-shot configuration seam modelled on the document store's, so that with server persistence enabled the pool is an `HttpAssetStore` carrying the same auth headers as the other stores. Three browser-specific mechanisms needed decisions rather than translation: `clearAssetPool` deletes an IndexedDB database today. In server mode there is nothing local to delete, and a "clear cache" action must not remove server-side assets -- that would destroy user data from a button promising to free space. It now revokes local object URLs and closes the client, and deletes nothing remote. The cross-tab replacement broadcast is unchanged. Its job is to make another tab drop a warm object URL when the bytes behind an id change, and that need is identical in server mode. The exclusivity proof that decides whether an in-place `replace` is safe reasons over this browser only: local documents, unflushed state, and a cross-tab presence probe. In server mode those cover a strict subset of the holders, since another device can reference the same id and no probe can see it -- and asking the server who else references an id would be an existence oracle over other principals. The proof therefore returns not exclusive unconditionally in server mode, and regeneration takes the existing fork path. That is the fail-closed direction and a graceful degradation rather than a break. The development authenticator had to grow the asset principal's required `key`; without it every asset request would have been correctly denied. It derives from the same learner partition, and a principal without one still gets no asset access. The route now states plainly that this key comes from a client-supplied header, so the cross-principal isolation the asset contract describes is not in force under this authenticator -- that warning previously lived only in the auth module's own header, where nobody mounting the route would see it. * fix(storage): close three defects in the asset app wiring (#1007) **The S3 branch could not load, and isolated nothing.** The route dynamically imported the storage package's S3 byte store, but that module statically imported the AWS SDK, so merely resolving it pulled the SDK into resolution and bundling. The route then separately performed an ignored native import of the SDK from the app's own scope, which fails with ERR_MODULE_NOT_FOUND because the SDK lives under the storage workspace package -- so a deployment that set the bucket got a broken store rather than an S3 one. There is now one lazy loader in one resolution scope, the storage module no longer imports the SDK at module load, and the bucket is validated at initialization rather than treated as valid because it is non-empty. A test mocks the SDK to throw on resolution and asserts that importing the byte-store module never resolves it. **A cross-tab replacement could leave a mounted lease stale.** The broadcast handler re-resolves, and `HttpAssetStore.resolve` coalesces concurrent calls for one id onto a single in-flight request -- correct on its own, and required by the contract. Together they meant a resolve triggered by the broadcast could merge into a GET that started before the peer's replacement, return the old revision, and leave the consumer on the old URL with nothing to dislodge it. The client gains an invalidation that retires its cached snapshot and advances the per-id generation without revoking any URL already issued; `release` would have been wrong here, since the contract keeps an issued URL valid until its holder releases it. Removing the invalidation call fails the new race test. **Clearing a server-backed pool left the closed client installed.** In server mode `clearAssetPool` closed the client and returned before clearing the module singleton, so every later `getAssetPool()` handed back a closed store until a page reload happened to intervene. Only the IndexedDB deletion is skipped in server mode now; the singleton is always cleared so the next call rebuilds through the configured factory. Also adds the tests whose absence let these through: a route round-trip carrying invalid UTF-8, the SDK-resolution assertion above, the broadcast race, and reopening the pool after a clear. The earlier tests were not wrong so much as aimed past the defect -- the route test used JSON text, and the S3 test injected a loader rather than exercising the imports. * chore(storage): bump package version to 0.3.0 * fix(persistence): make the route adapter faithful where it carries bytes The embedded persistence route bridges a Node `RequestListener` to the Fetch API by hand, and its response object is cast `as unknown as ServerResponse`, so the compiler checks none of that surface. Five divergences are closed here; the first two are the ones that corrupt data, the last three were raised in review. **`end` decoded bytes as text.** It ran `Buffer.from(chunk).toString()` on a `Uint8Array` chunk, which is UTF-8, so every byte outside that encoding became U+FFFD -- silent corruption with no error anywhere. The adapter now buffers chunks as bytes. **`write` was absent**, making any chunked handler a runtime `TypeError` rather than a compile error. **`write` callbacks ran inline.** Node invokes them after the chunk is handed off. A handler that writes its next chunk from each callback therefore recursed synchronously and could overflow the stack, and a throwing callback threw in the wrong execution phase. They are now deferred. **The encoding argument was ignored.** `write(s, 'latin1')` emitted `0xc3 0xa9` where Node emits `0xe9`, corrupting the body and potentially invalidating a `Content-Length` the handler had set. Both `write` and `end` now parse their full overload set and pass the encoding through. Node's own behaviour for an invalid encoding was established by probing `ServerResponse` on Node 22 rather than assumed: it throws `ERR_UNKNOWN_ENCODING` synchronously, so the adapter validates with `Buffer.isEncoding` and lets `Buffer.from` raise the native error. **Bodyless statuses were not suppressed.** Node discards writes after `writeHead(204)` or `writeHead(304)`; this passed the buffered bytes to the Fetch `Response` constructor, which throws for those statuses, so the request became a 500 instead of the intended 204 or 304. Buffered data is now discarded for 204, 205 and 304, and for `HEAD`. None of the five reaches a caller today -- the document and runtime handlers only ever end with JSON strings. They are traps laid for the first handler that carries bytes, which is why they surfaced while wiring the asset server backend, whose byte routes hit the first two immediately. Each is pinned by a test that fails when its fix is reverted. * fix(storage): keep { client, bucket } working for the S3 byte store `S3AssetByteStoreOptions.commands` was required and the store's methods called it directly, so the published `./asset/s3-bytes` export stopped accepting the `{ client, bucket }` construction it had always taken. Anyone not going through `loadS3AssetByteStore` either failed to compile or hit an undefined command factory at runtime. Make `commands` optional. Omitted, the store resolves the AWS SDK's own command constructors through a dynamic import on its first `write` / `read` / `delete` and caches them, so `{ client, bucket }` works again. Importing the module and constructing a store still never reach the optional peer dependency: a deployment that does not select S3 never resolves it. When the SDK cannot be resolved the call rejects naming `@aws-sdk/client-s3` rather than crashing on an undefined property, and the resolution is left uncached so installing the dependency fixes an already-constructed store. The lazily bound store is run against the byte-store contract, and the new isolation tests pin both that construction does not resolve the SDK and that an unresolvable SDK is reported by name. With the break gone this branch's change to the package is additive, so the version becomes a patch above main (0.2.4 -> 0.2.5) rather than the minor it was carrying. Refs #1007 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(persistence): run the asset collector in the shipped deployment `PgAssetStore.remove`, and a `replace` that changes content, only stamp `asset_blobs.unreferenced_at`. `AssetCollector.collect` is the only path that deletes anything, and the persistence route deliberately did not schedule it, on the grounds that scheduling is the deployment's decision. The deployment this repository ships is docker-compose.yml: the app and PostgreSQL, and nothing else that could make that decision. "The deployment decides" therefore meant "it never runs", and ordinary asset churn retained PostgreSQL bytes or S3 objects forever. Schedule it from instrumentation.ts, which Next runs once per server process — a route module can be instantiated more than once and has no shutdown hook, so a schedule started from one is really started per instantiation. Next 15 and later pick the file up with no config, so next.config.ts is unchanged. Collection is on by default with a 15-minute interval and the package's one-hour grace period, so the Compose stack is correct with no operator action. ASSET_COLLECTION_INTERVAL_MS, ASSET_COLLECTION_GRACE_MS, and ASSET_COLLECTION_ENABLED change or disable it, each documented where it is read. Without DATABASE_URL nothing is scheduled at all, and a failed pass is logged and retried on the next tick rather than ending the schedule or the process. Several instances may collect concurrently: each candidate blob row is re-checked and locked FOR UPDATE inside its own transaction, so collectors serialize on the row rather than racing. The comment says so, so nobody adds a distributed lock that is not needed. The byte-store selection moves to lib/persistence/asset-byte-store so the collector reclaims through the same layer the request path wrote through. A collector on the PostgreSQL byte store while the route wrote to S3 would drop the blob row and orphan the object permanently. Refs #1007 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(storage): bound each collector pass and report whether it filled An unbounded collect() is sized by however long a deployment ran before collection was scheduled: one statement selecting every eligible blob, then a transaction and a byte-layer delete each, in a loop nothing interrupts. The first pass over the Compose deployment that grew a backlog before collection was scheduled is exactly that pass. A pass now takes at most batchSize blobs (default one thousand) and returns. Ordinary churn between two scheduled passes is far below the cap, so a healthy deployment behaves as it did when a pass was unbounded; a capped pass costs the remainder one scheduling interval, which is what the interval is for. collect() still answers with the count, and the count alone cannot tell an empty backlog from a full batch -- a re-referenced or concurrently taken candidate is skipped, so even a full batch can collect less than batchSize. collectPass() returns the count together with capped, which is exactly "this batch was full, run again"; a caller draining a backlog loops while it is true. Candidates are taken oldest-unreferenced first, with the content hash as tiebreaker within one timestamp. Ordering by the hash alone would starve: digests are uniformly distributed, so a high-sorting blob waits behind every lower digest stamped after it, and under steady arrivals those keep coming. The tests pin the ordering against a planner that happens to answer the right rows without an ORDER BY. Refs #1007 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(persistence): keep asset backend faults off the shared handler Two review findings on the app wiring. A concrete asset pool instance is single-lifecycle, but resolveConfiguredAssetPoolStore handed the same object out again after clearAssetPool() had closed it, reinstalling a dead store as the live pool. The second handout now refuses loudly and points at the factory form, which rebuilds on every resolution. createPersistenceHandler eagerly awaited the asset byte store, so an invalid ASSET_S3_BUCKET or an unresolvable AWS SDK rejected the shared persistence handler and took document and runtime traffic down with an optional backend. Byte-store construction is now deferred to the first asset byte operation, and a failed construction is not cached, so the next asset request retries -- the route's own no-poisoned-singleton rule. Refs #1007 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> | 1 个月前 | |
feat(media): write generated media through the asset pool under server-backed persistence (#1392) * feat(media): store generated media in the asset pool when persistence is server-backed With server-backed persistence the document is durable and shared, but generated media stayed in the producing browser: the document kept its gen_img_* / gen_vid_* placeholder and narration kept a browser-derived audio id. Every new browser that opened such a course re-ran generation for every slide, and it never converged, because the address of the generated bytes was never written back into the document. Under server-backed persistence only, the classic generation chain now stores bytes in the asset pool first and writes the id the pool allocated into the document. - The client bootstrap configures the asset seam alongside the document and runtime seams: an HttpAssetStore over the persistence endpoint carrying the same credentials the document store carries, marked server-backed. The seam preflight now covers all three, so a failure still cannot half-configure persistence. - Image, video and TTS generation commit in one fixed order: provider, pool, document, local cache, task. A reference reaches the document only after put returned an id, so a document can never name bytes that were not stored. A failure before the write-back leaves the placeholder with the provider called exactly once; the retry happens on the next owner load. - The write-back is a per-slot rewrite through mutateDocument, which re-reads the current document under the per-stage lock, so it cannot clobber a newer scene. The open course is refreshed with the same rewrite without being marked dirty. - "Has this already been generated?" is answered by the document (the slide exists and no longer holds the placeholder) instead of by this browser's task table. - The classroom's resume effect fails closed on ownership: only a resolved owner starts generation, so a viewer opening a shared course spends nothing. - The local media and audio tables become a per-tab cache. A failed cache write costs a re-download, never the media. Browser-only mode is unchanged: every new call site sits behind the server-backed gate, the local tables stay authoritative there, and placeholders stay in the document. Rendering and export needed no changes. HttpAssetStore.resolve mints an object URL exactly as the browser store does, and the export byte resolver was already pool-first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(classroom): make the generation owner gate a three-outcome rule and apply it everywhere The gate refused everything but a resolved owner, which read a sidecar that answered "no ownership fact exists for this course" as a reason to block. That is the answer a deployment without the sidecar's server-side prerequisites gives for every course, and the answer a course with no ownership record gives: in both, there is nobody the operator's budget needs protecting from, and refusing strands the course's own author behind a question that can never be answered. Ownership is now four states over the sidecar's three outcomes. A definite answer splits into owner and not-owner. An absent record is its own answer, ownerless, and generation proceeds — the behaviour such a deployment had before the gate existed. Only the absence of an answer, a transport failure or a load that has not asked yet, stays unresolved and fails closed: "we could not ask" must never be read as "nobody owns this". One mapper turns a sidecar result into that state, and one predicate decides on it. The workbench classroom pane runs the same resume effect and had no ownership input at all, so a viewer opening a shared course there could still spend the budget. It now asks the sidecar once per course, in parallel with its load and feeding only the generation gate, so its read-only and edit behaviour is unchanged. The shared progressive-load policy carries the gate for it, with both new inputs required rather than defaulted so a future caller cannot omit them into an open budget. Its stale comment claiming ownership could not be expressed here is corrected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the write-back survive autosave, arrive before the scene does, and never leak Independent reviews of the write-back found three ways a durable document could still end up naming a placeholder, and two ways the gate that protects the operator's budget could be walked around. An autosave round captures the store synchronously and writes that capture, so a round already in flight when a rewrite landed wrote the placeholder straight back over the allocated id, and nothing marked the store dirty again to correct it. The rewrite now marks the units it changed, which leaves a corrective flush queued behind the stale one; re-saving a scene that already holds the id is idempotent, losing the id is not. Media is generated from outlines in parallel with scene content and usually finishes first, so the slide that will carry the placeholder does not exist yet and the write-back has nothing to rewrite. That was the ordinary path, not a tail case, and its result was discarded: the task was marked done, the scene was added afterwards with its placeholder intact, and a second pass in the same run could call the provider again. The allocation is now held under the placeholder — which also answers the skip test, so nothing pays twice — and applied when that scene is committed, before its first save. One complete pass now leaves no placeholder behind. A failed commit used to abandon what it had already allocated. A poster upload that failed threw away a stored video and sent the retry to submit the most expensive job in the system again; a rejected write-back left registry rows that name bytes nothing references, which the byte collector cannot reclaim because it only collects blobs no row names. A poster failure now costs the poster, and a write-back that reached nothing reclaims what it allocated. A partial write is left alone, because the document already names it. The ownership gate is fail-closed again. Treating the sidecar's 404 as permission was wrong: the client cannot tell "this course has no owner" from "this deployment told me nothing", so a visitor who opened a shared course could bill the operator. The root cause was the sidecar itself, which gated on the agent runtime although every persisted course has an owner regardless — the persistence route resolves one for every request. It now gates on server persistence, so the configuration that made 404 the universal answer has real ownership facts to report, and the gate can refuse everything but a named owner. Retry affordances answered to no gate at all. A viewer of a shared course with one failed image was shown a Retry button that called the provider. Both retry entry points and every surface that draws them now read one shared permission, so what is offered and what is allowed are the same value. Also: narration regeneration no longer pretends it can replace bytes behind a live id — the exclusivity proof that would allow it is refused by construction once references leave the browser, so it forks to a fresh id and says so; the "already generated" test lets a finished deck answer from the document alone, since scene order stops identifying an outline once slides are inserted or deleted; stored assets record a specific media type rather than a generic transfer type; the pane no longer asks the sidecar in browser-only mode; and the funnel's docstring now states what the per-stage lock actually guarantees, which is same-browser serialization and not a cross-browser compare-and-swap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): park allocations in the deciding turn, never reclaim on an ambiguous write A delta review of the write-back found the first-pass fix still had a window, and the reclamation it added could delete media the document already names. The allocation was parked after an awaited local cache write. A scene committed in that window reconciled against a registry that did not hold it yet, so the document kept the placeholder — and the entry recorded a moment later then answered the skip test as "already handled", so nothing could correct it. Parking now happens inside the write-back, in the same synchronous turn as the decision that nothing could take the reference; no await separates the live check from the park. Allocations parked by an earlier pass are handed to their slides at the start of the next one, before anything decides what still needs generating, so a held allocation whose scene has since arrived becomes a rewrite rather than an answer. Reclaiming on a rejected write was unsound: a rejection does not prove the server did not apply the write, so deleting the asset could break the scene that now names it. The funnel decides instead, and says so: it reclaims only when no store write was ever issued and nothing took the reference. Anything else is placed if its slide exists and parked if it does not, so the next pass reuses the bytes instead of paying for them again. When a write fails after part of it landed, the live store is brought up to the document before the error is rethrown — otherwise the next ordinary flush would overwrite the half that did land, with the ids deliberately not reclaimed. Parked allocations are now cleared with the course. Classic placeholders are reused across runs, so one surviving an interrupted run would be handed to a different slide of the next deck: the previous picture, on a slide whose provider was never asked. Both classroom surfaces clear the arriving course, the deletion cascade clears the deleted one, and clearing the database clears them all. Two more ways generation could start without asking the gate are closed. An overlapping pass — an outline retry re-enters generation with every outline while the first is still working — re-requested elements whose provider call was already in flight; a task that is not done is an answered request, not an unanswered one. And narration regeneration in the timeline editor called the TTS provider and allocated a pool asset with no ownership check at all; it now reads the same permission, which withholds both the per-line and whole-timeline controls and refuses the call. Finally, a pane opened during the stage-link availability gap recorded the sidecar's 404 for a course that was moments from existing and never asked again, leaving the real owner locked out of generation until it remounted. Ownership is re-fetched once the document becomes available; the gate stays closed until an answer arrives, so asking again can only open it for someone entitled to it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the asset routes reachable, and stale snapshots harmless A full-branch audit found that the deployment this project documents could not store a single generated asset, and that several routes into durable storage could still write a placeholder over a reference that had already landed. The persistence route sent asset requests through the development authenticator, which refuses outright in a production build that has not explicitly opted into it — and the documented server-persistence recipe produces exactly that build. Every store and every read answered 401, so images and video failed on every slide while re-billing the provider on each retry, and a narration failure stopped the deck at its first slide. Assets live in one shared partition by design, so there was never anything per-caller for that authenticator to decide: the route now resolves the asset principal itself, alongside the owner it already resolves for documents. Runtime sessions are genuinely per-learner and keep the development authenticator until real session verification replaces it. And narration that cannot be stored no longer fails its scene: the line stays unvoiced and retryable, which is what an image that cannot be stored does to its slide. Placeholders could also come back from behind. A queued autosave's snapshot, an editor-history entry replayed by an undo, the departing save a course switch flushes — each captures content at its own moment, and any of those moments can predate a write-back. Point fixes at each producer would leave the next producer to rediscover the bug, so the check lives at the write boundary every producer passes through, and the allocation record it consults now outlives the parked queue: a placeholder whose rewrite landed long ago is exactly the case it catches. Two ways generation could be lost or repeated are closed. A pass now claims the elements it will reach and releases them however it ends, so an overlapping pass stands down while an aborted one strands nothing — previously its tasks stayed `pending` and every later pass skipped them with no retry control to recover them. And the media abort controller is aborted before being replaced, so a superseded pass stops calling providers instead of running on for a course the user has left. The remaining two are narrower. The workbench pane asks for ownership only after a document load succeeds, and after every later one, mirroring the page route: the load is what creates the ownership row the first time a course is opened, so asking beforehand asked about a course that did not exist yet and locked its author out for the mount. And the ownership gate on the timeline editor now withholds narration regeneration alone; listening back to existing narration and seeing whether a line has any spend nothing and stay available. Known limitation, unchanged and now stated plainly in the comments that used to point at it as a solution: nothing reclaims an unreferenced pool asset. The registry sweep is written but not wired up, and the byte collector only reclaims blobs no registry row names, so every narration regeneration and every abandoned allocation leaves storage behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): gate asset mutations, and make claims and allocation records survive a handoff Opening the asset routes opened all of them. Reads and allocations are meant to be as open as document reads and creates already are, but no authorization hook was supplied, so the handler's default admitted PUT and DELETE too — and those scope by principal key alone, which is one shared constant. Any caller who learned an id, and a document read hands out every id its slides name, could overwrite or destroy another author's media. Mutations now require the deployment's credential, which in a production build without the development-auth opt-in means they are refused outright; reads and allocations stay open. The route comment says what the posture is and what it is not: the deployment-level fence is the access code, and no per-principal quota is configured. The client's own reclaim is best effort to match — losing an argument about deleting an asset must not cost a task its retry, and the bytes are left for server-side reclamation. The pass claim could not survive the handoff it was written for. A retry aborts the live media pass and starts its replacement in the same synchronous block, long before the aborted pass's cleanup runs, so the replacement saw every element still claimed, collected nothing, and returned — leaving each unreached element at pending with nobody coming back for it and no retry control to recover it, which is the exact failure the claim was introduced to prevent. A claim now carries its pass's signal and is retired the moment that signal aborts, and a pass releases only claims it still owns, so a late unwind cannot take its replacement's work. Claims are also acquired at the single point every request passes through, so a single-task retry participates too — previously a retry awaiting its provider was invisible to a pass starting alongside it and both called it. The allocation record could outlive the bytes it named. It was written before the write-back attempted anything and survived the reclaim that followed a failure, so when the slide finally arrived the write boundary stamped a deleted id into the document — and the placeholder it replaced was gone, which reads as already generated and stops anything from retrying. The record is now written only where the allocation is retained, and forgotten wherever a reclaim removes the bytes, including the narration rollback path. The tests follow. The route test drives the real storage handler against an in-memory registry instead of a stub, so it can see what the resolved principal is then allowed to do; the handoff test performs a real abort mid-pass rather than starting from an already-aborted signal; and the guards that could only assert file layout now assert the property they care about, or have been replaced by behaviour. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make media passes serial per course instead of tracking element ownership Three rounds of per-element claims each produced a new way to lose an element. Whole-pass reservations swallowed a Retry for an element the same pass had already failed, leaving it pending with the affordance gone. Retiring a claim by its signal freed an element whose commit was still uploading, so the replacement pass paid for it twice. A claim held for a failed element stranded its retry. The bookkeeping is the defect: every refinement of "who owns this element right now" answered the question at a moment when the answer was already stale. Passes for one course are now serial. A replacement aborts its predecessor, as before, and then waits for it to settle before collecting. That removes the question entirely: a commit already under way finishes — its bytes stored and its reference written, so the new pass sees a resolved slide and skips it — and an element the aborted pass never reached is still a placeholder and gets collected like any other. The claim set, the reservations, the signal retirement and the identity-checked release are all gone. The task table is consulted for one thing only: an element that is generating right now is a single-element retry running alongside the pass, and taking it too would pay twice. Pending is deliberately not a skip reason — it means a pass once intended to reach an element, which an abandoned pass leaves behind with nobody acting on it, and reading that as answered is what stranded elements before. A retry runs concurrently with a pass, because a pass never revisits an element it has processed, and it re-reads the task after its own await and refuses before touching it: marking first and refusing afterwards destroyed the failed state that draws the affordance. Browser-only mode is back to exactly what it was. The abort is now conditional, the waiting does not apply, and the original status-based skip is restored verbatim. Two baseline lines remain changed in each of the two files, and both are behind a server-backed fork whose else-branch is the original. Two smaller things. The allocation record becomes visible when a write goes on the wire rather than when the round trip ends, and the write boundary reconciles under the document lock rather than before it — a save queued during a write-back was otherwise captured with the placeholder and, for a course the user had left, had no corrective flush to follow. And the comments that said a refused reclaim leaves its bytes for server-side reclamation were wrong: nothing collects them, because the registry entry still names its blob and the sweep that would remove it is not wired up. They now say the bytes leak. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a deferred pass re-earn its right to run, and bound the commit it waits on Serializing passes moved their body out of the block that launched them, and three things followed from that. A pass now wakes when its predecessor settles, which can be after the user has left the course. It enqueued before it looked at its signal, into a task table keyed by element id alone — and placeholder ids are not unique across courses, which is why the classroom clears that table on arrival. So a departing course's pass seeded the arriving course's table with tasks carrying the wrong stage id, and a Retry routes by that id: the reference went into the wrong document. A pass now re-validates after the wait, before touching anything shared. The same lateness broke the skip test. `documentSkipIndex` answers only while the live store is on the pass's stage, and returning nothing put the collection loop on the browser-only rule — a silent demotion from "the document is the authority" to "this browser's task table is", on exactly the path where that table has just been cleared. Every element the predecessor had committed was collected again, paid for again, and its second write-back found no placeholder to rewrite, so its bytes were parked where nothing will ever reference them. In server-backed mode an unreadable document now means the pass stands down. And waiting was unbounded. A commit is uncancellable: the asset client takes no signal, and a document write cannot be half-undone. One stalled upload therefore froze the course's media generation for the session — the replacement never collected, the element sat on a skeleton that draws no Retry, and only a reload recovered. The pass's signal is now threaded into the media proxy fetch, and the commit is bounded by a deadline. The deadline is on the wait, not the work: the commit carries on, and if it lands late the document simply ends up correct, while the element becomes retryable and the queue moves on. The tests that were meant to pin the previous round were not sensitive to it. Two asserted end states where the mechanism only changes ordering, and one of them rigged the document read so the assertion held whether or not the pass had waited; a third covered half of what it claimed. They now observe the ordering directly — nothing is issued while another pass for the course is working; in browser-only mode a second pass reaches its provider immediately — and the reconciliation under the document lock has a test that fails when it moves back outside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * revert(media): drop the commit deadline and the abortable download The deadline bought less than it cost. Abandoning a commit after two minutes makes the element retryable while the real commit is still running, so a Retry starts a second commit for the same placeholder against the first: two provider calls, two allocations, and whichever lands second stamps its result over the other's task by element id. The allocation record is keyed by placeholder, so the loser's cleanup erases the winner's record, and the write boundary then puts the raw placeholder back into the document. That is the overlap serial passes were built to remove, reopened through the one door serialization never covered. So a stalled commit holds the course's media queue until it settles or the page is reloaded, and that is written down rather than papered over. The wait is unbounded on purpose: every ceiling on it turns out to be a way of running two commits for one element. Threading the pass signal into the download was also a mistake, in the other direction. The provider call that produced the URL has already been billed, so cancelling the download throws away work that is paid for — and the shared proxy cache records a cancelled request as a transient failure against that URL, which after three of them blocks it for every consumer in the session. Browser-only mode never asked for this: it had no way to observe an abort there, which is exactly why the bytes were kept. The signal is gone from the download again, and `fetchAsBlob` is byte-for-byte what it was before this branch. The regression guard for the stranded-element rule is restored alongside the timing test that was meant to supersede it. It catches a different rule — a task left pending being read as answered — and nothing else does: making the pass skip pending leaves every other suite green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): stop asking the pool for refs it never issued, bound it, and adopt cached bytes Four things a deployment found once this was running for real. A reference this application mints itself — a generation placeholder, a derived narration key — was never in the pool, because the pool allocates every id it holds. Asking anyway used to be an IndexedDB miss; once the pool is server-backed it is a request that answers 404, one per element per load, forever on a course that still holds placeholders. Every lease and probe now checks first. The check is a negative test on shapes this application owns, not an id validator: the pool's id domain stays unconstrained, and anything that is not one of ours is still asked about. The asset store can bound how much one principal holds, and enforces it inside the write transaction, but nothing ever passed the number. It does now, with a default rather than an opt-in: allocation is reachable by any caller a deployment admits, and with one shared principal an unbounded store is unbounded database growth with no operator-visible brake. Refusing asset mutations to unauthenticated callers was not enough, because every authenticated caller resolves to that same shared principal — so authentication decided nothing, and any signed-in visitor could delete any id they learned. Since this branch began storing media the registry is the only copy a course has. Replacing and deleting are now refused to everyone, and the browser no longer tries: an entry nothing references waits for server-side reclamation instead. What a browser must still do is forget its own record of an allocation that reached nothing, or a later save would stamp an id the document has no reason to trust. And a course generated before any of this holds placeholders in its document with its bytes only in the author's browser. Those bytes are paid for, so the author's next load converts them — stored to the pool and written back through the ordinary commit path, with no provider call — instead of buying them again. A row that records only a hosted URL is treated as absent: that URL is the provider's address, not something a document may hold. One renderer expectation moved with this. An untracked placeholder used to paint as pending on first render because asking the pool left a lease in flight; it settled to disabled a moment later either way, and now says so from the start. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): surface a full store as a refusal and convert legacy narration A quota refusal reached the browser as HTTP 500 with a generic message, which reads as a transient failure: the element kept a Retry that would pay a provider again and be refused again. The store raises the contract's own error and the handler maps it to 507, but the store answering a request is not always built by the same bundle as the handler -- the persistence provider is reached from the route bundle and from instrumentation, which is why its state lives on a Symbol.for global -- and `instanceof` is false across that boundary while the declared code is still right. Classify on the code as well as the class, and make the code a permanent, persisted refusal in the browser: recorded locally so it survives a reload, shown as "storage is full", and refused by the retry entry point so a stale button cannot buy a second generation. Every other storage failure stays retryable. Convert what a pre-server-backed course still holds. Generated media is adopted under either key this application has used for it -- the placeholder, and the allocated id of a course converted once and later rolled back -- instead of only the first. Narration is converted by a load-time pass over the open course's speech actions, since nothing re-enters generation for an action that already has an id: bytes to the pool, id written back through a funnel that mirrors the media one, owner-only and server-backed-only. A line whose bytes are in no browser is left alone rather than re-synthesized. Also: the pool guard is now a positive `ast_` test rather than an enumeration of the shapes we mint (imports never reach the pool, so this is safe in both modes); the slide ref collection is an exported pure function so its four lease sites are covered behaviourally; ASSET_QUOTA_BYTES treats every spelling of zero as opting out and refuses a malformed value at startup instead of falling back; the abort signal is re-checked after the cache read, before an uncancellable commit; and the unused `removeAsset` and pool `replace` surfaces are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * chore(storage): release 0.29.1 The asset HTTP handler now recognises a store refusal by the contract code it declares as well as by its class, so a quota refusal raised in another module realm answers 507 instead of 500. Same contract, stricter recognition, no API change: a patch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a full store recoverable and adoption course-safe Narration adoption read the local audio row by its derived key alone. That key carries no stage id and the table is keyed by id alone, so two courses can mint the same one -- a PPTX import numbers its scenes and actions deterministically, which gives every imported deck's first slide `tts_s1_speech-scene-p1`. Locally a collision only means one course plays another's clip in one browser; adopting it wrote that clip into the shared document permanently, for every device and every visitor. A row that names a course is now adopted only into that course, and a row from before that column existed only when the text it recorded is the text of the action being converted. A full asset store was made permanent last round, which was wrong three times over: it overwrote the refused bytes with an empty blob -- on the conversion path that row is a course's only copy of its own media -- it kept sending the rest of the deck to a provider against a ceiling it already knew was reached, and it left no way back once an operator raised that ceiling. A full store is neither the content's fault nor the configuration's, so it is now its own case: the bytes are kept, the pass stops at the first refusal, and the element shows the reason together with a Retry that re-attempts the upload from those bytes. Nothing retries automatically, so no one is re-billed. The narration write-back now reaches the write boundary every producer of a durable write passes through, not only the dirty mark: adoption never deletes the derived row, so a snapshot that reverts the rewrite is adopted again on the next load and allocates a fresh asset every time. Adoption is also mounted by both classroom surfaces rather than one, takes the course's abort signal, and re-validates that this browser still has the course open before each write. ASSET_QUOTA_BYTES is validated from instrumentation, where the README and the docstring already claimed it was: its only other consumer is lazy and memoised, so a malformed ceiling let the process boot and then failed every persistence request, documents and runtime included. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): remember a full store per course, and never lose retained bytes A stopped pass left the elements it never reached as placeholders with no persisted record -- deliberately, since nothing was attempted for them. But that left the next load with no reason not to try: it called a provider for the next placeholder and was refused at exactly the same point, once per reload, indefinitely. A full store is not a property of any slide. It belongs to the deployment and changes for reasons the document knows nothing about, so it is now remembered once per course in the browser's device KV. A pass that finds the marker stands down before spending anything and leaves every placeholder its "storage is full" state and its Retry; the first upload that succeeds clears it and the next pass runs normally. Narration adoption latched per course so it runs once per load, and the latch outlived the abort that leaving a course performs. On a surface that stays mounted across switches -- the workbench pane is one component for every course it shows -- owner course A, visitor course B, then back to A skipped exactly the clips the abort had cut off, and nothing else converts them. The latch is released with the abort now, and a course adopts one run at a time so a re-entry cannot hand a clip a second allocation while the previous run's uncancellable tail is still settling. A quota-blocked element retried into a network error or a 500 lost the bytes that were kept for it: the retry deleted the row before attempting the upload and wrote no replacement for an error carrying no structured code, so the next retry went back to a provider for media this browser had a moment earlier. The row now survives until an upload succeeds, the failure handler keeps whatever bytes the attempt was given, and the retry asks the question a pass asks -- does this browser already hold bytes for this element -- rather than reading an error code that a second failure has already overwritten. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): adopt real legacy narration, queue re-entries, report attempt outcomes Narration adoption admitted a stage-less row only when the text it recorded matched the action being converted. Both of those columns were added to the local audio table by the very change that moved narration onto allocated ids, so a row still carrying a derived key has neither: the rule refused every real pre-allocation course and passed only on fixtures built from post-allocation rows. What the row cannot say, the key can. A derived key names two clips only when two courses share a scene order and an action id, and an action id repeats only when something other than the generator minted it -- an import numbers them by slide position. So a key built from a generated action id is adopted on that basis, a key an import could have reproduced still needs matching text, and a row that names another course is refused however unique its key looks. Handing a re-entering caller the adoption run already in flight undid the latch release it was paired with: that run is bound to the signal the departure just aborted, so it stops at its next clip while the caller -- which has the course open and a live signal -- is told the work is done, and an effect replayed as mount, cleanup, mount adopts nothing at all. A later caller now waits for the uncancellable tail and scans again, which costs a lookup on a course that has nothing left and finishes the clips the abort cut off on one that does. One attempt at an element now reports both facts its callers need instead of a bare boolean: whether the store refused it for room, and whether bytes actually reached the store. Leaving a course clears the task table, so a retry that landed afterwards read "no failed task" as success and deleted the row holding the only copy of the media. Nothing is inferred from that table any more. Reading the localStorage property can throw where storage is denied by policy, typeof included, so the availability check moved inside the guard: this metadata is best-effort, and a rejection here strands a generation pass that has already enqueued its tasks. A retry is never blocked by the per-course "store is full" marker, but a retry that is refused again re-sets it, and adoption now reads and writes the same marker rather than issuing one refused upload per clip on every load. The two canvas element renderers and both thumbnail renderers show the reason beside the Retry, so a full store does not look like an ordinary failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): probe a full store instead of standing down, and pair notices with a Retry Narration adoption was given both halves of the per-course "the store is full" marker last round: it stood down when the marker was set, and it set the marker when its own upload was refused for room. Those halves are only safe together if something can lift the marker, and for adoption nothing could. It has no affordance of its own, it stood down before reaching its own clear, the media pass returns before its marker gate when there is nothing to generate -- so a narration-only deck, or one whose slides are already satisfied, painted no storage-full element and offered no Retry -- and narration generated rather than adopted allocates directly rather than through the media commit. The course's cached narration was then lost for good, where before it converted on the first load after the ceiling was raised. The gate is a probe now. A marked course attempts exactly one clip per load: refused, it stops and the marker stands, which costs what standing down cost; stored, it lifts the marker and finishes the course. Adoption spends no provider money, so the whole cost of probing a store that is still full is one refused upload. Generated narration lifts the marker too. The three surfaces that gained a failure notice last round drew it for any failure with a reason, including the one refusal that is reachable without server-backed persistence, so a browser-only deck painted something it had not painted before. The notice is drawn beside a Retry and nowhere else, which is what it was added for and what leaves browser-only output unchanged. Both are now asserted through the render harness the surface matrix already had. A caller arriving while a rescan is queued shares it rather than appending another. One rescan converts whatever the run in flight left and every later one would find an allocated id on every action, so a chain bought nothing and turned a single stalled upload into a course that never adopts again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): treat a refusal for room as a fact about one clip, not the deck The asset store checks each write against the headroom it has left, so a store that refuses a long opening clip can still hold every short clip behind it. Narration adoption assumed the opposite: it broke the deck at the first refusal and then re-attempted that same first clip on every later load, because the document names it first. A deck whose longest clip exceeds current headroom therefore never converted the clips that would have fit, with no affordance to recover it -- the state the probe was introduced to remove, reached through a narrower door. An unmarked load now attempts every clip, skipping the ones that do not fit, and remembers the condition only if the load ends with clips it still could not store. A marked load spends its single upload on the smallest clip left rather than the first one named: that is the clip that answers the question the marker asks, because if the smallest does not fit nothing does. The media pass keeps stopping at its first refusal, and for a reason adoption does not share -- every element it attempts costs a provider call. A rescan several callers share took the newest caller's signal, and the newest caller is not necessarily the one still there: a surface that opened a course and closed it again would stop work a surface still showing that course was waiting for, and that surface is latched, so it would never ask again. The shared run now takes a signal that is aborted only once every caller has left. The comment claiming the shared rescan contains a stalled upload was wrong -- the rescan is chained off the run in flight, so a stalled upload leaves every caller pending exactly as a chain would. It claims the bounded queue it actually provides, and the stall is recorded as a limitation. The failed-state containers took their stacking classes unconditionally, so markup differed in browser-only mode even though nothing moved on screen. Those classes are applied only when there is a notice to stack, and the tests assert the exact class attribute rather than a substring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): stop narration adoption writing the media pass's store-full marker The marker means "do not call a provider for this course". A path is entitled to write it only if its own refusal cost a provider call, and narration adoption's refusals cost nothing: it uploads bytes this browser already holds. The store also checks each write against the headroom it has left, so a clip that does not fit says nothing about whether a slide's image would. Adoption was writing it anyway, and one over-long narration clip was therefore enough to stand a course's entire image pass down on every later load -- on a store that had just accepted adoption's other clips. The author could still recover each element by hand, every load, for ever. Three rounds of narrowing this seam produced a finding each time, so it is removed rather than narrowed again. Gone: the marker read, the single-clip probe, the smallest-clip selection, and the up-front read of every row into an array -- which also retires a sampled-then-stale flag and the retention of a whole deck's blobs for the length of a run, and returns the loop to streaming one row at a time. Adoption's rule is now that every load attempts every clip it holds, once; any failure skips that clip and the load continues. The noise the coupling was meant to avoid does not arise, because after the first load the clips still outstanding are exactly the ones that did not fit -- normally none, or one. A successful write still clears the marker, and that is a different kind of statement: a write that went through is a fact this run established, where a refusal is an inference about what some other write would cost. For a course whose media needs nothing, adoption and generated narration are also the only paths that can establish it. The failure module still documented the deck-wide premise this contradicts. It now says what is true: the check is per write, and the media pass stops the deck as a judgement about cost rather than about certainty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bound a full store's cost from the store's own arithmetic Removing the store-full marker from narration adoption removed its bound too, and the code then asserted the bound was unnecessary. It is, on a store with room for most of a deck. On the store the whole mechanism exists for -- the ceiling reached, nothing fitting -- the outstanding set after every load is the entire deck, so a thirty-clip course posted thirty full blobs on every load, indefinitely. Each of those is not a cheap refusal: the bytes are uploaded, the server hashes the whole payload, and only then takes a per-principal lock and sums every entry that principal owns before saying no. The bound needs no flag, no key and nothing carried between loads. The store asks whether `used + addedBytes` exceeds the ceiling, and `used` only grows while a run is uploading, so a clip refused for want of room implies every clip at least that large is refused for the rest of that run. The run keeps the smallest size it has been refused and skips anything no smaller without uploading it; a smaller clip is still attempted, because it may fit. A deck the store refuses entirely now costs one upload per successive size minimum instead of one per clip, and a deck it has room for costs nothing extra, because nothing is refused. Only a refusal for room lowers the bar: a dropped connection says nothing about how much room there is. The deck-wide certainty premise the failure module retracted last round still stood verbatim at the site that implements the stand-down. Both copies now say the same thing: the check is per write, and the pass stops the deck as a judgement about cost rather than about certainty. The comment on adoption's marker clear now names its price. Narration of a few hundred bytes fits in headroom an image does not, so a proven write can let the next pass buy one more image that is refused again -- bounded at one, and the price of the alternative being a course whose media never generates again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(media): state the adoption bound exactly, and stop three comments describing the old rule The comment introducing the in-load bound gave its cost as "at most a handful, and the first load pays the most". Neither clause is a property of the rule. A clip is skipped only when something no larger was already refused, so a fully-refused deck costs one upload per successive size minimum in document order: one when the clips grow, about ln N for an arbitrary order, and one per clip when they only shrink -- a long opener followed by terser lines is exactly that shape. And no load is cheaper than the first, because the bound resets per run and a refused clip stays outstanding. The comment now says that, and points at what would make it exactly one for any ordering: the store returning its remaining headroom in the refusal's existing details channel, which the server leaves empty today. Two other comments still described the previous rule -- "attempts every clip it holds, every load" -- one of them twenty lines above the paragraph that introduces the bound, in the same block. Both now say what the code does. The bound's soundness is worth stating where a maintainer will look for it: quota is charged at full length with no discount for a duplicate, the sum it is checked against joins entries to blobs so the collector cannot lower it, the check takes a per-principal lock before summing, and replace and delete are refused to every browser. Nothing a run can do makes room appear inside it. One test installed a row implementation and replaced it wholesale a few lines later, so the first was dead and the survivor dropped the text the first clip's import-shaped key needs for the ownership rule -- it passed on the coincidence that the fixture's default text is the action's. Merged into one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): keep a Retry from re-buying parked media, and state the store seam once Five findings from an inline review. The standalone classroom route asked the ownership sidecar once per load and recorded only stage ownership when that ask failed. Every non-answer fails closed, so one transient 5xx left the genuine author with no resume, no Retry affordance and no legacy narration converted for the rest of the load, with nothing to change it short of a reload. The failure now records the fail-closed answer explicitly -- an answer an earlier load established must not outlive the failure that replaced it -- and an unresolved answer is asked for again, a few times over a few seconds. A real answer, however unwelcome, is final. A Retry could pay a provider for media the pool already held. When the bytes are stored and only the write-back fails in a way that keeps the allocation, it is parked and no local row exists, because that row is written only after a successful write-back. Retry now reads the parked queue exactly as the pass does and re-attempts the write-back: it re-keys the task done when the document takes it, leaves the entry parked when the slide still does not exist, and stays failed and retryable when the document refuses again. Object URLs a parked allocation owns are revoked when the entry is dropped. The commit path leaves them alone while the entry is parked, because it is then the only thing holding bytes this tab can render, so a course switch or a stage deletion was pinning the whole blob for the life of the tab. An entry a slide has already taken is left alone: the task table is displaying those URLs. The fallback lookup for cached bytes is a stage-scoped scan, and the keyed lookup misses for every row the commit path writes, so a pass was materializing and sorting the course's whole media table once per element. One scan per pass now, built on the first miss. It is sound and not merely cheaper: an element asks only for its own placeholder, and every row a pass writes carries the placeholder of the element that wrote it. "The store accepted a write, so it is not out of room" was enforced at three call sites under slightly different conditions, which made it a convention the next pool write path could silently break. It is stated once, in putAsset, for the course whose bytes it just stored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 16 天前 | |
fix(tests): make media tests pass on Windows path separators (#1615) | 7 天前 | |
feat(token-plan): add TokenDance one-key preset for every modality (#1525) * feat(token-plan): add TokenDance one-key preset for every modality TokenDance is a model gateway: chat and images are OpenAI-compatible at /gateway/v1, and the same key authenticates vendor-protocol routes on the same host (Ark, MiniMax, Bocha). The preset reuses the existing adapters with those route prefixes as base URLs, so one key lights up LLM, image, video, TTS and web search from Settings -> Token Plan. - providers: add a built-in `tokendance` OpenAI-compatible provider (TOKENDANCE_* env prefix, logo, provider name in all locales) - token-plan: add the TokenDance preset (Seedream image, MiniMax H3 video, MiniMax speech TTS, Bocha web search) - seedream: use a base URL that already ends in a version segment verbatim, so gateway routes like `/ark/v3` do not get `/api/v3` appended - minimax-video: route H3-family models through the v2 task API (content array submit, task-envelope poll); connectivity checks for H3 probe auth on the v2 query route instead of submitting a billable task - README: add a one-key quick example and replace the Gemini-specific model recommendation with a provider-agnostic setup recommendation Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qrrq9CPwb718mpouz8Y2KL * fix(token-plan): accept preset web-search base URLs and report H3 dimensions per ratio - web-search: the client base URL allowlist also accepts the exact base URL a built-in token plan preset writes for that provider, derived from TOKEN_PLAN_PRESETS. Applying a plan whose web-search route is not an official vendor host previously stored a URL that the route rejected with 400. Any other client URL is still rejected. - minimax-video: report H3 v2 clip dimensions for 16:9, 9:16, 4:3 and 1:1 instead of assuming landscape for every non-portrait ratio. - tests: pin the allowlist for every preset, the 1:1 H3 dimensions, and clear TOKENDANCE_* in the provider-config env isolation list. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qrrq9CPwb718mpouz8Y2KL --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> | 11 天前 | |
feat(media): write generated media through the asset pool under server-backed persistence (#1392) * feat(media): store generated media in the asset pool when persistence is server-backed With server-backed persistence the document is durable and shared, but generated media stayed in the producing browser: the document kept its gen_img_* / gen_vid_* placeholder and narration kept a browser-derived audio id. Every new browser that opened such a course re-ran generation for every slide, and it never converged, because the address of the generated bytes was never written back into the document. Under server-backed persistence only, the classic generation chain now stores bytes in the asset pool first and writes the id the pool allocated into the document. - The client bootstrap configures the asset seam alongside the document and runtime seams: an HttpAssetStore over the persistence endpoint carrying the same credentials the document store carries, marked server-backed. The seam preflight now covers all three, so a failure still cannot half-configure persistence. - Image, video and TTS generation commit in one fixed order: provider, pool, document, local cache, task. A reference reaches the document only after put returned an id, so a document can never name bytes that were not stored. A failure before the write-back leaves the placeholder with the provider called exactly once; the retry happens on the next owner load. - The write-back is a per-slot rewrite through mutateDocument, which re-reads the current document under the per-stage lock, so it cannot clobber a newer scene. The open course is refreshed with the same rewrite without being marked dirty. - "Has this already been generated?" is answered by the document (the slide exists and no longer holds the placeholder) instead of by this browser's task table. - The classroom's resume effect fails closed on ownership: only a resolved owner starts generation, so a viewer opening a shared course spends nothing. - The local media and audio tables become a per-tab cache. A failed cache write costs a re-download, never the media. Browser-only mode is unchanged: every new call site sits behind the server-backed gate, the local tables stay authoritative there, and placeholders stay in the document. Rendering and export needed no changes. HttpAssetStore.resolve mints an object URL exactly as the browser store does, and the export byte resolver was already pool-first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(classroom): make the generation owner gate a three-outcome rule and apply it everywhere The gate refused everything but a resolved owner, which read a sidecar that answered "no ownership fact exists for this course" as a reason to block. That is the answer a deployment without the sidecar's server-side prerequisites gives for every course, and the answer a course with no ownership record gives: in both, there is nobody the operator's budget needs protecting from, and refusing strands the course's own author behind a question that can never be answered. Ownership is now four states over the sidecar's three outcomes. A definite answer splits into owner and not-owner. An absent record is its own answer, ownerless, and generation proceeds — the behaviour such a deployment had before the gate existed. Only the absence of an answer, a transport failure or a load that has not asked yet, stays unresolved and fails closed: "we could not ask" must never be read as "nobody owns this". One mapper turns a sidecar result into that state, and one predicate decides on it. The workbench classroom pane runs the same resume effect and had no ownership input at all, so a viewer opening a shared course there could still spend the budget. It now asks the sidecar once per course, in parallel with its load and feeding only the generation gate, so its read-only and edit behaviour is unchanged. The shared progressive-load policy carries the gate for it, with both new inputs required rather than defaulted so a future caller cannot omit them into an open budget. Its stale comment claiming ownership could not be expressed here is corrected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the write-back survive autosave, arrive before the scene does, and never leak Independent reviews of the write-back found three ways a durable document could still end up naming a placeholder, and two ways the gate that protects the operator's budget could be walked around. An autosave round captures the store synchronously and writes that capture, so a round already in flight when a rewrite landed wrote the placeholder straight back over the allocated id, and nothing marked the store dirty again to correct it. The rewrite now marks the units it changed, which leaves a corrective flush queued behind the stale one; re-saving a scene that already holds the id is idempotent, losing the id is not. Media is generated from outlines in parallel with scene content and usually finishes first, so the slide that will carry the placeholder does not exist yet and the write-back has nothing to rewrite. That was the ordinary path, not a tail case, and its result was discarded: the task was marked done, the scene was added afterwards with its placeholder intact, and a second pass in the same run could call the provider again. The allocation is now held under the placeholder — which also answers the skip test, so nothing pays twice — and applied when that scene is committed, before its first save. One complete pass now leaves no placeholder behind. A failed commit used to abandon what it had already allocated. A poster upload that failed threw away a stored video and sent the retry to submit the most expensive job in the system again; a rejected write-back left registry rows that name bytes nothing references, which the byte collector cannot reclaim because it only collects blobs no row names. A poster failure now costs the poster, and a write-back that reached nothing reclaims what it allocated. A partial write is left alone, because the document already names it. The ownership gate is fail-closed again. Treating the sidecar's 404 as permission was wrong: the client cannot tell "this course has no owner" from "this deployment told me nothing", so a visitor who opened a shared course could bill the operator. The root cause was the sidecar itself, which gated on the agent runtime although every persisted course has an owner regardless — the persistence route resolves one for every request. It now gates on server persistence, so the configuration that made 404 the universal answer has real ownership facts to report, and the gate can refuse everything but a named owner. Retry affordances answered to no gate at all. A viewer of a shared course with one failed image was shown a Retry button that called the provider. Both retry entry points and every surface that draws them now read one shared permission, so what is offered and what is allowed are the same value. Also: narration regeneration no longer pretends it can replace bytes behind a live id — the exclusivity proof that would allow it is refused by construction once references leave the browser, so it forks to a fresh id and says so; the "already generated" test lets a finished deck answer from the document alone, since scene order stops identifying an outline once slides are inserted or deleted; stored assets record a specific media type rather than a generic transfer type; the pane no longer asks the sidecar in browser-only mode; and the funnel's docstring now states what the per-stage lock actually guarantees, which is same-browser serialization and not a cross-browser compare-and-swap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): park allocations in the deciding turn, never reclaim on an ambiguous write A delta review of the write-back found the first-pass fix still had a window, and the reclamation it added could delete media the document already names. The allocation was parked after an awaited local cache write. A scene committed in that window reconciled against a registry that did not hold it yet, so the document kept the placeholder — and the entry recorded a moment later then answered the skip test as "already handled", so nothing could correct it. Parking now happens inside the write-back, in the same synchronous turn as the decision that nothing could take the reference; no await separates the live check from the park. Allocations parked by an earlier pass are handed to their slides at the start of the next one, before anything decides what still needs generating, so a held allocation whose scene has since arrived becomes a rewrite rather than an answer. Reclaiming on a rejected write was unsound: a rejection does not prove the server did not apply the write, so deleting the asset could break the scene that now names it. The funnel decides instead, and says so: it reclaims only when no store write was ever issued and nothing took the reference. Anything else is placed if its slide exists and parked if it does not, so the next pass reuses the bytes instead of paying for them again. When a write fails after part of it landed, the live store is brought up to the document before the error is rethrown — otherwise the next ordinary flush would overwrite the half that did land, with the ids deliberately not reclaimed. Parked allocations are now cleared with the course. Classic placeholders are reused across runs, so one surviving an interrupted run would be handed to a different slide of the next deck: the previous picture, on a slide whose provider was never asked. Both classroom surfaces clear the arriving course, the deletion cascade clears the deleted one, and clearing the database clears them all. Two more ways generation could start without asking the gate are closed. An overlapping pass — an outline retry re-enters generation with every outline while the first is still working — re-requested elements whose provider call was already in flight; a task that is not done is an answered request, not an unanswered one. And narration regeneration in the timeline editor called the TTS provider and allocated a pool asset with no ownership check at all; it now reads the same permission, which withholds both the per-line and whole-timeline controls and refuses the call. Finally, a pane opened during the stage-link availability gap recorded the sidecar's 404 for a course that was moments from existing and never asked again, leaving the real owner locked out of generation until it remounted. Ownership is re-fetched once the document becomes available; the gate stays closed until an answer arrives, so asking again can only open it for someone entitled to it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the asset routes reachable, and stale snapshots harmless A full-branch audit found that the deployment this project documents could not store a single generated asset, and that several routes into durable storage could still write a placeholder over a reference that had already landed. The persistence route sent asset requests through the development authenticator, which refuses outright in a production build that has not explicitly opted into it — and the documented server-persistence recipe produces exactly that build. Every store and every read answered 401, so images and video failed on every slide while re-billing the provider on each retry, and a narration failure stopped the deck at its first slide. Assets live in one shared partition by design, so there was never anything per-caller for that authenticator to decide: the route now resolves the asset principal itself, alongside the owner it already resolves for documents. Runtime sessions are genuinely per-learner and keep the development authenticator until real session verification replaces it. And narration that cannot be stored no longer fails its scene: the line stays unvoiced and retryable, which is what an image that cannot be stored does to its slide. Placeholders could also come back from behind. A queued autosave's snapshot, an editor-history entry replayed by an undo, the departing save a course switch flushes — each captures content at its own moment, and any of those moments can predate a write-back. Point fixes at each producer would leave the next producer to rediscover the bug, so the check lives at the write boundary every producer passes through, and the allocation record it consults now outlives the parked queue: a placeholder whose rewrite landed long ago is exactly the case it catches. Two ways generation could be lost or repeated are closed. A pass now claims the elements it will reach and releases them however it ends, so an overlapping pass stands down while an aborted one strands nothing — previously its tasks stayed `pending` and every later pass skipped them with no retry control to recover them. And the media abort controller is aborted before being replaced, so a superseded pass stops calling providers instead of running on for a course the user has left. The remaining two are narrower. The workbench pane asks for ownership only after a document load succeeds, and after every later one, mirroring the page route: the load is what creates the ownership row the first time a course is opened, so asking beforehand asked about a course that did not exist yet and locked its author out for the mount. And the ownership gate on the timeline editor now withholds narration regeneration alone; listening back to existing narration and seeing whether a line has any spend nothing and stay available. Known limitation, unchanged and now stated plainly in the comments that used to point at it as a solution: nothing reclaims an unreferenced pool asset. The registry sweep is written but not wired up, and the byte collector only reclaims blobs no registry row names, so every narration regeneration and every abandoned allocation leaves storage behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): gate asset mutations, and make claims and allocation records survive a handoff Opening the asset routes opened all of them. Reads and allocations are meant to be as open as document reads and creates already are, but no authorization hook was supplied, so the handler's default admitted PUT and DELETE too — and those scope by principal key alone, which is one shared constant. Any caller who learned an id, and a document read hands out every id its slides name, could overwrite or destroy another author's media. Mutations now require the deployment's credential, which in a production build without the development-auth opt-in means they are refused outright; reads and allocations stay open. The route comment says what the posture is and what it is not: the deployment-level fence is the access code, and no per-principal quota is configured. The client's own reclaim is best effort to match — losing an argument about deleting an asset must not cost a task its retry, and the bytes are left for server-side reclamation. The pass claim could not survive the handoff it was written for. A retry aborts the live media pass and starts its replacement in the same synchronous block, long before the aborted pass's cleanup runs, so the replacement saw every element still claimed, collected nothing, and returned — leaving each unreached element at pending with nobody coming back for it and no retry control to recover it, which is the exact failure the claim was introduced to prevent. A claim now carries its pass's signal and is retired the moment that signal aborts, and a pass releases only claims it still owns, so a late unwind cannot take its replacement's work. Claims are also acquired at the single point every request passes through, so a single-task retry participates too — previously a retry awaiting its provider was invisible to a pass starting alongside it and both called it. The allocation record could outlive the bytes it named. It was written before the write-back attempted anything and survived the reclaim that followed a failure, so when the slide finally arrived the write boundary stamped a deleted id into the document — and the placeholder it replaced was gone, which reads as already generated and stops anything from retrying. The record is now written only where the allocation is retained, and forgotten wherever a reclaim removes the bytes, including the narration rollback path. The tests follow. The route test drives the real storage handler against an in-memory registry instead of a stub, so it can see what the resolved principal is then allowed to do; the handoff test performs a real abort mid-pass rather than starting from an already-aborted signal; and the guards that could only assert file layout now assert the property they care about, or have been replaced by behaviour. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make media passes serial per course instead of tracking element ownership Three rounds of per-element claims each produced a new way to lose an element. Whole-pass reservations swallowed a Retry for an element the same pass had already failed, leaving it pending with the affordance gone. Retiring a claim by its signal freed an element whose commit was still uploading, so the replacement pass paid for it twice. A claim held for a failed element stranded its retry. The bookkeeping is the defect: every refinement of "who owns this element right now" answered the question at a moment when the answer was already stale. Passes for one course are now serial. A replacement aborts its predecessor, as before, and then waits for it to settle before collecting. That removes the question entirely: a commit already under way finishes — its bytes stored and its reference written, so the new pass sees a resolved slide and skips it — and an element the aborted pass never reached is still a placeholder and gets collected like any other. The claim set, the reservations, the signal retirement and the identity-checked release are all gone. The task table is consulted for one thing only: an element that is generating right now is a single-element retry running alongside the pass, and taking it too would pay twice. Pending is deliberately not a skip reason — it means a pass once intended to reach an element, which an abandoned pass leaves behind with nobody acting on it, and reading that as answered is what stranded elements before. A retry runs concurrently with a pass, because a pass never revisits an element it has processed, and it re-reads the task after its own await and refuses before touching it: marking first and refusing afterwards destroyed the failed state that draws the affordance. Browser-only mode is back to exactly what it was. The abort is now conditional, the waiting does not apply, and the original status-based skip is restored verbatim. Two baseline lines remain changed in each of the two files, and both are behind a server-backed fork whose else-branch is the original. Two smaller things. The allocation record becomes visible when a write goes on the wire rather than when the round trip ends, and the write boundary reconciles under the document lock rather than before it — a save queued during a write-back was otherwise captured with the placeholder and, for a course the user had left, had no corrective flush to follow. And the comments that said a refused reclaim leaves its bytes for server-side reclamation were wrong: nothing collects them, because the registry entry still names its blob and the sweep that would remove it is not wired up. They now say the bytes leak. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a deferred pass re-earn its right to run, and bound the commit it waits on Serializing passes moved their body out of the block that launched them, and three things followed from that. A pass now wakes when its predecessor settles, which can be after the user has left the course. It enqueued before it looked at its signal, into a task table keyed by element id alone — and placeholder ids are not unique across courses, which is why the classroom clears that table on arrival. So a departing course's pass seeded the arriving course's table with tasks carrying the wrong stage id, and a Retry routes by that id: the reference went into the wrong document. A pass now re-validates after the wait, before touching anything shared. The same lateness broke the skip test. `documentSkipIndex` answers only while the live store is on the pass's stage, and returning nothing put the collection loop on the browser-only rule — a silent demotion from "the document is the authority" to "this browser's task table is", on exactly the path where that table has just been cleared. Every element the predecessor had committed was collected again, paid for again, and its second write-back found no placeholder to rewrite, so its bytes were parked where nothing will ever reference them. In server-backed mode an unreadable document now means the pass stands down. And waiting was unbounded. A commit is uncancellable: the asset client takes no signal, and a document write cannot be half-undone. One stalled upload therefore froze the course's media generation for the session — the replacement never collected, the element sat on a skeleton that draws no Retry, and only a reload recovered. The pass's signal is now threaded into the media proxy fetch, and the commit is bounded by a deadline. The deadline is on the wait, not the work: the commit carries on, and if it lands late the document simply ends up correct, while the element becomes retryable and the queue moves on. The tests that were meant to pin the previous round were not sensitive to it. Two asserted end states where the mechanism only changes ordering, and one of them rigged the document read so the assertion held whether or not the pass had waited; a third covered half of what it claimed. They now observe the ordering directly — nothing is issued while another pass for the course is working; in browser-only mode a second pass reaches its provider immediately — and the reconciliation under the document lock has a test that fails when it moves back outside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * revert(media): drop the commit deadline and the abortable download The deadline bought less than it cost. Abandoning a commit after two minutes makes the element retryable while the real commit is still running, so a Retry starts a second commit for the same placeholder against the first: two provider calls, two allocations, and whichever lands second stamps its result over the other's task by element id. The allocation record is keyed by placeholder, so the loser's cleanup erases the winner's record, and the write boundary then puts the raw placeholder back into the document. That is the overlap serial passes were built to remove, reopened through the one door serialization never covered. So a stalled commit holds the course's media queue until it settles or the page is reloaded, and that is written down rather than papered over. The wait is unbounded on purpose: every ceiling on it turns out to be a way of running two commits for one element. Threading the pass signal into the download was also a mistake, in the other direction. The provider call that produced the URL has already been billed, so cancelling the download throws away work that is paid for — and the shared proxy cache records a cancelled request as a transient failure against that URL, which after three of them blocks it for every consumer in the session. Browser-only mode never asked for this: it had no way to observe an abort there, which is exactly why the bytes were kept. The signal is gone from the download again, and `fetchAsBlob` is byte-for-byte what it was before this branch. The regression guard for the stranded-element rule is restored alongside the timing test that was meant to supersede it. It catches a different rule — a task left pending being read as answered — and nothing else does: making the pass skip pending leaves every other suite green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): stop asking the pool for refs it never issued, bound it, and adopt cached bytes Four things a deployment found once this was running for real. A reference this application mints itself — a generation placeholder, a derived narration key — was never in the pool, because the pool allocates every id it holds. Asking anyway used to be an IndexedDB miss; once the pool is server-backed it is a request that answers 404, one per element per load, forever on a course that still holds placeholders. Every lease and probe now checks first. The check is a negative test on shapes this application owns, not an id validator: the pool's id domain stays unconstrained, and anything that is not one of ours is still asked about. The asset store can bound how much one principal holds, and enforces it inside the write transaction, but nothing ever passed the number. It does now, with a default rather than an opt-in: allocation is reachable by any caller a deployment admits, and with one shared principal an unbounded store is unbounded database growth with no operator-visible brake. Refusing asset mutations to unauthenticated callers was not enough, because every authenticated caller resolves to that same shared principal — so authentication decided nothing, and any signed-in visitor could delete any id they learned. Since this branch began storing media the registry is the only copy a course has. Replacing and deleting are now refused to everyone, and the browser no longer tries: an entry nothing references waits for server-side reclamation instead. What a browser must still do is forget its own record of an allocation that reached nothing, or a later save would stamp an id the document has no reason to trust. And a course generated before any of this holds placeholders in its document with its bytes only in the author's browser. Those bytes are paid for, so the author's next load converts them — stored to the pool and written back through the ordinary commit path, with no provider call — instead of buying them again. A row that records only a hosted URL is treated as absent: that URL is the provider's address, not something a document may hold. One renderer expectation moved with this. An untracked placeholder used to paint as pending on first render because asking the pool left a lease in flight; it settled to disabled a moment later either way, and now says so from the start. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): surface a full store as a refusal and convert legacy narration A quota refusal reached the browser as HTTP 500 with a generic message, which reads as a transient failure: the element kept a Retry that would pay a provider again and be refused again. The store raises the contract's own error and the handler maps it to 507, but the store answering a request is not always built by the same bundle as the handler -- the persistence provider is reached from the route bundle and from instrumentation, which is why its state lives on a Symbol.for global -- and `instanceof` is false across that boundary while the declared code is still right. Classify on the code as well as the class, and make the code a permanent, persisted refusal in the browser: recorded locally so it survives a reload, shown as "storage is full", and refused by the retry entry point so a stale button cannot buy a second generation. Every other storage failure stays retryable. Convert what a pre-server-backed course still holds. Generated media is adopted under either key this application has used for it -- the placeholder, and the allocated id of a course converted once and later rolled back -- instead of only the first. Narration is converted by a load-time pass over the open course's speech actions, since nothing re-enters generation for an action that already has an id: bytes to the pool, id written back through a funnel that mirrors the media one, owner-only and server-backed-only. A line whose bytes are in no browser is left alone rather than re-synthesized. Also: the pool guard is now a positive `ast_` test rather than an enumeration of the shapes we mint (imports never reach the pool, so this is safe in both modes); the slide ref collection is an exported pure function so its four lease sites are covered behaviourally; ASSET_QUOTA_BYTES treats every spelling of zero as opting out and refuses a malformed value at startup instead of falling back; the abort signal is re-checked after the cache read, before an uncancellable commit; and the unused `removeAsset` and pool `replace` surfaces are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * chore(storage): release 0.29.1 The asset HTTP handler now recognises a store refusal by the contract code it declares as well as by its class, so a quota refusal raised in another module realm answers 507 instead of 500. Same contract, stricter recognition, no API change: a patch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a full store recoverable and adoption course-safe Narration adoption read the local audio row by its derived key alone. That key carries no stage id and the table is keyed by id alone, so two courses can mint the same one -- a PPTX import numbers its scenes and actions deterministically, which gives every imported deck's first slide `tts_s1_speech-scene-p1`. Locally a collision only means one course plays another's clip in one browser; adopting it wrote that clip into the shared document permanently, for every device and every visitor. A row that names a course is now adopted only into that course, and a row from before that column existed only when the text it recorded is the text of the action being converted. A full asset store was made permanent last round, which was wrong three times over: it overwrote the refused bytes with an empty blob -- on the conversion path that row is a course's only copy of its own media -- it kept sending the rest of the deck to a provider against a ceiling it already knew was reached, and it left no way back once an operator raised that ceiling. A full store is neither the content's fault nor the configuration's, so it is now its own case: the bytes are kept, the pass stops at the first refusal, and the element shows the reason together with a Retry that re-attempts the upload from those bytes. Nothing retries automatically, so no one is re-billed. The narration write-back now reaches the write boundary every producer of a durable write passes through, not only the dirty mark: adoption never deletes the derived row, so a snapshot that reverts the rewrite is adopted again on the next load and allocates a fresh asset every time. Adoption is also mounted by both classroom surfaces rather than one, takes the course's abort signal, and re-validates that this browser still has the course open before each write. ASSET_QUOTA_BYTES is validated from instrumentation, where the README and the docstring already claimed it was: its only other consumer is lazy and memoised, so a malformed ceiling let the process boot and then failed every persistence request, documents and runtime included. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): remember a full store per course, and never lose retained bytes A stopped pass left the elements it never reached as placeholders with no persisted record -- deliberately, since nothing was attempted for them. But that left the next load with no reason not to try: it called a provider for the next placeholder and was refused at exactly the same point, once per reload, indefinitely. A full store is not a property of any slide. It belongs to the deployment and changes for reasons the document knows nothing about, so it is now remembered once per course in the browser's device KV. A pass that finds the marker stands down before spending anything and leaves every placeholder its "storage is full" state and its Retry; the first upload that succeeds clears it and the next pass runs normally. Narration adoption latched per course so it runs once per load, and the latch outlived the abort that leaving a course performs. On a surface that stays mounted across switches -- the workbench pane is one component for every course it shows -- owner course A, visitor course B, then back to A skipped exactly the clips the abort had cut off, and nothing else converts them. The latch is released with the abort now, and a course adopts one run at a time so a re-entry cannot hand a clip a second allocation while the previous run's uncancellable tail is still settling. A quota-blocked element retried into a network error or a 500 lost the bytes that were kept for it: the retry deleted the row before attempting the upload and wrote no replacement for an error carrying no structured code, so the next retry went back to a provider for media this browser had a moment earlier. The row now survives until an upload succeeds, the failure handler keeps whatever bytes the attempt was given, and the retry asks the question a pass asks -- does this browser already hold bytes for this element -- rather than reading an error code that a second failure has already overwritten. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): adopt real legacy narration, queue re-entries, report attempt outcomes Narration adoption admitted a stage-less row only when the text it recorded matched the action being converted. Both of those columns were added to the local audio table by the very change that moved narration onto allocated ids, so a row still carrying a derived key has neither: the rule refused every real pre-allocation course and passed only on fixtures built from post-allocation rows. What the row cannot say, the key can. A derived key names two clips only when two courses share a scene order and an action id, and an action id repeats only when something other than the generator minted it -- an import numbers them by slide position. So a key built from a generated action id is adopted on that basis, a key an import could have reproduced still needs matching text, and a row that names another course is refused however unique its key looks. Handing a re-entering caller the adoption run already in flight undid the latch release it was paired with: that run is bound to the signal the departure just aborted, so it stops at its next clip while the caller -- which has the course open and a live signal -- is told the work is done, and an effect replayed as mount, cleanup, mount adopts nothing at all. A later caller now waits for the uncancellable tail and scans again, which costs a lookup on a course that has nothing left and finishes the clips the abort cut off on one that does. One attempt at an element now reports both facts its callers need instead of a bare boolean: whether the store refused it for room, and whether bytes actually reached the store. Leaving a course clears the task table, so a retry that landed afterwards read "no failed task" as success and deleted the row holding the only copy of the media. Nothing is inferred from that table any more. Reading the localStorage property can throw where storage is denied by policy, typeof included, so the availability check moved inside the guard: this metadata is best-effort, and a rejection here strands a generation pass that has already enqueued its tasks. A retry is never blocked by the per-course "store is full" marker, but a retry that is refused again re-sets it, and adoption now reads and writes the same marker rather than issuing one refused upload per clip on every load. The two canvas element renderers and both thumbnail renderers show the reason beside the Retry, so a full store does not look like an ordinary failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): probe a full store instead of standing down, and pair notices with a Retry Narration adoption was given both halves of the per-course "the store is full" marker last round: it stood down when the marker was set, and it set the marker when its own upload was refused for room. Those halves are only safe together if something can lift the marker, and for adoption nothing could. It has no affordance of its own, it stood down before reaching its own clear, the media pass returns before its marker gate when there is nothing to generate -- so a narration-only deck, or one whose slides are already satisfied, painted no storage-full element and offered no Retry -- and narration generated rather than adopted allocates directly rather than through the media commit. The course's cached narration was then lost for good, where before it converted on the first load after the ceiling was raised. The gate is a probe now. A marked course attempts exactly one clip per load: refused, it stops and the marker stands, which costs what standing down cost; stored, it lifts the marker and finishes the course. Adoption spends no provider money, so the whole cost of probing a store that is still full is one refused upload. Generated narration lifts the marker too. The three surfaces that gained a failure notice last round drew it for any failure with a reason, including the one refusal that is reachable without server-backed persistence, so a browser-only deck painted something it had not painted before. The notice is drawn beside a Retry and nowhere else, which is what it was added for and what leaves browser-only output unchanged. Both are now asserted through the render harness the surface matrix already had. A caller arriving while a rescan is queued shares it rather than appending another. One rescan converts whatever the run in flight left and every later one would find an allocated id on every action, so a chain bought nothing and turned a single stalled upload into a course that never adopts again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): treat a refusal for room as a fact about one clip, not the deck The asset store checks each write against the headroom it has left, so a store that refuses a long opening clip can still hold every short clip behind it. Narration adoption assumed the opposite: it broke the deck at the first refusal and then re-attempted that same first clip on every later load, because the document names it first. A deck whose longest clip exceeds current headroom therefore never converted the clips that would have fit, with no affordance to recover it -- the state the probe was introduced to remove, reached through a narrower door. An unmarked load now attempts every clip, skipping the ones that do not fit, and remembers the condition only if the load ends with clips it still could not store. A marked load spends its single upload on the smallest clip left rather than the first one named: that is the clip that answers the question the marker asks, because if the smallest does not fit nothing does. The media pass keeps stopping at its first refusal, and for a reason adoption does not share -- every element it attempts costs a provider call. A rescan several callers share took the newest caller's signal, and the newest caller is not necessarily the one still there: a surface that opened a course and closed it again would stop work a surface still showing that course was waiting for, and that surface is latched, so it would never ask again. The shared run now takes a signal that is aborted only once every caller has left. The comment claiming the shared rescan contains a stalled upload was wrong -- the rescan is chained off the run in flight, so a stalled upload leaves every caller pending exactly as a chain would. It claims the bounded queue it actually provides, and the stall is recorded as a limitation. The failed-state containers took their stacking classes unconditionally, so markup differed in browser-only mode even though nothing moved on screen. Those classes are applied only when there is a notice to stack, and the tests assert the exact class attribute rather than a substring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): stop narration adoption writing the media pass's store-full marker The marker means "do not call a provider for this course". A path is entitled to write it only if its own refusal cost a provider call, and narration adoption's refusals cost nothing: it uploads bytes this browser already holds. The store also checks each write against the headroom it has left, so a clip that does not fit says nothing about whether a slide's image would. Adoption was writing it anyway, and one over-long narration clip was therefore enough to stand a course's entire image pass down on every later load -- on a store that had just accepted adoption's other clips. The author could still recover each element by hand, every load, for ever. Three rounds of narrowing this seam produced a finding each time, so it is removed rather than narrowed again. Gone: the marker read, the single-clip probe, the smallest-clip selection, and the up-front read of every row into an array -- which also retires a sampled-then-stale flag and the retention of a whole deck's blobs for the length of a run, and returns the loop to streaming one row at a time. Adoption's rule is now that every load attempts every clip it holds, once; any failure skips that clip and the load continues. The noise the coupling was meant to avoid does not arise, because after the first load the clips still outstanding are exactly the ones that did not fit -- normally none, or one. A successful write still clears the marker, and that is a different kind of statement: a write that went through is a fact this run established, where a refusal is an inference about what some other write would cost. For a course whose media needs nothing, adoption and generated narration are also the only paths that can establish it. The failure module still documented the deck-wide premise this contradicts. It now says what is true: the check is per write, and the media pass stops the deck as a judgement about cost rather than about certainty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bound a full store's cost from the store's own arithmetic Removing the store-full marker from narration adoption removed its bound too, and the code then asserted the bound was unnecessary. It is, on a store with room for most of a deck. On the store the whole mechanism exists for -- the ceiling reached, nothing fitting -- the outstanding set after every load is the entire deck, so a thirty-clip course posted thirty full blobs on every load, indefinitely. Each of those is not a cheap refusal: the bytes are uploaded, the server hashes the whole payload, and only then takes a per-principal lock and sums every entry that principal owns before saying no. The bound needs no flag, no key and nothing carried between loads. The store asks whether `used + addedBytes` exceeds the ceiling, and `used` only grows while a run is uploading, so a clip refused for want of room implies every clip at least that large is refused for the rest of that run. The run keeps the smallest size it has been refused and skips anything no smaller without uploading it; a smaller clip is still attempted, because it may fit. A deck the store refuses entirely now costs one upload per successive size minimum instead of one per clip, and a deck it has room for costs nothing extra, because nothing is refused. Only a refusal for room lowers the bar: a dropped connection says nothing about how much room there is. The deck-wide certainty premise the failure module retracted last round still stood verbatim at the site that implements the stand-down. Both copies now say the same thing: the check is per write, and the pass stops the deck as a judgement about cost rather than about certainty. The comment on adoption's marker clear now names its price. Narration of a few hundred bytes fits in headroom an image does not, so a proven write can let the next pass buy one more image that is refused again -- bounded at one, and the price of the alternative being a course whose media never generates again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(media): state the adoption bound exactly, and stop three comments describing the old rule The comment introducing the in-load bound gave its cost as "at most a handful, and the first load pays the most". Neither clause is a property of the rule. A clip is skipped only when something no larger was already refused, so a fully-refused deck costs one upload per successive size minimum in document order: one when the clips grow, about ln N for an arbitrary order, and one per clip when they only shrink -- a long opener followed by terser lines is exactly that shape. And no load is cheaper than the first, because the bound resets per run and a refused clip stays outstanding. The comment now says that, and points at what would make it exactly one for any ordering: the store returning its remaining headroom in the refusal's existing details channel, which the server leaves empty today. Two other comments still described the previous rule -- "attempts every clip it holds, every load" -- one of them twenty lines above the paragraph that introduces the bound, in the same block. Both now say what the code does. The bound's soundness is worth stating where a maintainer will look for it: quota is charged at full length with no discount for a duplicate, the sum it is checked against joins entries to blobs so the collector cannot lower it, the check takes a per-principal lock before summing, and replace and delete are refused to every browser. Nothing a run can do makes room appear inside it. One test installed a row implementation and replaced it wholesale a few lines later, so the first was dead and the survivor dropped the text the first clip's import-shaped key needs for the ownership rule -- it passed on the coincidence that the fixture's default text is the action's. Merged into one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): keep a Retry from re-buying parked media, and state the store seam once Five findings from an inline review. The standalone classroom route asked the ownership sidecar once per load and recorded only stage ownership when that ask failed. Every non-answer fails closed, so one transient 5xx left the genuine author with no resume, no Retry affordance and no legacy narration converted for the rest of the load, with nothing to change it short of a reload. The failure now records the fail-closed answer explicitly -- an answer an earlier load established must not outlive the failure that replaced it -- and an unresolved answer is asked for again, a few times over a few seconds. A real answer, however unwelcome, is final. A Retry could pay a provider for media the pool already held. When the bytes are stored and only the write-back fails in a way that keeps the allocation, it is parked and no local row exists, because that row is written only after a successful write-back. Retry now reads the parked queue exactly as the pass does and re-attempts the write-back: it re-keys the task done when the document takes it, leaves the entry parked when the slide still does not exist, and stays failed and retryable when the document refuses again. Object URLs a parked allocation owns are revoked when the entry is dropped. The commit path leaves them alone while the entry is parked, because it is then the only thing holding bytes this tab can render, so a course switch or a stage deletion was pinning the whole blob for the life of the tab. An entry a slide has already taken is left alone: the task table is displaying those URLs. The fallback lookup for cached bytes is a stage-scoped scan, and the keyed lookup misses for every row the commit path writes, so a pass was materializing and sorting the course's whole media table once per element. One scan per pass now, built on the first miss. It is sound and not merely cheaper: an element asks only for its own placeholder, and every row a pass writes carries the placeholder of the element that wrote it. "The store accepted a write, so it is not out of room" was enforced at three call sites under slightly different conditions, which made it a convention the next pool write path could silently break. It is stated once, in putAsset, for the course whose bytes it just stored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 16 天前 | |
fix(tests): make media tests pass on Windows path separators (#1615) | 7 天前 | |
refactor(media): one client-side pool commit primitive; keep refused narration instead of re-billing it (#1523) Extracts commitToPool, the single client-side sequence for storing bytes in the asset pool, writing the allocated id back, and mirroring locally; routes the media pass, narration adoption and fresh TTS through it. A store-full refusal during TTS now retains the already-billed clip so the next load adopts it with zero provider calls. Closes #1467. | 12 天前 | |
feat(media): write generated media through the asset pool under server-backed persistence (#1392) * feat(media): store generated media in the asset pool when persistence is server-backed With server-backed persistence the document is durable and shared, but generated media stayed in the producing browser: the document kept its gen_img_* / gen_vid_* placeholder and narration kept a browser-derived audio id. Every new browser that opened such a course re-ran generation for every slide, and it never converged, because the address of the generated bytes was never written back into the document. Under server-backed persistence only, the classic generation chain now stores bytes in the asset pool first and writes the id the pool allocated into the document. - The client bootstrap configures the asset seam alongside the document and runtime seams: an HttpAssetStore over the persistence endpoint carrying the same credentials the document store carries, marked server-backed. The seam preflight now covers all three, so a failure still cannot half-configure persistence. - Image, video and TTS generation commit in one fixed order: provider, pool, document, local cache, task. A reference reaches the document only after put returned an id, so a document can never name bytes that were not stored. A failure before the write-back leaves the placeholder with the provider called exactly once; the retry happens on the next owner load. - The write-back is a per-slot rewrite through mutateDocument, which re-reads the current document under the per-stage lock, so it cannot clobber a newer scene. The open course is refreshed with the same rewrite without being marked dirty. - "Has this already been generated?" is answered by the document (the slide exists and no longer holds the placeholder) instead of by this browser's task table. - The classroom's resume effect fails closed on ownership: only a resolved owner starts generation, so a viewer opening a shared course spends nothing. - The local media and audio tables become a per-tab cache. A failed cache write costs a re-download, never the media. Browser-only mode is unchanged: every new call site sits behind the server-backed gate, the local tables stay authoritative there, and placeholders stay in the document. Rendering and export needed no changes. HttpAssetStore.resolve mints an object URL exactly as the browser store does, and the export byte resolver was already pool-first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(classroom): make the generation owner gate a three-outcome rule and apply it everywhere The gate refused everything but a resolved owner, which read a sidecar that answered "no ownership fact exists for this course" as a reason to block. That is the answer a deployment without the sidecar's server-side prerequisites gives for every course, and the answer a course with no ownership record gives: in both, there is nobody the operator's budget needs protecting from, and refusing strands the course's own author behind a question that can never be answered. Ownership is now four states over the sidecar's three outcomes. A definite answer splits into owner and not-owner. An absent record is its own answer, ownerless, and generation proceeds — the behaviour such a deployment had before the gate existed. Only the absence of an answer, a transport failure or a load that has not asked yet, stays unresolved and fails closed: "we could not ask" must never be read as "nobody owns this". One mapper turns a sidecar result into that state, and one predicate decides on it. The workbench classroom pane runs the same resume effect and had no ownership input at all, so a viewer opening a shared course there could still spend the budget. It now asks the sidecar once per course, in parallel with its load and feeding only the generation gate, so its read-only and edit behaviour is unchanged. The shared progressive-load policy carries the gate for it, with both new inputs required rather than defaulted so a future caller cannot omit them into an open budget. Its stale comment claiming ownership could not be expressed here is corrected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the write-back survive autosave, arrive before the scene does, and never leak Independent reviews of the write-back found three ways a durable document could still end up naming a placeholder, and two ways the gate that protects the operator's budget could be walked around. An autosave round captures the store synchronously and writes that capture, so a round already in flight when a rewrite landed wrote the placeholder straight back over the allocated id, and nothing marked the store dirty again to correct it. The rewrite now marks the units it changed, which leaves a corrective flush queued behind the stale one; re-saving a scene that already holds the id is idempotent, losing the id is not. Media is generated from outlines in parallel with scene content and usually finishes first, so the slide that will carry the placeholder does not exist yet and the write-back has nothing to rewrite. That was the ordinary path, not a tail case, and its result was discarded: the task was marked done, the scene was added afterwards with its placeholder intact, and a second pass in the same run could call the provider again. The allocation is now held under the placeholder — which also answers the skip test, so nothing pays twice — and applied when that scene is committed, before its first save. One complete pass now leaves no placeholder behind. A failed commit used to abandon what it had already allocated. A poster upload that failed threw away a stored video and sent the retry to submit the most expensive job in the system again; a rejected write-back left registry rows that name bytes nothing references, which the byte collector cannot reclaim because it only collects blobs no row names. A poster failure now costs the poster, and a write-back that reached nothing reclaims what it allocated. A partial write is left alone, because the document already names it. The ownership gate is fail-closed again. Treating the sidecar's 404 as permission was wrong: the client cannot tell "this course has no owner" from "this deployment told me nothing", so a visitor who opened a shared course could bill the operator. The root cause was the sidecar itself, which gated on the agent runtime although every persisted course has an owner regardless — the persistence route resolves one for every request. It now gates on server persistence, so the configuration that made 404 the universal answer has real ownership facts to report, and the gate can refuse everything but a named owner. Retry affordances answered to no gate at all. A viewer of a shared course with one failed image was shown a Retry button that called the provider. Both retry entry points and every surface that draws them now read one shared permission, so what is offered and what is allowed are the same value. Also: narration regeneration no longer pretends it can replace bytes behind a live id — the exclusivity proof that would allow it is refused by construction once references leave the browser, so it forks to a fresh id and says so; the "already generated" test lets a finished deck answer from the document alone, since scene order stops identifying an outline once slides are inserted or deleted; stored assets record a specific media type rather than a generic transfer type; the pane no longer asks the sidecar in browser-only mode; and the funnel's docstring now states what the per-stage lock actually guarantees, which is same-browser serialization and not a cross-browser compare-and-swap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): park allocations in the deciding turn, never reclaim on an ambiguous write A delta review of the write-back found the first-pass fix still had a window, and the reclamation it added could delete media the document already names. The allocation was parked after an awaited local cache write. A scene committed in that window reconciled against a registry that did not hold it yet, so the document kept the placeholder — and the entry recorded a moment later then answered the skip test as "already handled", so nothing could correct it. Parking now happens inside the write-back, in the same synchronous turn as the decision that nothing could take the reference; no await separates the live check from the park. Allocations parked by an earlier pass are handed to their slides at the start of the next one, before anything decides what still needs generating, so a held allocation whose scene has since arrived becomes a rewrite rather than an answer. Reclaiming on a rejected write was unsound: a rejection does not prove the server did not apply the write, so deleting the asset could break the scene that now names it. The funnel decides instead, and says so: it reclaims only when no store write was ever issued and nothing took the reference. Anything else is placed if its slide exists and parked if it does not, so the next pass reuses the bytes instead of paying for them again. When a write fails after part of it landed, the live store is brought up to the document before the error is rethrown — otherwise the next ordinary flush would overwrite the half that did land, with the ids deliberately not reclaimed. Parked allocations are now cleared with the course. Classic placeholders are reused across runs, so one surviving an interrupted run would be handed to a different slide of the next deck: the previous picture, on a slide whose provider was never asked. Both classroom surfaces clear the arriving course, the deletion cascade clears the deleted one, and clearing the database clears them all. Two more ways generation could start without asking the gate are closed. An overlapping pass — an outline retry re-enters generation with every outline while the first is still working — re-requested elements whose provider call was already in flight; a task that is not done is an answered request, not an unanswered one. And narration regeneration in the timeline editor called the TTS provider and allocated a pool asset with no ownership check at all; it now reads the same permission, which withholds both the per-line and whole-timeline controls and refuses the call. Finally, a pane opened during the stage-link availability gap recorded the sidecar's 404 for a course that was moments from existing and never asked again, leaving the real owner locked out of generation until it remounted. Ownership is re-fetched once the document becomes available; the gate stays closed until an answer arrives, so asking again can only open it for someone entitled to it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the asset routes reachable, and stale snapshots harmless A full-branch audit found that the deployment this project documents could not store a single generated asset, and that several routes into durable storage could still write a placeholder over a reference that had already landed. The persistence route sent asset requests through the development authenticator, which refuses outright in a production build that has not explicitly opted into it — and the documented server-persistence recipe produces exactly that build. Every store and every read answered 401, so images and video failed on every slide while re-billing the provider on each retry, and a narration failure stopped the deck at its first slide. Assets live in one shared partition by design, so there was never anything per-caller for that authenticator to decide: the route now resolves the asset principal itself, alongside the owner it already resolves for documents. Runtime sessions are genuinely per-learner and keep the development authenticator until real session verification replaces it. And narration that cannot be stored no longer fails its scene: the line stays unvoiced and retryable, which is what an image that cannot be stored does to its slide. Placeholders could also come back from behind. A queued autosave's snapshot, an editor-history entry replayed by an undo, the departing save a course switch flushes — each captures content at its own moment, and any of those moments can predate a write-back. Point fixes at each producer would leave the next producer to rediscover the bug, so the check lives at the write boundary every producer passes through, and the allocation record it consults now outlives the parked queue: a placeholder whose rewrite landed long ago is exactly the case it catches. Two ways generation could be lost or repeated are closed. A pass now claims the elements it will reach and releases them however it ends, so an overlapping pass stands down while an aborted one strands nothing — previously its tasks stayed `pending` and every later pass skipped them with no retry control to recover them. And the media abort controller is aborted before being replaced, so a superseded pass stops calling providers instead of running on for a course the user has left. The remaining two are narrower. The workbench pane asks for ownership only after a document load succeeds, and after every later one, mirroring the page route: the load is what creates the ownership row the first time a course is opened, so asking beforehand asked about a course that did not exist yet and locked its author out for the mount. And the ownership gate on the timeline editor now withholds narration regeneration alone; listening back to existing narration and seeing whether a line has any spend nothing and stay available. Known limitation, unchanged and now stated plainly in the comments that used to point at it as a solution: nothing reclaims an unreferenced pool asset. The registry sweep is written but not wired up, and the byte collector only reclaims blobs no registry row names, so every narration regeneration and every abandoned allocation leaves storage behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): gate asset mutations, and make claims and allocation records survive a handoff Opening the asset routes opened all of them. Reads and allocations are meant to be as open as document reads and creates already are, but no authorization hook was supplied, so the handler's default admitted PUT and DELETE too — and those scope by principal key alone, which is one shared constant. Any caller who learned an id, and a document read hands out every id its slides name, could overwrite or destroy another author's media. Mutations now require the deployment's credential, which in a production build without the development-auth opt-in means they are refused outright; reads and allocations stay open. The route comment says what the posture is and what it is not: the deployment-level fence is the access code, and no per-principal quota is configured. The client's own reclaim is best effort to match — losing an argument about deleting an asset must not cost a task its retry, and the bytes are left for server-side reclamation. The pass claim could not survive the handoff it was written for. A retry aborts the live media pass and starts its replacement in the same synchronous block, long before the aborted pass's cleanup runs, so the replacement saw every element still claimed, collected nothing, and returned — leaving each unreached element at pending with nobody coming back for it and no retry control to recover it, which is the exact failure the claim was introduced to prevent. A claim now carries its pass's signal and is retired the moment that signal aborts, and a pass releases only claims it still owns, so a late unwind cannot take its replacement's work. Claims are also acquired at the single point every request passes through, so a single-task retry participates too — previously a retry awaiting its provider was invisible to a pass starting alongside it and both called it. The allocation record could outlive the bytes it named. It was written before the write-back attempted anything and survived the reclaim that followed a failure, so when the slide finally arrived the write boundary stamped a deleted id into the document — and the placeholder it replaced was gone, which reads as already generated and stops anything from retrying. The record is now written only where the allocation is retained, and forgotten wherever a reclaim removes the bytes, including the narration rollback path. The tests follow. The route test drives the real storage handler against an in-memory registry instead of a stub, so it can see what the resolved principal is then allowed to do; the handoff test performs a real abort mid-pass rather than starting from an already-aborted signal; and the guards that could only assert file layout now assert the property they care about, or have been replaced by behaviour. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make media passes serial per course instead of tracking element ownership Three rounds of per-element claims each produced a new way to lose an element. Whole-pass reservations swallowed a Retry for an element the same pass had already failed, leaving it pending with the affordance gone. Retiring a claim by its signal freed an element whose commit was still uploading, so the replacement pass paid for it twice. A claim held for a failed element stranded its retry. The bookkeeping is the defect: every refinement of "who owns this element right now" answered the question at a moment when the answer was already stale. Passes for one course are now serial. A replacement aborts its predecessor, as before, and then waits for it to settle before collecting. That removes the question entirely: a commit already under way finishes — its bytes stored and its reference written, so the new pass sees a resolved slide and skips it — and an element the aborted pass never reached is still a placeholder and gets collected like any other. The claim set, the reservations, the signal retirement and the identity-checked release are all gone. The task table is consulted for one thing only: an element that is generating right now is a single-element retry running alongside the pass, and taking it too would pay twice. Pending is deliberately not a skip reason — it means a pass once intended to reach an element, which an abandoned pass leaves behind with nobody acting on it, and reading that as answered is what stranded elements before. A retry runs concurrently with a pass, because a pass never revisits an element it has processed, and it re-reads the task after its own await and refuses before touching it: marking first and refusing afterwards destroyed the failed state that draws the affordance. Browser-only mode is back to exactly what it was. The abort is now conditional, the waiting does not apply, and the original status-based skip is restored verbatim. Two baseline lines remain changed in each of the two files, and both are behind a server-backed fork whose else-branch is the original. Two smaller things. The allocation record becomes visible when a write goes on the wire rather than when the round trip ends, and the write boundary reconciles under the document lock rather than before it — a save queued during a write-back was otherwise captured with the placeholder and, for a course the user had left, had no corrective flush to follow. And the comments that said a refused reclaim leaves its bytes for server-side reclamation were wrong: nothing collects them, because the registry entry still names its blob and the sweep that would remove it is not wired up. They now say the bytes leak. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a deferred pass re-earn its right to run, and bound the commit it waits on Serializing passes moved their body out of the block that launched them, and three things followed from that. A pass now wakes when its predecessor settles, which can be after the user has left the course. It enqueued before it looked at its signal, into a task table keyed by element id alone — and placeholder ids are not unique across courses, which is why the classroom clears that table on arrival. So a departing course's pass seeded the arriving course's table with tasks carrying the wrong stage id, and a Retry routes by that id: the reference went into the wrong document. A pass now re-validates after the wait, before touching anything shared. The same lateness broke the skip test. `documentSkipIndex` answers only while the live store is on the pass's stage, and returning nothing put the collection loop on the browser-only rule — a silent demotion from "the document is the authority" to "this browser's task table is", on exactly the path where that table has just been cleared. Every element the predecessor had committed was collected again, paid for again, and its second write-back found no placeholder to rewrite, so its bytes were parked where nothing will ever reference them. In server-backed mode an unreadable document now means the pass stands down. And waiting was unbounded. A commit is uncancellable: the asset client takes no signal, and a document write cannot be half-undone. One stalled upload therefore froze the course's media generation for the session — the replacement never collected, the element sat on a skeleton that draws no Retry, and only a reload recovered. The pass's signal is now threaded into the media proxy fetch, and the commit is bounded by a deadline. The deadline is on the wait, not the work: the commit carries on, and if it lands late the document simply ends up correct, while the element becomes retryable and the queue moves on. The tests that were meant to pin the previous round were not sensitive to it. Two asserted end states where the mechanism only changes ordering, and one of them rigged the document read so the assertion held whether or not the pass had waited; a third covered half of what it claimed. They now observe the ordering directly — nothing is issued while another pass for the course is working; in browser-only mode a second pass reaches its provider immediately — and the reconciliation under the document lock has a test that fails when it moves back outside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * revert(media): drop the commit deadline and the abortable download The deadline bought less than it cost. Abandoning a commit after two minutes makes the element retryable while the real commit is still running, so a Retry starts a second commit for the same placeholder against the first: two provider calls, two allocations, and whichever lands second stamps its result over the other's task by element id. The allocation record is keyed by placeholder, so the loser's cleanup erases the winner's record, and the write boundary then puts the raw placeholder back into the document. That is the overlap serial passes were built to remove, reopened through the one door serialization never covered. So a stalled commit holds the course's media queue until it settles or the page is reloaded, and that is written down rather than papered over. The wait is unbounded on purpose: every ceiling on it turns out to be a way of running two commits for one element. Threading the pass signal into the download was also a mistake, in the other direction. The provider call that produced the URL has already been billed, so cancelling the download throws away work that is paid for — and the shared proxy cache records a cancelled request as a transient failure against that URL, which after three of them blocks it for every consumer in the session. Browser-only mode never asked for this: it had no way to observe an abort there, which is exactly why the bytes were kept. The signal is gone from the download again, and `fetchAsBlob` is byte-for-byte what it was before this branch. The regression guard for the stranded-element rule is restored alongside the timing test that was meant to supersede it. It catches a different rule — a task left pending being read as answered — and nothing else does: making the pass skip pending leaves every other suite green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): stop asking the pool for refs it never issued, bound it, and adopt cached bytes Four things a deployment found once this was running for real. A reference this application mints itself — a generation placeholder, a derived narration key — was never in the pool, because the pool allocates every id it holds. Asking anyway used to be an IndexedDB miss; once the pool is server-backed it is a request that answers 404, one per element per load, forever on a course that still holds placeholders. Every lease and probe now checks first. The check is a negative test on shapes this application owns, not an id validator: the pool's id domain stays unconstrained, and anything that is not one of ours is still asked about. The asset store can bound how much one principal holds, and enforces it inside the write transaction, but nothing ever passed the number. It does now, with a default rather than an opt-in: allocation is reachable by any caller a deployment admits, and with one shared principal an unbounded store is unbounded database growth with no operator-visible brake. Refusing asset mutations to unauthenticated callers was not enough, because every authenticated caller resolves to that same shared principal — so authentication decided nothing, and any signed-in visitor could delete any id they learned. Since this branch began storing media the registry is the only copy a course has. Replacing and deleting are now refused to everyone, and the browser no longer tries: an entry nothing references waits for server-side reclamation instead. What a browser must still do is forget its own record of an allocation that reached nothing, or a later save would stamp an id the document has no reason to trust. And a course generated before any of this holds placeholders in its document with its bytes only in the author's browser. Those bytes are paid for, so the author's next load converts them — stored to the pool and written back through the ordinary commit path, with no provider call — instead of buying them again. A row that records only a hosted URL is treated as absent: that URL is the provider's address, not something a document may hold. One renderer expectation moved with this. An untracked placeholder used to paint as pending on first render because asking the pool left a lease in flight; it settled to disabled a moment later either way, and now says so from the start. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): surface a full store as a refusal and convert legacy narration A quota refusal reached the browser as HTTP 500 with a generic message, which reads as a transient failure: the element kept a Retry that would pay a provider again and be refused again. The store raises the contract's own error and the handler maps it to 507, but the store answering a request is not always built by the same bundle as the handler -- the persistence provider is reached from the route bundle and from instrumentation, which is why its state lives on a Symbol.for global -- and `instanceof` is false across that boundary while the declared code is still right. Classify on the code as well as the class, and make the code a permanent, persisted refusal in the browser: recorded locally so it survives a reload, shown as "storage is full", and refused by the retry entry point so a stale button cannot buy a second generation. Every other storage failure stays retryable. Convert what a pre-server-backed course still holds. Generated media is adopted under either key this application has used for it -- the placeholder, and the allocated id of a course converted once and later rolled back -- instead of only the first. Narration is converted by a load-time pass over the open course's speech actions, since nothing re-enters generation for an action that already has an id: bytes to the pool, id written back through a funnel that mirrors the media one, owner-only and server-backed-only. A line whose bytes are in no browser is left alone rather than re-synthesized. Also: the pool guard is now a positive `ast_` test rather than an enumeration of the shapes we mint (imports never reach the pool, so this is safe in both modes); the slide ref collection is an exported pure function so its four lease sites are covered behaviourally; ASSET_QUOTA_BYTES treats every spelling of zero as opting out and refuses a malformed value at startup instead of falling back; the abort signal is re-checked after the cache read, before an uncancellable commit; and the unused `removeAsset` and pool `replace` surfaces are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * chore(storage): release 0.29.1 The asset HTTP handler now recognises a store refusal by the contract code it declares as well as by its class, so a quota refusal raised in another module realm answers 507 instead of 500. Same contract, stricter recognition, no API change: a patch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a full store recoverable and adoption course-safe Narration adoption read the local audio row by its derived key alone. That key carries no stage id and the table is keyed by id alone, so two courses can mint the same one -- a PPTX import numbers its scenes and actions deterministically, which gives every imported deck's first slide `tts_s1_speech-scene-p1`. Locally a collision only means one course plays another's clip in one browser; adopting it wrote that clip into the shared document permanently, for every device and every visitor. A row that names a course is now adopted only into that course, and a row from before that column existed only when the text it recorded is the text of the action being converted. A full asset store was made permanent last round, which was wrong three times over: it overwrote the refused bytes with an empty blob -- on the conversion path that row is a course's only copy of its own media -- it kept sending the rest of the deck to a provider against a ceiling it already knew was reached, and it left no way back once an operator raised that ceiling. A full store is neither the content's fault nor the configuration's, so it is now its own case: the bytes are kept, the pass stops at the first refusal, and the element shows the reason together with a Retry that re-attempts the upload from those bytes. Nothing retries automatically, so no one is re-billed. The narration write-back now reaches the write boundary every producer of a durable write passes through, not only the dirty mark: adoption never deletes the derived row, so a snapshot that reverts the rewrite is adopted again on the next load and allocates a fresh asset every time. Adoption is also mounted by both classroom surfaces rather than one, takes the course's abort signal, and re-validates that this browser still has the course open before each write. ASSET_QUOTA_BYTES is validated from instrumentation, where the README and the docstring already claimed it was: its only other consumer is lazy and memoised, so a malformed ceiling let the process boot and then failed every persistence request, documents and runtime included. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): remember a full store per course, and never lose retained bytes A stopped pass left the elements it never reached as placeholders with no persisted record -- deliberately, since nothing was attempted for them. But that left the next load with no reason not to try: it called a provider for the next placeholder and was refused at exactly the same point, once per reload, indefinitely. A full store is not a property of any slide. It belongs to the deployment and changes for reasons the document knows nothing about, so it is now remembered once per course in the browser's device KV. A pass that finds the marker stands down before spending anything and leaves every placeholder its "storage is full" state and its Retry; the first upload that succeeds clears it and the next pass runs normally. Narration adoption latched per course so it runs once per load, and the latch outlived the abort that leaving a course performs. On a surface that stays mounted across switches -- the workbench pane is one component for every course it shows -- owner course A, visitor course B, then back to A skipped exactly the clips the abort had cut off, and nothing else converts them. The latch is released with the abort now, and a course adopts one run at a time so a re-entry cannot hand a clip a second allocation while the previous run's uncancellable tail is still settling. A quota-blocked element retried into a network error or a 500 lost the bytes that were kept for it: the retry deleted the row before attempting the upload and wrote no replacement for an error carrying no structured code, so the next retry went back to a provider for media this browser had a moment earlier. The row now survives until an upload succeeds, the failure handler keeps whatever bytes the attempt was given, and the retry asks the question a pass asks -- does this browser already hold bytes for this element -- rather than reading an error code that a second failure has already overwritten. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): adopt real legacy narration, queue re-entries, report attempt outcomes Narration adoption admitted a stage-less row only when the text it recorded matched the action being converted. Both of those columns were added to the local audio table by the very change that moved narration onto allocated ids, so a row still carrying a derived key has neither: the rule refused every real pre-allocation course and passed only on fixtures built from post-allocation rows. What the row cannot say, the key can. A derived key names two clips only when two courses share a scene order and an action id, and an action id repeats only when something other than the generator minted it -- an import numbers them by slide position. So a key built from a generated action id is adopted on that basis, a key an import could have reproduced still needs matching text, and a row that names another course is refused however unique its key looks. Handing a re-entering caller the adoption run already in flight undid the latch release it was paired with: that run is bound to the signal the departure just aborted, so it stops at its next clip while the caller -- which has the course open and a live signal -- is told the work is done, and an effect replayed as mount, cleanup, mount adopts nothing at all. A later caller now waits for the uncancellable tail and scans again, which costs a lookup on a course that has nothing left and finishes the clips the abort cut off on one that does. One attempt at an element now reports both facts its callers need instead of a bare boolean: whether the store refused it for room, and whether bytes actually reached the store. Leaving a course clears the task table, so a retry that landed afterwards read "no failed task" as success and deleted the row holding the only copy of the media. Nothing is inferred from that table any more. Reading the localStorage property can throw where storage is denied by policy, typeof included, so the availability check moved inside the guard: this metadata is best-effort, and a rejection here strands a generation pass that has already enqueued its tasks. A retry is never blocked by the per-course "store is full" marker, but a retry that is refused again re-sets it, and adoption now reads and writes the same marker rather than issuing one refused upload per clip on every load. The two canvas element renderers and both thumbnail renderers show the reason beside the Retry, so a full store does not look like an ordinary failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): probe a full store instead of standing down, and pair notices with a Retry Narration adoption was given both halves of the per-course "the store is full" marker last round: it stood down when the marker was set, and it set the marker when its own upload was refused for room. Those halves are only safe together if something can lift the marker, and for adoption nothing could. It has no affordance of its own, it stood down before reaching its own clear, the media pass returns before its marker gate when there is nothing to generate -- so a narration-only deck, or one whose slides are already satisfied, painted no storage-full element and offered no Retry -- and narration generated rather than adopted allocates directly rather than through the media commit. The course's cached narration was then lost for good, where before it converted on the first load after the ceiling was raised. The gate is a probe now. A marked course attempts exactly one clip per load: refused, it stops and the marker stands, which costs what standing down cost; stored, it lifts the marker and finishes the course. Adoption spends no provider money, so the whole cost of probing a store that is still full is one refused upload. Generated narration lifts the marker too. The three surfaces that gained a failure notice last round drew it for any failure with a reason, including the one refusal that is reachable without server-backed persistence, so a browser-only deck painted something it had not painted before. The notice is drawn beside a Retry and nowhere else, which is what it was added for and what leaves browser-only output unchanged. Both are now asserted through the render harness the surface matrix already had. A caller arriving while a rescan is queued shares it rather than appending another. One rescan converts whatever the run in flight left and every later one would find an allocated id on every action, so a chain bought nothing and turned a single stalled upload into a course that never adopts again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): treat a refusal for room as a fact about one clip, not the deck The asset store checks each write against the headroom it has left, so a store that refuses a long opening clip can still hold every short clip behind it. Narration adoption assumed the opposite: it broke the deck at the first refusal and then re-attempted that same first clip on every later load, because the document names it first. A deck whose longest clip exceeds current headroom therefore never converted the clips that would have fit, with no affordance to recover it -- the state the probe was introduced to remove, reached through a narrower door. An unmarked load now attempts every clip, skipping the ones that do not fit, and remembers the condition only if the load ends with clips it still could not store. A marked load spends its single upload on the smallest clip left rather than the first one named: that is the clip that answers the question the marker asks, because if the smallest does not fit nothing does. The media pass keeps stopping at its first refusal, and for a reason adoption does not share -- every element it attempts costs a provider call. A rescan several callers share took the newest caller's signal, and the newest caller is not necessarily the one still there: a surface that opened a course and closed it again would stop work a surface still showing that course was waiting for, and that surface is latched, so it would never ask again. The shared run now takes a signal that is aborted only once every caller has left. The comment claiming the shared rescan contains a stalled upload was wrong -- the rescan is chained off the run in flight, so a stalled upload leaves every caller pending exactly as a chain would. It claims the bounded queue it actually provides, and the stall is recorded as a limitation. The failed-state containers took their stacking classes unconditionally, so markup differed in browser-only mode even though nothing moved on screen. Those classes are applied only when there is a notice to stack, and the tests assert the exact class attribute rather than a substring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): stop narration adoption writing the media pass's store-full marker The marker means "do not call a provider for this course". A path is entitled to write it only if its own refusal cost a provider call, and narration adoption's refusals cost nothing: it uploads bytes this browser already holds. The store also checks each write against the headroom it has left, so a clip that does not fit says nothing about whether a slide's image would. Adoption was writing it anyway, and one over-long narration clip was therefore enough to stand a course's entire image pass down on every later load -- on a store that had just accepted adoption's other clips. The author could still recover each element by hand, every load, for ever. Three rounds of narrowing this seam produced a finding each time, so it is removed rather than narrowed again. Gone: the marker read, the single-clip probe, the smallest-clip selection, and the up-front read of every row into an array -- which also retires a sampled-then-stale flag and the retention of a whole deck's blobs for the length of a run, and returns the loop to streaming one row at a time. Adoption's rule is now that every load attempts every clip it holds, once; any failure skips that clip and the load continues. The noise the coupling was meant to avoid does not arise, because after the first load the clips still outstanding are exactly the ones that did not fit -- normally none, or one. A successful write still clears the marker, and that is a different kind of statement: a write that went through is a fact this run established, where a refusal is an inference about what some other write would cost. For a course whose media needs nothing, adoption and generated narration are also the only paths that can establish it. The failure module still documented the deck-wide premise this contradicts. It now says what is true: the check is per write, and the media pass stops the deck as a judgement about cost rather than about certainty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bound a full store's cost from the store's own arithmetic Removing the store-full marker from narration adoption removed its bound too, and the code then asserted the bound was unnecessary. It is, on a store with room for most of a deck. On the store the whole mechanism exists for -- the ceiling reached, nothing fitting -- the outstanding set after every load is the entire deck, so a thirty-clip course posted thirty full blobs on every load, indefinitely. Each of those is not a cheap refusal: the bytes are uploaded, the server hashes the whole payload, and only then takes a per-principal lock and sums every entry that principal owns before saying no. The bound needs no flag, no key and nothing carried between loads. The store asks whether `used + addedBytes` exceeds the ceiling, and `used` only grows while a run is uploading, so a clip refused for want of room implies every clip at least that large is refused for the rest of that run. The run keeps the smallest size it has been refused and skips anything no smaller without uploading it; a smaller clip is still attempted, because it may fit. A deck the store refuses entirely now costs one upload per successive size minimum instead of one per clip, and a deck it has room for costs nothing extra, because nothing is refused. Only a refusal for room lowers the bar: a dropped connection says nothing about how much room there is. The deck-wide certainty premise the failure module retracted last round still stood verbatim at the site that implements the stand-down. Both copies now say the same thing: the check is per write, and the pass stops the deck as a judgement about cost rather than about certainty. The comment on adoption's marker clear now names its price. Narration of a few hundred bytes fits in headroom an image does not, so a proven write can let the next pass buy one more image that is refused again -- bounded at one, and the price of the alternative being a course whose media never generates again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(media): state the adoption bound exactly, and stop three comments describing the old rule The comment introducing the in-load bound gave its cost as "at most a handful, and the first load pays the most". Neither clause is a property of the rule. A clip is skipped only when something no larger was already refused, so a fully-refused deck costs one upload per successive size minimum in document order: one when the clips grow, about ln N for an arbitrary order, and one per clip when they only shrink -- a long opener followed by terser lines is exactly that shape. And no load is cheaper than the first, because the bound resets per run and a refused clip stays outstanding. The comment now says that, and points at what would make it exactly one for any ordering: the store returning its remaining headroom in the refusal's existing details channel, which the server leaves empty today. Two other comments still described the previous rule -- "attempts every clip it holds, every load" -- one of them twenty lines above the paragraph that introduces the bound, in the same block. Both now say what the code does. The bound's soundness is worth stating where a maintainer will look for it: quota is charged at full length with no discount for a duplicate, the sum it is checked against joins entries to blobs so the collector cannot lower it, the check takes a per-principal lock before summing, and replace and delete are refused to every browser. Nothing a run can do makes room appear inside it. One test installed a row implementation and replaced it wholesale a few lines later, so the first was dead and the survivor dropped the text the first clip's import-shaped key needs for the ownership rule -- it passed on the coincidence that the fixture's default text is the action's. Merged into one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): keep a Retry from re-buying parked media, and state the store seam once Five findings from an inline review. The standalone classroom route asked the ownership sidecar once per load and recorded only stage ownership when that ask failed. Every non-answer fails closed, so one transient 5xx left the genuine author with no resume, no Retry affordance and no legacy narration converted for the rest of the load, with nothing to change it short of a reload. The failure now records the fail-closed answer explicitly -- an answer an earlier load established must not outlive the failure that replaced it -- and an unresolved answer is asked for again, a few times over a few seconds. A real answer, however unwelcome, is final. A Retry could pay a provider for media the pool already held. When the bytes are stored and only the write-back fails in a way that keeps the allocation, it is parked and no local row exists, because that row is written only after a successful write-back. Retry now reads the parked queue exactly as the pass does and re-attempts the write-back: it re-keys the task done when the document takes it, leaves the entry parked when the slide still does not exist, and stays failed and retryable when the document refuses again. Object URLs a parked allocation owns are revoked when the entry is dropped. The commit path leaves them alone while the entry is parked, because it is then the only thing holding bytes this tab can render, so a course switch or a stage deletion was pinning the whole blob for the life of the tab. An entry a slide has already taken is left alone: the task table is displaying those URLs. The fallback lookup for cached bytes is a stage-scoped scan, and the keyed lookup misses for every row the commit path writes, so a pass was materializing and sorting the course's whole media table once per element. One scan per pass now, built on the first miss. It is sound and not merely cheaper: an element asks only for its own placeholder, and every row a pass writes carries the placeholder of the element that wrote it. "The store accepted a write, so it is not out of room" was enforced at three call sites under slightly different conditions, which made it a convention the next pool write path could silently break. It is stated once, in putAsset, for the course whose bytes it just stored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 16 天前 | |
feat(media): write generated media through the asset pool under server-backed persistence (#1392) * feat(media): store generated media in the asset pool when persistence is server-backed With server-backed persistence the document is durable and shared, but generated media stayed in the producing browser: the document kept its gen_img_* / gen_vid_* placeholder and narration kept a browser-derived audio id. Every new browser that opened such a course re-ran generation for every slide, and it never converged, because the address of the generated bytes was never written back into the document. Under server-backed persistence only, the classic generation chain now stores bytes in the asset pool first and writes the id the pool allocated into the document. - The client bootstrap configures the asset seam alongside the document and runtime seams: an HttpAssetStore over the persistence endpoint carrying the same credentials the document store carries, marked server-backed. The seam preflight now covers all three, so a failure still cannot half-configure persistence. - Image, video and TTS generation commit in one fixed order: provider, pool, document, local cache, task. A reference reaches the document only after put returned an id, so a document can never name bytes that were not stored. A failure before the write-back leaves the placeholder with the provider called exactly once; the retry happens on the next owner load. - The write-back is a per-slot rewrite through mutateDocument, which re-reads the current document under the per-stage lock, so it cannot clobber a newer scene. The open course is refreshed with the same rewrite without being marked dirty. - "Has this already been generated?" is answered by the document (the slide exists and no longer holds the placeholder) instead of by this browser's task table. - The classroom's resume effect fails closed on ownership: only a resolved owner starts generation, so a viewer opening a shared course spends nothing. - The local media and audio tables become a per-tab cache. A failed cache write costs a re-download, never the media. Browser-only mode is unchanged: every new call site sits behind the server-backed gate, the local tables stay authoritative there, and placeholders stay in the document. Rendering and export needed no changes. HttpAssetStore.resolve mints an object URL exactly as the browser store does, and the export byte resolver was already pool-first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(classroom): make the generation owner gate a three-outcome rule and apply it everywhere The gate refused everything but a resolved owner, which read a sidecar that answered "no ownership fact exists for this course" as a reason to block. That is the answer a deployment without the sidecar's server-side prerequisites gives for every course, and the answer a course with no ownership record gives: in both, there is nobody the operator's budget needs protecting from, and refusing strands the course's own author behind a question that can never be answered. Ownership is now four states over the sidecar's three outcomes. A definite answer splits into owner and not-owner. An absent record is its own answer, ownerless, and generation proceeds — the behaviour such a deployment had before the gate existed. Only the absence of an answer, a transport failure or a load that has not asked yet, stays unresolved and fails closed: "we could not ask" must never be read as "nobody owns this". One mapper turns a sidecar result into that state, and one predicate decides on it. The workbench classroom pane runs the same resume effect and had no ownership input at all, so a viewer opening a shared course there could still spend the budget. It now asks the sidecar once per course, in parallel with its load and feeding only the generation gate, so its read-only and edit behaviour is unchanged. The shared progressive-load policy carries the gate for it, with both new inputs required rather than defaulted so a future caller cannot omit them into an open budget. Its stale comment claiming ownership could not be expressed here is corrected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the write-back survive autosave, arrive before the scene does, and never leak Independent reviews of the write-back found three ways a durable document could still end up naming a placeholder, and two ways the gate that protects the operator's budget could be walked around. An autosave round captures the store synchronously and writes that capture, so a round already in flight when a rewrite landed wrote the placeholder straight back over the allocated id, and nothing marked the store dirty again to correct it. The rewrite now marks the units it changed, which leaves a corrective flush queued behind the stale one; re-saving a scene that already holds the id is idempotent, losing the id is not. Media is generated from outlines in parallel with scene content and usually finishes first, so the slide that will carry the placeholder does not exist yet and the write-back has nothing to rewrite. That was the ordinary path, not a tail case, and its result was discarded: the task was marked done, the scene was added afterwards with its placeholder intact, and a second pass in the same run could call the provider again. The allocation is now held under the placeholder — which also answers the skip test, so nothing pays twice — and applied when that scene is committed, before its first save. One complete pass now leaves no placeholder behind. A failed commit used to abandon what it had already allocated. A poster upload that failed threw away a stored video and sent the retry to submit the most expensive job in the system again; a rejected write-back left registry rows that name bytes nothing references, which the byte collector cannot reclaim because it only collects blobs no row names. A poster failure now costs the poster, and a write-back that reached nothing reclaims what it allocated. A partial write is left alone, because the document already names it. The ownership gate is fail-closed again. Treating the sidecar's 404 as permission was wrong: the client cannot tell "this course has no owner" from "this deployment told me nothing", so a visitor who opened a shared course could bill the operator. The root cause was the sidecar itself, which gated on the agent runtime although every persisted course has an owner regardless — the persistence route resolves one for every request. It now gates on server persistence, so the configuration that made 404 the universal answer has real ownership facts to report, and the gate can refuse everything but a named owner. Retry affordances answered to no gate at all. A viewer of a shared course with one failed image was shown a Retry button that called the provider. Both retry entry points and every surface that draws them now read one shared permission, so what is offered and what is allowed are the same value. Also: narration regeneration no longer pretends it can replace bytes behind a live id — the exclusivity proof that would allow it is refused by construction once references leave the browser, so it forks to a fresh id and says so; the "already generated" test lets a finished deck answer from the document alone, since scene order stops identifying an outline once slides are inserted or deleted; stored assets record a specific media type rather than a generic transfer type; the pane no longer asks the sidecar in browser-only mode; and the funnel's docstring now states what the per-stage lock actually guarantees, which is same-browser serialization and not a cross-browser compare-and-swap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): park allocations in the deciding turn, never reclaim on an ambiguous write A delta review of the write-back found the first-pass fix still had a window, and the reclamation it added could delete media the document already names. The allocation was parked after an awaited local cache write. A scene committed in that window reconciled against a registry that did not hold it yet, so the document kept the placeholder — and the entry recorded a moment later then answered the skip test as "already handled", so nothing could correct it. Parking now happens inside the write-back, in the same synchronous turn as the decision that nothing could take the reference; no await separates the live check from the park. Allocations parked by an earlier pass are handed to their slides at the start of the next one, before anything decides what still needs generating, so a held allocation whose scene has since arrived becomes a rewrite rather than an answer. Reclaiming on a rejected write was unsound: a rejection does not prove the server did not apply the write, so deleting the asset could break the scene that now names it. The funnel decides instead, and says so: it reclaims only when no store write was ever issued and nothing took the reference. Anything else is placed if its slide exists and parked if it does not, so the next pass reuses the bytes instead of paying for them again. When a write fails after part of it landed, the live store is brought up to the document before the error is rethrown — otherwise the next ordinary flush would overwrite the half that did land, with the ids deliberately not reclaimed. Parked allocations are now cleared with the course. Classic placeholders are reused across runs, so one surviving an interrupted run would be handed to a different slide of the next deck: the previous picture, on a slide whose provider was never asked. Both classroom surfaces clear the arriving course, the deletion cascade clears the deleted one, and clearing the database clears them all. Two more ways generation could start without asking the gate are closed. An overlapping pass — an outline retry re-enters generation with every outline while the first is still working — re-requested elements whose provider call was already in flight; a task that is not done is an answered request, not an unanswered one. And narration regeneration in the timeline editor called the TTS provider and allocated a pool asset with no ownership check at all; it now reads the same permission, which withholds both the per-line and whole-timeline controls and refuses the call. Finally, a pane opened during the stage-link availability gap recorded the sidecar's 404 for a course that was moments from existing and never asked again, leaving the real owner locked out of generation until it remounted. Ownership is re-fetched once the document becomes available; the gate stays closed until an answer arrives, so asking again can only open it for someone entitled to it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the asset routes reachable, and stale snapshots harmless A full-branch audit found that the deployment this project documents could not store a single generated asset, and that several routes into durable storage could still write a placeholder over a reference that had already landed. The persistence route sent asset requests through the development authenticator, which refuses outright in a production build that has not explicitly opted into it — and the documented server-persistence recipe produces exactly that build. Every store and every read answered 401, so images and video failed on every slide while re-billing the provider on each retry, and a narration failure stopped the deck at its first slide. Assets live in one shared partition by design, so there was never anything per-caller for that authenticator to decide: the route now resolves the asset principal itself, alongside the owner it already resolves for documents. Runtime sessions are genuinely per-learner and keep the development authenticator until real session verification replaces it. And narration that cannot be stored no longer fails its scene: the line stays unvoiced and retryable, which is what an image that cannot be stored does to its slide. Placeholders could also come back from behind. A queued autosave's snapshot, an editor-history entry replayed by an undo, the departing save a course switch flushes — each captures content at its own moment, and any of those moments can predate a write-back. Point fixes at each producer would leave the next producer to rediscover the bug, so the check lives at the write boundary every producer passes through, and the allocation record it consults now outlives the parked queue: a placeholder whose rewrite landed long ago is exactly the case it catches. Two ways generation could be lost or repeated are closed. A pass now claims the elements it will reach and releases them however it ends, so an overlapping pass stands down while an aborted one strands nothing — previously its tasks stayed `pending` and every later pass skipped them with no retry control to recover them. And the media abort controller is aborted before being replaced, so a superseded pass stops calling providers instead of running on for a course the user has left. The remaining two are narrower. The workbench pane asks for ownership only after a document load succeeds, and after every later one, mirroring the page route: the load is what creates the ownership row the first time a course is opened, so asking beforehand asked about a course that did not exist yet and locked its author out for the mount. And the ownership gate on the timeline editor now withholds narration regeneration alone; listening back to existing narration and seeing whether a line has any spend nothing and stay available. Known limitation, unchanged and now stated plainly in the comments that used to point at it as a solution: nothing reclaims an unreferenced pool asset. The registry sweep is written but not wired up, and the byte collector only reclaims blobs no registry row names, so every narration regeneration and every abandoned allocation leaves storage behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): gate asset mutations, and make claims and allocation records survive a handoff Opening the asset routes opened all of them. Reads and allocations are meant to be as open as document reads and creates already are, but no authorization hook was supplied, so the handler's default admitted PUT and DELETE too — and those scope by principal key alone, which is one shared constant. Any caller who learned an id, and a document read hands out every id its slides name, could overwrite or destroy another author's media. Mutations now require the deployment's credential, which in a production build without the development-auth opt-in means they are refused outright; reads and allocations stay open. The route comment says what the posture is and what it is not: the deployment-level fence is the access code, and no per-principal quota is configured. The client's own reclaim is best effort to match — losing an argument about deleting an asset must not cost a task its retry, and the bytes are left for server-side reclamation. The pass claim could not survive the handoff it was written for. A retry aborts the live media pass and starts its replacement in the same synchronous block, long before the aborted pass's cleanup runs, so the replacement saw every element still claimed, collected nothing, and returned — leaving each unreached element at pending with nobody coming back for it and no retry control to recover it, which is the exact failure the claim was introduced to prevent. A claim now carries its pass's signal and is retired the moment that signal aborts, and a pass releases only claims it still owns, so a late unwind cannot take its replacement's work. Claims are also acquired at the single point every request passes through, so a single-task retry participates too — previously a retry awaiting its provider was invisible to a pass starting alongside it and both called it. The allocation record could outlive the bytes it named. It was written before the write-back attempted anything and survived the reclaim that followed a failure, so when the slide finally arrived the write boundary stamped a deleted id into the document — and the placeholder it replaced was gone, which reads as already generated and stops anything from retrying. The record is now written only where the allocation is retained, and forgotten wherever a reclaim removes the bytes, including the narration rollback path. The tests follow. The route test drives the real storage handler against an in-memory registry instead of a stub, so it can see what the resolved principal is then allowed to do; the handoff test performs a real abort mid-pass rather than starting from an already-aborted signal; and the guards that could only assert file layout now assert the property they care about, or have been replaced by behaviour. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make media passes serial per course instead of tracking element ownership Three rounds of per-element claims each produced a new way to lose an element. Whole-pass reservations swallowed a Retry for an element the same pass had already failed, leaving it pending with the affordance gone. Retiring a claim by its signal freed an element whose commit was still uploading, so the replacement pass paid for it twice. A claim held for a failed element stranded its retry. The bookkeeping is the defect: every refinement of "who owns this element right now" answered the question at a moment when the answer was already stale. Passes for one course are now serial. A replacement aborts its predecessor, as before, and then waits for it to settle before collecting. That removes the question entirely: a commit already under way finishes — its bytes stored and its reference written, so the new pass sees a resolved slide and skips it — and an element the aborted pass never reached is still a placeholder and gets collected like any other. The claim set, the reservations, the signal retirement and the identity-checked release are all gone. The task table is consulted for one thing only: an element that is generating right now is a single-element retry running alongside the pass, and taking it too would pay twice. Pending is deliberately not a skip reason — it means a pass once intended to reach an element, which an abandoned pass leaves behind with nobody acting on it, and reading that as answered is what stranded elements before. A retry runs concurrently with a pass, because a pass never revisits an element it has processed, and it re-reads the task after its own await and refuses before touching it: marking first and refusing afterwards destroyed the failed state that draws the affordance. Browser-only mode is back to exactly what it was. The abort is now conditional, the waiting does not apply, and the original status-based skip is restored verbatim. Two baseline lines remain changed in each of the two files, and both are behind a server-backed fork whose else-branch is the original. Two smaller things. The allocation record becomes visible when a write goes on the wire rather than when the round trip ends, and the write boundary reconciles under the document lock rather than before it — a save queued during a write-back was otherwise captured with the placeholder and, for a course the user had left, had no corrective flush to follow. And the comments that said a refused reclaim leaves its bytes for server-side reclamation were wrong: nothing collects them, because the registry entry still names its blob and the sweep that would remove it is not wired up. They now say the bytes leak. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a deferred pass re-earn its right to run, and bound the commit it waits on Serializing passes moved their body out of the block that launched them, and three things followed from that. A pass now wakes when its predecessor settles, which can be after the user has left the course. It enqueued before it looked at its signal, into a task table keyed by element id alone — and placeholder ids are not unique across courses, which is why the classroom clears that table on arrival. So a departing course's pass seeded the arriving course's table with tasks carrying the wrong stage id, and a Retry routes by that id: the reference went into the wrong document. A pass now re-validates after the wait, before touching anything shared. The same lateness broke the skip test. `documentSkipIndex` answers only while the live store is on the pass's stage, and returning nothing put the collection loop on the browser-only rule — a silent demotion from "the document is the authority" to "this browser's task table is", on exactly the path where that table has just been cleared. Every element the predecessor had committed was collected again, paid for again, and its second write-back found no placeholder to rewrite, so its bytes were parked where nothing will ever reference them. In server-backed mode an unreadable document now means the pass stands down. And waiting was unbounded. A commit is uncancellable: the asset client takes no signal, and a document write cannot be half-undone. One stalled upload therefore froze the course's media generation for the session — the replacement never collected, the element sat on a skeleton that draws no Retry, and only a reload recovered. The pass's signal is now threaded into the media proxy fetch, and the commit is bounded by a deadline. The deadline is on the wait, not the work: the commit carries on, and if it lands late the document simply ends up correct, while the element becomes retryable and the queue moves on. The tests that were meant to pin the previous round were not sensitive to it. Two asserted end states where the mechanism only changes ordering, and one of them rigged the document read so the assertion held whether or not the pass had waited; a third covered half of what it claimed. They now observe the ordering directly — nothing is issued while another pass for the course is working; in browser-only mode a second pass reaches its provider immediately — and the reconciliation under the document lock has a test that fails when it moves back outside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * revert(media): drop the commit deadline and the abortable download The deadline bought less than it cost. Abandoning a commit after two minutes makes the element retryable while the real commit is still running, so a Retry starts a second commit for the same placeholder against the first: two provider calls, two allocations, and whichever lands second stamps its result over the other's task by element id. The allocation record is keyed by placeholder, so the loser's cleanup erases the winner's record, and the write boundary then puts the raw placeholder back into the document. That is the overlap serial passes were built to remove, reopened through the one door serialization never covered. So a stalled commit holds the course's media queue until it settles or the page is reloaded, and that is written down rather than papered over. The wait is unbounded on purpose: every ceiling on it turns out to be a way of running two commits for one element. Threading the pass signal into the download was also a mistake, in the other direction. The provider call that produced the URL has already been billed, so cancelling the download throws away work that is paid for — and the shared proxy cache records a cancelled request as a transient failure against that URL, which after three of them blocks it for every consumer in the session. Browser-only mode never asked for this: it had no way to observe an abort there, which is exactly why the bytes were kept. The signal is gone from the download again, and `fetchAsBlob` is byte-for-byte what it was before this branch. The regression guard for the stranded-element rule is restored alongside the timing test that was meant to supersede it. It catches a different rule — a task left pending being read as answered — and nothing else does: making the pass skip pending leaves every other suite green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): stop asking the pool for refs it never issued, bound it, and adopt cached bytes Four things a deployment found once this was running for real. A reference this application mints itself — a generation placeholder, a derived narration key — was never in the pool, because the pool allocates every id it holds. Asking anyway used to be an IndexedDB miss; once the pool is server-backed it is a request that answers 404, one per element per load, forever on a course that still holds placeholders. Every lease and probe now checks first. The check is a negative test on shapes this application owns, not an id validator: the pool's id domain stays unconstrained, and anything that is not one of ours is still asked about. The asset store can bound how much one principal holds, and enforces it inside the write transaction, but nothing ever passed the number. It does now, with a default rather than an opt-in: allocation is reachable by any caller a deployment admits, and with one shared principal an unbounded store is unbounded database growth with no operator-visible brake. Refusing asset mutations to unauthenticated callers was not enough, because every authenticated caller resolves to that same shared principal — so authentication decided nothing, and any signed-in visitor could delete any id they learned. Since this branch began storing media the registry is the only copy a course has. Replacing and deleting are now refused to everyone, and the browser no longer tries: an entry nothing references waits for server-side reclamation instead. What a browser must still do is forget its own record of an allocation that reached nothing, or a later save would stamp an id the document has no reason to trust. And a course generated before any of this holds placeholders in its document with its bytes only in the author's browser. Those bytes are paid for, so the author's next load converts them — stored to the pool and written back through the ordinary commit path, with no provider call — instead of buying them again. A row that records only a hosted URL is treated as absent: that URL is the provider's address, not something a document may hold. One renderer expectation moved with this. An untracked placeholder used to paint as pending on first render because asking the pool left a lease in flight; it settled to disabled a moment later either way, and now says so from the start. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): surface a full store as a refusal and convert legacy narration A quota refusal reached the browser as HTTP 500 with a generic message, which reads as a transient failure: the element kept a Retry that would pay a provider again and be refused again. The store raises the contract's own error and the handler maps it to 507, but the store answering a request is not always built by the same bundle as the handler -- the persistence provider is reached from the route bundle and from instrumentation, which is why its state lives on a Symbol.for global -- and `instanceof` is false across that boundary while the declared code is still right. Classify on the code as well as the class, and make the code a permanent, persisted refusal in the browser: recorded locally so it survives a reload, shown as "storage is full", and refused by the retry entry point so a stale button cannot buy a second generation. Every other storage failure stays retryable. Convert what a pre-server-backed course still holds. Generated media is adopted under either key this application has used for it -- the placeholder, and the allocated id of a course converted once and later rolled back -- instead of only the first. Narration is converted by a load-time pass over the open course's speech actions, since nothing re-enters generation for an action that already has an id: bytes to the pool, id written back through a funnel that mirrors the media one, owner-only and server-backed-only. A line whose bytes are in no browser is left alone rather than re-synthesized. Also: the pool guard is now a positive `ast_` test rather than an enumeration of the shapes we mint (imports never reach the pool, so this is safe in both modes); the slide ref collection is an exported pure function so its four lease sites are covered behaviourally; ASSET_QUOTA_BYTES treats every spelling of zero as opting out and refuses a malformed value at startup instead of falling back; the abort signal is re-checked after the cache read, before an uncancellable commit; and the unused `removeAsset` and pool `replace` surfaces are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * chore(storage): release 0.29.1 The asset HTTP handler now recognises a store refusal by the contract code it declares as well as by its class, so a quota refusal raised in another module realm answers 507 instead of 500. Same contract, stricter recognition, no API change: a patch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a full store recoverable and adoption course-safe Narration adoption read the local audio row by its derived key alone. That key carries no stage id and the table is keyed by id alone, so two courses can mint the same one -- a PPTX import numbers its scenes and actions deterministically, which gives every imported deck's first slide `tts_s1_speech-scene-p1`. Locally a collision only means one course plays another's clip in one browser; adopting it wrote that clip into the shared document permanently, for every device and every visitor. A row that names a course is now adopted only into that course, and a row from before that column existed only when the text it recorded is the text of the action being converted. A full asset store was made permanent last round, which was wrong three times over: it overwrote the refused bytes with an empty blob -- on the conversion path that row is a course's only copy of its own media -- it kept sending the rest of the deck to a provider against a ceiling it already knew was reached, and it left no way back once an operator raised that ceiling. A full store is neither the content's fault nor the configuration's, so it is now its own case: the bytes are kept, the pass stops at the first refusal, and the element shows the reason together with a Retry that re-attempts the upload from those bytes. Nothing retries automatically, so no one is re-billed. The narration write-back now reaches the write boundary every producer of a durable write passes through, not only the dirty mark: adoption never deletes the derived row, so a snapshot that reverts the rewrite is adopted again on the next load and allocates a fresh asset every time. Adoption is also mounted by both classroom surfaces rather than one, takes the course's abort signal, and re-validates that this browser still has the course open before each write. ASSET_QUOTA_BYTES is validated from instrumentation, where the README and the docstring already claimed it was: its only other consumer is lazy and memoised, so a malformed ceiling let the process boot and then failed every persistence request, documents and runtime included. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): remember a full store per course, and never lose retained bytes A stopped pass left the elements it never reached as placeholders with no persisted record -- deliberately, since nothing was attempted for them. But that left the next load with no reason not to try: it called a provider for the next placeholder and was refused at exactly the same point, once per reload, indefinitely. A full store is not a property of any slide. It belongs to the deployment and changes for reasons the document knows nothing about, so it is now remembered once per course in the browser's device KV. A pass that finds the marker stands down before spending anything and leaves every placeholder its "storage is full" state and its Retry; the first upload that succeeds clears it and the next pass runs normally. Narration adoption latched per course so it runs once per load, and the latch outlived the abort that leaving a course performs. On a surface that stays mounted across switches -- the workbench pane is one component for every course it shows -- owner course A, visitor course B, then back to A skipped exactly the clips the abort had cut off, and nothing else converts them. The latch is released with the abort now, and a course adopts one run at a time so a re-entry cannot hand a clip a second allocation while the previous run's uncancellable tail is still settling. A quota-blocked element retried into a network error or a 500 lost the bytes that were kept for it: the retry deleted the row before attempting the upload and wrote no replacement for an error carrying no structured code, so the next retry went back to a provider for media this browser had a moment earlier. The row now survives until an upload succeeds, the failure handler keeps whatever bytes the attempt was given, and the retry asks the question a pass asks -- does this browser already hold bytes for this element -- rather than reading an error code that a second failure has already overwritten. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): adopt real legacy narration, queue re-entries, report attempt outcomes Narration adoption admitted a stage-less row only when the text it recorded matched the action being converted. Both of those columns were added to the local audio table by the very change that moved narration onto allocated ids, so a row still carrying a derived key has neither: the rule refused every real pre-allocation course and passed only on fixtures built from post-allocation rows. What the row cannot say, the key can. A derived key names two clips only when two courses share a scene order and an action id, and an action id repeats only when something other than the generator minted it -- an import numbers them by slide position. So a key built from a generated action id is adopted on that basis, a key an import could have reproduced still needs matching text, and a row that names another course is refused however unique its key looks. Handing a re-entering caller the adoption run already in flight undid the latch release it was paired with: that run is bound to the signal the departure just aborted, so it stops at its next clip while the caller -- which has the course open and a live signal -- is told the work is done, and an effect replayed as mount, cleanup, mount adopts nothing at all. A later caller now waits for the uncancellable tail and scans again, which costs a lookup on a course that has nothing left and finishes the clips the abort cut off on one that does. One attempt at an element now reports both facts its callers need instead of a bare boolean: whether the store refused it for room, and whether bytes actually reached the store. Leaving a course clears the task table, so a retry that landed afterwards read "no failed task" as success and deleted the row holding the only copy of the media. Nothing is inferred from that table any more. Reading the localStorage property can throw where storage is denied by policy, typeof included, so the availability check moved inside the guard: this metadata is best-effort, and a rejection here strands a generation pass that has already enqueued its tasks. A retry is never blocked by the per-course "store is full" marker, but a retry that is refused again re-sets it, and adoption now reads and writes the same marker rather than issuing one refused upload per clip on every load. The two canvas element renderers and both thumbnail renderers show the reason beside the Retry, so a full store does not look like an ordinary failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): probe a full store instead of standing down, and pair notices with a Retry Narration adoption was given both halves of the per-course "the store is full" marker last round: it stood down when the marker was set, and it set the marker when its own upload was refused for room. Those halves are only safe together if something can lift the marker, and for adoption nothing could. It has no affordance of its own, it stood down before reaching its own clear, the media pass returns before its marker gate when there is nothing to generate -- so a narration-only deck, or one whose slides are already satisfied, painted no storage-full element and offered no Retry -- and narration generated rather than adopted allocates directly rather than through the media commit. The course's cached narration was then lost for good, where before it converted on the first load after the ceiling was raised. The gate is a probe now. A marked course attempts exactly one clip per load: refused, it stops and the marker stands, which costs what standing down cost; stored, it lifts the marker and finishes the course. Adoption spends no provider money, so the whole cost of probing a store that is still full is one refused upload. Generated narration lifts the marker too. The three surfaces that gained a failure notice last round drew it for any failure with a reason, including the one refusal that is reachable without server-backed persistence, so a browser-only deck painted something it had not painted before. The notice is drawn beside a Retry and nowhere else, which is what it was added for and what leaves browser-only output unchanged. Both are now asserted through the render harness the surface matrix already had. A caller arriving while a rescan is queued shares it rather than appending another. One rescan converts whatever the run in flight left and every later one would find an allocated id on every action, so a chain bought nothing and turned a single stalled upload into a course that never adopts again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): treat a refusal for room as a fact about one clip, not the deck The asset store checks each write against the headroom it has left, so a store that refuses a long opening clip can still hold every short clip behind it. Narration adoption assumed the opposite: it broke the deck at the first refusal and then re-attempted that same first clip on every later load, because the document names it first. A deck whose longest clip exceeds current headroom therefore never converted the clips that would have fit, with no affordance to recover it -- the state the probe was introduced to remove, reached through a narrower door. An unmarked load now attempts every clip, skipping the ones that do not fit, and remembers the condition only if the load ends with clips it still could not store. A marked load spends its single upload on the smallest clip left rather than the first one named: that is the clip that answers the question the marker asks, because if the smallest does not fit nothing does. The media pass keeps stopping at its first refusal, and for a reason adoption does not share -- every element it attempts costs a provider call. A rescan several callers share took the newest caller's signal, and the newest caller is not necessarily the one still there: a surface that opened a course and closed it again would stop work a surface still showing that course was waiting for, and that surface is latched, so it would never ask again. The shared run now takes a signal that is aborted only once every caller has left. The comment claiming the shared rescan contains a stalled upload was wrong -- the rescan is chained off the run in flight, so a stalled upload leaves every caller pending exactly as a chain would. It claims the bounded queue it actually provides, and the stall is recorded as a limitation. The failed-state containers took their stacking classes unconditionally, so markup differed in browser-only mode even though nothing moved on screen. Those classes are applied only when there is a notice to stack, and the tests assert the exact class attribute rather than a substring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): stop narration adoption writing the media pass's store-full marker The marker means "do not call a provider for this course". A path is entitled to write it only if its own refusal cost a provider call, and narration adoption's refusals cost nothing: it uploads bytes this browser already holds. The store also checks each write against the headroom it has left, so a clip that does not fit says nothing about whether a slide's image would. Adoption was writing it anyway, and one over-long narration clip was therefore enough to stand a course's entire image pass down on every later load -- on a store that had just accepted adoption's other clips. The author could still recover each element by hand, every load, for ever. Three rounds of narrowing this seam produced a finding each time, so it is removed rather than narrowed again. Gone: the marker read, the single-clip probe, the smallest-clip selection, and the up-front read of every row into an array -- which also retires a sampled-then-stale flag and the retention of a whole deck's blobs for the length of a run, and returns the loop to streaming one row at a time. Adoption's rule is now that every load attempts every clip it holds, once; any failure skips that clip and the load continues. The noise the coupling was meant to avoid does not arise, because after the first load the clips still outstanding are exactly the ones that did not fit -- normally none, or one. A successful write still clears the marker, and that is a different kind of statement: a write that went through is a fact this run established, where a refusal is an inference about what some other write would cost. For a course whose media needs nothing, adoption and generated narration are also the only paths that can establish it. The failure module still documented the deck-wide premise this contradicts. It now says what is true: the check is per write, and the media pass stops the deck as a judgement about cost rather than about certainty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bound a full store's cost from the store's own arithmetic Removing the store-full marker from narration adoption removed its bound too, and the code then asserted the bound was unnecessary. It is, on a store with room for most of a deck. On the store the whole mechanism exists for -- the ceiling reached, nothing fitting -- the outstanding set after every load is the entire deck, so a thirty-clip course posted thirty full blobs on every load, indefinitely. Each of those is not a cheap refusal: the bytes are uploaded, the server hashes the whole payload, and only then takes a per-principal lock and sums every entry that principal owns before saying no. The bound needs no flag, no key and nothing carried between loads. The store asks whether `used + addedBytes` exceeds the ceiling, and `used` only grows while a run is uploading, so a clip refused for want of room implies every clip at least that large is refused for the rest of that run. The run keeps the smallest size it has been refused and skips anything no smaller without uploading it; a smaller clip is still attempted, because it may fit. A deck the store refuses entirely now costs one upload per successive size minimum instead of one per clip, and a deck it has room for costs nothing extra, because nothing is refused. Only a refusal for room lowers the bar: a dropped connection says nothing about how much room there is. The deck-wide certainty premise the failure module retracted last round still stood verbatim at the site that implements the stand-down. Both copies now say the same thing: the check is per write, and the pass stops the deck as a judgement about cost rather than about certainty. The comment on adoption's marker clear now names its price. Narration of a few hundred bytes fits in headroom an image does not, so a proven write can let the next pass buy one more image that is refused again -- bounded at one, and the price of the alternative being a course whose media never generates again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(media): state the adoption bound exactly, and stop three comments describing the old rule The comment introducing the in-load bound gave its cost as "at most a handful, and the first load pays the most". Neither clause is a property of the rule. A clip is skipped only when something no larger was already refused, so a fully-refused deck costs one upload per successive size minimum in document order: one when the clips grow, about ln N for an arbitrary order, and one per clip when they only shrink -- a long opener followed by terser lines is exactly that shape. And no load is cheaper than the first, because the bound resets per run and a refused clip stays outstanding. The comment now says that, and points at what would make it exactly one for any ordering: the store returning its remaining headroom in the refusal's existing details channel, which the server leaves empty today. Two other comments still described the previous rule -- "attempts every clip it holds, every load" -- one of them twenty lines above the paragraph that introduces the bound, in the same block. Both now say what the code does. The bound's soundness is worth stating where a maintainer will look for it: quota is charged at full length with no discount for a duplicate, the sum it is checked against joins entries to blobs so the collector cannot lower it, the check takes a per-principal lock before summing, and replace and delete are refused to every browser. Nothing a run can do makes room appear inside it. One test installed a row implementation and replaced it wholesale a few lines later, so the first was dead and the survivor dropped the text the first clip's import-shaped key needs for the ownership rule -- it passed on the coincidence that the fixture's default text is the action's. Merged into one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): keep a Retry from re-buying parked media, and state the store seam once Five findings from an inline review. The standalone classroom route asked the ownership sidecar once per load and recorded only stage ownership when that ask failed. Every non-answer fails closed, so one transient 5xx left the genuine author with no resume, no Retry affordance and no legacy narration converted for the rest of the load, with nothing to change it short of a reload. The failure now records the fail-closed answer explicitly -- an answer an earlier load established must not outlive the failure that replaced it -- and an unresolved answer is asked for again, a few times over a few seconds. A real answer, however unwelcome, is final. A Retry could pay a provider for media the pool already held. When the bytes are stored and only the write-back fails in a way that keeps the allocation, it is parked and no local row exists, because that row is written only after a successful write-back. Retry now reads the parked queue exactly as the pass does and re-attempts the write-back: it re-keys the task done when the document takes it, leaves the entry parked when the slide still does not exist, and stays failed and retryable when the document refuses again. Object URLs a parked allocation owns are revoked when the entry is dropped. The commit path leaves them alone while the entry is parked, because it is then the only thing holding bytes this tab can render, so a course switch or a stage deletion was pinning the whole blob for the life of the tab. An entry a slide has already taken is left alone: the task table is displaying those URLs. The fallback lookup for cached bytes is a stage-scoped scan, and the keyed lookup misses for every row the commit path writes, so a pass was materializing and sorting the course's whole media table once per element. One scan per pass now, built on the first miss. It is sound and not merely cheaper: an element asks only for its own placeholder, and every row a pass writes carries the placeholder of the element that wrote it. "The store accepted a write, so it is not out of room" was enforced at three call sites under slightly different conditions, which made it a convention the next pool write path could silently break. It is stated once, in putAsset, for the course whose bytes it just stored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 16 天前 | |
fix: keep long Grok relay generations alive and inline image bytes (#1364) Two independent Grok failures seen when the provider is reached through a relay (a custom base URL) instead of api.x.ai directly: - lib/ai/providers.ts: a long non-streaming chat completion was cut off by the relay with a 504 at its idle timeout (~5 min), because nothing is sent upstream until the model has the whole answer. Adding 'grok' to the existing streaming-compat path (OPENAI_COMPAT_USE_STREAMING_CHAT=true) keeps bytes flowing across the idle window; the SSE is buffered back into a normal JSON response for the caller. The path is for relays only. usesCustomOpenAIBaseUrl recognises OpenAI's origin alone, so Grok's own api.x.ai also read as "custom" and was forced onto the compat transport; the provider's native endpoint is now excluded. - lib/media/adapters/grok-image-adapter.ts: response_format 'url' returns a link on the relay's CDN host (imgen.x.ai), which may be unreachable from the server's network. The generation then failed at the follow-up fetch through /api/proxy-media even though the image had been produced successfully. 'b64_json' inlines the bytes and removes that second hop. Inline bytes declare no media type, so the adapter reports one on ImageGenerationResult and returns a typed data URL, which is the shape openrouter-image-adapter already uses. Consumers take the type from there: agent image persistence records it, the client's stored media row keeps the type its data URL states, and classroom-media-generation names the file with the matching extension. A JPEG is no longer stored, served or named as a PNG. The process-wide undici timeout that previously accompanied these changes is dropped: upstream #1404 now gives LLM calls their own undici headers/body timeouts, which covers the same failure without raising the defaults globally. Co-authored-by: ciclou1 <ciclou1@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> | 5 天前 | |
fix(media): refuse redirects on adapter generation and poll calls (#1636) #930 made the connectivity probes in the media adapters pass `redirect: 'manual'`. The generation and poll calls in the same 14 files were left following redirects. Those requests carry the provider credential and go to a base URL that comes from provider settings a caller can supply, so a 3xx would replay the credential at a host the caller chose — and the redirect target can be an address the outbound guard already refused. Every such call now passes `redirect: 'manual'` and rejects a 3xx through a shared `assertNotRedirected` helper, which reports it as "<provider>: Redirects are not allowed (HTTP <status>)" instead of letting the generic failure path describe it as a provider error. - 26 call sites across the image adapters (seedream, openai, qwen, grok, lemonade, minimax, nano-banana), the video adapters (seedance, kling, grok, happyhorse, minimax, veo) and ComfyUI's submit and image fetch. - ComfyUI's `pollHistory` keeps its contract of handing the caller a retryable failure rather than aborting the generation: it logs the refusal and returns null. - ComfyUI's same-origin workflow load is deliberately untouched — it reads the app's own public/ asset, carries no credential and is not provider-influenced. - The two adapters added since #930 (OpenRouter image and video) already did this. tests/media/adapter-redirects.test.ts covers one case per adapter family. Each serves a 302 and asserts that the call rejects with the redirect message and that every request carrying an init object asked fetch not to follow redirects; each case fails if its adapter stops passing `redirect: 'manual'`. The HappyHorse test asserted the exact request init, so it now includes the new option. AI-assisted commit Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> | 5 天前 | |
feat(media): resolve image, video, and ASR models from server config (#1175) * feat(media): resolve image, video, and ASR models from server config The image generation chain never consulted the server-side IMAGE_<PREFIX>_MODELS config: the route only forwarded the client x-image-model header and each adapter silently fell back to a hardcoded vendor default, so a managed provider could generate with a model the operator never chose. The provider itself also defaulted to a hardcoded vendor id when the client sent no preference. - Add resolveImageModel / resolveVideoModel / resolveASRModel and resolveServerImageProviderId / resolveServerVideoProviderId to the server provider config, mirroring how TTS already resolves server-managed settings. - Routes resolve provider and model server-side when the client sends no preference and fail loud with a clear 400 when nothing resolves; no more hardcoded vendor defaults. - Adapters require an explicit model via requireModel; connectivity probes keep a fixed probe payload since they are auth handshakes, not user-facing generation. - Tests cover the resolver precedence and the fail-loud paths. * fix(media): align model resolution semantics across capabilities resolveImageModel and resolveASRModel let the server pin win unconditionally, while resolveVideoModel implements allowlist semantics (client choice wins when it is in the pinned list, otherwise the first pinned entry). Align image and ASR to the video allowlist semantics so a client picking any server-pinned model is honored, and update the resolver docstrings and tests to pin the aligned behavior for all three capabilities. The classroom media path has no client model and no HTTP response to fail loud with: it now falls back to the first catalog model when the operator pins no _MODELS list, so key-only deployments keep generating media instead of silently skipping every element via the adapter's requireModel backstop. * fix(media): normalize configured model lists and client model ids | 1 个月前 | |
fix: keep long Grok relay generations alive and inline image bytes (#1364) Two independent Grok failures seen when the provider is reached through a relay (a custom base URL) instead of api.x.ai directly: - lib/ai/providers.ts: a long non-streaming chat completion was cut off by the relay with a 504 at its idle timeout (~5 min), because nothing is sent upstream until the model has the whole answer. Adding 'grok' to the existing streaming-compat path (OPENAI_COMPAT_USE_STREAMING_CHAT=true) keeps bytes flowing across the idle window; the SSE is buffered back into a normal JSON response for the caller. The path is for relays only. usesCustomOpenAIBaseUrl recognises OpenAI's origin alone, so Grok's own api.x.ai also read as "custom" and was forced onto the compat transport; the provider's native endpoint is now excluded. - lib/media/adapters/grok-image-adapter.ts: response_format 'url' returns a link on the relay's CDN host (imgen.x.ai), which may be unreachable from the server's network. The generation then failed at the follow-up fetch through /api/proxy-media even though the image had been produced successfully. 'b64_json' inlines the bytes and removes that second hop. Inline bytes declare no media type, so the adapter reports one on ImageGenerationResult and returns a typed data URL, which is the shape openrouter-image-adapter already uses. Consumers take the type from there: agent image persistence records it, the client's stored media row keeps the type its data URL states, and classroom-media-generation names the file with the matching extension. A JPEG is no longer stored, served or named as a PNG. The process-wide undici timeout that previously accompanied these changes is dropped: upstream #1404 now gives LLM calls their own undici headers/body timeouts, which covers the same failure without raising the defaults globally. Co-authored-by: ciclou1 <ciclou1@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> | 5 天前 | |
feat(agent-runtime): workbench generate_image / generate_video write through the asset pool (#1007 part 6) (#1524) The workbench tools store generated bytes in the asset pool under the shared principal and write the allocated id into the document, the same discipline as the classic chain since #1392; the runner's putScene creates the reference rows and commits the allocations. The video completion patch rewrites every placeholder slot and retires anything that would shadow the new id; immediate render is preserved by leasing the id at the render boundary. A store-full refusal fails the tool with a model-readable error and writes nothing. Legacy /api/classroom-media documents keep rendering. Closes #1522. | 11 天前 | |
feat(media): allocate generated assets through the registry (#1007 part 2, step b) (#1039) * feat(media): establish shared asset ownership primitives Introduce global browser asset-pool ownership, asset-reference collection, stage reclamation planning, and lease-based URL access. Carry allocated media identity through storage and generation boundaries with regression coverage. * fix(media): enforce safe resolution across every consumer Route image, video, thumbnail, presentation, and video-export consumers through one resolution state machine. Prevent opaque allocated or generated references from reaching render and export sinks, with fallback and ownership tests. * fix(media): protect document-owned assets across mutations Allocate pool bytes before compatibility writes and document commits, then roll back uncommitted generations safely. Preserve document ownership across edits, retries, imports, speech generation, scene changes, and stage deletion. * fix(media): scope retries and tighten ownership guard Scope retries to the target scene and slide and refuse ambiguous shared-reference mutations. Expand retry rendering coverage and keep direct pool URL resolution behind the shared lease owner. * fix(storage): avoid nested lock during stage cleanup Execute prepared reclamation plans against an explicitly deleted document so the compatibility cascade cannot re-enter the per-document lock. Cover deletion of stage-owned media rows even when the document has no references. * refactor(media): confine reclamation to stage deletion Remove inline pool and compatibility-row cleanup from element, speech, scene, and audio replacement flows. Keep whole-stage reclamation behind explicit stage ownership, preserve stage-less legacy audio rows, and document deferred document-truth sweeping. * fix(media): close retry and resolution gaps Restore shared source tasks after successful forks, scope retries across both whiteboard locations, and wait for parallel TTS workers before rollback. Resolve background media through import, rendering, and PPTX export paths while keeping retry controls visible over last-good bytes. * test(media): execute the consumer safety matrix Replace source-substring checks with resolver seam execution across all six UI consumers, stage hydration, and both exporters. Pin each rollback layer independently and harden the ownership guard against aliased pool imports. * chore(packages): publish the additive DSL field Bump the DSL patch version for the optional speech-action field. Keep the transitional reclamation policy app-owned and document the legitimate transaction rollback removals there. * fix(media): make retry rollback task-safe Delay allocation task re-keying until final document reconciliation succeeds, restore shared source tasks on failed forks, and require exact stage-whiteboard targets before falling through from a missed scene. * fix(media): clear private assets and refresh leases Delete the asset-pool database during the confirmed local-data wipe. Notify the app lease layer after same-id replacement so mounted consumers re-resolve current bytes without reaching into storage internals. * test(media): pin closure safety guards Exercise the CSS allocation boundary, unfiltered legacy-row ownership, delete-and-undo byte survival, and real distinct consumer seams. Restore the video-only manifest overwrite condition and share the direct video resolution hook across both element variants. * test(media): narrow rollback element assertion Narrow the reconciled slide element to an image before checking its source so the rollback regression remains type-safe under the full root compiler configuration. * fix(media): close lease refresh races Gate the first batch publication by unique resolved refs, then publish every replacement snapshot without mutating the prior React state object. Serialize invalidation behind pending releases, evict rejected refreshes, and register replacement observation at the pool boundary. * fix(media): reopen pool after clear failures Always evict the singleton once its store has been closed, including blocked and failed database deletion paths. Report blocked deletion as deferred and prove a later write uses a fresh live store. * fix(media): isolate shared retry progress Track forked regeneration under the selected element until it receives a fresh asset identity, leaving the shared source task and bytes untouched. Surface targeted failures through renderer task lookup and skip the redundant fork reconciliation lock. * fix(media): close final asset retry gaps Keep blocked asset clears fail-loud until a successful retry, with actionable settings guidance. Clear failed shared-fork state across durable and live key spaces, and pin renderer lookups, lease publication identity, and committed-ref rollback protection. * fix: preserve actionable retry failures Localize the blocked cache-clear recovery hint across every supported locale and pin the deferred-error mapping. Retain durable fork failure rows while retries run, deleting them only after successful generation so unstructured failures survive reload. * fix(media): hydrate legacy stored video thumbnails Home-page recent-video thumbnails regressed for legacy Dexie mediaFiles rows keyed by gen_vid placeholders: the reworked hydration resolved the row's bytes through the sealed resolver but never surfaced the stored blob (and its poster) as object URLs for the preview card, so the CI recent-video-thumbnail e2e specs found no visible element. Hydration now materializes legacy stored video rows into blob URLs for both the element src and poster while keeping the resolver invariants: opaque refs still never reach a DOM src, and concrete addresses are never blanked. Unit pins cover the seam so the vitest suite catches this class without a browser. * fix(media): resolve sole restored legacy video Classroom playback restored tasks by exact document media references. Legacy gen_vid references can outlive the key used by the one persisted video row, leaving the player on a placeholder even though bytes were restored. Select the sole completed stage video only for legacy sequential refs after exact and reconciled matches. Keep exact failures authoritative, refuse ambiguous candidates, and cover success, ambiguity, and failure precedence in unit tests. * fix(media): scope legacy video recovery Decide restored legacy video recovery once from the complete document and record it through the shared task lookup consumed by playback, editing, and resolved slides. Keep ambiguous documents as placeholders, apply the same decision to thumbnail hydration, and preserve exact failure precedence. * fix(media): exclude claimed video recovery tasks Model restored legacy recovery as a two-pass match across document video elements and task rows. Remove tasks claimed by exact, targeted, or placeholder lookup before applying the sole-candidate fallback, and cover the ownership/cardinality matrix. * fix(media): unify video element resolution Centralize source, task, poster, and legacy recovery decisions for every video consumer. Ensure direct URLs win over opaque refs and play_video waits on element-targeted retry tasks. * refactor(media): route legacy recovery through resolver Let document-aware consumers request legacy video recovery through the unified element binding API, keeping thumbnail hydration on a single decision path. * fix(media): preserve import refs and prefer pool bytes Recognize unambiguous extensionless relative media addresses during classroom import. Resolve allocated export and thumbnail assets from the shared pool before falling back to lagging compatibility rows. * fix(export): preserve concrete video sources Route PPTX video elements through the shared media binding resolver so an unresolved opaque reference cannot replace a playable source. Complete browser and persisted-store cleanup when asset-pool deletion is deferred, while retaining distinct hard-failure behavior. * fix(media): fork retries without exclusive ownership Enumerate logical asset owners across every persisted document before allowing global pool replacement. Thread explicit targets through fresh-id rewrites and cover cross-document aliases plus unreadable ownership. * fix(media): guard global asset reclamation Share a fail-closed persisted-document liveness check between stage deletion and retry replacement. Preserve cross-document pool aliases while deleting stage-owned compatibility rows and cover enumeration failures. * fix(media): cover complete slide asset references Route slide media traversal through a shared mutable slot contract so backgrounds participate in export, thumbnail hydration, collection, and rewrite lifecycles. Snapshot complete surviving-document refs once per reclamation and preserve manifest-only owners while retaining fail-closed behavior. * fix(media): preserve exclusive retry asset ids Allow targeted retries to replace exclusively owned pool assets in place. Keep shared and unprovable ownership paths on fresh allocations, and pin production-shaped retries plus compatibility-row cleanup. * fix(media): revalidate asset bindings at completion Recheck repository-wide ownership before replacing generated media and fork scoped retries when exclusivity changed. Route video export selection through the unified resolver and keep concrete posters independent of task state. * fix(media): count unflushed owners and broadcast replacements The completion-time exclusivity proof read only the persisted document, but slide duplication updates the Zustand aggregate synchronously and schedules persistence behind a debounce. A retry finishing inside that window saw a single persisted owner and replaced the bytes behind a reference the duplicate also held. The proof now also counts owners in the live stage snapshot when that snapshot represents the stage being retried, so an unflushed duplicate forks instead. Same-id replacement notifications were realm-local, so a second tab showing the same classroom kept its lease pinned to the superseded blob URL. The notification now travels over a BroadcastChannel; each receiving realm runs its own observers against its own pool, so a spoofed message can at most force a re-resolve. A missing or failing channel never fails the replacement. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bind replacement listeners and spare shared audio rows A realm that only renders never sends a replacement, so binding the channel from the sender path left passive tabs deaf to peers. Binding now happens where the observer is registered, when the asset-pool module loads, and the receiving realm resolves its own pool lazily so a cleared or unavailable pool degrades to the next resolve instead of throwing. Observer notifications ran under Promise.all, so a rejection surfaced after BrowserAssetStore.replace had already committed and turned a durable success into a reported failure. They are settled individually now; the callback is wrapped because a synchronous throw would otherwise escape before allSettled sees the array. audioFiles rows are keyed globally by audioId, so deleting a stage removed the sole row for an id a surviving document still referenced — playback and both export paths read that table directly and cannot fall back to the preserved pool blob. Rows are now filtered against surviving references, while a failed enumeration still withholds only the irreversible pool removal. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): serve replaced bytes from the pool everywhere Unknown survivor liveness deleted every planned audioFiles row. Those rows are keyed globally by audioId, and playback plus both export paths read the table directly, so losing one is as irreversible for them as removing the pool entry. Unknown liveness now preserves the rows too, leaving bounded garbage for a later pass that can prove exclusivity. The earlier pool-first change covered PPTX, video collection and thumbnail hydration but missed classroom ZIP export, which still serialized the stale compatibility row after a lagged same-id replacement, shipping media the classroom no longer renders. Auditing every direct reader of the media and audio tables surfaced the same gap in playback: speech regeneration also replaces bytes under a stable id and does not roll the pool back when the compatibility write fails, so the player kept serving superseded narration. It now resolves the pool first and falls back to stored rows for legacy and imported audio. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(audio): replace exclusively owned speech clips in place regenerateSpeechAudio accepted the action's audioId but always passed undefined as replaceAssetId, so regenerating an exclusively owned pool-backed clip allocated a new asset every time, rewrote the action and orphaned the previous pool entry and compatibility row until stage reclamation — contradicting the stable-id path generateAndStoreTTS already implements for media. Ownership is now established before synthesis, and the rule itself moved to the shared reference module so media retries, poster replacement and speech regeneration consume one implementation instead of restating it. An exclusively owned clip keeps its id and has its bytes replaced; a shared clip, a legacy id with no pool entry, or unprovable ownership still gets a fresh allocation so other holders keep their audio. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(media): drop the import left behind by the ownership move Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): resolve audio pool-first and fence peer realms Stable-id TTS regeneration commits replaced narration to the pool before the audioFiles mirror write, so a failed mirror leaves the row stale. AudioPlayer already resolved pool-first, but audioObjectUrl, collectAudioFiles and the video timeline dependencies still read Dexie directly and would serve the superseded clip. All allocated-audio readers now share one resolver, with Dexie kept as the fallback for legacy and imported rows. The exclusivity proof modelled unflushed owners in the active realm only, so another tab duplicating the same asset during its save debounce could still be observed as a single owner and have its bytes replaced globally. A peer's pending state cannot be read across realms, so presence is probed instead: any realm holding the stage forces the fork path. A deferred asset-pool deletion no longer reloads the page. The database is still on disk and the guidance asks the user to close the other tab and retry, which the reload discarded. The decision moved into a shared helper so the rule is pinned rather than living inline in the component. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): fail closed when presence cannot be probed The presence helper documented that an unanswerable probe must count as a peer, but every unavailable path — no BroadcastChannel, a constructor that threw, a send that failed, and the window before the pool's asynchronous binding completes — returned false, so the ownership proof cleared a single local owner and replaced globally shared bytes in place. Probing now returns present, absent or unknown, and only a probe that was actually sent and went unanswered is absent; the ownership decision treats unknown exactly like present. The pool declares its binding intent synchronously so a probe issued during the load-time window waits for the bind instead of concluding that presence is unavailable, and releases that gate if the import fails. Coverage reaches the write boundary: with presence unknown, a production-shaped targeted retry forks to a fresh id instead of calling replace, and the original bytes stay intact for a peer's unflushed owner. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(dsl): bump to 0.6.3 after the release dedupe took 0.6.2 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(i18n): add the deferred-clear guidance to fr-FR The locale landed on main after this branch added the key, so the alignment check flagged it as the one missing translation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 1 个月前 | |
feat(agent-runtime): workbench generate_image / generate_video write through the asset pool (#1007 part 6) (#1524) The workbench tools store generated bytes in the asset pool under the shared principal and write the allocated id into the document, the same discipline as the classic chain since #1392; the runner's putScene creates the reference rows and commits the allocations. The video completion patch rewrites every placeholder slot and retires anything that would shadow the new id; immediate render is preserved by leasing the id at the render boundary. A store-full refusal fails the tool with a model-readable error and writes nothing. Legacy /api/classroom-media documents keep rendering. Closes #1522. | 11 天前 | |
feat(dsl): standardize the asset manifest and converge the export paths (#1007 part 3) (#1117) * feat(dsl): standardize the document asset manifest Add asset-manifest.ts to @openmaic/dsl: the canonical AssetManifestEntry shape (ref, kind, and byteSize/mimeType/duration/voice/prompt metadata where available) plus enumerateAssetManifest, the pure document-to-manifest enumeration. An entry's ref is the reference exactly as the document holds it -- the manifest is the id-based reference enumeration with metadata, not a content hash and not a resolution result. The traversal walks the stage whiteboard, each scene's canvas/whiteboards/speech actions, and the stage video-manifest keys in document order, with logical-owner reference counts that match the accounting duplication-safe replacement uses. This settles the media-ref + asset-manifest schema question (#779 open question 4) on the side the asset-pool RFC already implied: the schema is a function of the id semantics decided there. The type lives in the dsl rather than a new @openmaic/exporter package because the enumeration is pure over document types the dsl already owns (Stage/Scene/Slide/Action), so a separate package would add a published artifact and release-workflow surface without adding a capability; the storage contract comment now points at the module. Refs #1007 * refactor(export): drive the classroom ZIP from the asset manifest collectMediaFiles used to scan the whole mediaFiles table for the stage, so any row the document no longer references -- an orphan left by an edit or a superseded regeneration -- rode along into the archive. Both ZIP collectors now take their reference sets from the standardized asset manifest (buildStageAssetManifest wraps the dsl enumeration with the compatibility rows' metadata): only referenced assets are archived, and a referenced asset whose bytes exist only in the pool is still collected via a synthesized record. Byte resolution is unchanged: pool first through resolveStoredBytes / resolveAudioBlob, with the compatibility row kept as the legacy byte fallback and as the metadata source. mediaIndex is now a serialized view of the manifest, and the missing-audio report derives from the manifest's audio entries instead of a second action walk. The audioRef mapping and the legacy audioUrl fetch path (collectLegacyAudioForExport) are untouched. Refs #1007 * refactor(video-export): take the timeline's reference sets from the manifest createVideoTimelineDeps scanned the whole mediaFiles table for the stage and derived its audio id set from its own action walk -- a third, independent answer to "which media does this course use?". Both record loads now key off the standardized asset manifest: media rows are read per manifest ref by compound key instead of by table scan, and the audio id set is the manifest's audio entries. Orphan rows were never reachable through the scene-scoped elementId-to-mediaRef bridge; now they are not even read. The bridge itself is untouched: element ids recur across scenes, so the elementId-to-mediaRef mapping stays scoped per scene, and the legacy audioUrl fallback keeps its own action walk because a URL is not a manifest ref. AssetPlan remains the video IR's view of the same references. Refs #1007 * refactor(export): resolve PPTX media through the shared resolver only Each PPTX element branch carried its own resolution chain: a task-state renderable-URL lookup first, then -- gated on the legacy placeholder predicate -- a stored-bytes override, with the poster block repeating the pattern. One helper now owns resolution for backgrounds, images, video / audio sources, and posters: opaque refs (allocated ids and legacy placeholders alike, no placeholder-pattern gate) resolve pool-first through resolveStoredBytes and embed as data URLs, concrete addresses resolve through the media state machine and keep the caller's fetch path. exportMediaResolution and the resolveStoredMediaBlob wrapper fold into the helper; resolvePptxMediaBinding stays as the state-machine entry the resolution-surface test matrix drives. Refs #1007 * refactor(export): retire the export-side Dexie byte fallbacks Export call sites no longer read bytes off compatibility rows directly. The ZIP collectors and the video timeline's audio load resolve bytes only through the shared resolvers (resolveStoredBytes / resolveAudioBlob), which answer pool-first and keep the compatibility row as their internal legacy fallback level; the row reads that remain at the call sites supply metadata (format/duration/voice/mime/size/prompt) only. The rows themselves stay for legacy and regeneration readers -- what goes is the export paths' own fallback logic. One observable tightening: a failed media row (error set, empty placeholder blob) no longer ships a 0-byte file into the classroom ZIP, and an evicted row no longer ships its empty local blob; referenced-but- byteless assets are simply absent from the archive, as they already were when no row existed. Refs #1007 * test(media): cover the enriched stage asset manifest builder Pins the join between the pure dsl enumeration and the compatibility rows: metadata attaches by ref, rows no document reference names never appear, and a referenced asset with no row keeps a metadata-free entry. Refs #1007 * fix(video-export): widen the deps stage input for the manifest enumeration enumerateAssetManifest reads the stage's whiteboard and videoManifest, so createVideoTimelineDeps declares them on its input instead of the bare id; callers pass only the id today and the optional fields stay absent. Also applies the repo prettier formatting to the files this branch touched. Refs #1007 * fix(dsl): enumerate slide audio elements in the asset manifest Slide audio elements carry their own src, and the manifest skipped them, so a manifest-driven collector could never archive their bytes. The audio slot maps to kind 'audio' alongside narration ids. Refs #1007 * fix(media): harden ref-keyed lookups against prototype-named asset refs AssetRef is an unconstrained string alias, so a media reference can legitimately be "__proto__", "constructor", or any other Object.prototype member. Plain objects keyed by such refs silently drop assignments or answer lookups with the prototype object, which rewrite paths then accept as a mapped id. Convert the remaining ref-keyed lookup tables introduced by the export convergence to prototype-safe structures: the classroom import media/poster alias maps and the legacy-conversion video-manifest reconstruction now use Map / null-prototype containers with explicit membership checks, and every consumed value is validated as a string before it is written into a src / mediaRef / audioId slot. The shared media-task lookup receives the same treatment: one centralized own-property-checked lookupMediaTask now serves the stored-bytes resolver, the PPTX embeddable-src path, the video collection path, and the element/background task resolution, so a prototype-named placeholderRef can no longer hide a re-keyed task from the fallback chain. Adversarial tests drive "__proto__" and "constructor" refs through the import round trip, the PPTX fallback path, the legacy conversion commit path, and the media-task fallback end to end, including a buildPptxBlob regression with a task re-keyed to an allocated id while retaining a prototype-named placeholderRef. * fix(export): use safe archive asset paths * refactor(dsl): centralize slide media slot roles * refactor(export): derive consumer refs from manifest * fix(export): sanitize classroom archive extensions * fix(video-export): preserve narration speech order * fix(export): enforce kind-coherent archive media * fix(export): define media coherence boundary * fix(export): carry task-owned poster binding for PPTX export A video element with no explicit poster falls back to its media task's generated poster URL, but resolveVideoMediaForElement left posterTask undefined for that case, so the PPTX manifest guard saw a foreign URL with no task-ownership exemption and dropped the video element instead of using the established runtime poster fallback. Carry the poster task binding whenever the task poster is the effective poster: the task-owned URL then satisfies the guard's objectUrl exemption end to end. A concrete explicit element poster still stays element-owned and never borrows the binding, and the guard's foreign-ref rejection is preserved (and exported as a directly testable predicate). Coverage: an element with no poster plus a task-provided poster embeds the task poster as the PPTX cover (red at the pre-fix head, green now), and a genuinely unrelated URL with no task ownership is still rejected by the guard. * fix(export): preserve legacy narration source refs in the media index The explicit sourceRef contract was partial: primary audio and generated media entries carried it, but legacy URL narration serialized no source ref. The legacy URL itself is the natural source ref — it is known at fetch time — so wire it through the collected blob into the mediaIndex entry. Import already registers serialized sourceRefs as aliases, so the URL now round-trips as an explicit mapping instead of being reconstructed only from the action's audioRef. Poster siblings are deliberately NOT given their own mediaIndex entry: a sibling poster (media/asset-<n>.poster.<ext>) is a legacy byte copy written from the video record and is not an independently referenced document asset — when the poster is a real document asset it already has its own indexed entry with a sourceRef, and import reconstructs the sibling by path derivation from its parent video entry, reusing the poster's own indexed allocation when one exists. The PR description is narrowed to match; corrected paragraph: "Archive names never interpolate refs — sequential safe paths (media/asset-<n>.<ext>, audio/audio-<n>.<ext>) with the original ref preserved through an explicit sourceRef mapping on every independently indexed media entry: generated media assets, poster assets, primary narration, and legacy URL narration (the legacy URL itself is the entry's sourceRef). Extensions are allowlisted per kind. The one exception is the legacy sibling poster byte copy (media/asset-<n>.poster.<ext>, written next to its video when the video record still carries the pre-pool poster bytes): it is not an independently referenced document asset, so it has no mediaIndex entry or sourceRef of its own — its identity is derivable from its parent video entry (same index), and import reconstructs it by sibling-path derivation from that video entry, reusing the poster's own indexed allocation when one exists." --------- Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 1 个月前 | |
release: OpenMAIC 1.0.0 — the agent workbench (#1228) * feat(storage): add an agent-session store with PG backend and layered contracts (#1163) * feat(storage): add agent-session store with PG backend and layered contracts * test(storage): avoid BigInt literals for pre-ES2020 root typecheck * fix(storage): close agent-session store review findings * docs(storage): align hook ordering and contention-probe claims with the code * ci: run on the agent-workbench integration branch * chore(storage): bump to 0.5.0 for the agent-session store * fix(storage): carry replay compaction across page boundaries * feat(agent): add the driver model contract and stage route dialect (#1165) * feat(agent): add the driver model contract and stage route dialect * fix(agent): validate route context windows and clarify dialect precedence * feat(agent): adapt the agent-session store and runtime foundations (#1167) * feat(agent): adapt the agent-session store and runtime foundations * feat(agent): resolve request owner identity via an anonymous cookie * docs(agent): document the opt-in compaction default and harden edge cases * feat(agent): add the background session runner (#1169) * feat(agent): add the background session runner * feat(agent): wire the runner into startup behind feature flags * fix(agent): stop clean interruptions from consuming the attempt budget * fix(storage): charge the attempt budget for abandoned leases but not clean parks * docs(storage): document the attempt-charging contract and decouple its tests * feat(agent): add agent session and owner event streams (#1170) * feat(agent): add agent session and owner event streams * fix(agent): close the session-existence oracle and document the owner seam * feat(agent): add agent session lifecycle routes (#1171) * feat(agent): add agent session lifecycle routes * fix(agent): validate session-create input and preserve the owner cookie on errors * refactor(storage): drop the unused active-stage API from the agent-session contract (#1174) * refactor(storage): drop the unused active-stage API from the agent-session contract Tools address stages explicitly on every call, so the store keeps no mutable session-level stage pointer. Removes resolveActiveStage and setActiveStage from the store interface, their PG implementations, the active_stage_changed lifecycle event, the session_active_stage owner event variant, and the contract tests pinning them. The active_stage_id column and the DDL check constraint stay untouched for schema compatibility. * chore(storage): bump @openmaic/storage to 0.7.0 for the contract removal * docs: document the agent runtime configuration surface (#1176) * fix(agent): repair orphaned and late tool results across interruption boundaries (#1180) * fix(agent): repair orphaned and late tool results across interruption boundaries A crash, shutdown, or provider failure can leave the durable transcript with tool calls that have no result, or with results ordered illegally for the provider. Three failure modes were fixed: - Orphaned tool calls: a run that died between an assistant tool-call frame and its result left a dangling call in the entry tree. Resume no longer synthesizes and persists receipts for it: interrupted results are a read-time provider view owned by a shared read-boundary repair, which returns the original array for a healthy transcript and never mutates the tree. - Late parallel results: a parallel tool can finish while pi unwinds an aborted assistant frame, leaving result(A), assistant(aborted), result(B) in durable order. Strict providers reject non-contiguous results, so the read-boundary repair moves existing results next to their owning assistant frame (in call order), omits incomplete unwind frames, and synthesizes receipts only for genuinely missing calls. - Interrupted calls at the write boundary: a call still in flight when the run winds down (shutdown, lease loss, cancellation, provider failure) had no receipt at all. The runner now tracks in-flight calls from their assistant frames and, before the terminal flush, appends an interrupted-result receipt for each still-orphaned call through the same attempt-fenced write chain, so a lease-stealing zombie never writes and the next claim sees a provider-safe transcript. * test(agent): pin the runner wiring for interruption-boundary tool repair * feat(agent): add neutral tool foundation libraries (#1184) * feat(agent): register a web_search tool on the session runner (#1185) * feat(storage): add a per-session URL trust gate (#1186) * feat(agent): add the skills system (#1189) * feat(agent): add the skills system (builtin directories and durable user skills) * fix(storage): serialize the user-skill quota check-and-insert per owner Two concurrent creates at the 50-skill boundary both counted 49 rows and both inserted (READ COMMITTED, no lock), overshooting the quota contract. The create transaction now takes a per-owner pg_advisory_xact_lock first, and the same-name idempotency check runs before the count check so an at-least-once retry of the create that committed as the owner's 50th row still returns its durable receipt instead of a quota error. The 23505 backstop is retained for writes that do not take the lock. * fix(agent): share unstorable-character validation and align skill lookup * feat(agent): add session materials and a fetch_url tool behind the URL trust gate (#1190) * feat(agent): add session materials and a fetch_url tool behind the URL trust gate * fix(agent): harden session material fetching * feat(storage): add an ownership scope to stage documents (#1191) * feat(agent): add material read and search tools (#1192) * feat(agent): add stage read and patch tools (#1194) * feat(agent): add page generation and deck editing tools (#1198) * test(storage): keep the PG contract suite order-independent (#1200) * fix(agent): revoke deleted-session URL authority and reject private ISATAP endpoints (#1199) * fix(storage): revoke deleted session URL authority * fix(ssrf): reject private ISATAP endpoints in strict fetches * chore(storage): bump to 0.11.1 for the session-URL authority fix * feat(agent): add roster and voice registration tools (#1201) * feat(agent): add folder organisation tools (#1202) * feat(api): add stage and material HTTP routes (#1203) * feat(workbench): add the client data layer (#1204) * feat(workbench): add the client data layer * docs(workbench): write the ported comments in English * chore(edit): remove the in-editor agent panel (#1210) * chore(edit): remove the in-editor agent panel * style: apply prettier formatting * fix(agent): report the runtime as unusable without a database (#1207) * fix(agent): report the runtime as unusable without a database * style: apply prettier formatting * feat(agent): add image, video and pptx import tools (#1211) * feat(workbench): add the agent chat surface (#1205) * feat(workbench): add the agent chat surface * docs(workbench): write the ported comments in English * fix(workbench): label the folder and rename tools on the timeline * fix(workbench): label the roster and voice tools on the timeline The reconciliation test iterates every tool the runner registers and requires a display label of its own. The roster and voice-clone tools (list_voices, set_roster, clip_audio, register_voice) reached the integration base with the roster/voice-registration tools but never gained presentation rows, so they fell through to the default branch and rendered their wire names. Port their rows from the reference implementation (labels and i18n keys verbatim) and extend the reconciliation allowlist with ROSTER_TOOL_NAMES and VOICE_CLONE_TOOL_NAMES, so a future tool cannot enter the product without a label. * feat(agent): add the material extraction lifecycle (#1212) * feat(storage): add material extraction lifecycle * feat(agent): execute queued material extraction * style: apply prettier formatting * style: satisfy prefer-const in the extraction runner * test: give material fixtures the extraction lifecycle fields The media-tools slice and the extraction lifecycle slice were each green in isolation but never compiled together: the lifecycle made derivedFrom and extraction required on AgentSessionMaterial while the media-tool fixtures predate them. * chore: remove stray task notes * fix(workbench): label the extraction lifecycle tools on the timeline * feat(workbench): add the workspace shell (#1206) * feat(workbench): add the workspace shell * docs(workbench): write the ported comments in English * i18n(workbench): align workspace keys across locales * fix(workbench): adopt the landed data layer and label the extraction tools - replace the sibling-slice seam stubs with the real data-layer modules - drop ambient declarations now shadowed by landed files - port timeline labels for the extraction lifecycle tools from the reference - align the new i18n keys across all locales * ci: retrigger * feat(api): folder routes, stage-meta viewer surfaces, and the material upload contract (#1215) * fix(storage): restore capability-based stage access * fix(api): bind document access to request owner * fix(agent): restore three-state stage access on the tool layer Port probeStageAccess and the three-state StageAccess (owned / foreign / missing / tombstoned) and gate every stageId-bearing stage tool on an owned probe, mirroring the reference per tool: - move_to_folder, rename_stage, read_stage_outline refuse a non-owned stage with the single not-yours message before touching the store. - The course/DSL toolset and the roster toolset are wrapped by withOwnerStageAuthorization: read_stage, patch_stage, grep_stage and every writer refuse a foreign stage with the same message and refusal shape. - Scene preview keeps its own probe and its own refusal text, and is registered beside the course toolset (never double-gated). - The runner injects one probe factory at the three call sites. Tests: the dsl cross-owner test premise (a foreign stage is readable by id) encoded an invented capability-read policy that the reference does not have at the tool layer; it now asserts foreign read/patch/grep are all refused while the owner still reads. Curriculum cross-owner assertions were already the reference's and now pass with the probes in place. * docs: correct per-file test counts in the fidelity report * test: fix type errors in stage-access fidelity test * test: adapt media-tool and gate suites to the owner-scoped store seam * feat(api): add owner-scoped course-folder HTTP routes Port the reference implementation's /api/folders family (list, create, rename, delete with ungroup/remove modes, and folder membership) onto the owner-bound document store, replacing its provider-based auth with the existing withRequestOwnerId / owner-scoped store seams. The storage package's folder store grows the pieces the routes need: DocumentFolder.order (schema column + max+1 assignment + ordering), renameFolder, deleteFolder(mode) with captured member ids, and setStageFolder(stageId, folderId | null) with idempotent un-filing. FolderNameError moves into folder-name-validation.ts (stage-storage re-exports it, keeping import sites intact). Every route gates on the configured agent runtime (plain 404 when off or unconfigured), keeps the reference's machine codes and envelopes, and is covered by gate tests plus a behavior suite. * feat(api): add stage-meta viewer surfaces for the classroom Port the reference implementation's viewer-facing stage state — can-edit / collected / published / generation-complete — on top of the stage-access base (stage_meta + tombstones). stage_meta gains published_at and generation_complete columns plus a stage_bookmarks table; the reference's deployment-specific origin/claimed_at columns are stripped. New gated routes: GET /api/stage-meta/[stageId] (per-viewer facts, 404 for absent/tombstoned, never returns the owner id), GET /api/stages/[id]/status, POST generation-complete / publish / unpublish (owner-only), POST /api/bookmarks. The resolver lives in lib/server/stage-access.ts. Wiring: a fetchStageMeta client with the reference's three-outcome contract, stage-store isOwner/isBookmarked/readOnly fields (upstream single-user defaults, no-op until the sidecar answers) plus setViewerAccess, the classroom apply path computing readOnly = !(isOwner || isBookmarked), the Stage editability gate, and a sidecar probe after each classroom load. A sidecar 'absent' answer keeps the editable default here because the classroom also serves local-only courses; server writes stay owner-enforced. * feat(api): port the reference material upload contract Rewrite POST /api/materials to the reference implementation's upload shape so the workbench uploader (uploadWorkbenchMaterial, which posts no session id and expects a flat 201 view) works unchanged: owner-scoped upload with mime normalization/validation (415), per-class size caps checked on the declared content-length and the streamed body (413), empty body (400), quota (429), sha256 reserve->store->finalize lifecycle with abandon on failure, flat { materialId, originalName, bytes, mime, extraction } 201, and an x-request-id echo. Adds the owner-scoped material library (owner_material table + quota + 24h lazy sweep, bytes in the host's asset registry as the neutral replacement for the reference's object-storage byte path) and the material cap configuration. The session-scoped GET list is left as-is; the reference's owner-material extraction worker is not ported (the branch's session-material extraction lifecycle already covers extraction). Gate tests now cover all 23 persistence routes across the three runtime env states; the materials behavior suite pins the new contract. * feat(media): add an optional local ffmpeg media extractor (#1213) Adds a local ffmpeg/ffprobe pipeline as a second media extraction provider behind the extractor registry, ported faithfully from the reference implementation: duration probing, keyframe-safe chunking, per-chunk ASR with timeout and deadline budgets, and timestamped transcript assembly. - Availability probing feeds the registry's candidate selection: the provider simply is not a candidate when ffmpeg/ffprobe are absent. - With neither ffmpeg nor a cloud provider configured, extraction fails with an actionable message naming both enablement paths. - Media materials route through the same extraction lifecycle and lease fence as documents; no parallel queue. - Tests inject the executable resolver so the missing-ffmpeg path is the default-tested one; the real pipeline test is skip-if-unavailable. - @openmaic/storage 0.13.0 -> 0.14.0 (media routing in the material lifecycle surface). * feat(storage): per-scene monotonic revisions via database triggers (#1214) * feat(storage): per-scene monotonic revisions via database triggers Restore the reference implementation's freshness granularity: a per-scene monotonic revision maintained by database triggers, so every writer (HTTP routes, agent tools, jobs, manual SQL) bumps it without application cooperation. - Companion revision tables + trigger functions in the storage package's idempotent schema bootstrap, with the lock-order invariant, pg_notify wakeup and the suppression switch for batch writers. - ensureDocumentSchema gained a dollar-quote-aware statement splitter. - The freshness and manifest routes serve per-scene revisions. - Mutation-verified: dropping the triggers turns the revision tests red. - @openmaic/storage 0.13.0 -> 0.14.0. * fix: forward the freshness manifest through the owner-bound store * feat(workbench): add the Pro entry points and preserve the mode-transition semantics (#1208) * feat(workbench): add the Pro entry points * feat(workbench): preserve Pro mode transition semantics * fix(workbench): drop ambient declarations shadowed by landed slices * fix(workbench): drop ambient declarations shadowed by the landed shell * feat: port workspace shell sibling modules Port the 16 leaf modules the Pro workspace shell imports but that were only ambient-declared, replacing the compile-time bridge with real implementations adapted from the sibling-slice reference: pure workbench helpers (session title, rail tab, course-chat bootstrap, created-course tabs, course-tabs memory, workspace navigation, pane navigation, pro-edit sizing, existing-course minting, first-message session), the neutral brand context and course-rename server API, the server-action session delete, the home discovery hook, the classroom pane host with its load-policy leaf, the theme toggle and floating-layer owner, plus the floating-layer-owner wiring the dialog/dropdown/tooltip portals stamp. Also add the workbench-shell locale copy for all 12 locales, port the reference tests for the ported modules, and drop types/workbench-sibling-slices.d.ts now that every declaration has a real implementation. * docs: keep ported comments in English and deployment-neutral * docs: announce 1.0.0 and refresh the feature overview (#1216) * docs: announce 1.0.0 and refresh the feature overview * docs: finalize 1.0.0 README after feature merge * fix(agent): control-plane routes answer 404, not 500, without a database The agent control-plane routes gated only on the runtime flag, so an enabled-but-unconfigured deployment (flag on, DATABASE_URL empty) answered 500 from a store that cannot connect. Gate them on the configured check instead, matching the stage/material routes: the whole surface is cleanly absent until both the flag and the database are present. The status probe keeps reporting both bits. * test: mock both runtime gate exports in the control-plane route suites * fix(agent): abort in-flight TTS on cancel and bound each provider request with a timeout (#1217) The generate_tts / scene-tts path checked the runner's AbortSignal between actions but never created the provider HTTP requests with it, so a session cancel left a hung synthesis fetch in flight until a restart repaired the tool result. Thread the signal end-to-end: TTSModelConfig carries an optional signal, generateTTS combines it with a per-request timeout (TTS_REQUEST_TIMEOUT_MS, default 30s, ported from the reference runtime's TTS bounds) via AbortSignal.any, and every provider fetch (openai, azure, glm, qwen incl. voice-clone + audio download, voxcpm, minimax, doubao, elevenlabs, lemonade) is created with that signal. A timeout now fails the tool call with TTSRequestTimeoutError (a clear retryable error) instead of wedging the session; a caller cancel propagates as the interruption so the runner settles the session as cancelled without a restart. Tests: hung-provider simulation rejects at the timeout with the retryable error; abort mid-flight aborts the captured request signal and surfaces the interrupted shape; removing the signal wiring makes the abort tests fail (red), restoring them turns green. * fix(workbench): PG-mode home listing via owner stages; keep the interrupted terminal course card (#1218) Finding 1: with server persistence on, listStages resolved to the generic GET /api/persistence/documents listing, which the capability model deliberately answers 403 FORBIDDEN_DOCUMENTS for (reads by id, listings owner-only). The home/workspace library now lists through the owner-scoped GET /api/stages surface (same anonymous-owner cookie the workbench uses) when server persistence is enabled; the server-side 403 is untouched. Finding 2: a run interrupted (session_interrupted) and repaired (session_resumed) that ends cancelled before agent_end stranded its pending classroom sightings, so the timeline's terminal card lost the course the answer produced. session_end (cancelled) now flushes the pending sightings into the same course card set agent_end paints, before the stopped caption. * chore(workbench): remove the bookmark concept and the saved-courses drawer (#1219) * chore(classroom): remove the bookmark ('collected') concept entirely The stage-meta viewer port introduced a bookmark surface (stage_bookmarks table, POST /api/bookmarks, the isBookmarked sidecar field, and a readOnly rule that let a saved course stay editable). The product has no such concept, so remove it as a closure: - delete the /api/bookmarks route and the stage_bookmarks table plus its query helpers from the persistence bootstrap - drop isBookmarked from GET /api/stage-meta/[stageId] - simplify the classroom read-only rule to readOnly = !isOwner across the sidecar client, ownership signal, classroom load, stage store and the classroom page - keep publish/unpublish, generation-complete, isOwner and isPublic exactly as they were - update the gate and stage-meta route suites and the README mentions The workspace rail's Bookmark glyphs and comments describe the upstream saved-courses (favorites) section, which is driven by isOwner and renders no collect affordance; they are kept as unrelated homonyms. * chore(workbench): remove the saved-courses drawer UI The first pass removed the bookmark data model but kept the rail's "Saved courses" drawer, judging it a separate surface driven by `isOwner === false`. The home/workspace listing is owner-scoped, so that flag can never occur: `allSaved` is permanently empty and the drawer (plus the collapsed-rail Bookmark mini-button) is a dead affordance. Remove it: the SavedDrawer component and its mount, the savedOpen / savedSection state, the allSaved / matchedSaved derivations, the 'saved' variant of the course-list renderers, the mini Bookmark glyph, the drawer-only CSS, and the drawer's i18n keys from all 12 locales. The courses tab is now exactly one folders tree. The authored/favorites split in workspace-tree.ts goes with it; the tree module no longer reads `isOwner`. The discovery course type keeps the field — the shell still reads it for read-only gating. Upstream has no collect concept; the drawer could only ever render empty here. The reference implementation HAS this drawer (its favorites come from its account system), so this removal is a deliberate upstream product decision, not a fidelity bug. * fix(workbench): restore the attach entry, add the rail settings entry, pin all three entry points (#1221) * fix(workbench): restore the composer attach entry by gating it on the live runtime The AttachButton's rollout probe read a `materialsEnabled` field that this branch's /api/agent/runtime never answers (the materials routes gate on the runtime itself, like the stages), so the gate could never pass and the attach button never rendered — the Pro launch and chat composers showed only the @-mention and enhance glyphs. Substitute the field with the runtime's `enabled` value, which IS the upload action's precondition: POST /api/materials answers 404 whenever it is false, so the render condition now equals the action precondition (no dead button). The button's label (`proMode.attach`) is a user-visible string that becomes visible again; port the reference implementation's own translations verbatim into the 11 locales that still carried the Chinese copy. * feat(workbench): add the settings entry to the rail's bottom-left cluster The reference's rail foot carries a cluster of utilities (its saved-courses drawer, the language switcher, the display toggle). This branch removed the drawer — it could only ever render empty here — and the product decision is to fill that freed spot with the settings entry. Add a settings trigger to the foot cluster (expanded rail, beside the language and display toggles, and on the collapsed strip) and mount the model/provider SettingsDialog in the rail, wired to the trigger. It is the same dialog the classic home opens from its header pill; the workspace had no settings entry of its own, so nothing is duplicated within a surface. * test(workbench): pin the restored upload, attach, and settings entry points Covers the three restored entry points: - the courses-tab upload control: rendered beside the course name filter, wired to the discovery hook's ZIP import trigger, disabled while an import runs, and gated by the same condition as its action (the courses tab); - the composer attach control: an actual render of AttachButton under both probe answers (visible when the runtime says the upload path is live, hidden otherwise), its mounts in the launch and chat composers, the branch's runtime-field substitution in the probe, and the reference's own `proMode.attach` copy in all 12 locales; - the settings entry: the trigger in the rail's foot cluster (expanded and collapsed), beside the language and display toggles, opening the SettingsDialog the rail mounts. * chore(config): the Pro workbench flag implies the MAIC Editor gate (#1223) A workbench build without the editor toggle has no way to edit a course: enabling NEXT_PUBLIC_PRO_WORKBENCH_ENABLED while forgetting NEXT_PUBLIC_MAIC_EDITOR_ENABLED produced exactly that split-brain bundle. The workbench IS Pro mode, so its flag now implies the editor gate; the standalone flag remains for deployments that want the classroom editor without the workbench. Documents both flags in .env.example. * fix(agent): wake SSE tails and the runner on durable deltas (streaming fidelity) (#1222) The Pro workbench chat did not stream: the session/owner SSE routes polled the durable event log on a 5s/30s clock with no wakeup, so message_update deltas (written at 150ms cadence) reached the browser in poll-sized blocks and the thinking strip only mounted after the whole reasoning text had accumulated. Port the reference's LISTEN/NOTIFY delta path: - storage: add in-transaction wake hooks (onSessionEventAppended, onOwnerEventAppended, onCancelRequested) so a host queues pg_notify in the same transaction as the durable append; align readEventsAfterForReplay to rank the bounded page so the first delta after the cursor is always kept (the live tail can never starve). Bump @openmaic/storage to 0.18.0. - app: port the process-wide event-notify bus (dedicated LISTEN client, self-check probe, reconnect backoff; notify through the storage transaction surface), wire the store hooks, subscribe both SSE routes before the initial read with the reference's initializing gate, and give the runner one {kind:'session'} subscription whose wake runs the cancel check and the message drain. Polls stay as the lossy-NOTIFY backstop. - lifecycle: start/stop the bus from instrumentation. Tests: storage hook + compaction contract; route wakeup latency; runner wakeup wiring with a fake agent; bus unit tests; PG contracts proving a real append wakes the routes and a live SSE route forwards a message_update on the wakeup, and that a rolled-back append never wakes. Also fix the pre-existing park-attempt-budget PG test TRUNCATE (missing CASCADE against newer FK tables). * fix(storage): asset writes self-deadlocked against pooled PostgreSQL (#1225) * fix(storage): refuse the non-transactional byte-write deadlock configuration A byte store whose plain write() runs on its own pooled connection cannot be invoked from inside a registry write transaction: after the transaction has claimed the blob-row lock, that write blocks on the lock the transaction just took while the transaction waits on the write - a self-deadlock PostgreSQL cannot detect (one side is idle in transaction). There is no lock-safe ordering for such a writer: bytes must be written after the row claim (writing before it lets the collector delete the bytes while the upsert waits), and any second-connection write after the claim is the deadlock. The configuration is therefore detected and refused: - AssetByteStore gains writesOutsideRegistryDatabase?: true, declaring that the layer's plain byte operations cannot contend for the registry's row locks. - PgAssetStore refuses put()/replace() up front (and defends coordinatedWrite) when the byte store has no writeWith and does not declare the flag, throwing a clear configuration error before any row is claimed. - The collector mirrors the guard on its delete path (deleteWith or a declared out-of-registry layer, else a configuration error). - The object store declares the flag (its out-of-transaction write remains legitimate); the in-registry PostgreSQL byte column provides writeWith / deleteWith instead. - Write transactions (put/replace/remove) set SET LOCAL lock_timeout = 30s so any future lock-contention variant fails loudly instead of hanging. Bumps @openmaic/storage to 0.18.0. * fix(persistence): forward the transactional byte methods through the lazy asset byte-store wrapper The no-bucket case of lazyAssetByteStore returned a bare { write, read, delete } and dropped writeWith/readWith even though the underlying PgAssetByteStore has them. The registry's hasTransactionalWriter duck check then failed and put() fell back to the byte store's own pooled connection, which blocks forever on the blob-row lock the registry transaction just took when the bytes live in the same PostgreSQL - the production self-deadlock. The no-bucket layer is statically PgAssetByteStore, so its transaction-pinned methods are forwarded eagerly (typed against the real signatures via PgForwardedByteStore). The bucket case keeps its lazy-probing semantics: no transactional writer exists there, the signed-URL method stays absent or lazy exactly as documented, and the wrapper now declares writesOutsideRegistryDatabase so the registry may run the plain write inside its transaction. New tests pin the wrapper's transactional capability red-to-green and assert put()/resolve() route byte traffic through the transaction-pinned queryable. * fix(home): cap the generate-prep ingest drain at 3s so Generate never waits the full server budget The classic home flow's Generate click drained in-flight ingests for the full 15s server budget. Cap the wait at GENERATE_DRAIN_CAP_MS (3000ms, documented as a UX bound) and reuse the existing timeout fallback: sources that miss the cap proceed on the legacy byte path and each late-resolving id is released. * chore(storage): bump to 0.19.0 over the concurrently landed 0.18.0 * fix(agent): bound every tool call with a timeout; never resurrect a cancelled session (#1226) * fix(agent): bound every tool call with a global timeout and settle it on cancel A tool await that neither resolves nor rejects wedges the session forever: the lease keeps heartbeating and the driver never reaches its next cancel checkpoint. Race every tool execution (in buildAgent) against a hard budget (OPENMAIC_AGENT_TOOL_TIMEOUT_MS, default 10 min, per-tool overrides for known long runners) and against the caller's AbortSignal, so even a signal-ignoring await cannot keep a cancelled session running. On timeout the call rejects with AgentToolTimeoutError; the agent loop turns the rejection into a structured error tool-result the agent can retry or proceed from, and the abort signal is delivered to the tool's in-flight work through a derived controller. Zombie-tool updates after settlement are dropped. * fix(storage): never re-lease a cancel-requested session; settle it as cancelled on claim The claim scan treated a session with cancel_requested_at set as a normal claim candidate: after a restart it re-leased the same session for attempt N+1 and resumed generating despite the pending cancel. claimNextSession now settles such candidates as cancelled under the claim lock (status cancelled, attempt reset, lease and cancel request cleared, terminal session_end event and owner projection) instead of leasing them, then keeps scanning. Bump @openmaic/storage to 0.18.0. * docs: takeaway-style 1.0.0 announcement with bilingual guide links The 1.0.0 head is now a short takeaway block — badge links to the official user guides (English and Chinese), five one-line highlights, and pointers into Features and the workbench setup section — instead of six dense paragraphs. The detailed provider-neutrality and freshness notes move into the Features workbench section, phrased database- neutrally (the announcement no longer names a specific database). Release date corrected to August 27. * fix(workbench): restore editor chrome, mode transition, streaming, materials, mentions, folders (#1229) * fix(workbench): wire workspace folder routes * fix(editor): restore reference workbench chrome * fix(workbench): persist composer materials and course refs * fix(workbench): preserve live reasoning frames * fix(persistence): back off failed streaming saves * chore(workbench): retire stale slice seams * test(editor): cover element pin layer * chore(storage): bump to 0.21.0 for the user-message ref/material fields * chore(editor): translate ported code comments to English * fix(agent): fence durable tool writes and consume cancel requests atomically (#1230) * fix(agent): enforce provider force-off in agent tools and scrub vendor identity from tool results (#1231) * fix(materials): serialize per-owner quota reservations and make crashed uploads reclaimable (#1232) * fix(editor): resolve dock-bar i18n keys, remove dock height drag, wire element referencing (#1233) * fix(workbench): send the opening session message exactly once with refs intact (#1234) * feat(editor): port timeline TTS preview single-flight and voice-all state latching (#1235) * fix(media): restore the reference classic media chain (#1236) * fix(import): adapt imported PPTX canvas size so decks render without overflow (#1237) * fix(editor): complete element referencing — renderer DOM contract and GenUI picking aligned with the reference (#1238) * test(providers): reconcile the provider-config vendor-token debt count after the main merge The integration line's AK/SK fallback for the managed document provider adds occurrences that main's allowlist snapshot predates. Same mixed-composition debt category the group already documents; no new vendor behavior. * test(providers): reconcile vendor-token debt counts with the integration line The main-merge brought main's neutrality-guard snapshot next to integration features it predates (media-extractor fallback chain, local voice-profile deletion semantics, the enabled-TTS helper). Same debt categories the guard already documents; counts updated to the guard's own tally and two grouped entries added. No new vendor behavior. * fix(agent): carry reasoning through the completions dialect so the thinking strip renders (#1239) * feat(skills): add Feynman and spiral curriculum methods (#1240) * feat(agent): port missing reference tools and skills (parity audit) (#1241) * feat(media): retire asset-registry wiring; media and materials follow the reference byte model (#1242) * fix(classroom): center adapted canvases in the stage and send back navigation home during generation (#1243) * feat(settings): skill management with real list, download, delete, and upload (#1244) * feat(settings): skill management section with real list, detail, and zip download * feat(skills): owner skill delete and upload across storage, API, and settings * fixup! feat(settings): skill management section with real list, detail, and zip download chore: neutralize a reference note in the settings header comment * fix(media): persist origin-independent classroom-media references from the agent runtime (#1245) * feat(editor): float the insert toolbar in the outer frame with collapse (#1246) The insert strip was bounded to the slide card, so it could only ever sit on top of slide content: the card's overflow clipped it and it could not be parked in the padding beside the slide. Move it into the studio frame the element picker's panel already roams (CanvasOverlayPortal + the frame selector), so both canvas overlays share one bounding container and their handles behave the same. While picking, the strip rises over the picker and goes inert, which is the z-order CANVAS_OVERLAY_Z already documents. Add a fold beside the grip: the chevron collapses the strip to that grip row and back, with the buttons unmounted rather than hidden. The fold is session-local state owned by EditShell, next to the drag offset, so a surface swap keeps it; nothing is persisted. Expanding a strip parked at the bottom edge re-clamps through the same bounds rule the keyboard move uses. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * fix(workbench): align the chat timeline's left edge with the composer (#1247) * fix(agent): fence session claims while an ask_user question is outstanding (#1248) * fix(agent): settle-time rescue tracks real delivery instead of a count offset (#1249) * fix(persistence): migrate owner_material to oss_key and drop legacy asset_id (#1250) * docs(readme): surface the 1.0.0 user guide badges at the top (#1253) * fix(workbench): show newly created folders in the sidebar without reload (#1254) * docs(readme): add the release version prefix and drop the opt-in framing * fix(workbench): single-source the chat gutter so timeline and composer share a left edge (#1255) The transcript and the composer each established their own column: their own `px-*` gutter and their own `mx-auto w-full max-w-*` centering wrapper. Equal padding values were never enough, because the two columns are centered inside different containing blocks — the transcript's is a scroll container, whose content box is narrower than the composer footer's by the scrollbar's width: transcript text left = pad + (pane - 2*pad - scrollbar - measure) / 2 composer box left = pad + (pane - 2*pad - measure) / 2 The padding cancels out of the difference and what remains is `-scrollbar/2` at every padding value, so the transcript sat half a scrollbar to the left of the composer and tuning the two paddings against each other could not move it. The column is now established once, by the nearest common ancestor of both (`chatColumn`), and the scroll viewport and the composer footer are siblings inside it that add no horizontal inset of their own. The cap carries the gutter on top of the 760px reading measure, so the text column keeps its width. The handed-over question row drops the padding that indented it past the agent's prose; framed rows keep their own inner padding, which is what a card's border sitting on the column edge means. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * fix(workbench): lock pane-embedded classroom to edit mode (#1256) The workspace right pane painted the full learning chrome — speed control, play button, learner avatars, mic bar — for a course the agent had just created, then flipped to edit once the first scene landed. resolveStageChromeMode treated playback as the DEFAULT branch for a hosted classroom, so every shortfall fell into it: a course whose tab opens at stage_link time has no scenes yet, so currentSceneId is null and isHostedSceneEditable is false. A folded pane parked the playback root behind the fold and cross-faded it out over the pane on unfold, and a failed editor chunk dropped into playback permanently. Lock it at the pane instead of defaulting per entry path: - WorkbenchPanelProvider — the single element that mounts a classroom into the workspace — publishes editPinned (visible && !playback). Every entry path passes through it, so none of them decides. - The hosted resolution can no longer degrade to playback. Start Learning (workbenchLearning, new input, split out from pane visibility) is the one door; everything else resolves between the neutral loading shell and edit. - Stage's chrome dispatch is exhaustive on chromeMode, so the playback root is no longer the else-branch of a condition about the current scene. No flicker: chromeMode is resolved during render, and preloadEditor now answers synchronously (isEditorPreloaded) so a remount with the chunk already registered paints edit on the first frame. A failed import is no longer cached forever, so the lock cannot strand the pane. Standalone classrooms keep their stored mode unchanged. --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> | 1 个月前 | |
feat(token-plan): one-click token-plan setup + deployment usage dashboard (#784) * feat(usage): add usage normalization, pricing, model-fetch, balance, storage Foundational layer for token-plan usage tracking (cc-switch-modeled): - lib/usage/normalize.ts: AI SDK v6 usage → four-class token shape - lib/usage/pricing.ts + defaults: per-class USD pricing table - lib/server/model-fetch.ts: /models candidate-URL multi-fallback (ported) - lib/usage/balance-providers.ts: built-in balance queries + detection - lib/server/usage-storage.ts: fire-and-forget jsonl logging to data/usage All pure/storage logic covered by vitest (32 tests). * feat(usage): capture token usage at the callLLM/streamLLM chokepoint - callLLM records result.usage before returning - streamLLM wraps onFinish to record totalUsage on stream completion, preserving any caller-supplied onFinish - provider/model derived from the model instance (no route changes) - fs-backed storage imported dynamically; fire-and-forget, never throws - include_usage is already sent by @ai-sdk/openai, so streaming is covered * feat(usage): add probe-models, balance, and usage API routes - POST /api/provider/probe-models: discover chat models via /models with candidate fallback; SSRF-guarded; filters non-chat ids; 401/404 typed - POST /api/provider/balance: built-in balance detection + billing fallback - GET /api/usage: aggregate jsonl by model/day/source with costIncomplete flag Verified e2e against the live MAIC gateway: 16 chat models (6 filtered), balance detected, and a real callLLM writes a costed usage row. * feat(settings): add token-plan preset picker to provider dialog - lib/config/token-plan-presets.ts: data-driven vendor presets (Huawei/MiniMax/ Xiaomi token plans, OpenRouter/SiliconFlow gateways, DeepSeek/GLM/Qwen/Hunyuan/ Doubao direct) with baseURL, protocol, optional modelsUrl, category - add-provider-dialog: '选择厂商/自定义' tabs; picking a preset auto-fills baseURL+protocol+modelsUrl, custom tab unchanged - ProviderSettings.modelsUrl carries the optional /models override - i18n keys added across all 8 locales Verified in-browser: preset picker renders grouped by category. * feat(settings): add Fetch Models button and balance bar to provider panel - 拉取模型: probes /models, merges discovered ids into the model list (dedupe, keeps manual additions), with success/no-endpoint/auth messages - 查询余额: queries /api/provider/balance, renders a balance bar or a 'check console' hint when unsupported - index.tsx: handleModelsFetched merges probe results into provider config - i18n keys across all 8 locales; removed a now-unused eslint-disable Verified in-browser against MAIC gateway: 16 models fetched, balance shown. * feat(settings): add usage dashboard to System Settings - usage-dashboard.tsx: echarts dual-axis daily trend (tokens + cost), totals cards, by-model table, refresh; reads GET /api/usage - mounted at the top of GeneralSettings (系统设置) - honest disclaimer + costIncomplete marker when a model lacks pricing - i18n keys across all 8 locales Verified in-browser: shows 1 request / 31 tokens / $0.0003 from a prior call. * feat(token-plan): multi-modal one-click setup in System Settings - token-plan-presets.ts: presets now declare per-modality targets (llm/image/video/tts/webSearch); MiniMax is the full-set template, others LLM-only — extend by adding entries (at-our-best adaptation) - apply-token-plan.ts: fills one key into every declared modality via injected store setters, isolating per-modality failures (TDD, 4 tests) - token-plan-settings.tsx: new sidebar page — pick plan, enter key, one-click apply lights up adapted modalities (+LLM model probe), shows 'not adapted yet' for the rest, balance bar reused - reverted add-provider dialog to plain custom form (preset picker moved here) - i18n across all 8 locales Verified in-browser: Token Plan page renders, MiniMax shows all 5 modalities, apply/balance UI wired. 859 tests pass, build clean. * feat(token-plan): add custom token plan entry to Token Plan page - Custom card at the bottom of the provider list: pick it to manually enter name + protocol + baseURL (LLM-only), then key + one-click apply via the same applyTokenPlan flow (mirrors cc-switch's custom provider) - effectivePreset unifies preset vs custom for apply/balance/probe - i18n keys (customGroup/customName/customHint) across all 8 locales Verified in-browser: custom card expands the manual form; preset flow intact. * feat(token-plan): reflect persisted config on the Token Plan page Other settings panels read the store directly, so they survive section switches; the Token Plan page only wrote the store, so it looked blank on return. Now it also reads providersConfig: - selecting a preset prefills its saved API key - configured presets show a '已配置/Configured' badge State persistence was never broken (apply writes zustand→localStorage); this fixes the missing read-back so the page reflects it. * fix(settings): isLLMProviderConfigured crashed on providers without models Root cause of 'token plan config disappears': applyTokenPlan writes a new LLM provider with no models yet (probe fills them later). isLLMProviderConfigured did config.models.length unguarded → threw inside setProviderConfig's resolver → the whole set() aborted → key/baseUrl/type never persisted. - guard models in isLLMProviderConfigured (shared validator; affects any provider written without a models array) - seed models:[] in applyTokenPlan's LLM write for a valid initial shape - regression tests: validator no longer throws; apply+probe write keeps apiKey Verified in-browser: DeepSeek persists key and shows 已配置 after section switch. * feat(token-plan): add remove/teardown for a configured token plan Apply had no inverse. Add removeTokenPlan: clears the API key and disables (enabled:false) every modality the plan declared; a custom LLM provider is deleted entirely (removeProvider), built-ins keep the cleared shape. - removeTokenPlan in apply-token-plan.ts (mirrors applyTokenPlan, isolated per-modality, injected setters; 3 tests) - trash button on configured preset cards in the Token Plan page; resets page state if the removed plan was selected - i18n 'remove' across all 8 locales Verified in-browser: removing DeepSeek empties its key and drops the 已配置 badge. * feat(usage): track multimodal usage, drop cost/pricing Reframe usage stats as pure usage (no cost), per user decision: - usage-storage: drop all cost fields; add kind (llm/image/video/tts/asr) + quantity + unit; LLM keeps token counts, others store quantity - delete lib/usage/pricing.ts + pricing-defaults.json + its test - instrument image/video/tts at the server-only API routes (not in the provider dispatch — those files are in the client graph and importing fs-backed usage-storage broke the client bundle) - /api/usage: aggregate by model/day/modality, no cost - dashboard: per-modality usage (tokens/images/seconds/chars), token-only daily trend, no '$'; i18n cost keys replaced with modality/unit keys - backward compatible: legacy rows (no kind, stray cost fields) read as llm 912 tests pass, build clean. * feat(usage): per-modality dashboard layout + softer dark-mode chart - group usage into per-modality sections (LLM/image/video/tts/asr), each table's usage column uses one consistent unit (token/image/sec/char) — no more mixed units in a single column - summary chips per modality with their own unit - trend chart now plots daily REQUESTS (unit-agnostic, works for any modality) instead of LLM-only tokens - theme-aware chart: faint thin line + soft gradient area, muted axis/grid colors via useTheme — fixes the harsh solid stroke in dark mode Verified in-browser (dark): TTS shows '字符', LLM shows 'Token', separate sections; chart no longer has a hard line. * refactor(usage): dedupe model/fetch/usage helpers, parallel balance probe Review cleanups on the token-plan/usage branch: - extract modelInfoFromId() (shared vision heuristic + ModelInfo shape) - extract fetchWithTimeout() shared by model-fetch and balance-providers - extract recordGenerationUsage() to dedupe the image/tts/video routes - queryBalance: fetch billing subscription + usage in parallel - parseOneApiBilling: report quota without remaining when usage endpoint is unavailable, instead of implying zero spend (full balance) - split the two jammed imports in settings/index.tsx * feat(token-plan): drop custom token-plan support Custom token plans (manual baseURL/protocol entry) added complexity for little gain — a one-off provider is better configured directly on the Providers page. Token Plan is now preset-only: - remove custom mode, manual fields, and the custom card from the UI - collapse effectivePreset back to the selected preset - drop the now-dead removeProvider action + custom-id branch in removeTokenPlan - remove orphaned customGroup/customName/customHint i18n keys (8 locales) * feat(token-plan): add Volcengine/Tencent/Bailian plans, drop balance feature Add three vendor token-plan presets (all map to existing built-in LLM providers, so it's data-only — no new adapters): - 火山方舟 Volcengine Ark → doubao, OpenAI /api/v3 - 腾讯 TokenHub Token Plan → tencent-hunyuan, OpenAI /plan/v3 (the plan-specific base; /v1 is the pay-as-you-go gateway) - 阿里百炼 Token Plan → qwen, cross-model plan (Qwen + DeepSeek/Kimi/ GLM/MiniMax) on one key; model list is probed/entered Remove the balance/quota feature entirely — we now track usage, not cost, and every vendor's balance query needs its own cloud AK/SK + signature (Volcengine SigV4 / Tencent TC3 / Aliyun BSS), which the Bearer-key billing-endpoint probe never supported anyway: - delete lib/usage/balance-providers.ts, /api/provider/balance, its test - strip the Check Balance button + balance bar from the token-plan page and the provider config panel - remove the 4 balance i18n keys across all 8 locales - restore an eslint-disable the branch had dropped in provider-config-panel * feat(token-plan): use cloud-brand logos for vendor token plans The three vendor plans are cloud offerings, not single-model products, so icon them with the cloud brand rather than a model logo: - 火山方舟 → volcengine.svg (was doubao.svg) - 腾讯 TokenHub → tencentcloud.svg (was hunyuan.svg) - 阿里百炼 → alibabacloud.svg (was bailian.svg) Logos are the colored brand variants from lobehub/lobe-icons, matching the existing colored-logo style (plain <img>, no dark:invert needed). * feat(token-plan): keep only MiniMax and Volcengine presets Trim the token-plan list to the two we want to ship: MiniMax (full-set template) and 火山方舟 Volcengine Ark. Drop the Tencent/Bailian plans and the OpenRouter/SiliconFlow/DeepSeek/GLM/Qwen entries. - remove the now-unused tencentcloud.svg / alibabacloud.svg logos (volcengine.svg stays; the other logos are still used by the provider registry) - retarget the LLM-only apply test from the deleted deepseek preset to volcengine-ark * feat(token-plan): restore aggregator/third-party presets Previous commit over-trimmed: the intent was to drop only the Tencent and Bailian token plans, not the OpenRouter/SiliconFlow/DeepSeek/GLM/Qwen entries. Bring those back; keep only Tencent/Bailian removed. - token_plan: MiniMax, 火山方舟 Volcengine Ark - aggregator: OpenRouter, SiliconFlow - third_party: DeepSeek, GLM, Qwen Revert the apply test back to the deepseek fixture (restored). tencentcloud.svg / alibabacloud.svg stay deleted (their plans are gone). * fix(token-plan): point Volcengine plan at the Coding Plan endpoint The plan's ark--prefixed API keys authenticate only against /api/coding/v3, not the general /api/v3 endpoint — the latter rejects them with "The API key format is incorrect", so model probing returned nothing. Switch the base URL to https://ark.cn-beijing.volces.com/api/coding/v3. * fix(token-plan): Volcengine is an Agent Plan (Anthropic /api/plan) Per the Ark Agent Plan docs, the ark--prefixed keys authenticate ONLY against the dedicated Anthropic-compatible base https://ark.cn-beijing. volces.com/api/plan ("其他 Base URL 无法在 Agent Plan 中使用"). The general /api/v3 and the Coding Plan /api/coding endpoints both reject the key as "API key format is incorrect", which is why model probing kept returning 0. - baseUrl → https://ark.cn-beijing.volces.com/api/plan/v1 (the /v1 lets the Anthropic SDK land on /api/plan/v1/messages) - apiFormat → anthropic - rename to 火山方舟 Agent Plan Probe still targets /api/plan/v1/models (the path exists); if the Anthropic gateway doesn't return an OpenAI-shaped list, users fall back to typing a model id like ark-code-latest. * fix(token-plan): Volcengine Agent Plan = OpenAI /api/plan/v3 + ark-code-latest Settled after probing the real key and reading cc-switch's approach: - The ark- plan key works on the OpenAI-compatible /api/plan/v3 endpoint (chat/completions returns 200); switch apiFormat back to openai. - The plan exposes NO /models list (every /api/plan/*/models is 404), which is why probing kept returning 0. cc-switch handles this by hardcoding a single ark-code-latest (an auto-routing alias valid on any tier) and does NOT use AK/SK for model listing — so we do the same. - Seed defaultModels: ['ark-code-latest'] only; users add specific ids by hand. Supporting machinery (kept, general-purpose): - applyTokenPlan seeds models from defaultModels instead of wiping to [] - handleApply uses defaultModels and skips the doomed probe when present - drop stray .playwright-mcp/ debug artifacts and gitignore them * style: fix prettier formatting in usage files CI runs prettier on the whole repo (prettier . --check); these four files predate this branch's formatting pass and tripped the check. * feat(token-plan): verify Volcengine Agent Plan's published model set The Agent Plan publishes a fixed model set but exposes no /models endpoint, so carry the documented models as CANDIDATES and verify each on apply: - add verifyModels flag to TokenPlanModalityTarget - new /api/provider/probe-chat-models route: sends a minimal chat request per candidate (OpenAI /chat/completions or Anthropic /messages) in parallel, returns the subset that succeeds; SSRF-guarded, auth-failure short-circuits - handleApply gains a verify branch (before the fixed-defaultModels fast path), falling back to the seeded list if verification fails - Volcengine preset now carries the 12 published Agent Plan text models (doubao-seed-2.0-*/deepseek-v4-*/minimax-m*/glm-5.2/kimi-k2.*) as candidates This auto-prunes retired (docs flag deepseek-v3.2/glm-5.1 as 即将下线) and tier-gated models without code changes. Verified all 12 resolve against a real plan key. * feat(token-plan): wire Volcengine Agent Plan image + video modalities Make the Ark seedream/seedance adapters path-configurable and light up the image/video modalities on the Volcengine plan: - seedream/seedance adapters: resolveArkRoot() uses baseUrl verbatim when it already carries an /api/... path (token plan's /api/plan/v3), else appends the standard /api/v3 — no regression for the pay-as-you-go default host. - applyTokenPlan: image/video branches inject a modality's defaultModels as customModels and set them as the active provider+model, so generation works out of the box. New optional setImageProvider/ModelId + setVideoProvider/ ModelId actions (UI passes the store setters; tests omit them). - Volcengine preset declares image (doubao-seedream-5.0-lite, verified 200 on /api/plan/v3/images/generations) and video (doubao-seedance-2.0/1.5-pro — Medium+ tiers only; lower tiers reject at call time, no code change needed to upgrade). Applying the plan overwrites the shared seedream/seedance slot with the plan config (same overwrite model as LLM); switching back to pay-as-you-go is a manual edit or plan removal. Verified image end-to-end with a real plan key. * feat(token-plan): verify image/video models on apply, disable unsupported tiers The Volcengine plan lit up video optimistically, but lower tiers (Small) don't include video — so using it 404'd with UnsupportedModel. Probe media models on apply and only keep what the tier actually supports: - generalize /api/provider/probe-chat-models with a `kind` (chat|image|video): image hits /images/generations, video hits /contents/generations/tasks with empty content. The model-support check (404 UnsupportedModel) runs before any billable work, so probing never starts a real image/video job; for media, "supported" = any non-404 response. - handleApply: after lighting up image/video, probe each verifyModels modality; prune to the verified model set + re-select a working model, or disable the modality entirely if none pass (no false "available"). - Volcengine preset: image/video targets gain verifyModels: true. - add settings.tokenPlan.tierUnsupported across 8 locales. Verified with a real Small-tier key: image (seedream-5.0-lite) passes and is kept; video (seedance-2.0/1.5-pro) 404s and is disabled. * feat(web-search): add Doubao (豆包搜索) provider Doubao Search (Custom 版) over its REST endpoint POST open.feedcoopapi.com/search_api/web_search with Bearer auth — the same endpoint the askecho-search-infinity MCP server wraps, so the Volcengine Agent Plan key authenticates directly. Mirrors the MiniMax adapter: maps Result.WebResults to WebSearchSource (prefers Summary, the query-relevant excerpt, over Snippet for LLM use) and surfaces errors from ResponseMetadata.Error. - register 'doubao' in WebSearchProviderId + WEB_SEARCH_PROVIDERS - searchWithDoubao adapter, searchWeb dispatch, store default config - SSRF allowlist entry for the search host * feat(audio): support Agent Plan single-key auth for Doubao TTS generateDoubaoTTS now picks auth + endpoint from the key shape, since Volcengine exposes Seed-TTS as two products with separate credentials (verified: a plan key 401s on the normal endpoint, and the plan endpoint rejects appId-style auth): - single key (no colon) -> X-Api-Key, for the Agent Plan /plan endpoint - appId:accessKey -> X-Api-App-Id + X-Api-Access-Key (unchanged) A malformed pair (empty half) fails clearly instead of sending an empty header. Reuses the existing NDJSON/base64-mp3 parsing and voice list. * fix(media): map MiniMax video 720p to its real 768P tier normalizeVideoOptions defaults minimax-video to '720p' (the first supported resolution), but Hailuo 2.3 only accepts 768P/1080P and rejects 720P with '2013 ... does not support resolution 720P'. MiniMax's mid tier is 768P, not 720P (the adapter already falls back to 768P, as does the connectivity test), so map the shared enum's '720p' to 768P. Regression tests lock the mapping. * feat(token-plan): add web search + TTS to Volcengine Agent Plan, widen image tiers Extend the volcengine-ark preset now that the adapters exist: - webSearch -> doubao (own host open.feedcoopapi.com, not the ark endpoint) - tts -> doubao-tts on the /api/plan/tts endpoint (single-key auth) - image defaultModels widened to a best-first Seedream 5.0/4.5/4.0 list so a higher tier keeps the strongest model while verifyModels prunes the rest; video keeps the 2.0 + 1.5-pro candidates Comments record the verified host/auth quirks of each modality. * feat(token-plan): show result panel only after probing, with two clear states Addresses review feedback that a green check implied generation works when it only meant 'configured'. The panel now renders after probing finishes (gated on results && !applying) so it reflects the final set, and uses two states: green when the modality is configured/usable, muted when a live probe proved it unavailable (e.g. video on a tier without it). * feat(token-plan): scope presets to true multi-modal token plans Drop the single-modality LLM presets (OpenRouter, SiliconFlow, DeepSeek, GLM, Qwen) from Token Plan. A token plan's defining trait is one key spanning many modalities; those entries are ordinary LLM API providers already covered by the add-provider flow, and listing them here muddied the 'one key, every modality' promise. Only MiniMax and the Volcengine Ark Agent Plan remain. The UI already hides categories with no entries. apply-token-plan's LLM-only test now uses a local fixture instead of the removed deepseek preset. * feat(token-plan): progressive reveal of probe results on apply The result panel previously rendered all at once after probing finished, reading as dead air during the model probe. Now rows appear immediately on Apply: modalities with a live probe in flight show a spinner ('pending') and resolve to lit/failed independently as each probe returns, while non-probe modalities show lit right away. Probes run in parallel (Promise.all) instead of sequentially. A row only turns green once its own probe confirms, so this reveals structure + live progress without a premature green — complementing the earlier 'render only after probing' intent rather than reverting it. Per review feedback from @wyuc on #784. * fix(token-plan): enrich seeded models with built-in thinking capability Token Plan built ModelInfo objects from probed ids via modelInfoFromId(), filling only streaming/tools/vision — so a model that supports configurable thinking lost capabilities.thinking and InlineThinkingControl was hidden. modelInfoFromId now takes an optional providerId and overlays the catalog thinking capability for that (provider, model) pair; applyTokenPlan does the same for its synchronously-seeded list. Added the Ark Agent Plan's dotted aliases to the metadata table: - native Doubao Seed 2.0 family (doubao-seed-2.0-pro/code/lite/mini) - cross-vendor models the plan serves through its OpenAI-compatible endpoint (deepseek-v4-pro/flash, glm-5.2, kimi-k2.7-code/k2.6, minimax-m3/m2.7, ark-code-latest) All verified against a live plan key: each accepts the gateway's unified reasoning_effort field (low/medium/high) and actually reasons. They share the doubao effort adapter, which disables via 'minimal' (not 'none') — matching what the plan endpoint accepts (it rejects reasoning_effort:'none'). Addresses review point #1 from @wyuc on #784. * Improve token plan capability setup UI * fix token plan setup flow * chore: prettier format tts-providers.ts | 2 个月前 | |
feat(media): write generated media through the asset pool under server-backed persistence (#1392) * feat(media): store generated media in the asset pool when persistence is server-backed With server-backed persistence the document is durable and shared, but generated media stayed in the producing browser: the document kept its gen_img_* / gen_vid_* placeholder and narration kept a browser-derived audio id. Every new browser that opened such a course re-ran generation for every slide, and it never converged, because the address of the generated bytes was never written back into the document. Under server-backed persistence only, the classic generation chain now stores bytes in the asset pool first and writes the id the pool allocated into the document. - The client bootstrap configures the asset seam alongside the document and runtime seams: an HttpAssetStore over the persistence endpoint carrying the same credentials the document store carries, marked server-backed. The seam preflight now covers all three, so a failure still cannot half-configure persistence. - Image, video and TTS generation commit in one fixed order: provider, pool, document, local cache, task. A reference reaches the document only after put returned an id, so a document can never name bytes that were not stored. A failure before the write-back leaves the placeholder with the provider called exactly once; the retry happens on the next owner load. - The write-back is a per-slot rewrite through mutateDocument, which re-reads the current document under the per-stage lock, so it cannot clobber a newer scene. The open course is refreshed with the same rewrite without being marked dirty. - "Has this already been generated?" is answered by the document (the slide exists and no longer holds the placeholder) instead of by this browser's task table. - The classroom's resume effect fails closed on ownership: only a resolved owner starts generation, so a viewer opening a shared course spends nothing. - The local media and audio tables become a per-tab cache. A failed cache write costs a re-download, never the media. Browser-only mode is unchanged: every new call site sits behind the server-backed gate, the local tables stay authoritative there, and placeholders stay in the document. Rendering and export needed no changes. HttpAssetStore.resolve mints an object URL exactly as the browser store does, and the export byte resolver was already pool-first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(classroom): make the generation owner gate a three-outcome rule and apply it everywhere The gate refused everything but a resolved owner, which read a sidecar that answered "no ownership fact exists for this course" as a reason to block. That is the answer a deployment without the sidecar's server-side prerequisites gives for every course, and the answer a course with no ownership record gives: in both, there is nobody the operator's budget needs protecting from, and refusing strands the course's own author behind a question that can never be answered. Ownership is now four states over the sidecar's three outcomes. A definite answer splits into owner and not-owner. An absent record is its own answer, ownerless, and generation proceeds — the behaviour such a deployment had before the gate existed. Only the absence of an answer, a transport failure or a load that has not asked yet, stays unresolved and fails closed: "we could not ask" must never be read as "nobody owns this". One mapper turns a sidecar result into that state, and one predicate decides on it. The workbench classroom pane runs the same resume effect and had no ownership input at all, so a viewer opening a shared course there could still spend the budget. It now asks the sidecar once per course, in parallel with its load and feeding only the generation gate, so its read-only and edit behaviour is unchanged. The shared progressive-load policy carries the gate for it, with both new inputs required rather than defaulted so a future caller cannot omit them into an open budget. Its stale comment claiming ownership could not be expressed here is corrected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the write-back survive autosave, arrive before the scene does, and never leak Independent reviews of the write-back found three ways a durable document could still end up naming a placeholder, and two ways the gate that protects the operator's budget could be walked around. An autosave round captures the store synchronously and writes that capture, so a round already in flight when a rewrite landed wrote the placeholder straight back over the allocated id, and nothing marked the store dirty again to correct it. The rewrite now marks the units it changed, which leaves a corrective flush queued behind the stale one; re-saving a scene that already holds the id is idempotent, losing the id is not. Media is generated from outlines in parallel with scene content and usually finishes first, so the slide that will carry the placeholder does not exist yet and the write-back has nothing to rewrite. That was the ordinary path, not a tail case, and its result was discarded: the task was marked done, the scene was added afterwards with its placeholder intact, and a second pass in the same run could call the provider again. The allocation is now held under the placeholder — which also answers the skip test, so nothing pays twice — and applied when that scene is committed, before its first save. One complete pass now leaves no placeholder behind. A failed commit used to abandon what it had already allocated. A poster upload that failed threw away a stored video and sent the retry to submit the most expensive job in the system again; a rejected write-back left registry rows that name bytes nothing references, which the byte collector cannot reclaim because it only collects blobs no row names. A poster failure now costs the poster, and a write-back that reached nothing reclaims what it allocated. A partial write is left alone, because the document already names it. The ownership gate is fail-closed again. Treating the sidecar's 404 as permission was wrong: the client cannot tell "this course has no owner" from "this deployment told me nothing", so a visitor who opened a shared course could bill the operator. The root cause was the sidecar itself, which gated on the agent runtime although every persisted course has an owner regardless — the persistence route resolves one for every request. It now gates on server persistence, so the configuration that made 404 the universal answer has real ownership facts to report, and the gate can refuse everything but a named owner. Retry affordances answered to no gate at all. A viewer of a shared course with one failed image was shown a Retry button that called the provider. Both retry entry points and every surface that draws them now read one shared permission, so what is offered and what is allowed are the same value. Also: narration regeneration no longer pretends it can replace bytes behind a live id — the exclusivity proof that would allow it is refused by construction once references leave the browser, so it forks to a fresh id and says so; the "already generated" test lets a finished deck answer from the document alone, since scene order stops identifying an outline once slides are inserted or deleted; stored assets record a specific media type rather than a generic transfer type; the pane no longer asks the sidecar in browser-only mode; and the funnel's docstring now states what the per-stage lock actually guarantees, which is same-browser serialization and not a cross-browser compare-and-swap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): park allocations in the deciding turn, never reclaim on an ambiguous write A delta review of the write-back found the first-pass fix still had a window, and the reclamation it added could delete media the document already names. The allocation was parked after an awaited local cache write. A scene committed in that window reconciled against a registry that did not hold it yet, so the document kept the placeholder — and the entry recorded a moment later then answered the skip test as "already handled", so nothing could correct it. Parking now happens inside the write-back, in the same synchronous turn as the decision that nothing could take the reference; no await separates the live check from the park. Allocations parked by an earlier pass are handed to their slides at the start of the next one, before anything decides what still needs generating, so a held allocation whose scene has since arrived becomes a rewrite rather than an answer. Reclaiming on a rejected write was unsound: a rejection does not prove the server did not apply the write, so deleting the asset could break the scene that now names it. The funnel decides instead, and says so: it reclaims only when no store write was ever issued and nothing took the reference. Anything else is placed if its slide exists and parked if it does not, so the next pass reuses the bytes instead of paying for them again. When a write fails after part of it landed, the live store is brought up to the document before the error is rethrown — otherwise the next ordinary flush would overwrite the half that did land, with the ids deliberately not reclaimed. Parked allocations are now cleared with the course. Classic placeholders are reused across runs, so one surviving an interrupted run would be handed to a different slide of the next deck: the previous picture, on a slide whose provider was never asked. Both classroom surfaces clear the arriving course, the deletion cascade clears the deleted one, and clearing the database clears them all. Two more ways generation could start without asking the gate are closed. An overlapping pass — an outline retry re-enters generation with every outline while the first is still working — re-requested elements whose provider call was already in flight; a task that is not done is an answered request, not an unanswered one. And narration regeneration in the timeline editor called the TTS provider and allocated a pool asset with no ownership check at all; it now reads the same permission, which withholds both the per-line and whole-timeline controls and refuses the call. Finally, a pane opened during the stage-link availability gap recorded the sidecar's 404 for a course that was moments from existing and never asked again, leaving the real owner locked out of generation until it remounted. Ownership is re-fetched once the document becomes available; the gate stays closed until an answer arrives, so asking again can only open it for someone entitled to it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the asset routes reachable, and stale snapshots harmless A full-branch audit found that the deployment this project documents could not store a single generated asset, and that several routes into durable storage could still write a placeholder over a reference that had already landed. The persistence route sent asset requests through the development authenticator, which refuses outright in a production build that has not explicitly opted into it — and the documented server-persistence recipe produces exactly that build. Every store and every read answered 401, so images and video failed on every slide while re-billing the provider on each retry, and a narration failure stopped the deck at its first slide. Assets live in one shared partition by design, so there was never anything per-caller for that authenticator to decide: the route now resolves the asset principal itself, alongside the owner it already resolves for documents. Runtime sessions are genuinely per-learner and keep the development authenticator until real session verification replaces it. And narration that cannot be stored no longer fails its scene: the line stays unvoiced and retryable, which is what an image that cannot be stored does to its slide. Placeholders could also come back from behind. A queued autosave's snapshot, an editor-history entry replayed by an undo, the departing save a course switch flushes — each captures content at its own moment, and any of those moments can predate a write-back. Point fixes at each producer would leave the next producer to rediscover the bug, so the check lives at the write boundary every producer passes through, and the allocation record it consults now outlives the parked queue: a placeholder whose rewrite landed long ago is exactly the case it catches. Two ways generation could be lost or repeated are closed. A pass now claims the elements it will reach and releases them however it ends, so an overlapping pass stands down while an aborted one strands nothing — previously its tasks stayed `pending` and every later pass skipped them with no retry control to recover them. And the media abort controller is aborted before being replaced, so a superseded pass stops calling providers instead of running on for a course the user has left. The remaining two are narrower. The workbench pane asks for ownership only after a document load succeeds, and after every later one, mirroring the page route: the load is what creates the ownership row the first time a course is opened, so asking beforehand asked about a course that did not exist yet and locked its author out for the mount. And the ownership gate on the timeline editor now withholds narration regeneration alone; listening back to existing narration and seeing whether a line has any spend nothing and stay available. Known limitation, unchanged and now stated plainly in the comments that used to point at it as a solution: nothing reclaims an unreferenced pool asset. The registry sweep is written but not wired up, and the byte collector only reclaims blobs no registry row names, so every narration regeneration and every abandoned allocation leaves storage behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): gate asset mutations, and make claims and allocation records survive a handoff Opening the asset routes opened all of them. Reads and allocations are meant to be as open as document reads and creates already are, but no authorization hook was supplied, so the handler's default admitted PUT and DELETE too — and those scope by principal key alone, which is one shared constant. Any caller who learned an id, and a document read hands out every id its slides name, could overwrite or destroy another author's media. Mutations now require the deployment's credential, which in a production build without the development-auth opt-in means they are refused outright; reads and allocations stay open. The route comment says what the posture is and what it is not: the deployment-level fence is the access code, and no per-principal quota is configured. The client's own reclaim is best effort to match — losing an argument about deleting an asset must not cost a task its retry, and the bytes are left for server-side reclamation. The pass claim could not survive the handoff it was written for. A retry aborts the live media pass and starts its replacement in the same synchronous block, long before the aborted pass's cleanup runs, so the replacement saw every element still claimed, collected nothing, and returned — leaving each unreached element at pending with nobody coming back for it and no retry control to recover it, which is the exact failure the claim was introduced to prevent. A claim now carries its pass's signal and is retired the moment that signal aborts, and a pass releases only claims it still owns, so a late unwind cannot take its replacement's work. Claims are also acquired at the single point every request passes through, so a single-task retry participates too — previously a retry awaiting its provider was invisible to a pass starting alongside it and both called it. The allocation record could outlive the bytes it named. It was written before the write-back attempted anything and survived the reclaim that followed a failure, so when the slide finally arrived the write boundary stamped a deleted id into the document — and the placeholder it replaced was gone, which reads as already generated and stops anything from retrying. The record is now written only where the allocation is retained, and forgotten wherever a reclaim removes the bytes, including the narration rollback path. The tests follow. The route test drives the real storage handler against an in-memory registry instead of a stub, so it can see what the resolved principal is then allowed to do; the handoff test performs a real abort mid-pass rather than starting from an already-aborted signal; and the guards that could only assert file layout now assert the property they care about, or have been replaced by behaviour. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make media passes serial per course instead of tracking element ownership Three rounds of per-element claims each produced a new way to lose an element. Whole-pass reservations swallowed a Retry for an element the same pass had already failed, leaving it pending with the affordance gone. Retiring a claim by its signal freed an element whose commit was still uploading, so the replacement pass paid for it twice. A claim held for a failed element stranded its retry. The bookkeeping is the defect: every refinement of "who owns this element right now" answered the question at a moment when the answer was already stale. Passes for one course are now serial. A replacement aborts its predecessor, as before, and then waits for it to settle before collecting. That removes the question entirely: a commit already under way finishes — its bytes stored and its reference written, so the new pass sees a resolved slide and skips it — and an element the aborted pass never reached is still a placeholder and gets collected like any other. The claim set, the reservations, the signal retirement and the identity-checked release are all gone. The task table is consulted for one thing only: an element that is generating right now is a single-element retry running alongside the pass, and taking it too would pay twice. Pending is deliberately not a skip reason — it means a pass once intended to reach an element, which an abandoned pass leaves behind with nobody acting on it, and reading that as answered is what stranded elements before. A retry runs concurrently with a pass, because a pass never revisits an element it has processed, and it re-reads the task after its own await and refuses before touching it: marking first and refusing afterwards destroyed the failed state that draws the affordance. Browser-only mode is back to exactly what it was. The abort is now conditional, the waiting does not apply, and the original status-based skip is restored verbatim. Two baseline lines remain changed in each of the two files, and both are behind a server-backed fork whose else-branch is the original. Two smaller things. The allocation record becomes visible when a write goes on the wire rather than when the round trip ends, and the write boundary reconciles under the document lock rather than before it — a save queued during a write-back was otherwise captured with the placeholder and, for a course the user had left, had no corrective flush to follow. And the comments that said a refused reclaim leaves its bytes for server-side reclamation were wrong: nothing collects them, because the registry entry still names its blob and the sweep that would remove it is not wired up. They now say the bytes leak. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a deferred pass re-earn its right to run, and bound the commit it waits on Serializing passes moved their body out of the block that launched them, and three things followed from that. A pass now wakes when its predecessor settles, which can be after the user has left the course. It enqueued before it looked at its signal, into a task table keyed by element id alone — and placeholder ids are not unique across courses, which is why the classroom clears that table on arrival. So a departing course's pass seeded the arriving course's table with tasks carrying the wrong stage id, and a Retry routes by that id: the reference went into the wrong document. A pass now re-validates after the wait, before touching anything shared. The same lateness broke the skip test. `documentSkipIndex` answers only while the live store is on the pass's stage, and returning nothing put the collection loop on the browser-only rule — a silent demotion from "the document is the authority" to "this browser's task table is", on exactly the path where that table has just been cleared. Every element the predecessor had committed was collected again, paid for again, and its second write-back found no placeholder to rewrite, so its bytes were parked where nothing will ever reference them. In server-backed mode an unreadable document now means the pass stands down. And waiting was unbounded. A commit is uncancellable: the asset client takes no signal, and a document write cannot be half-undone. One stalled upload therefore froze the course's media generation for the session — the replacement never collected, the element sat on a skeleton that draws no Retry, and only a reload recovered. The pass's signal is now threaded into the media proxy fetch, and the commit is bounded by a deadline. The deadline is on the wait, not the work: the commit carries on, and if it lands late the document simply ends up correct, while the element becomes retryable and the queue moves on. The tests that were meant to pin the previous round were not sensitive to it. Two asserted end states where the mechanism only changes ordering, and one of them rigged the document read so the assertion held whether or not the pass had waited; a third covered half of what it claimed. They now observe the ordering directly — nothing is issued while another pass for the course is working; in browser-only mode a second pass reaches its provider immediately — and the reconciliation under the document lock has a test that fails when it moves back outside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * revert(media): drop the commit deadline and the abortable download The deadline bought less than it cost. Abandoning a commit after two minutes makes the element retryable while the real commit is still running, so a Retry starts a second commit for the same placeholder against the first: two provider calls, two allocations, and whichever lands second stamps its result over the other's task by element id. The allocation record is keyed by placeholder, so the loser's cleanup erases the winner's record, and the write boundary then puts the raw placeholder back into the document. That is the overlap serial passes were built to remove, reopened through the one door serialization never covered. So a stalled commit holds the course's media queue until it settles or the page is reloaded, and that is written down rather than papered over. The wait is unbounded on purpose: every ceiling on it turns out to be a way of running two commits for one element. Threading the pass signal into the download was also a mistake, in the other direction. The provider call that produced the URL has already been billed, so cancelling the download throws away work that is paid for — and the shared proxy cache records a cancelled request as a transient failure against that URL, which after three of them blocks it for every consumer in the session. Browser-only mode never asked for this: it had no way to observe an abort there, which is exactly why the bytes were kept. The signal is gone from the download again, and `fetchAsBlob` is byte-for-byte what it was before this branch. The regression guard for the stranded-element rule is restored alongside the timing test that was meant to supersede it. It catches a different rule — a task left pending being read as answered — and nothing else does: making the pass skip pending leaves every other suite green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): stop asking the pool for refs it never issued, bound it, and adopt cached bytes Four things a deployment found once this was running for real. A reference this application mints itself — a generation placeholder, a derived narration key — was never in the pool, because the pool allocates every id it holds. Asking anyway used to be an IndexedDB miss; once the pool is server-backed it is a request that answers 404, one per element per load, forever on a course that still holds placeholders. Every lease and probe now checks first. The check is a negative test on shapes this application owns, not an id validator: the pool's id domain stays unconstrained, and anything that is not one of ours is still asked about. The asset store can bound how much one principal holds, and enforces it inside the write transaction, but nothing ever passed the number. It does now, with a default rather than an opt-in: allocation is reachable by any caller a deployment admits, and with one shared principal an unbounded store is unbounded database growth with no operator-visible brake. Refusing asset mutations to unauthenticated callers was not enough, because every authenticated caller resolves to that same shared principal — so authentication decided nothing, and any signed-in visitor could delete any id they learned. Since this branch began storing media the registry is the only copy a course has. Replacing and deleting are now refused to everyone, and the browser no longer tries: an entry nothing references waits for server-side reclamation instead. What a browser must still do is forget its own record of an allocation that reached nothing, or a later save would stamp an id the document has no reason to trust. And a course generated before any of this holds placeholders in its document with its bytes only in the author's browser. Those bytes are paid for, so the author's next load converts them — stored to the pool and written back through the ordinary commit path, with no provider call — instead of buying them again. A row that records only a hosted URL is treated as absent: that URL is the provider's address, not something a document may hold. One renderer expectation moved with this. An untracked placeholder used to paint as pending on first render because asking the pool left a lease in flight; it settled to disabled a moment later either way, and now says so from the start. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): surface a full store as a refusal and convert legacy narration A quota refusal reached the browser as HTTP 500 with a generic message, which reads as a transient failure: the element kept a Retry that would pay a provider again and be refused again. The store raises the contract's own error and the handler maps it to 507, but the store answering a request is not always built by the same bundle as the handler -- the persistence provider is reached from the route bundle and from instrumentation, which is why its state lives on a Symbol.for global -- and `instanceof` is false across that boundary while the declared code is still right. Classify on the code as well as the class, and make the code a permanent, persisted refusal in the browser: recorded locally so it survives a reload, shown as "storage is full", and refused by the retry entry point so a stale button cannot buy a second generation. Every other storage failure stays retryable. Convert what a pre-server-backed course still holds. Generated media is adopted under either key this application has used for it -- the placeholder, and the allocated id of a course converted once and later rolled back -- instead of only the first. Narration is converted by a load-time pass over the open course's speech actions, since nothing re-enters generation for an action that already has an id: bytes to the pool, id written back through a funnel that mirrors the media one, owner-only and server-backed-only. A line whose bytes are in no browser is left alone rather than re-synthesized. Also: the pool guard is now a positive `ast_` test rather than an enumeration of the shapes we mint (imports never reach the pool, so this is safe in both modes); the slide ref collection is an exported pure function so its four lease sites are covered behaviourally; ASSET_QUOTA_BYTES treats every spelling of zero as opting out and refuses a malformed value at startup instead of falling back; the abort signal is re-checked after the cache read, before an uncancellable commit; and the unused `removeAsset` and pool `replace` surfaces are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * chore(storage): release 0.29.1 The asset HTTP handler now recognises a store refusal by the contract code it declares as well as by its class, so a quota refusal raised in another module realm answers 507 instead of 500. Same contract, stricter recognition, no API change: a patch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a full store recoverable and adoption course-safe Narration adoption read the local audio row by its derived key alone. That key carries no stage id and the table is keyed by id alone, so two courses can mint the same one -- a PPTX import numbers its scenes and actions deterministically, which gives every imported deck's first slide `tts_s1_speech-scene-p1`. Locally a collision only means one course plays another's clip in one browser; adopting it wrote that clip into the shared document permanently, for every device and every visitor. A row that names a course is now adopted only into that course, and a row from before that column existed only when the text it recorded is the text of the action being converted. A full asset store was made permanent last round, which was wrong three times over: it overwrote the refused bytes with an empty blob -- on the conversion path that row is a course's only copy of its own media -- it kept sending the rest of the deck to a provider against a ceiling it already knew was reached, and it left no way back once an operator raised that ceiling. A full store is neither the content's fault nor the configuration's, so it is now its own case: the bytes are kept, the pass stops at the first refusal, and the element shows the reason together with a Retry that re-attempts the upload from those bytes. Nothing retries automatically, so no one is re-billed. The narration write-back now reaches the write boundary every producer of a durable write passes through, not only the dirty mark: adoption never deletes the derived row, so a snapshot that reverts the rewrite is adopted again on the next load and allocates a fresh asset every time. Adoption is also mounted by both classroom surfaces rather than one, takes the course's abort signal, and re-validates that this browser still has the course open before each write. ASSET_QUOTA_BYTES is validated from instrumentation, where the README and the docstring already claimed it was: its only other consumer is lazy and memoised, so a malformed ceiling let the process boot and then failed every persistence request, documents and runtime included. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): remember a full store per course, and never lose retained bytes A stopped pass left the elements it never reached as placeholders with no persisted record -- deliberately, since nothing was attempted for them. But that left the next load with no reason not to try: it called a provider for the next placeholder and was refused at exactly the same point, once per reload, indefinitely. A full store is not a property of any slide. It belongs to the deployment and changes for reasons the document knows nothing about, so it is now remembered once per course in the browser's device KV. A pass that finds the marker stands down before spending anything and leaves every placeholder its "storage is full" state and its Retry; the first upload that succeeds clears it and the next pass runs normally. Narration adoption latched per course so it runs once per load, and the latch outlived the abort that leaving a course performs. On a surface that stays mounted across switches -- the workbench pane is one component for every course it shows -- owner course A, visitor course B, then back to A skipped exactly the clips the abort had cut off, and nothing else converts them. The latch is released with the abort now, and a course adopts one run at a time so a re-entry cannot hand a clip a second allocation while the previous run's uncancellable tail is still settling. A quota-blocked element retried into a network error or a 500 lost the bytes that were kept for it: the retry deleted the row before attempting the upload and wrote no replacement for an error carrying no structured code, so the next retry went back to a provider for media this browser had a moment earlier. The row now survives until an upload succeeds, the failure handler keeps whatever bytes the attempt was given, and the retry asks the question a pass asks -- does this browser already hold bytes for this element -- rather than reading an error code that a second failure has already overwritten. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): adopt real legacy narration, queue re-entries, report attempt outcomes Narration adoption admitted a stage-less row only when the text it recorded matched the action being converted. Both of those columns were added to the local audio table by the very change that moved narration onto allocated ids, so a row still carrying a derived key has neither: the rule refused every real pre-allocation course and passed only on fixtures built from post-allocation rows. What the row cannot say, the key can. A derived key names two clips only when two courses share a scene order and an action id, and an action id repeats only when something other than the generator minted it -- an import numbers them by slide position. So a key built from a generated action id is adopted on that basis, a key an import could have reproduced still needs matching text, and a row that names another course is refused however unique its key looks. Handing a re-entering caller the adoption run already in flight undid the latch release it was paired with: that run is bound to the signal the departure just aborted, so it stops at its next clip while the caller -- which has the course open and a live signal -- is told the work is done, and an effect replayed as mount, cleanup, mount adopts nothing at all. A later caller now waits for the uncancellable tail and scans again, which costs a lookup on a course that has nothing left and finishes the clips the abort cut off on one that does. One attempt at an element now reports both facts its callers need instead of a bare boolean: whether the store refused it for room, and whether bytes actually reached the store. Leaving a course clears the task table, so a retry that landed afterwards read "no failed task" as success and deleted the row holding the only copy of the media. Nothing is inferred from that table any more. Reading the localStorage property can throw where storage is denied by policy, typeof included, so the availability check moved inside the guard: this metadata is best-effort, and a rejection here strands a generation pass that has already enqueued its tasks. A retry is never blocked by the per-course "store is full" marker, but a retry that is refused again re-sets it, and adoption now reads and writes the same marker rather than issuing one refused upload per clip on every load. The two canvas element renderers and both thumbnail renderers show the reason beside the Retry, so a full store does not look like an ordinary failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): probe a full store instead of standing down, and pair notices with a Retry Narration adoption was given both halves of the per-course "the store is full" marker last round: it stood down when the marker was set, and it set the marker when its own upload was refused for room. Those halves are only safe together if something can lift the marker, and for adoption nothing could. It has no affordance of its own, it stood down before reaching its own clear, the media pass returns before its marker gate when there is nothing to generate -- so a narration-only deck, or one whose slides are already satisfied, painted no storage-full element and offered no Retry -- and narration generated rather than adopted allocates directly rather than through the media commit. The course's cached narration was then lost for good, where before it converted on the first load after the ceiling was raised. The gate is a probe now. A marked course attempts exactly one clip per load: refused, it stops and the marker stands, which costs what standing down cost; stored, it lifts the marker and finishes the course. Adoption spends no provider money, so the whole cost of probing a store that is still full is one refused upload. Generated narration lifts the marker too. The three surfaces that gained a failure notice last round drew it for any failure with a reason, including the one refusal that is reachable without server-backed persistence, so a browser-only deck painted something it had not painted before. The notice is drawn beside a Retry and nowhere else, which is what it was added for and what leaves browser-only output unchanged. Both are now asserted through the render harness the surface matrix already had. A caller arriving while a rescan is queued shares it rather than appending another. One rescan converts whatever the run in flight left and every later one would find an allocated id on every action, so a chain bought nothing and turned a single stalled upload into a course that never adopts again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): treat a refusal for room as a fact about one clip, not the deck The asset store checks each write against the headroom it has left, so a store that refuses a long opening clip can still hold every short clip behind it. Narration adoption assumed the opposite: it broke the deck at the first refusal and then re-attempted that same first clip on every later load, because the document names it first. A deck whose longest clip exceeds current headroom therefore never converted the clips that would have fit, with no affordance to recover it -- the state the probe was introduced to remove, reached through a narrower door. An unmarked load now attempts every clip, skipping the ones that do not fit, and remembers the condition only if the load ends with clips it still could not store. A marked load spends its single upload on the smallest clip left rather than the first one named: that is the clip that answers the question the marker asks, because if the smallest does not fit nothing does. The media pass keeps stopping at its first refusal, and for a reason adoption does not share -- every element it attempts costs a provider call. A rescan several callers share took the newest caller's signal, and the newest caller is not necessarily the one still there: a surface that opened a course and closed it again would stop work a surface still showing that course was waiting for, and that surface is latched, so it would never ask again. The shared run now takes a signal that is aborted only once every caller has left. The comment claiming the shared rescan contains a stalled upload was wrong -- the rescan is chained off the run in flight, so a stalled upload leaves every caller pending exactly as a chain would. It claims the bounded queue it actually provides, and the stall is recorded as a limitation. The failed-state containers took their stacking classes unconditionally, so markup differed in browser-only mode even though nothing moved on screen. Those classes are applied only when there is a notice to stack, and the tests assert the exact class attribute rather than a substring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): stop narration adoption writing the media pass's store-full marker The marker means "do not call a provider for this course". A path is entitled to write it only if its own refusal cost a provider call, and narration adoption's refusals cost nothing: it uploads bytes this browser already holds. The store also checks each write against the headroom it has left, so a clip that does not fit says nothing about whether a slide's image would. Adoption was writing it anyway, and one over-long narration clip was therefore enough to stand a course's entire image pass down on every later load -- on a store that had just accepted adoption's other clips. The author could still recover each element by hand, every load, for ever. Three rounds of narrowing this seam produced a finding each time, so it is removed rather than narrowed again. Gone: the marker read, the single-clip probe, the smallest-clip selection, and the up-front read of every row into an array -- which also retires a sampled-then-stale flag and the retention of a whole deck's blobs for the length of a run, and returns the loop to streaming one row at a time. Adoption's rule is now that every load attempts every clip it holds, once; any failure skips that clip and the load continues. The noise the coupling was meant to avoid does not arise, because after the first load the clips still outstanding are exactly the ones that did not fit -- normally none, or one. A successful write still clears the marker, and that is a different kind of statement: a write that went through is a fact this run established, where a refusal is an inference about what some other write would cost. For a course whose media needs nothing, adoption and generated narration are also the only paths that can establish it. The failure module still documented the deck-wide premise this contradicts. It now says what is true: the check is per write, and the media pass stops the deck as a judgement about cost rather than about certainty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bound a full store's cost from the store's own arithmetic Removing the store-full marker from narration adoption removed its bound too, and the code then asserted the bound was unnecessary. It is, on a store with room for most of a deck. On the store the whole mechanism exists for -- the ceiling reached, nothing fitting -- the outstanding set after every load is the entire deck, so a thirty-clip course posted thirty full blobs on every load, indefinitely. Each of those is not a cheap refusal: the bytes are uploaded, the server hashes the whole payload, and only then takes a per-principal lock and sums every entry that principal owns before saying no. The bound needs no flag, no key and nothing carried between loads. The store asks whether `used + addedBytes` exceeds the ceiling, and `used` only grows while a run is uploading, so a clip refused for want of room implies every clip at least that large is refused for the rest of that run. The run keeps the smallest size it has been refused and skips anything no smaller without uploading it; a smaller clip is still attempted, because it may fit. A deck the store refuses entirely now costs one upload per successive size minimum instead of one per clip, and a deck it has room for costs nothing extra, because nothing is refused. Only a refusal for room lowers the bar: a dropped connection says nothing about how much room there is. The deck-wide certainty premise the failure module retracted last round still stood verbatim at the site that implements the stand-down. Both copies now say the same thing: the check is per write, and the pass stops the deck as a judgement about cost rather than about certainty. The comment on adoption's marker clear now names its price. Narration of a few hundred bytes fits in headroom an image does not, so a proven write can let the next pass buy one more image that is refused again -- bounded at one, and the price of the alternative being a course whose media never generates again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(media): state the adoption bound exactly, and stop three comments describing the old rule The comment introducing the in-load bound gave its cost as "at most a handful, and the first load pays the most". Neither clause is a property of the rule. A clip is skipped only when something no larger was already refused, so a fully-refused deck costs one upload per successive size minimum in document order: one when the clips grow, about ln N for an arbitrary order, and one per clip when they only shrink -- a long opener followed by terser lines is exactly that shape. And no load is cheaper than the first, because the bound resets per run and a refused clip stays outstanding. The comment now says that, and points at what would make it exactly one for any ordering: the store returning its remaining headroom in the refusal's existing details channel, which the server leaves empty today. Two other comments still described the previous rule -- "attempts every clip it holds, every load" -- one of them twenty lines above the paragraph that introduces the bound, in the same block. Both now say what the code does. The bound's soundness is worth stating where a maintainer will look for it: quota is charged at full length with no discount for a duplicate, the sum it is checked against joins entries to blobs so the collector cannot lower it, the check takes a per-principal lock before summing, and replace and delete are refused to every browser. Nothing a run can do makes room appear inside it. One test installed a row implementation and replaced it wholesale a few lines later, so the first was dead and the survivor dropped the text the first clip's import-shaped key needs for the ownership rule -- it passed on the coincidence that the fixture's default text is the action's. Merged into one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): keep a Retry from re-buying parked media, and state the store seam once Five findings from an inline review. The standalone classroom route asked the ownership sidecar once per load and recorded only stage ownership when that ask failed. Every non-answer fails closed, so one transient 5xx left the genuine author with no resume, no Retry affordance and no legacy narration converted for the rest of the load, with nothing to change it short of a reload. The failure now records the fail-closed answer explicitly -- an answer an earlier load established must not outlive the failure that replaced it -- and an unresolved answer is asked for again, a few times over a few seconds. A real answer, however unwelcome, is final. A Retry could pay a provider for media the pool already held. When the bytes are stored and only the write-back fails in a way that keeps the allocation, it is parked and no local row exists, because that row is written only after a successful write-back. Retry now reads the parked queue exactly as the pass does and re-attempts the write-back: it re-keys the task done when the document takes it, leaves the entry parked when the slide still does not exist, and stays failed and retryable when the document refuses again. Object URLs a parked allocation owns are revoked when the entry is dropped. The commit path leaves them alone while the entry is parked, because it is then the only thing holding bytes this tab can render, so a course switch or a stage deletion was pinning the whole blob for the life of the tab. An entry a slide has already taken is left alone: the task table is displaying those URLs. The fallback lookup for cached bytes is a stage-scoped scan, and the keyed lookup misses for every row the commit path writes, so a pass was materializing and sorting the course's whole media table once per element. One scan per pass now, built on the first miss. It is sound and not merely cheaper: an element asks only for its own placeholder, and every row a pass writes carries the placeholder of the element that wrote it. "The store accepted a write, so it is not out of room" was enforced at three call sites under slightly different conditions, which made it a convention the next pool write path could silently break. It is stated once, in putAsset, for the course whose bytes it just stored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 16 天前 | |
feat(media): resolve image, video, and ASR models from server config (#1175) * feat(media): resolve image, video, and ASR models from server config The image generation chain never consulted the server-side IMAGE_<PREFIX>_MODELS config: the route only forwarded the client x-image-model header and each adapter silently fell back to a hardcoded vendor default, so a managed provider could generate with a model the operator never chose. The provider itself also defaulted to a hardcoded vendor id when the client sent no preference. - Add resolveImageModel / resolveVideoModel / resolveASRModel and resolveServerImageProviderId / resolveServerVideoProviderId to the server provider config, mirroring how TTS already resolves server-managed settings. - Routes resolve provider and model server-side when the client sends no preference and fail loud with a clear 400 when nothing resolves; no more hardcoded vendor defaults. - Adapters require an explicit model via requireModel; connectivity probes keep a fixed probe payload since they are auth handshakes, not user-facing generation. - Tests cover the resolver precedence and the fail-loud paths. * fix(media): align model resolution semantics across capabilities resolveImageModel and resolveASRModel let the server pin win unconditionally, while resolveVideoModel implements allowlist semantics (client choice wins when it is in the pinned list, otherwise the first pinned entry). Align image and ASR to the video allowlist semantics so a client picking any server-pinned model is honored, and update the resolver docstrings and tests to pin the aligned behavior for all three capabilities. The classroom media path has no client model and no HTTP response to fail loud with: it now falls back to the first catalog model when the operator pins no _MODELS list, so key-only deployments keep generating media instead of silently skipping every element via the adapter's requireModel backstop. * fix(media): normalize configured model lists and client model ids | 1 个月前 | |
feat(media): add OpenRouter image and video providers (#1356) * feat(media): add OpenRouter image and video providers OpenMAIC ships six separate video providers (Veo, Kling, Seedance, MiniMax, Grok, HappyHorse) and seven image providers, each needing its own key. OpenRouter fronts those same model families behind one key and one account, so this adds it as a provider on both sides. Both use OpenRouter's dedicated media endpoints, not chat-completions: - Image: POST /images -> { data: [{ b64_json }] } - Video: POST /videos -> 202 { id, status }, poll GET /videos/{id}, then GET /videos/{id}/content for the mp4 bytes The model list is fetched live from GET /images/models and GET /videos/models through /api/openrouter-models rather than pinned in the registry: OpenRouter hosts 48 image and 28 video models today and adds more, so a hardcoded shortlist would decide for the operator which models exist. The registry keeps a three-entry seed as an offline fallback, and the existing custom-model UI still accepts any model id. Both catalogs answer unauthenticated, so the picker fills before a key is pasted; a key is forwarded when present for proxied base URLs. Adapter contracts are covered by stubbed-fetch tests (request shape, empty-response handling, and the video job state machine including terminal failure). No test performs a billable call. Closes #1355 * fix(media): validate the key and tolerate a pasted endpoint URL Three fixes found while configuring the new provider: 1. Both connectivity probes hit the model catalogs, which answer 200 unauthenticated — so "Test Connection" reported success for any string, including an invalid key. Probe GET /key instead: equally cheap, and it actually rejects a bad key. 2. The settings field is labelled "Base URL" but the panel echoes it back as "Request URL", so pasting the full endpoint (https://openrouter.ai/api/v1/images) is the natural mistake. That built /api/v1/images/images and 404'd. Trim a trailing slash and a trailing /images or /videos so both forms work; a proxy path that merely contains the word is left alone. 3. The image and video settings panels read `data.message` on a failed test, but failures answer with `error` (apiError) and only successes carry `message`. Every failing connectivity test — for any provider, not just OpenRouter — rendered "connection failed: undefined" instead of the reason. Pre-existing; surfaced by 1 and 2 above. Closes #1355 * fix(media): make every OpenRouter model selectable, and always select a provider Two gaps found while configuring the new provider. The settings Models list is a read-only catalog for every provider; the actual model picker is the media popover. That picker built its groups from the static registry array, so OpenRouter offered only the three-entry seed while settings listed the full live catalog — the models were visible but not choosable. Feed the same live catalog into the popover, fetched only once the provider is usable so an unconfigured install makes no request. Separately, `imageProviderId`/`videoProviderId` are empty until a provider is chosen (first-run auto-config leaves them blank when the server reports no media provider). Opening the settings panel on an empty id selected nothing: the header rendered the missing name key as "settings.undefined", and Test Connection posted a blank x-image-provider/x-video-provider, so it failed with "No image/video provider configured" whatever key was typed. Fall back to the first catalog entry so the panel always has a selection. Pre-existing and not specific to OpenRouter. Closes #1355 * fix(tts): request a browser-playable format from custom providers `generateOpenAITTS` serves every custom OpenAI-compatible TTS provider but never sent `response_format`, so it inherited whatever each provider defaults to. OpenAI defaults to mp3; OpenRouter's /audio/speech defaults to raw `pcm`. The unknown content type then fell through to the `'mp3'` default below, the client built `data:audio/mp3;base64,…` from headerless PCM samples, and playback failed with "no supported source was found" — while the server logged a clean 200, because the audio really was generated. Name the format instead of inheriting it. Also stop mislabelling an unrecognised body: `pcm`/`l16` now raises a message naming the cause, and `aac`/`opus` are recognised. Two supporting fixes: - /api/openrouter-models normalises its base URL the way the adapters do and falls back to the public catalog when a custom base URL fails, so a typo in a free-text settings field cannot empty the model picker. Also types the headers object so tsc accepts the conditional. - provider-neutrality-guard pins exact per-vendor occurrence counts in lib/server/provider-config.ts. Adding the image and video env entries raises "openrouter" from 2 to 6 (each entry contributes both its key and its value); CI failed without the bump. Closes #1355 * fix(security): never send the operator key to a client-chosen host Review found `/api/openrouter-models` was an SSRF and key-exfiltration path, and the finding is correct. The route took `x-base-url` from the caller at highest precedence while preferring the *server* env key, so any caller could make the server send the operator's OpenRouter credential as an `Authorization: Bearer` header to an arbitrary URL. The route's own comment claimed it followed `/api/verify-image-provider`; that pattern runs `validateUrlForSSRF` on client base URLs, and this route did not. The boundary is now explicit: the server key travels only to the operator's own base URL. A client-supplied URL is SSRF-validated and carries only that caller's own `x-api-key` — the server key is dropped — and the unauthenticated public-catalog fallback never forwards a credential chosen for a different host. Redirects are no longer followed (`redirect: 'manual'`), since a redirect would carry the Authorization header off-host and reopen the same hole, and upstream reads are bounded by a timeout. The per-URL cache is now keyed by destination *and* a hash of the credential, and bounded to 64 entries with oldest-first eviction, so client-supplied URLs cannot grow it without limit and one caller's key-authorised catalog is never served to another. Also from the review: - The image adapter discarded the reported `media_type`. The orchestration layer wraps a bare `base64` as `data:image/png` unconditionally, so jpeg/webp results were mislabelled; the adapter now returns a data URL carrying the real type. - Adapter generation and poll requests set `redirect: 'manual'`, matching the `/key` probe that already did. - `runPolledTask` accepts an `AbortSignal` so the sleep between polls is cancellable; the video adapter passes the caller's signal. Without it a cancelled generation still slept out a full 10s interval. Tests cover the highest-risk paths the review named: which credential reaches which URL, that an SSRF-rejected destination is never contacted, that the fallback is unauthenticated, cache isolation between callers, and MIME preservation. Findings 2 (base-URL normalisation) and 4 (neutrality-guard debt) were already fixed in d553a08, pushed after the review was submitted; CI is green on that commit. Closes #1355 * ci: retry flaky voice clone timeout * fix(vercel): keep OpenRouter catalogs within Hobby function limit * fix(vercel): avoid tracing self-hosted sharp binaries --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> | 10 天前 | |
feat(media): write generated media through the asset pool under server-backed persistence (#1392) * feat(media): store generated media in the asset pool when persistence is server-backed With server-backed persistence the document is durable and shared, but generated media stayed in the producing browser: the document kept its gen_img_* / gen_vid_* placeholder and narration kept a browser-derived audio id. Every new browser that opened such a course re-ran generation for every slide, and it never converged, because the address of the generated bytes was never written back into the document. Under server-backed persistence only, the classic generation chain now stores bytes in the asset pool first and writes the id the pool allocated into the document. - The client bootstrap configures the asset seam alongside the document and runtime seams: an HttpAssetStore over the persistence endpoint carrying the same credentials the document store carries, marked server-backed. The seam preflight now covers all three, so a failure still cannot half-configure persistence. - Image, video and TTS generation commit in one fixed order: provider, pool, document, local cache, task. A reference reaches the document only after put returned an id, so a document can never name bytes that were not stored. A failure before the write-back leaves the placeholder with the provider called exactly once; the retry happens on the next owner load. - The write-back is a per-slot rewrite through mutateDocument, which re-reads the current document under the per-stage lock, so it cannot clobber a newer scene. The open course is refreshed with the same rewrite without being marked dirty. - "Has this already been generated?" is answered by the document (the slide exists and no longer holds the placeholder) instead of by this browser's task table. - The classroom's resume effect fails closed on ownership: only a resolved owner starts generation, so a viewer opening a shared course spends nothing. - The local media and audio tables become a per-tab cache. A failed cache write costs a re-download, never the media. Browser-only mode is unchanged: every new call site sits behind the server-backed gate, the local tables stay authoritative there, and placeholders stay in the document. Rendering and export needed no changes. HttpAssetStore.resolve mints an object URL exactly as the browser store does, and the export byte resolver was already pool-first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(classroom): make the generation owner gate a three-outcome rule and apply it everywhere The gate refused everything but a resolved owner, which read a sidecar that answered "no ownership fact exists for this course" as a reason to block. That is the answer a deployment without the sidecar's server-side prerequisites gives for every course, and the answer a course with no ownership record gives: in both, there is nobody the operator's budget needs protecting from, and refusing strands the course's own author behind a question that can never be answered. Ownership is now four states over the sidecar's three outcomes. A definite answer splits into owner and not-owner. An absent record is its own answer, ownerless, and generation proceeds — the behaviour such a deployment had before the gate existed. Only the absence of an answer, a transport failure or a load that has not asked yet, stays unresolved and fails closed: "we could not ask" must never be read as "nobody owns this". One mapper turns a sidecar result into that state, and one predicate decides on it. The workbench classroom pane runs the same resume effect and had no ownership input at all, so a viewer opening a shared course there could still spend the budget. It now asks the sidecar once per course, in parallel with its load and feeding only the generation gate, so its read-only and edit behaviour is unchanged. The shared progressive-load policy carries the gate for it, with both new inputs required rather than defaulted so a future caller cannot omit them into an open budget. Its stale comment claiming ownership could not be expressed here is corrected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the write-back survive autosave, arrive before the scene does, and never leak Independent reviews of the write-back found three ways a durable document could still end up naming a placeholder, and two ways the gate that protects the operator's budget could be walked around. An autosave round captures the store synchronously and writes that capture, so a round already in flight when a rewrite landed wrote the placeholder straight back over the allocated id, and nothing marked the store dirty again to correct it. The rewrite now marks the units it changed, which leaves a corrective flush queued behind the stale one; re-saving a scene that already holds the id is idempotent, losing the id is not. Media is generated from outlines in parallel with scene content and usually finishes first, so the slide that will carry the placeholder does not exist yet and the write-back has nothing to rewrite. That was the ordinary path, not a tail case, and its result was discarded: the task was marked done, the scene was added afterwards with its placeholder intact, and a second pass in the same run could call the provider again. The allocation is now held under the placeholder — which also answers the skip test, so nothing pays twice — and applied when that scene is committed, before its first save. One complete pass now leaves no placeholder behind. A failed commit used to abandon what it had already allocated. A poster upload that failed threw away a stored video and sent the retry to submit the most expensive job in the system again; a rejected write-back left registry rows that name bytes nothing references, which the byte collector cannot reclaim because it only collects blobs no row names. A poster failure now costs the poster, and a write-back that reached nothing reclaims what it allocated. A partial write is left alone, because the document already names it. The ownership gate is fail-closed again. Treating the sidecar's 404 as permission was wrong: the client cannot tell "this course has no owner" from "this deployment told me nothing", so a visitor who opened a shared course could bill the operator. The root cause was the sidecar itself, which gated on the agent runtime although every persisted course has an owner regardless — the persistence route resolves one for every request. It now gates on server persistence, so the configuration that made 404 the universal answer has real ownership facts to report, and the gate can refuse everything but a named owner. Retry affordances answered to no gate at all. A viewer of a shared course with one failed image was shown a Retry button that called the provider. Both retry entry points and every surface that draws them now read one shared permission, so what is offered and what is allowed are the same value. Also: narration regeneration no longer pretends it can replace bytes behind a live id — the exclusivity proof that would allow it is refused by construction once references leave the browser, so it forks to a fresh id and says so; the "already generated" test lets a finished deck answer from the document alone, since scene order stops identifying an outline once slides are inserted or deleted; stored assets record a specific media type rather than a generic transfer type; the pane no longer asks the sidecar in browser-only mode; and the funnel's docstring now states what the per-stage lock actually guarantees, which is same-browser serialization and not a cross-browser compare-and-swap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): park allocations in the deciding turn, never reclaim on an ambiguous write A delta review of the write-back found the first-pass fix still had a window, and the reclamation it added could delete media the document already names. The allocation was parked after an awaited local cache write. A scene committed in that window reconciled against a registry that did not hold it yet, so the document kept the placeholder — and the entry recorded a moment later then answered the skip test as "already handled", so nothing could correct it. Parking now happens inside the write-back, in the same synchronous turn as the decision that nothing could take the reference; no await separates the live check from the park. Allocations parked by an earlier pass are handed to their slides at the start of the next one, before anything decides what still needs generating, so a held allocation whose scene has since arrived becomes a rewrite rather than an answer. Reclaiming on a rejected write was unsound: a rejection does not prove the server did not apply the write, so deleting the asset could break the scene that now names it. The funnel decides instead, and says so: it reclaims only when no store write was ever issued and nothing took the reference. Anything else is placed if its slide exists and parked if it does not, so the next pass reuses the bytes instead of paying for them again. When a write fails after part of it landed, the live store is brought up to the document before the error is rethrown — otherwise the next ordinary flush would overwrite the half that did land, with the ids deliberately not reclaimed. Parked allocations are now cleared with the course. Classic placeholders are reused across runs, so one surviving an interrupted run would be handed to a different slide of the next deck: the previous picture, on a slide whose provider was never asked. Both classroom surfaces clear the arriving course, the deletion cascade clears the deleted one, and clearing the database clears them all. Two more ways generation could start without asking the gate are closed. An overlapping pass — an outline retry re-enters generation with every outline while the first is still working — re-requested elements whose provider call was already in flight; a task that is not done is an answered request, not an unanswered one. And narration regeneration in the timeline editor called the TTS provider and allocated a pool asset with no ownership check at all; it now reads the same permission, which withholds both the per-line and whole-timeline controls and refuses the call. Finally, a pane opened during the stage-link availability gap recorded the sidecar's 404 for a course that was moments from existing and never asked again, leaving the real owner locked out of generation until it remounted. Ownership is re-fetched once the document becomes available; the gate stays closed until an answer arrives, so asking again can only open it for someone entitled to it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the asset routes reachable, and stale snapshots harmless A full-branch audit found that the deployment this project documents could not store a single generated asset, and that several routes into durable storage could still write a placeholder over a reference that had already landed. The persistence route sent asset requests through the development authenticator, which refuses outright in a production build that has not explicitly opted into it — and the documented server-persistence recipe produces exactly that build. Every store and every read answered 401, so images and video failed on every slide while re-billing the provider on each retry, and a narration failure stopped the deck at its first slide. Assets live in one shared partition by design, so there was never anything per-caller for that authenticator to decide: the route now resolves the asset principal itself, alongside the owner it already resolves for documents. Runtime sessions are genuinely per-learner and keep the development authenticator until real session verification replaces it. And narration that cannot be stored no longer fails its scene: the line stays unvoiced and retryable, which is what an image that cannot be stored does to its slide. Placeholders could also come back from behind. A queued autosave's snapshot, an editor-history entry replayed by an undo, the departing save a course switch flushes — each captures content at its own moment, and any of those moments can predate a write-back. Point fixes at each producer would leave the next producer to rediscover the bug, so the check lives at the write boundary every producer passes through, and the allocation record it consults now outlives the parked queue: a placeholder whose rewrite landed long ago is exactly the case it catches. Two ways generation could be lost or repeated are closed. A pass now claims the elements it will reach and releases them however it ends, so an overlapping pass stands down while an aborted one strands nothing — previously its tasks stayed `pending` and every later pass skipped them with no retry control to recover them. And the media abort controller is aborted before being replaced, so a superseded pass stops calling providers instead of running on for a course the user has left. The remaining two are narrower. The workbench pane asks for ownership only after a document load succeeds, and after every later one, mirroring the page route: the load is what creates the ownership row the first time a course is opened, so asking beforehand asked about a course that did not exist yet and locked its author out for the mount. And the ownership gate on the timeline editor now withholds narration regeneration alone; listening back to existing narration and seeing whether a line has any spend nothing and stay available. Known limitation, unchanged and now stated plainly in the comments that used to point at it as a solution: nothing reclaims an unreferenced pool asset. The registry sweep is written but not wired up, and the byte collector only reclaims blobs no registry row names, so every narration regeneration and every abandoned allocation leaves storage behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): gate asset mutations, and make claims and allocation records survive a handoff Opening the asset routes opened all of them. Reads and allocations are meant to be as open as document reads and creates already are, but no authorization hook was supplied, so the handler's default admitted PUT and DELETE too — and those scope by principal key alone, which is one shared constant. Any caller who learned an id, and a document read hands out every id its slides name, could overwrite or destroy another author's media. Mutations now require the deployment's credential, which in a production build without the development-auth opt-in means they are refused outright; reads and allocations stay open. The route comment says what the posture is and what it is not: the deployment-level fence is the access code, and no per-principal quota is configured. The client's own reclaim is best effort to match — losing an argument about deleting an asset must not cost a task its retry, and the bytes are left for server-side reclamation. The pass claim could not survive the handoff it was written for. A retry aborts the live media pass and starts its replacement in the same synchronous block, long before the aborted pass's cleanup runs, so the replacement saw every element still claimed, collected nothing, and returned — leaving each unreached element at pending with nobody coming back for it and no retry control to recover it, which is the exact failure the claim was introduced to prevent. A claim now carries its pass's signal and is retired the moment that signal aborts, and a pass releases only claims it still owns, so a late unwind cannot take its replacement's work. Claims are also acquired at the single point every request passes through, so a single-task retry participates too — previously a retry awaiting its provider was invisible to a pass starting alongside it and both called it. The allocation record could outlive the bytes it named. It was written before the write-back attempted anything and survived the reclaim that followed a failure, so when the slide finally arrived the write boundary stamped a deleted id into the document — and the placeholder it replaced was gone, which reads as already generated and stops anything from retrying. The record is now written only where the allocation is retained, and forgotten wherever a reclaim removes the bytes, including the narration rollback path. The tests follow. The route test drives the real storage handler against an in-memory registry instead of a stub, so it can see what the resolved principal is then allowed to do; the handoff test performs a real abort mid-pass rather than starting from an already-aborted signal; and the guards that could only assert file layout now assert the property they care about, or have been replaced by behaviour. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make media passes serial per course instead of tracking element ownership Three rounds of per-element claims each produced a new way to lose an element. Whole-pass reservations swallowed a Retry for an element the same pass had already failed, leaving it pending with the affordance gone. Retiring a claim by its signal freed an element whose commit was still uploading, so the replacement pass paid for it twice. A claim held for a failed element stranded its retry. The bookkeeping is the defect: every refinement of "who owns this element right now" answered the question at a moment when the answer was already stale. Passes for one course are now serial. A replacement aborts its predecessor, as before, and then waits for it to settle before collecting. That removes the question entirely: a commit already under way finishes — its bytes stored and its reference written, so the new pass sees a resolved slide and skips it — and an element the aborted pass never reached is still a placeholder and gets collected like any other. The claim set, the reservations, the signal retirement and the identity-checked release are all gone. The task table is consulted for one thing only: an element that is generating right now is a single-element retry running alongside the pass, and taking it too would pay twice. Pending is deliberately not a skip reason — it means a pass once intended to reach an element, which an abandoned pass leaves behind with nobody acting on it, and reading that as answered is what stranded elements before. A retry runs concurrently with a pass, because a pass never revisits an element it has processed, and it re-reads the task after its own await and refuses before touching it: marking first and refusing afterwards destroyed the failed state that draws the affordance. Browser-only mode is back to exactly what it was. The abort is now conditional, the waiting does not apply, and the original status-based skip is restored verbatim. Two baseline lines remain changed in each of the two files, and both are behind a server-backed fork whose else-branch is the original. Two smaller things. The allocation record becomes visible when a write goes on the wire rather than when the round trip ends, and the write boundary reconciles under the document lock rather than before it — a save queued during a write-back was otherwise captured with the placeholder and, for a course the user had left, had no corrective flush to follow. And the comments that said a refused reclaim leaves its bytes for server-side reclamation were wrong: nothing collects them, because the registry entry still names its blob and the sweep that would remove it is not wired up. They now say the bytes leak. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a deferred pass re-earn its right to run, and bound the commit it waits on Serializing passes moved their body out of the block that launched them, and three things followed from that. A pass now wakes when its predecessor settles, which can be after the user has left the course. It enqueued before it looked at its signal, into a task table keyed by element id alone — and placeholder ids are not unique across courses, which is why the classroom clears that table on arrival. So a departing course's pass seeded the arriving course's table with tasks carrying the wrong stage id, and a Retry routes by that id: the reference went into the wrong document. A pass now re-validates after the wait, before touching anything shared. The same lateness broke the skip test. `documentSkipIndex` answers only while the live store is on the pass's stage, and returning nothing put the collection loop on the browser-only rule — a silent demotion from "the document is the authority" to "this browser's task table is", on exactly the path where that table has just been cleared. Every element the predecessor had committed was collected again, paid for again, and its second write-back found no placeholder to rewrite, so its bytes were parked where nothing will ever reference them. In server-backed mode an unreadable document now means the pass stands down. And waiting was unbounded. A commit is uncancellable: the asset client takes no signal, and a document write cannot be half-undone. One stalled upload therefore froze the course's media generation for the session — the replacement never collected, the element sat on a skeleton that draws no Retry, and only a reload recovered. The pass's signal is now threaded into the media proxy fetch, and the commit is bounded by a deadline. The deadline is on the wait, not the work: the commit carries on, and if it lands late the document simply ends up correct, while the element becomes retryable and the queue moves on. The tests that were meant to pin the previous round were not sensitive to it. Two asserted end states where the mechanism only changes ordering, and one of them rigged the document read so the assertion held whether or not the pass had waited; a third covered half of what it claimed. They now observe the ordering directly — nothing is issued while another pass for the course is working; in browser-only mode a second pass reaches its provider immediately — and the reconciliation under the document lock has a test that fails when it moves back outside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * revert(media): drop the commit deadline and the abortable download The deadline bought less than it cost. Abandoning a commit after two minutes makes the element retryable while the real commit is still running, so a Retry starts a second commit for the same placeholder against the first: two provider calls, two allocations, and whichever lands second stamps its result over the other's task by element id. The allocation record is keyed by placeholder, so the loser's cleanup erases the winner's record, and the write boundary then puts the raw placeholder back into the document. That is the overlap serial passes were built to remove, reopened through the one door serialization never covered. So a stalled commit holds the course's media queue until it settles or the page is reloaded, and that is written down rather than papered over. The wait is unbounded on purpose: every ceiling on it turns out to be a way of running two commits for one element. Threading the pass signal into the download was also a mistake, in the other direction. The provider call that produced the URL has already been billed, so cancelling the download throws away work that is paid for — and the shared proxy cache records a cancelled request as a transient failure against that URL, which after three of them blocks it for every consumer in the session. Browser-only mode never asked for this: it had no way to observe an abort there, which is exactly why the bytes were kept. The signal is gone from the download again, and `fetchAsBlob` is byte-for-byte what it was before this branch. The regression guard for the stranded-element rule is restored alongside the timing test that was meant to supersede it. It catches a different rule — a task left pending being read as answered — and nothing else does: making the pass skip pending leaves every other suite green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): stop asking the pool for refs it never issued, bound it, and adopt cached bytes Four things a deployment found once this was running for real. A reference this application mints itself — a generation placeholder, a derived narration key — was never in the pool, because the pool allocates every id it holds. Asking anyway used to be an IndexedDB miss; once the pool is server-backed it is a request that answers 404, one per element per load, forever on a course that still holds placeholders. Every lease and probe now checks first. The check is a negative test on shapes this application owns, not an id validator: the pool's id domain stays unconstrained, and anything that is not one of ours is still asked about. The asset store can bound how much one principal holds, and enforces it inside the write transaction, but nothing ever passed the number. It does now, with a default rather than an opt-in: allocation is reachable by any caller a deployment admits, and with one shared principal an unbounded store is unbounded database growth with no operator-visible brake. Refusing asset mutations to unauthenticated callers was not enough, because every authenticated caller resolves to that same shared principal — so authentication decided nothing, and any signed-in visitor could delete any id they learned. Since this branch began storing media the registry is the only copy a course has. Replacing and deleting are now refused to everyone, and the browser no longer tries: an entry nothing references waits for server-side reclamation instead. What a browser must still do is forget its own record of an allocation that reached nothing, or a later save would stamp an id the document has no reason to trust. And a course generated before any of this holds placeholders in its document with its bytes only in the author's browser. Those bytes are paid for, so the author's next load converts them — stored to the pool and written back through the ordinary commit path, with no provider call — instead of buying them again. A row that records only a hosted URL is treated as absent: that URL is the provider's address, not something a document may hold. One renderer expectation moved with this. An untracked placeholder used to paint as pending on first render because asking the pool left a lease in flight; it settled to disabled a moment later either way, and now says so from the start. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): surface a full store as a refusal and convert legacy narration A quota refusal reached the browser as HTTP 500 with a generic message, which reads as a transient failure: the element kept a Retry that would pay a provider again and be refused again. The store raises the contract's own error and the handler maps it to 507, but the store answering a request is not always built by the same bundle as the handler -- the persistence provider is reached from the route bundle and from instrumentation, which is why its state lives on a Symbol.for global -- and `instanceof` is false across that boundary while the declared code is still right. Classify on the code as well as the class, and make the code a permanent, persisted refusal in the browser: recorded locally so it survives a reload, shown as "storage is full", and refused by the retry entry point so a stale button cannot buy a second generation. Every other storage failure stays retryable. Convert what a pre-server-backed course still holds. Generated media is adopted under either key this application has used for it -- the placeholder, and the allocated id of a course converted once and later rolled back -- instead of only the first. Narration is converted by a load-time pass over the open course's speech actions, since nothing re-enters generation for an action that already has an id: bytes to the pool, id written back through a funnel that mirrors the media one, owner-only and server-backed-only. A line whose bytes are in no browser is left alone rather than re-synthesized. Also: the pool guard is now a positive `ast_` test rather than an enumeration of the shapes we mint (imports never reach the pool, so this is safe in both modes); the slide ref collection is an exported pure function so its four lease sites are covered behaviourally; ASSET_QUOTA_BYTES treats every spelling of zero as opting out and refuses a malformed value at startup instead of falling back; the abort signal is re-checked after the cache read, before an uncancellable commit; and the unused `removeAsset` and pool `replace` surfaces are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * chore(storage): release 0.29.1 The asset HTTP handler now recognises a store refusal by the contract code it declares as well as by its class, so a quota refusal raised in another module realm answers 507 instead of 500. Same contract, stricter recognition, no API change: a patch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a full store recoverable and adoption course-safe Narration adoption read the local audio row by its derived key alone. That key carries no stage id and the table is keyed by id alone, so two courses can mint the same one -- a PPTX import numbers its scenes and actions deterministically, which gives every imported deck's first slide `tts_s1_speech-scene-p1`. Locally a collision only means one course plays another's clip in one browser; adopting it wrote that clip into the shared document permanently, for every device and every visitor. A row that names a course is now adopted only into that course, and a row from before that column existed only when the text it recorded is the text of the action being converted. A full asset store was made permanent last round, which was wrong three times over: it overwrote the refused bytes with an empty blob -- on the conversion path that row is a course's only copy of its own media -- it kept sending the rest of the deck to a provider against a ceiling it already knew was reached, and it left no way back once an operator raised that ceiling. A full store is neither the content's fault nor the configuration's, so it is now its own case: the bytes are kept, the pass stops at the first refusal, and the element shows the reason together with a Retry that re-attempts the upload from those bytes. Nothing retries automatically, so no one is re-billed. The narration write-back now reaches the write boundary every producer of a durable write passes through, not only the dirty mark: adoption never deletes the derived row, so a snapshot that reverts the rewrite is adopted again on the next load and allocates a fresh asset every time. Adoption is also mounted by both classroom surfaces rather than one, takes the course's abort signal, and re-validates that this browser still has the course open before each write. ASSET_QUOTA_BYTES is validated from instrumentation, where the README and the docstring already claimed it was: its only other consumer is lazy and memoised, so a malformed ceiling let the process boot and then failed every persistence request, documents and runtime included. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): remember a full store per course, and never lose retained bytes A stopped pass left the elements it never reached as placeholders with no persisted record -- deliberately, since nothing was attempted for them. But that left the next load with no reason not to try: it called a provider for the next placeholder and was refused at exactly the same point, once per reload, indefinitely. A full store is not a property of any slide. It belongs to the deployment and changes for reasons the document knows nothing about, so it is now remembered once per course in the browser's device KV. A pass that finds the marker stands down before spending anything and leaves every placeholder its "storage is full" state and its Retry; the first upload that succeeds clears it and the next pass runs normally. Narration adoption latched per course so it runs once per load, and the latch outlived the abort that leaving a course performs. On a surface that stays mounted across switches -- the workbench pane is one component for every course it shows -- owner course A, visitor course B, then back to A skipped exactly the clips the abort had cut off, and nothing else converts them. The latch is released with the abort now, and a course adopts one run at a time so a re-entry cannot hand a clip a second allocation while the previous run's uncancellable tail is still settling. A quota-blocked element retried into a network error or a 500 lost the bytes that were kept for it: the retry deleted the row before attempting the upload and wrote no replacement for an error carrying no structured code, so the next retry went back to a provider for media this browser had a moment earlier. The row now survives until an upload succeeds, the failure handler keeps whatever bytes the attempt was given, and the retry asks the question a pass asks -- does this browser already hold bytes for this element -- rather than reading an error code that a second failure has already overwritten. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): adopt real legacy narration, queue re-entries, report attempt outcomes Narration adoption admitted a stage-less row only when the text it recorded matched the action being converted. Both of those columns were added to the local audio table by the very change that moved narration onto allocated ids, so a row still carrying a derived key has neither: the rule refused every real pre-allocation course and passed only on fixtures built from post-allocation rows. What the row cannot say, the key can. A derived key names two clips only when two courses share a scene order and an action id, and an action id repeats only when something other than the generator minted it -- an import numbers them by slide position. So a key built from a generated action id is adopted on that basis, a key an import could have reproduced still needs matching text, and a row that names another course is refused however unique its key looks. Handing a re-entering caller the adoption run already in flight undid the latch release it was paired with: that run is bound to the signal the departure just aborted, so it stops at its next clip while the caller -- which has the course open and a live signal -- is told the work is done, and an effect replayed as mount, cleanup, mount adopts nothing at all. A later caller now waits for the uncancellable tail and scans again, which costs a lookup on a course that has nothing left and finishes the clips the abort cut off on one that does. One attempt at an element now reports both facts its callers need instead of a bare boolean: whether the store refused it for room, and whether bytes actually reached the store. Leaving a course clears the task table, so a retry that landed afterwards read "no failed task" as success and deleted the row holding the only copy of the media. Nothing is inferred from that table any more. Reading the localStorage property can throw where storage is denied by policy, typeof included, so the availability check moved inside the guard: this metadata is best-effort, and a rejection here strands a generation pass that has already enqueued its tasks. A retry is never blocked by the per-course "store is full" marker, but a retry that is refused again re-sets it, and adoption now reads and writes the same marker rather than issuing one refused upload per clip on every load. The two canvas element renderers and both thumbnail renderers show the reason beside the Retry, so a full store does not look like an ordinary failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): probe a full store instead of standing down, and pair notices with a Retry Narration adoption was given both halves of the per-course "the store is full" marker last round: it stood down when the marker was set, and it set the marker when its own upload was refused for room. Those halves are only safe together if something can lift the marker, and for adoption nothing could. It has no affordance of its own, it stood down before reaching its own clear, the media pass returns before its marker gate when there is nothing to generate -- so a narration-only deck, or one whose slides are already satisfied, painted no storage-full element and offered no Retry -- and narration generated rather than adopted allocates directly rather than through the media commit. The course's cached narration was then lost for good, where before it converted on the first load after the ceiling was raised. The gate is a probe now. A marked course attempts exactly one clip per load: refused, it stops and the marker stands, which costs what standing down cost; stored, it lifts the marker and finishes the course. Adoption spends no provider money, so the whole cost of probing a store that is still full is one refused upload. Generated narration lifts the marker too. The three surfaces that gained a failure notice last round drew it for any failure with a reason, including the one refusal that is reachable without server-backed persistence, so a browser-only deck painted something it had not painted before. The notice is drawn beside a Retry and nowhere else, which is what it was added for and what leaves browser-only output unchanged. Both are now asserted through the render harness the surface matrix already had. A caller arriving while a rescan is queued shares it rather than appending another. One rescan converts whatever the run in flight left and every later one would find an allocated id on every action, so a chain bought nothing and turned a single stalled upload into a course that never adopts again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): treat a refusal for room as a fact about one clip, not the deck The asset store checks each write against the headroom it has left, so a store that refuses a long opening clip can still hold every short clip behind it. Narration adoption assumed the opposite: it broke the deck at the first refusal and then re-attempted that same first clip on every later load, because the document names it first. A deck whose longest clip exceeds current headroom therefore never converted the clips that would have fit, with no affordance to recover it -- the state the probe was introduced to remove, reached through a narrower door. An unmarked load now attempts every clip, skipping the ones that do not fit, and remembers the condition only if the load ends with clips it still could not store. A marked load spends its single upload on the smallest clip left rather than the first one named: that is the clip that answers the question the marker asks, because if the smallest does not fit nothing does. The media pass keeps stopping at its first refusal, and for a reason adoption does not share -- every element it attempts costs a provider call. A rescan several callers share took the newest caller's signal, and the newest caller is not necessarily the one still there: a surface that opened a course and closed it again would stop work a surface still showing that course was waiting for, and that surface is latched, so it would never ask again. The shared run now takes a signal that is aborted only once every caller has left. The comment claiming the shared rescan contains a stalled upload was wrong -- the rescan is chained off the run in flight, so a stalled upload leaves every caller pending exactly as a chain would. It claims the bounded queue it actually provides, and the stall is recorded as a limitation. The failed-state containers took their stacking classes unconditionally, so markup differed in browser-only mode even though nothing moved on screen. Those classes are applied only when there is a notice to stack, and the tests assert the exact class attribute rather than a substring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): stop narration adoption writing the media pass's store-full marker The marker means "do not call a provider for this course". A path is entitled to write it only if its own refusal cost a provider call, and narration adoption's refusals cost nothing: it uploads bytes this browser already holds. The store also checks each write against the headroom it has left, so a clip that does not fit says nothing about whether a slide's image would. Adoption was writing it anyway, and one over-long narration clip was therefore enough to stand a course's entire image pass down on every later load -- on a store that had just accepted adoption's other clips. The author could still recover each element by hand, every load, for ever. Three rounds of narrowing this seam produced a finding each time, so it is removed rather than narrowed again. Gone: the marker read, the single-clip probe, the smallest-clip selection, and the up-front read of every row into an array -- which also retires a sampled-then-stale flag and the retention of a whole deck's blobs for the length of a run, and returns the loop to streaming one row at a time. Adoption's rule is now that every load attempts every clip it holds, once; any failure skips that clip and the load continues. The noise the coupling was meant to avoid does not arise, because after the first load the clips still outstanding are exactly the ones that did not fit -- normally none, or one. A successful write still clears the marker, and that is a different kind of statement: a write that went through is a fact this run established, where a refusal is an inference about what some other write would cost. For a course whose media needs nothing, adoption and generated narration are also the only paths that can establish it. The failure module still documented the deck-wide premise this contradicts. It now says what is true: the check is per write, and the media pass stops the deck as a judgement about cost rather than about certainty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bound a full store's cost from the store's own arithmetic Removing the store-full marker from narration adoption removed its bound too, and the code then asserted the bound was unnecessary. It is, on a store with room for most of a deck. On the store the whole mechanism exists for -- the ceiling reached, nothing fitting -- the outstanding set after every load is the entire deck, so a thirty-clip course posted thirty full blobs on every load, indefinitely. Each of those is not a cheap refusal: the bytes are uploaded, the server hashes the whole payload, and only then takes a per-principal lock and sums every entry that principal owns before saying no. The bound needs no flag, no key and nothing carried between loads. The store asks whether `used + addedBytes` exceeds the ceiling, and `used` only grows while a run is uploading, so a clip refused for want of room implies every clip at least that large is refused for the rest of that run. The run keeps the smallest size it has been refused and skips anything no smaller without uploading it; a smaller clip is still attempted, because it may fit. A deck the store refuses entirely now costs one upload per successive size minimum instead of one per clip, and a deck it has room for costs nothing extra, because nothing is refused. Only a refusal for room lowers the bar: a dropped connection says nothing about how much room there is. The deck-wide certainty premise the failure module retracted last round still stood verbatim at the site that implements the stand-down. Both copies now say the same thing: the check is per write, and the pass stops the deck as a judgement about cost rather than about certainty. The comment on adoption's marker clear now names its price. Narration of a few hundred bytes fits in headroom an image does not, so a proven write can let the next pass buy one more image that is refused again -- bounded at one, and the price of the alternative being a course whose media never generates again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(media): state the adoption bound exactly, and stop three comments describing the old rule The comment introducing the in-load bound gave its cost as "at most a handful, and the first load pays the most". Neither clause is a property of the rule. A clip is skipped only when something no larger was already refused, so a fully-refused deck costs one upload per successive size minimum in document order: one when the clips grow, about ln N for an arbitrary order, and one per clip when they only shrink -- a long opener followed by terser lines is exactly that shape. And no load is cheaper than the first, because the bound resets per run and a refused clip stays outstanding. The comment now says that, and points at what would make it exactly one for any ordering: the store returning its remaining headroom in the refusal's existing details channel, which the server leaves empty today. Two other comments still described the previous rule -- "attempts every clip it holds, every load" -- one of them twenty lines above the paragraph that introduces the bound, in the same block. Both now say what the code does. The bound's soundness is worth stating where a maintainer will look for it: quota is charged at full length with no discount for a duplicate, the sum it is checked against joins entries to blobs so the collector cannot lower it, the check takes a per-principal lock before summing, and replace and delete are refused to every browser. Nothing a run can do makes room appear inside it. One test installed a row implementation and replaced it wholesale a few lines later, so the first was dead and the survivor dropped the text the first clip's import-shaped key needs for the ownership rule -- it passed on the coincidence that the fixture's default text is the action's. Merged into one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): keep a Retry from re-buying parked media, and state the store seam once Five findings from an inline review. The standalone classroom route asked the ownership sidecar once per load and recorded only stage ownership when that ask failed. Every non-answer fails closed, so one transient 5xx left the genuine author with no resume, no Retry affordance and no legacy narration converted for the rest of the load, with nothing to change it short of a reload. The failure now records the fail-closed answer explicitly -- an answer an earlier load established must not outlive the failure that replaced it -- and an unresolved answer is asked for again, a few times over a few seconds. A real answer, however unwelcome, is final. A Retry could pay a provider for media the pool already held. When the bytes are stored and only the write-back fails in a way that keeps the allocation, it is parked and no local row exists, because that row is written only after a successful write-back. Retry now reads the parked queue exactly as the pass does and re-attempts the write-back: it re-keys the task done when the document takes it, leaves the entry parked when the slide still does not exist, and stays failed and retryable when the document refuses again. Object URLs a parked allocation owns are revoked when the entry is dropped. The commit path leaves them alone while the entry is parked, because it is then the only thing holding bytes this tab can render, so a course switch or a stage deletion was pinning the whole blob for the life of the tab. An entry a slide has already taken is left alone: the task table is displaying those URLs. The fallback lookup for cached bytes is a stage-scoped scan, and the keyed lookup misses for every row the commit path writes, so a pass was materializing and sorting the course's whole media table once per element. One scan per pass now, built on the first miss. It is sound and not merely cheaper: an element asks only for its own placeholder, and every row a pass writes carries the placeholder of the element that wrote it. "The store accepted a write, so it is not out of room" was enforced at three call sites under slightly different conditions, which made it a convention the next pool write path could silently break. It is stated once, in putAsset, for the course whose bytes it just stored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 16 天前 | |
feat(media): write generated media through the asset pool under server-backed persistence (#1392) * feat(media): store generated media in the asset pool when persistence is server-backed With server-backed persistence the document is durable and shared, but generated media stayed in the producing browser: the document kept its gen_img_* / gen_vid_* placeholder and narration kept a browser-derived audio id. Every new browser that opened such a course re-ran generation for every slide, and it never converged, because the address of the generated bytes was never written back into the document. Under server-backed persistence only, the classic generation chain now stores bytes in the asset pool first and writes the id the pool allocated into the document. - The client bootstrap configures the asset seam alongside the document and runtime seams: an HttpAssetStore over the persistence endpoint carrying the same credentials the document store carries, marked server-backed. The seam preflight now covers all three, so a failure still cannot half-configure persistence. - Image, video and TTS generation commit in one fixed order: provider, pool, document, local cache, task. A reference reaches the document only after put returned an id, so a document can never name bytes that were not stored. A failure before the write-back leaves the placeholder with the provider called exactly once; the retry happens on the next owner load. - The write-back is a per-slot rewrite through mutateDocument, which re-reads the current document under the per-stage lock, so it cannot clobber a newer scene. The open course is refreshed with the same rewrite without being marked dirty. - "Has this already been generated?" is answered by the document (the slide exists and no longer holds the placeholder) instead of by this browser's task table. - The classroom's resume effect fails closed on ownership: only a resolved owner starts generation, so a viewer opening a shared course spends nothing. - The local media and audio tables become a per-tab cache. A failed cache write costs a re-download, never the media. Browser-only mode is unchanged: every new call site sits behind the server-backed gate, the local tables stay authoritative there, and placeholders stay in the document. Rendering and export needed no changes. HttpAssetStore.resolve mints an object URL exactly as the browser store does, and the export byte resolver was already pool-first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(classroom): make the generation owner gate a three-outcome rule and apply it everywhere The gate refused everything but a resolved owner, which read a sidecar that answered "no ownership fact exists for this course" as a reason to block. That is the answer a deployment without the sidecar's server-side prerequisites gives for every course, and the answer a course with no ownership record gives: in both, there is nobody the operator's budget needs protecting from, and refusing strands the course's own author behind a question that can never be answered. Ownership is now four states over the sidecar's three outcomes. A definite answer splits into owner and not-owner. An absent record is its own answer, ownerless, and generation proceeds — the behaviour such a deployment had before the gate existed. Only the absence of an answer, a transport failure or a load that has not asked yet, stays unresolved and fails closed: "we could not ask" must never be read as "nobody owns this". One mapper turns a sidecar result into that state, and one predicate decides on it. The workbench classroom pane runs the same resume effect and had no ownership input at all, so a viewer opening a shared course there could still spend the budget. It now asks the sidecar once per course, in parallel with its load and feeding only the generation gate, so its read-only and edit behaviour is unchanged. The shared progressive-load policy carries the gate for it, with both new inputs required rather than defaulted so a future caller cannot omit them into an open budget. Its stale comment claiming ownership could not be expressed here is corrected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the write-back survive autosave, arrive before the scene does, and never leak Independent reviews of the write-back found three ways a durable document could still end up naming a placeholder, and two ways the gate that protects the operator's budget could be walked around. An autosave round captures the store synchronously and writes that capture, so a round already in flight when a rewrite landed wrote the placeholder straight back over the allocated id, and nothing marked the store dirty again to correct it. The rewrite now marks the units it changed, which leaves a corrective flush queued behind the stale one; re-saving a scene that already holds the id is idempotent, losing the id is not. Media is generated from outlines in parallel with scene content and usually finishes first, so the slide that will carry the placeholder does not exist yet and the write-back has nothing to rewrite. That was the ordinary path, not a tail case, and its result was discarded: the task was marked done, the scene was added afterwards with its placeholder intact, and a second pass in the same run could call the provider again. The allocation is now held under the placeholder — which also answers the skip test, so nothing pays twice — and applied when that scene is committed, before its first save. One complete pass now leaves no placeholder behind. A failed commit used to abandon what it had already allocated. A poster upload that failed threw away a stored video and sent the retry to submit the most expensive job in the system again; a rejected write-back left registry rows that name bytes nothing references, which the byte collector cannot reclaim because it only collects blobs no row names. A poster failure now costs the poster, and a write-back that reached nothing reclaims what it allocated. A partial write is left alone, because the document already names it. The ownership gate is fail-closed again. Treating the sidecar's 404 as permission was wrong: the client cannot tell "this course has no owner" from "this deployment told me nothing", so a visitor who opened a shared course could bill the operator. The root cause was the sidecar itself, which gated on the agent runtime although every persisted course has an owner regardless — the persistence route resolves one for every request. It now gates on server persistence, so the configuration that made 404 the universal answer has real ownership facts to report, and the gate can refuse everything but a named owner. Retry affordances answered to no gate at all. A viewer of a shared course with one failed image was shown a Retry button that called the provider. Both retry entry points and every surface that draws them now read one shared permission, so what is offered and what is allowed are the same value. Also: narration regeneration no longer pretends it can replace bytes behind a live id — the exclusivity proof that would allow it is refused by construction once references leave the browser, so it forks to a fresh id and says so; the "already generated" test lets a finished deck answer from the document alone, since scene order stops identifying an outline once slides are inserted or deleted; stored assets record a specific media type rather than a generic transfer type; the pane no longer asks the sidecar in browser-only mode; and the funnel's docstring now states what the per-stage lock actually guarantees, which is same-browser serialization and not a cross-browser compare-and-swap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): park allocations in the deciding turn, never reclaim on an ambiguous write A delta review of the write-back found the first-pass fix still had a window, and the reclamation it added could delete media the document already names. The allocation was parked after an awaited local cache write. A scene committed in that window reconciled against a registry that did not hold it yet, so the document kept the placeholder — and the entry recorded a moment later then answered the skip test as "already handled", so nothing could correct it. Parking now happens inside the write-back, in the same synchronous turn as the decision that nothing could take the reference; no await separates the live check from the park. Allocations parked by an earlier pass are handed to their slides at the start of the next one, before anything decides what still needs generating, so a held allocation whose scene has since arrived becomes a rewrite rather than an answer. Reclaiming on a rejected write was unsound: a rejection does not prove the server did not apply the write, so deleting the asset could break the scene that now names it. The funnel decides instead, and says so: it reclaims only when no store write was ever issued and nothing took the reference. Anything else is placed if its slide exists and parked if it does not, so the next pass reuses the bytes instead of paying for them again. When a write fails after part of it landed, the live store is brought up to the document before the error is rethrown — otherwise the next ordinary flush would overwrite the half that did land, with the ids deliberately not reclaimed. Parked allocations are now cleared with the course. Classic placeholders are reused across runs, so one surviving an interrupted run would be handed to a different slide of the next deck: the previous picture, on a slide whose provider was never asked. Both classroom surfaces clear the arriving course, the deletion cascade clears the deleted one, and clearing the database clears them all. Two more ways generation could start without asking the gate are closed. An overlapping pass — an outline retry re-enters generation with every outline while the first is still working — re-requested elements whose provider call was already in flight; a task that is not done is an answered request, not an unanswered one. And narration regeneration in the timeline editor called the TTS provider and allocated a pool asset with no ownership check at all; it now reads the same permission, which withholds both the per-line and whole-timeline controls and refuses the call. Finally, a pane opened during the stage-link availability gap recorded the sidecar's 404 for a course that was moments from existing and never asked again, leaving the real owner locked out of generation until it remounted. Ownership is re-fetched once the document becomes available; the gate stays closed until an answer arrives, so asking again can only open it for someone entitled to it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the asset routes reachable, and stale snapshots harmless A full-branch audit found that the deployment this project documents could not store a single generated asset, and that several routes into durable storage could still write a placeholder over a reference that had already landed. The persistence route sent asset requests through the development authenticator, which refuses outright in a production build that has not explicitly opted into it — and the documented server-persistence recipe produces exactly that build. Every store and every read answered 401, so images and video failed on every slide while re-billing the provider on each retry, and a narration failure stopped the deck at its first slide. Assets live in one shared partition by design, so there was never anything per-caller for that authenticator to decide: the route now resolves the asset principal itself, alongside the owner it already resolves for documents. Runtime sessions are genuinely per-learner and keep the development authenticator until real session verification replaces it. And narration that cannot be stored no longer fails its scene: the line stays unvoiced and retryable, which is what an image that cannot be stored does to its slide. Placeholders could also come back from behind. A queued autosave's snapshot, an editor-history entry replayed by an undo, the departing save a course switch flushes — each captures content at its own moment, and any of those moments can predate a write-back. Point fixes at each producer would leave the next producer to rediscover the bug, so the check lives at the write boundary every producer passes through, and the allocation record it consults now outlives the parked queue: a placeholder whose rewrite landed long ago is exactly the case it catches. Two ways generation could be lost or repeated are closed. A pass now claims the elements it will reach and releases them however it ends, so an overlapping pass stands down while an aborted one strands nothing — previously its tasks stayed `pending` and every later pass skipped them with no retry control to recover them. And the media abort controller is aborted before being replaced, so a superseded pass stops calling providers instead of running on for a course the user has left. The remaining two are narrower. The workbench pane asks for ownership only after a document load succeeds, and after every later one, mirroring the page route: the load is what creates the ownership row the first time a course is opened, so asking beforehand asked about a course that did not exist yet and locked its author out for the mount. And the ownership gate on the timeline editor now withholds narration regeneration alone; listening back to existing narration and seeing whether a line has any spend nothing and stay available. Known limitation, unchanged and now stated plainly in the comments that used to point at it as a solution: nothing reclaims an unreferenced pool asset. The registry sweep is written but not wired up, and the byte collector only reclaims blobs no registry row names, so every narration regeneration and every abandoned allocation leaves storage behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): gate asset mutations, and make claims and allocation records survive a handoff Opening the asset routes opened all of them. Reads and allocations are meant to be as open as document reads and creates already are, but no authorization hook was supplied, so the handler's default admitted PUT and DELETE too — and those scope by principal key alone, which is one shared constant. Any caller who learned an id, and a document read hands out every id its slides name, could overwrite or destroy another author's media. Mutations now require the deployment's credential, which in a production build without the development-auth opt-in means they are refused outright; reads and allocations stay open. The route comment says what the posture is and what it is not: the deployment-level fence is the access code, and no per-principal quota is configured. The client's own reclaim is best effort to match — losing an argument about deleting an asset must not cost a task its retry, and the bytes are left for server-side reclamation. The pass claim could not survive the handoff it was written for. A retry aborts the live media pass and starts its replacement in the same synchronous block, long before the aborted pass's cleanup runs, so the replacement saw every element still claimed, collected nothing, and returned — leaving each unreached element at pending with nobody coming back for it and no retry control to recover it, which is the exact failure the claim was introduced to prevent. A claim now carries its pass's signal and is retired the moment that signal aborts, and a pass releases only claims it still owns, so a late unwind cannot take its replacement's work. Claims are also acquired at the single point every request passes through, so a single-task retry participates too — previously a retry awaiting its provider was invisible to a pass starting alongside it and both called it. The allocation record could outlive the bytes it named. It was written before the write-back attempted anything and survived the reclaim that followed a failure, so when the slide finally arrived the write boundary stamped a deleted id into the document — and the placeholder it replaced was gone, which reads as already generated and stops anything from retrying. The record is now written only where the allocation is retained, and forgotten wherever a reclaim removes the bytes, including the narration rollback path. The tests follow. The route test drives the real storage handler against an in-memory registry instead of a stub, so it can see what the resolved principal is then allowed to do; the handoff test performs a real abort mid-pass rather than starting from an already-aborted signal; and the guards that could only assert file layout now assert the property they care about, or have been replaced by behaviour. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make media passes serial per course instead of tracking element ownership Three rounds of per-element claims each produced a new way to lose an element. Whole-pass reservations swallowed a Retry for an element the same pass had already failed, leaving it pending with the affordance gone. Retiring a claim by its signal freed an element whose commit was still uploading, so the replacement pass paid for it twice. A claim held for a failed element stranded its retry. The bookkeeping is the defect: every refinement of "who owns this element right now" answered the question at a moment when the answer was already stale. Passes for one course are now serial. A replacement aborts its predecessor, as before, and then waits for it to settle before collecting. That removes the question entirely: a commit already under way finishes — its bytes stored and its reference written, so the new pass sees a resolved slide and skips it — and an element the aborted pass never reached is still a placeholder and gets collected like any other. The claim set, the reservations, the signal retirement and the identity-checked release are all gone. The task table is consulted for one thing only: an element that is generating right now is a single-element retry running alongside the pass, and taking it too would pay twice. Pending is deliberately not a skip reason — it means a pass once intended to reach an element, which an abandoned pass leaves behind with nobody acting on it, and reading that as answered is what stranded elements before. A retry runs concurrently with a pass, because a pass never revisits an element it has processed, and it re-reads the task after its own await and refuses before touching it: marking first and refusing afterwards destroyed the failed state that draws the affordance. Browser-only mode is back to exactly what it was. The abort is now conditional, the waiting does not apply, and the original status-based skip is restored verbatim. Two baseline lines remain changed in each of the two files, and both are behind a server-backed fork whose else-branch is the original. Two smaller things. The allocation record becomes visible when a write goes on the wire rather than when the round trip ends, and the write boundary reconciles under the document lock rather than before it — a save queued during a write-back was otherwise captured with the placeholder and, for a course the user had left, had no corrective flush to follow. And the comments that said a refused reclaim leaves its bytes for server-side reclamation were wrong: nothing collects them, because the registry entry still names its blob and the sweep that would remove it is not wired up. They now say the bytes leak. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a deferred pass re-earn its right to run, and bound the commit it waits on Serializing passes moved their body out of the block that launched them, and three things followed from that. A pass now wakes when its predecessor settles, which can be after the user has left the course. It enqueued before it looked at its signal, into a task table keyed by element id alone — and placeholder ids are not unique across courses, which is why the classroom clears that table on arrival. So a departing course's pass seeded the arriving course's table with tasks carrying the wrong stage id, and a Retry routes by that id: the reference went into the wrong document. A pass now re-validates after the wait, before touching anything shared. The same lateness broke the skip test. `documentSkipIndex` answers only while the live store is on the pass's stage, and returning nothing put the collection loop on the browser-only rule — a silent demotion from "the document is the authority" to "this browser's task table is", on exactly the path where that table has just been cleared. Every element the predecessor had committed was collected again, paid for again, and its second write-back found no placeholder to rewrite, so its bytes were parked where nothing will ever reference them. In server-backed mode an unreadable document now means the pass stands down. And waiting was unbounded. A commit is uncancellable: the asset client takes no signal, and a document write cannot be half-undone. One stalled upload therefore froze the course's media generation for the session — the replacement never collected, the element sat on a skeleton that draws no Retry, and only a reload recovered. The pass's signal is now threaded into the media proxy fetch, and the commit is bounded by a deadline. The deadline is on the wait, not the work: the commit carries on, and if it lands late the document simply ends up correct, while the element becomes retryable and the queue moves on. The tests that were meant to pin the previous round were not sensitive to it. Two asserted end states where the mechanism only changes ordering, and one of them rigged the document read so the assertion held whether or not the pass had waited; a third covered half of what it claimed. They now observe the ordering directly — nothing is issued while another pass for the course is working; in browser-only mode a second pass reaches its provider immediately — and the reconciliation under the document lock has a test that fails when it moves back outside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * revert(media): drop the commit deadline and the abortable download The deadline bought less than it cost. Abandoning a commit after two minutes makes the element retryable while the real commit is still running, so a Retry starts a second commit for the same placeholder against the first: two provider calls, two allocations, and whichever lands second stamps its result over the other's task by element id. The allocation record is keyed by placeholder, so the loser's cleanup erases the winner's record, and the write boundary then puts the raw placeholder back into the document. That is the overlap serial passes were built to remove, reopened through the one door serialization never covered. So a stalled commit holds the course's media queue until it settles or the page is reloaded, and that is written down rather than papered over. The wait is unbounded on purpose: every ceiling on it turns out to be a way of running two commits for one element. Threading the pass signal into the download was also a mistake, in the other direction. The provider call that produced the URL has already been billed, so cancelling the download throws away work that is paid for — and the shared proxy cache records a cancelled request as a transient failure against that URL, which after three of them blocks it for every consumer in the session. Browser-only mode never asked for this: it had no way to observe an abort there, which is exactly why the bytes were kept. The signal is gone from the download again, and `fetchAsBlob` is byte-for-byte what it was before this branch. The regression guard for the stranded-element rule is restored alongside the timing test that was meant to supersede it. It catches a different rule — a task left pending being read as answered — and nothing else does: making the pass skip pending leaves every other suite green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): stop asking the pool for refs it never issued, bound it, and adopt cached bytes Four things a deployment found once this was running for real. A reference this application mints itself — a generation placeholder, a derived narration key — was never in the pool, because the pool allocates every id it holds. Asking anyway used to be an IndexedDB miss; once the pool is server-backed it is a request that answers 404, one per element per load, forever on a course that still holds placeholders. Every lease and probe now checks first. The check is a negative test on shapes this application owns, not an id validator: the pool's id domain stays unconstrained, and anything that is not one of ours is still asked about. The asset store can bound how much one principal holds, and enforces it inside the write transaction, but nothing ever passed the number. It does now, with a default rather than an opt-in: allocation is reachable by any caller a deployment admits, and with one shared principal an unbounded store is unbounded database growth with no operator-visible brake. Refusing asset mutations to unauthenticated callers was not enough, because every authenticated caller resolves to that same shared principal — so authentication decided nothing, and any signed-in visitor could delete any id they learned. Since this branch began storing media the registry is the only copy a course has. Replacing and deleting are now refused to everyone, and the browser no longer tries: an entry nothing references waits for server-side reclamation instead. What a browser must still do is forget its own record of an allocation that reached nothing, or a later save would stamp an id the document has no reason to trust. And a course generated before any of this holds placeholders in its document with its bytes only in the author's browser. Those bytes are paid for, so the author's next load converts them — stored to the pool and written back through the ordinary commit path, with no provider call — instead of buying them again. A row that records only a hosted URL is treated as absent: that URL is the provider's address, not something a document may hold. One renderer expectation moved with this. An untracked placeholder used to paint as pending on first render because asking the pool left a lease in flight; it settled to disabled a moment later either way, and now says so from the start. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): surface a full store as a refusal and convert legacy narration A quota refusal reached the browser as HTTP 500 with a generic message, which reads as a transient failure: the element kept a Retry that would pay a provider again and be refused again. The store raises the contract's own error and the handler maps it to 507, but the store answering a request is not always built by the same bundle as the handler -- the persistence provider is reached from the route bundle and from instrumentation, which is why its state lives on a Symbol.for global -- and `instanceof` is false across that boundary while the declared code is still right. Classify on the code as well as the class, and make the code a permanent, persisted refusal in the browser: recorded locally so it survives a reload, shown as "storage is full", and refused by the retry entry point so a stale button cannot buy a second generation. Every other storage failure stays retryable. Convert what a pre-server-backed course still holds. Generated media is adopted under either key this application has used for it -- the placeholder, and the allocated id of a course converted once and later rolled back -- instead of only the first. Narration is converted by a load-time pass over the open course's speech actions, since nothing re-enters generation for an action that already has an id: bytes to the pool, id written back through a funnel that mirrors the media one, owner-only and server-backed-only. A line whose bytes are in no browser is left alone rather than re-synthesized. Also: the pool guard is now a positive `ast_` test rather than an enumeration of the shapes we mint (imports never reach the pool, so this is safe in both modes); the slide ref collection is an exported pure function so its four lease sites are covered behaviourally; ASSET_QUOTA_BYTES treats every spelling of zero as opting out and refuses a malformed value at startup instead of falling back; the abort signal is re-checked after the cache read, before an uncancellable commit; and the unused `removeAsset` and pool `replace` surfaces are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * chore(storage): release 0.29.1 The asset HTTP handler now recognises a store refusal by the contract code it declares as well as by its class, so a quota refusal raised in another module realm answers 507 instead of 500. Same contract, stricter recognition, no API change: a patch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a full store recoverable and adoption course-safe Narration adoption read the local audio row by its derived key alone. That key carries no stage id and the table is keyed by id alone, so two courses can mint the same one -- a PPTX import numbers its scenes and actions deterministically, which gives every imported deck's first slide `tts_s1_speech-scene-p1`. Locally a collision only means one course plays another's clip in one browser; adopting it wrote that clip into the shared document permanently, for every device and every visitor. A row that names a course is now adopted only into that course, and a row from before that column existed only when the text it recorded is the text of the action being converted. A full asset store was made permanent last round, which was wrong three times over: it overwrote the refused bytes with an empty blob -- on the conversion path that row is a course's only copy of its own media -- it kept sending the rest of the deck to a provider against a ceiling it already knew was reached, and it left no way back once an operator raised that ceiling. A full store is neither the content's fault nor the configuration's, so it is now its own case: the bytes are kept, the pass stops at the first refusal, and the element shows the reason together with a Retry that re-attempts the upload from those bytes. Nothing retries automatically, so no one is re-billed. The narration write-back now reaches the write boundary every producer of a durable write passes through, not only the dirty mark: adoption never deletes the derived row, so a snapshot that reverts the rewrite is adopted again on the next load and allocates a fresh asset every time. Adoption is also mounted by both classroom surfaces rather than one, takes the course's abort signal, and re-validates that this browser still has the course open before each write. ASSET_QUOTA_BYTES is validated from instrumentation, where the README and the docstring already claimed it was: its only other consumer is lazy and memoised, so a malformed ceiling let the process boot and then failed every persistence request, documents and runtime included. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): remember a full store per course, and never lose retained bytes A stopped pass left the elements it never reached as placeholders with no persisted record -- deliberately, since nothing was attempted for them. But that left the next load with no reason not to try: it called a provider for the next placeholder and was refused at exactly the same point, once per reload, indefinitely. A full store is not a property of any slide. It belongs to the deployment and changes for reasons the document knows nothing about, so it is now remembered once per course in the browser's device KV. A pass that finds the marker stands down before spending anything and leaves every placeholder its "storage is full" state and its Retry; the first upload that succeeds clears it and the next pass runs normally. Narration adoption latched per course so it runs once per load, and the latch outlived the abort that leaving a course performs. On a surface that stays mounted across switches -- the workbench pane is one component for every course it shows -- owner course A, visitor course B, then back to A skipped exactly the clips the abort had cut off, and nothing else converts them. The latch is released with the abort now, and a course adopts one run at a time so a re-entry cannot hand a clip a second allocation while the previous run's uncancellable tail is still settling. A quota-blocked element retried into a network error or a 500 lost the bytes that were kept for it: the retry deleted the row before attempting the upload and wrote no replacement for an error carrying no structured code, so the next retry went back to a provider for media this browser had a moment earlier. The row now survives until an upload succeeds, the failure handler keeps whatever bytes the attempt was given, and the retry asks the question a pass asks -- does this browser already hold bytes for this element -- rather than reading an error code that a second failure has already overwritten. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): adopt real legacy narration, queue re-entries, report attempt outcomes Narration adoption admitted a stage-less row only when the text it recorded matched the action being converted. Both of those columns were added to the local audio table by the very change that moved narration onto allocated ids, so a row still carrying a derived key has neither: the rule refused every real pre-allocation course and passed only on fixtures built from post-allocation rows. What the row cannot say, the key can. A derived key names two clips only when two courses share a scene order and an action id, and an action id repeats only when something other than the generator minted it -- an import numbers them by slide position. So a key built from a generated action id is adopted on that basis, a key an import could have reproduced still needs matching text, and a row that names another course is refused however unique its key looks. Handing a re-entering caller the adoption run already in flight undid the latch release it was paired with: that run is bound to the signal the departure just aborted, so it stops at its next clip while the caller -- which has the course open and a live signal -- is told the work is done, and an effect replayed as mount, cleanup, mount adopts nothing at all. A later caller now waits for the uncancellable tail and scans again, which costs a lookup on a course that has nothing left and finishes the clips the abort cut off on one that does. One attempt at an element now reports both facts its callers need instead of a bare boolean: whether the store refused it for room, and whether bytes actually reached the store. Leaving a course clears the task table, so a retry that landed afterwards read "no failed task" as success and deleted the row holding the only copy of the media. Nothing is inferred from that table any more. Reading the localStorage property can throw where storage is denied by policy, typeof included, so the availability check moved inside the guard: this metadata is best-effort, and a rejection here strands a generation pass that has already enqueued its tasks. A retry is never blocked by the per-course "store is full" marker, but a retry that is refused again re-sets it, and adoption now reads and writes the same marker rather than issuing one refused upload per clip on every load. The two canvas element renderers and both thumbnail renderers show the reason beside the Retry, so a full store does not look like an ordinary failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): probe a full store instead of standing down, and pair notices with a Retry Narration adoption was given both halves of the per-course "the store is full" marker last round: it stood down when the marker was set, and it set the marker when its own upload was refused for room. Those halves are only safe together if something can lift the marker, and for adoption nothing could. It has no affordance of its own, it stood down before reaching its own clear, the media pass returns before its marker gate when there is nothing to generate -- so a narration-only deck, or one whose slides are already satisfied, painted no storage-full element and offered no Retry -- and narration generated rather than adopted allocates directly rather than through the media commit. The course's cached narration was then lost for good, where before it converted on the first load after the ceiling was raised. The gate is a probe now. A marked course attempts exactly one clip per load: refused, it stops and the marker stands, which costs what standing down cost; stored, it lifts the marker and finishes the course. Adoption spends no provider money, so the whole cost of probing a store that is still full is one refused upload. Generated narration lifts the marker too. The three surfaces that gained a failure notice last round drew it for any failure with a reason, including the one refusal that is reachable without server-backed persistence, so a browser-only deck painted something it had not painted before. The notice is drawn beside a Retry and nowhere else, which is what it was added for and what leaves browser-only output unchanged. Both are now asserted through the render harness the surface matrix already had. A caller arriving while a rescan is queued shares it rather than appending another. One rescan converts whatever the run in flight left and every later one would find an allocated id on every action, so a chain bought nothing and turned a single stalled upload into a course that never adopts again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): treat a refusal for room as a fact about one clip, not the deck The asset store checks each write against the headroom it has left, so a store that refuses a long opening clip can still hold every short clip behind it. Narration adoption assumed the opposite: it broke the deck at the first refusal and then re-attempted that same first clip on every later load, because the document names it first. A deck whose longest clip exceeds current headroom therefore never converted the clips that would have fit, with no affordance to recover it -- the state the probe was introduced to remove, reached through a narrower door. An unmarked load now attempts every clip, skipping the ones that do not fit, and remembers the condition only if the load ends with clips it still could not store. A marked load spends its single upload on the smallest clip left rather than the first one named: that is the clip that answers the question the marker asks, because if the smallest does not fit nothing does. The media pass keeps stopping at its first refusal, and for a reason adoption does not share -- every element it attempts costs a provider call. A rescan several callers share took the newest caller's signal, and the newest caller is not necessarily the one still there: a surface that opened a course and closed it again would stop work a surface still showing that course was waiting for, and that surface is latched, so it would never ask again. The shared run now takes a signal that is aborted only once every caller has left. The comment claiming the shared rescan contains a stalled upload was wrong -- the rescan is chained off the run in flight, so a stalled upload leaves every caller pending exactly as a chain would. It claims the bounded queue it actually provides, and the stall is recorded as a limitation. The failed-state containers took their stacking classes unconditionally, so markup differed in browser-only mode even though nothing moved on screen. Those classes are applied only when there is a notice to stack, and the tests assert the exact class attribute rather than a substring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): stop narration adoption writing the media pass's store-full marker The marker means "do not call a provider for this course". A path is entitled to write it only if its own refusal cost a provider call, and narration adoption's refusals cost nothing: it uploads bytes this browser already holds. The store also checks each write against the headroom it has left, so a clip that does not fit says nothing about whether a slide's image would. Adoption was writing it anyway, and one over-long narration clip was therefore enough to stand a course's entire image pass down on every later load -- on a store that had just accepted adoption's other clips. The author could still recover each element by hand, every load, for ever. Three rounds of narrowing this seam produced a finding each time, so it is removed rather than narrowed again. Gone: the marker read, the single-clip probe, the smallest-clip selection, and the up-front read of every row into an array -- which also retires a sampled-then-stale flag and the retention of a whole deck's blobs for the length of a run, and returns the loop to streaming one row at a time. Adoption's rule is now that every load attempts every clip it holds, once; any failure skips that clip and the load continues. The noise the coupling was meant to avoid does not arise, because after the first load the clips still outstanding are exactly the ones that did not fit -- normally none, or one. A successful write still clears the marker, and that is a different kind of statement: a write that went through is a fact this run established, where a refusal is an inference about what some other write would cost. For a course whose media needs nothing, adoption and generated narration are also the only paths that can establish it. The failure module still documented the deck-wide premise this contradicts. It now says what is true: the check is per write, and the media pass stops the deck as a judgement about cost rather than about certainty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bound a full store's cost from the store's own arithmetic Removing the store-full marker from narration adoption removed its bound too, and the code then asserted the bound was unnecessary. It is, on a store with room for most of a deck. On the store the whole mechanism exists for -- the ceiling reached, nothing fitting -- the outstanding set after every load is the entire deck, so a thirty-clip course posted thirty full blobs on every load, indefinitely. Each of those is not a cheap refusal: the bytes are uploaded, the server hashes the whole payload, and only then takes a per-principal lock and sums every entry that principal owns before saying no. The bound needs no flag, no key and nothing carried between loads. The store asks whether `used + addedBytes` exceeds the ceiling, and `used` only grows while a run is uploading, so a clip refused for want of room implies every clip at least that large is refused for the rest of that run. The run keeps the smallest size it has been refused and skips anything no smaller without uploading it; a smaller clip is still attempted, because it may fit. A deck the store refuses entirely now costs one upload per successive size minimum instead of one per clip, and a deck it has room for costs nothing extra, because nothing is refused. Only a refusal for room lowers the bar: a dropped connection says nothing about how much room there is. The deck-wide certainty premise the failure module retracted last round still stood verbatim at the site that implements the stand-down. Both copies now say the same thing: the check is per write, and the pass stops the deck as a judgement about cost rather than about certainty. The comment on adoption's marker clear now names its price. Narration of a few hundred bytes fits in headroom an image does not, so a proven write can let the next pass buy one more image that is refused again -- bounded at one, and the price of the alternative being a course whose media never generates again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(media): state the adoption bound exactly, and stop three comments describing the old rule The comment introducing the in-load bound gave its cost as "at most a handful, and the first load pays the most". Neither clause is a property of the rule. A clip is skipped only when something no larger was already refused, so a fully-refused deck costs one upload per successive size minimum in document order: one when the clips grow, about ln N for an arbitrary order, and one per clip when they only shrink -- a long opener followed by terser lines is exactly that shape. And no load is cheaper than the first, because the bound resets per run and a refused clip stays outstanding. The comment now says that, and points at what would make it exactly one for any ordering: the store returning its remaining headroom in the refusal's existing details channel, which the server leaves empty today. Two other comments still described the previous rule -- "attempts every clip it holds, every load" -- one of them twenty lines above the paragraph that introduces the bound, in the same block. Both now say what the code does. The bound's soundness is worth stating where a maintainer will look for it: quota is charged at full length with no discount for a duplicate, the sum it is checked against joins entries to blobs so the collector cannot lower it, the check takes a per-principal lock before summing, and replace and delete are refused to every browser. Nothing a run can do makes room appear inside it. One test installed a row implementation and replaced it wholesale a few lines later, so the first was dead and the survivor dropped the text the first clip's import-shaped key needs for the ownership rule -- it passed on the coincidence that the fixture's default text is the action's. Merged into one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): keep a Retry from re-buying parked media, and state the store seam once Five findings from an inline review. The standalone classroom route asked the ownership sidecar once per load and recorded only stage ownership when that ask failed. Every non-answer fails closed, so one transient 5xx left the genuine author with no resume, no Retry affordance and no legacy narration converted for the rest of the load, with nothing to change it short of a reload. The failure now records the fail-closed answer explicitly -- an answer an earlier load established must not outlive the failure that replaced it -- and an unresolved answer is asked for again, a few times over a few seconds. A real answer, however unwelcome, is final. A Retry could pay a provider for media the pool already held. When the bytes are stored and only the write-back fails in a way that keeps the allocation, it is parked and no local row exists, because that row is written only after a successful write-back. Retry now reads the parked queue exactly as the pass does and re-attempts the write-back: it re-keys the task done when the document takes it, leaves the entry parked when the slide still does not exist, and stays failed and retryable when the document refuses again. Object URLs a parked allocation owns are revoked when the entry is dropped. The commit path leaves them alone while the entry is parked, because it is then the only thing holding bytes this tab can render, so a course switch or a stage deletion was pinning the whole blob for the life of the tab. An entry a slide has already taken is left alone: the task table is displaying those URLs. The fallback lookup for cached bytes is a stage-scoped scan, and the keyed lookup misses for every row the commit path writes, so a pass was materializing and sorting the course's whole media table once per element. One scan per pass now, built on the first miss. It is sound and not merely cheaper: an element asks only for its own placeholder, and every row a pass writes carries the placeholder of the element that wrote it. "The store accepted a write, so it is not out of room" was enforced at three call sites under slightly different conditions, which made it a convention the next pool write path could silently break. It is stated once, in putAsset, for the course whose bytes it just stored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 16 天前 | |
refactor(media): share submit-poll task driver (#900) * refactor(media): add shared polled task driver * refactor(media): migrate video adapters to shared polling * docs(media): remove obsolete task adapter guidance --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> | 2 个月前 | |
fix(ssrf): disable redirects in media connectivity probes (#930) * fix(ssrf): reject redirects in media auth probes * fix(ssrf): stop redirects in strict media probes * fix(ssrf): disable redirects in Google media probes --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> | 2 个月前 | |
feat(dsl): standardize the asset manifest and converge the export paths (#1007 part 3) (#1117) * feat(dsl): standardize the document asset manifest Add asset-manifest.ts to @openmaic/dsl: the canonical AssetManifestEntry shape (ref, kind, and byteSize/mimeType/duration/voice/prompt metadata where available) plus enumerateAssetManifest, the pure document-to-manifest enumeration. An entry's ref is the reference exactly as the document holds it -- the manifest is the id-based reference enumeration with metadata, not a content hash and not a resolution result. The traversal walks the stage whiteboard, each scene's canvas/whiteboards/speech actions, and the stage video-manifest keys in document order, with logical-owner reference counts that match the accounting duplication-safe replacement uses. This settles the media-ref + asset-manifest schema question (#779 open question 4) on the side the asset-pool RFC already implied: the schema is a function of the id semantics decided there. The type lives in the dsl rather than a new @openmaic/exporter package because the enumeration is pure over document types the dsl already owns (Stage/Scene/Slide/Action), so a separate package would add a published artifact and release-workflow surface without adding a capability; the storage contract comment now points at the module. Refs #1007 * refactor(export): drive the classroom ZIP from the asset manifest collectMediaFiles used to scan the whole mediaFiles table for the stage, so any row the document no longer references -- an orphan left by an edit or a superseded regeneration -- rode along into the archive. Both ZIP collectors now take their reference sets from the standardized asset manifest (buildStageAssetManifest wraps the dsl enumeration with the compatibility rows' metadata): only referenced assets are archived, and a referenced asset whose bytes exist only in the pool is still collected via a synthesized record. Byte resolution is unchanged: pool first through resolveStoredBytes / resolveAudioBlob, with the compatibility row kept as the legacy byte fallback and as the metadata source. mediaIndex is now a serialized view of the manifest, and the missing-audio report derives from the manifest's audio entries instead of a second action walk. The audioRef mapping and the legacy audioUrl fetch path (collectLegacyAudioForExport) are untouched. Refs #1007 * refactor(video-export): take the timeline's reference sets from the manifest createVideoTimelineDeps scanned the whole mediaFiles table for the stage and derived its audio id set from its own action walk -- a third, independent answer to "which media does this course use?". Both record loads now key off the standardized asset manifest: media rows are read per manifest ref by compound key instead of by table scan, and the audio id set is the manifest's audio entries. Orphan rows were never reachable through the scene-scoped elementId-to-mediaRef bridge; now they are not even read. The bridge itself is untouched: element ids recur across scenes, so the elementId-to-mediaRef mapping stays scoped per scene, and the legacy audioUrl fallback keeps its own action walk because a URL is not a manifest ref. AssetPlan remains the video IR's view of the same references. Refs #1007 * refactor(export): resolve PPTX media through the shared resolver only Each PPTX element branch carried its own resolution chain: a task-state renderable-URL lookup first, then -- gated on the legacy placeholder predicate -- a stored-bytes override, with the poster block repeating the pattern. One helper now owns resolution for backgrounds, images, video / audio sources, and posters: opaque refs (allocated ids and legacy placeholders alike, no placeholder-pattern gate) resolve pool-first through resolveStoredBytes and embed as data URLs, concrete addresses resolve through the media state machine and keep the caller's fetch path. exportMediaResolution and the resolveStoredMediaBlob wrapper fold into the helper; resolvePptxMediaBinding stays as the state-machine entry the resolution-surface test matrix drives. Refs #1007 * refactor(export): retire the export-side Dexie byte fallbacks Export call sites no longer read bytes off compatibility rows directly. The ZIP collectors and the video timeline's audio load resolve bytes only through the shared resolvers (resolveStoredBytes / resolveAudioBlob), which answer pool-first and keep the compatibility row as their internal legacy fallback level; the row reads that remain at the call sites supply metadata (format/duration/voice/mime/size/prompt) only. The rows themselves stay for legacy and regeneration readers -- what goes is the export paths' own fallback logic. One observable tightening: a failed media row (error set, empty placeholder blob) no longer ships a 0-byte file into the classroom ZIP, and an evicted row no longer ships its empty local blob; referenced-but- byteless assets are simply absent from the archive, as they already were when no row existed. Refs #1007 * test(media): cover the enriched stage asset manifest builder Pins the join between the pure dsl enumeration and the compatibility rows: metadata attaches by ref, rows no document reference names never appear, and a referenced asset with no row keeps a metadata-free entry. Refs #1007 * fix(video-export): widen the deps stage input for the manifest enumeration enumerateAssetManifest reads the stage's whiteboard and videoManifest, so createVideoTimelineDeps declares them on its input instead of the bare id; callers pass only the id today and the optional fields stay absent. Also applies the repo prettier formatting to the files this branch touched. Refs #1007 * fix(dsl): enumerate slide audio elements in the asset manifest Slide audio elements carry their own src, and the manifest skipped them, so a manifest-driven collector could never archive their bytes. The audio slot maps to kind 'audio' alongside narration ids. Refs #1007 * fix(media): harden ref-keyed lookups against prototype-named asset refs AssetRef is an unconstrained string alias, so a media reference can legitimately be "__proto__", "constructor", or any other Object.prototype member. Plain objects keyed by such refs silently drop assignments or answer lookups with the prototype object, which rewrite paths then accept as a mapped id. Convert the remaining ref-keyed lookup tables introduced by the export convergence to prototype-safe structures: the classroom import media/poster alias maps and the legacy-conversion video-manifest reconstruction now use Map / null-prototype containers with explicit membership checks, and every consumed value is validated as a string before it is written into a src / mediaRef / audioId slot. The shared media-task lookup receives the same treatment: one centralized own-property-checked lookupMediaTask now serves the stored-bytes resolver, the PPTX embeddable-src path, the video collection path, and the element/background task resolution, so a prototype-named placeholderRef can no longer hide a re-keyed task from the fallback chain. Adversarial tests drive "__proto__" and "constructor" refs through the import round trip, the PPTX fallback path, the legacy conversion commit path, and the media-task fallback end to end, including a buildPptxBlob regression with a task re-keyed to an allocated id while retaining a prototype-named placeholderRef. * fix(export): use safe archive asset paths * refactor(dsl): centralize slide media slot roles * refactor(export): derive consumer refs from manifest * fix(export): sanitize classroom archive extensions * fix(video-export): preserve narration speech order * fix(export): enforce kind-coherent archive media * fix(export): define media coherence boundary * fix(export): carry task-owned poster binding for PPTX export A video element with no explicit poster falls back to its media task's generated poster URL, but resolveVideoMediaForElement left posterTask undefined for that case, so the PPTX manifest guard saw a foreign URL with no task-ownership exemption and dropped the video element instead of using the established runtime poster fallback. Carry the poster task binding whenever the task poster is the effective poster: the task-owned URL then satisfies the guard's objectUrl exemption end to end. A concrete explicit element poster still stays element-owned and never borrows the binding, and the guard's foreign-ref rejection is preserved (and exported as a directly testable predicate). Coverage: an element with no poster plus a task-provided poster embeds the task poster as the PPTX cover (red at the pre-fix head, green now), and a genuinely unrelated URL with no task ownership is still rejected by the guard. * fix(export): preserve legacy narration source refs in the media index The explicit sourceRef contract was partial: primary audio and generated media entries carried it, but legacy URL narration serialized no source ref. The legacy URL itself is the natural source ref — it is known at fetch time — so wire it through the collected blob into the mediaIndex entry. Import already registers serialized sourceRefs as aliases, so the URL now round-trips as an explicit mapping instead of being reconstructed only from the action's audioRef. Poster siblings are deliberately NOT given their own mediaIndex entry: a sibling poster (media/asset-<n>.poster.<ext>) is a legacy byte copy written from the video record and is not an independently referenced document asset — when the poster is a real document asset it already has its own indexed entry with a sourceRef, and import reconstructs the sibling by path derivation from its parent video entry, reusing the poster's own indexed allocation when one exists. The PR description is narrowed to match; corrected paragraph: "Archive names never interpolate refs — sequential safe paths (media/asset-<n>.<ext>, audio/audio-<n>.<ext>) with the original ref preserved through an explicit sourceRef mapping on every independently indexed media entry: generated media assets, poster assets, primary narration, and legacy URL narration (the legacy URL itself is the entry's sourceRef). Extensions are allowlisted per kind. The one exception is the legacy sibling poster byte copy (media/asset-<n>.poster.<ext>, written next to its video when the video record still carries the pre-pool poster bytes): it is not an independently referenced document asset, so it has no mediaIndex entry or sourceRef of its own — its identity is derivable from its parent video entry (same index), and import reconstructs it by sibling-path derivation from that video entry, reusing the poster's own indexed allocation when one exists." --------- Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 1 个月前 | |
release: OpenMAIC 1.0.0 — the agent workbench (#1228) * feat(storage): add an agent-session store with PG backend and layered contracts (#1163) * feat(storage): add agent-session store with PG backend and layered contracts * test(storage): avoid BigInt literals for pre-ES2020 root typecheck * fix(storage): close agent-session store review findings * docs(storage): align hook ordering and contention-probe claims with the code * ci: run on the agent-workbench integration branch * chore(storage): bump to 0.5.0 for the agent-session store * fix(storage): carry replay compaction across page boundaries * feat(agent): add the driver model contract and stage route dialect (#1165) * feat(agent): add the driver model contract and stage route dialect * fix(agent): validate route context windows and clarify dialect precedence * feat(agent): adapt the agent-session store and runtime foundations (#1167) * feat(agent): adapt the agent-session store and runtime foundations * feat(agent): resolve request owner identity via an anonymous cookie * docs(agent): document the opt-in compaction default and harden edge cases * feat(agent): add the background session runner (#1169) * feat(agent): add the background session runner * feat(agent): wire the runner into startup behind feature flags * fix(agent): stop clean interruptions from consuming the attempt budget * fix(storage): charge the attempt budget for abandoned leases but not clean parks * docs(storage): document the attempt-charging contract and decouple its tests * feat(agent): add agent session and owner event streams (#1170) * feat(agent): add agent session and owner event streams * fix(agent): close the session-existence oracle and document the owner seam * feat(agent): add agent session lifecycle routes (#1171) * feat(agent): add agent session lifecycle routes * fix(agent): validate session-create input and preserve the owner cookie on errors * refactor(storage): drop the unused active-stage API from the agent-session contract (#1174) * refactor(storage): drop the unused active-stage API from the agent-session contract Tools address stages explicitly on every call, so the store keeps no mutable session-level stage pointer. Removes resolveActiveStage and setActiveStage from the store interface, their PG implementations, the active_stage_changed lifecycle event, the session_active_stage owner event variant, and the contract tests pinning them. The active_stage_id column and the DDL check constraint stay untouched for schema compatibility. * chore(storage): bump @openmaic/storage to 0.7.0 for the contract removal * docs: document the agent runtime configuration surface (#1176) * fix(agent): repair orphaned and late tool results across interruption boundaries (#1180) * fix(agent): repair orphaned and late tool results across interruption boundaries A crash, shutdown, or provider failure can leave the durable transcript with tool calls that have no result, or with results ordered illegally for the provider. Three failure modes were fixed: - Orphaned tool calls: a run that died between an assistant tool-call frame and its result left a dangling call in the entry tree. Resume no longer synthesizes and persists receipts for it: interrupted results are a read-time provider view owned by a shared read-boundary repair, which returns the original array for a healthy transcript and never mutates the tree. - Late parallel results: a parallel tool can finish while pi unwinds an aborted assistant frame, leaving result(A), assistant(aborted), result(B) in durable order. Strict providers reject non-contiguous results, so the read-boundary repair moves existing results next to their owning assistant frame (in call order), omits incomplete unwind frames, and synthesizes receipts only for genuinely missing calls. - Interrupted calls at the write boundary: a call still in flight when the run winds down (shutdown, lease loss, cancellation, provider failure) had no receipt at all. The runner now tracks in-flight calls from their assistant frames and, before the terminal flush, appends an interrupted-result receipt for each still-orphaned call through the same attempt-fenced write chain, so a lease-stealing zombie never writes and the next claim sees a provider-safe transcript. * test(agent): pin the runner wiring for interruption-boundary tool repair * feat(agent): add neutral tool foundation libraries (#1184) * feat(agent): register a web_search tool on the session runner (#1185) * feat(storage): add a per-session URL trust gate (#1186) * feat(agent): add the skills system (#1189) * feat(agent): add the skills system (builtin directories and durable user skills) * fix(storage): serialize the user-skill quota check-and-insert per owner Two concurrent creates at the 50-skill boundary both counted 49 rows and both inserted (READ COMMITTED, no lock), overshooting the quota contract. The create transaction now takes a per-owner pg_advisory_xact_lock first, and the same-name idempotency check runs before the count check so an at-least-once retry of the create that committed as the owner's 50th row still returns its durable receipt instead of a quota error. The 23505 backstop is retained for writes that do not take the lock. * fix(agent): share unstorable-character validation and align skill lookup * feat(agent): add session materials and a fetch_url tool behind the URL trust gate (#1190) * feat(agent): add session materials and a fetch_url tool behind the URL trust gate * fix(agent): harden session material fetching * feat(storage): add an ownership scope to stage documents (#1191) * feat(agent): add material read and search tools (#1192) * feat(agent): add stage read and patch tools (#1194) * feat(agent): add page generation and deck editing tools (#1198) * test(storage): keep the PG contract suite order-independent (#1200) * fix(agent): revoke deleted-session URL authority and reject private ISATAP endpoints (#1199) * fix(storage): revoke deleted session URL authority * fix(ssrf): reject private ISATAP endpoints in strict fetches * chore(storage): bump to 0.11.1 for the session-URL authority fix * feat(agent): add roster and voice registration tools (#1201) * feat(agent): add folder organisation tools (#1202) * feat(api): add stage and material HTTP routes (#1203) * feat(workbench): add the client data layer (#1204) * feat(workbench): add the client data layer * docs(workbench): write the ported comments in English * chore(edit): remove the in-editor agent panel (#1210) * chore(edit): remove the in-editor agent panel * style: apply prettier formatting * fix(agent): report the runtime as unusable without a database (#1207) * fix(agent): report the runtime as unusable without a database * style: apply prettier formatting * feat(agent): add image, video and pptx import tools (#1211) * feat(workbench): add the agent chat surface (#1205) * feat(workbench): add the agent chat surface * docs(workbench): write the ported comments in English * fix(workbench): label the folder and rename tools on the timeline * fix(workbench): label the roster and voice tools on the timeline The reconciliation test iterates every tool the runner registers and requires a display label of its own. The roster and voice-clone tools (list_voices, set_roster, clip_audio, register_voice) reached the integration base with the roster/voice-registration tools but never gained presentation rows, so they fell through to the default branch and rendered their wire names. Port their rows from the reference implementation (labels and i18n keys verbatim) and extend the reconciliation allowlist with ROSTER_TOOL_NAMES and VOICE_CLONE_TOOL_NAMES, so a future tool cannot enter the product without a label. * feat(agent): add the material extraction lifecycle (#1212) * feat(storage): add material extraction lifecycle * feat(agent): execute queued material extraction * style: apply prettier formatting * style: satisfy prefer-const in the extraction runner * test: give material fixtures the extraction lifecycle fields The media-tools slice and the extraction lifecycle slice were each green in isolation but never compiled together: the lifecycle made derivedFrom and extraction required on AgentSessionMaterial while the media-tool fixtures predate them. * chore: remove stray task notes * fix(workbench): label the extraction lifecycle tools on the timeline * feat(workbench): add the workspace shell (#1206) * feat(workbench): add the workspace shell * docs(workbench): write the ported comments in English * i18n(workbench): align workspace keys across locales * fix(workbench): adopt the landed data layer and label the extraction tools - replace the sibling-slice seam stubs with the real data-layer modules - drop ambient declarations now shadowed by landed files - port timeline labels for the extraction lifecycle tools from the reference - align the new i18n keys across all locales * ci: retrigger * feat(api): folder routes, stage-meta viewer surfaces, and the material upload contract (#1215) * fix(storage): restore capability-based stage access * fix(api): bind document access to request owner * fix(agent): restore three-state stage access on the tool layer Port probeStageAccess and the three-state StageAccess (owned / foreign / missing / tombstoned) and gate every stageId-bearing stage tool on an owned probe, mirroring the reference per tool: - move_to_folder, rename_stage, read_stage_outline refuse a non-owned stage with the single not-yours message before touching the store. - The course/DSL toolset and the roster toolset are wrapped by withOwnerStageAuthorization: read_stage, patch_stage, grep_stage and every writer refuse a foreign stage with the same message and refusal shape. - Scene preview keeps its own probe and its own refusal text, and is registered beside the course toolset (never double-gated). - The runner injects one probe factory at the three call sites. Tests: the dsl cross-owner test premise (a foreign stage is readable by id) encoded an invented capability-read policy that the reference does not have at the tool layer; it now asserts foreign read/patch/grep are all refused while the owner still reads. Curriculum cross-owner assertions were already the reference's and now pass with the probes in place. * docs: correct per-file test counts in the fidelity report * test: fix type errors in stage-access fidelity test * test: adapt media-tool and gate suites to the owner-scoped store seam * feat(api): add owner-scoped course-folder HTTP routes Port the reference implementation's /api/folders family (list, create, rename, delete with ungroup/remove modes, and folder membership) onto the owner-bound document store, replacing its provider-based auth with the existing withRequestOwnerId / owner-scoped store seams. The storage package's folder store grows the pieces the routes need: DocumentFolder.order (schema column + max+1 assignment + ordering), renameFolder, deleteFolder(mode) with captured member ids, and setStageFolder(stageId, folderId | null) with idempotent un-filing. FolderNameError moves into folder-name-validation.ts (stage-storage re-exports it, keeping import sites intact). Every route gates on the configured agent runtime (plain 404 when off or unconfigured), keeps the reference's machine codes and envelopes, and is covered by gate tests plus a behavior suite. * feat(api): add stage-meta viewer surfaces for the classroom Port the reference implementation's viewer-facing stage state — can-edit / collected / published / generation-complete — on top of the stage-access base (stage_meta + tombstones). stage_meta gains published_at and generation_complete columns plus a stage_bookmarks table; the reference's deployment-specific origin/claimed_at columns are stripped. New gated routes: GET /api/stage-meta/[stageId] (per-viewer facts, 404 for absent/tombstoned, never returns the owner id), GET /api/stages/[id]/status, POST generation-complete / publish / unpublish (owner-only), POST /api/bookmarks. The resolver lives in lib/server/stage-access.ts. Wiring: a fetchStageMeta client with the reference's three-outcome contract, stage-store isOwner/isBookmarked/readOnly fields (upstream single-user defaults, no-op until the sidecar answers) plus setViewerAccess, the classroom apply path computing readOnly = !(isOwner || isBookmarked), the Stage editability gate, and a sidecar probe after each classroom load. A sidecar 'absent' answer keeps the editable default here because the classroom also serves local-only courses; server writes stay owner-enforced. * feat(api): port the reference material upload contract Rewrite POST /api/materials to the reference implementation's upload shape so the workbench uploader (uploadWorkbenchMaterial, which posts no session id and expects a flat 201 view) works unchanged: owner-scoped upload with mime normalization/validation (415), per-class size caps checked on the declared content-length and the streamed body (413), empty body (400), quota (429), sha256 reserve->store->finalize lifecycle with abandon on failure, flat { materialId, originalName, bytes, mime, extraction } 201, and an x-request-id echo. Adds the owner-scoped material library (owner_material table + quota + 24h lazy sweep, bytes in the host's asset registry as the neutral replacement for the reference's object-storage byte path) and the material cap configuration. The session-scoped GET list is left as-is; the reference's owner-material extraction worker is not ported (the branch's session-material extraction lifecycle already covers extraction). Gate tests now cover all 23 persistence routes across the three runtime env states; the materials behavior suite pins the new contract. * feat(media): add an optional local ffmpeg media extractor (#1213) Adds a local ffmpeg/ffprobe pipeline as a second media extraction provider behind the extractor registry, ported faithfully from the reference implementation: duration probing, keyframe-safe chunking, per-chunk ASR with timeout and deadline budgets, and timestamped transcript assembly. - Availability probing feeds the registry's candidate selection: the provider simply is not a candidate when ffmpeg/ffprobe are absent. - With neither ffmpeg nor a cloud provider configured, extraction fails with an actionable message naming both enablement paths. - Media materials route through the same extraction lifecycle and lease fence as documents; no parallel queue. - Tests inject the executable resolver so the missing-ffmpeg path is the default-tested one; the real pipeline test is skip-if-unavailable. - @openmaic/storage 0.13.0 -> 0.14.0 (media routing in the material lifecycle surface). * feat(storage): per-scene monotonic revisions via database triggers (#1214) * feat(storage): per-scene monotonic revisions via database triggers Restore the reference implementation's freshness granularity: a per-scene monotonic revision maintained by database triggers, so every writer (HTTP routes, agent tools, jobs, manual SQL) bumps it without application cooperation. - Companion revision tables + trigger functions in the storage package's idempotent schema bootstrap, with the lock-order invariant, pg_notify wakeup and the suppression switch for batch writers. - ensureDocumentSchema gained a dollar-quote-aware statement splitter. - The freshness and manifest routes serve per-scene revisions. - Mutation-verified: dropping the triggers turns the revision tests red. - @openmaic/storage 0.13.0 -> 0.14.0. * fix: forward the freshness manifest through the owner-bound store * feat(workbench): add the Pro entry points and preserve the mode-transition semantics (#1208) * feat(workbench): add the Pro entry points * feat(workbench): preserve Pro mode transition semantics * fix(workbench): drop ambient declarations shadowed by landed slices * fix(workbench): drop ambient declarations shadowed by the landed shell * feat: port workspace shell sibling modules Port the 16 leaf modules the Pro workspace shell imports but that were only ambient-declared, replacing the compile-time bridge with real implementations adapted from the sibling-slice reference: pure workbench helpers (session title, rail tab, course-chat bootstrap, created-course tabs, course-tabs memory, workspace navigation, pane navigation, pro-edit sizing, existing-course minting, first-message session), the neutral brand context and course-rename server API, the server-action session delete, the home discovery hook, the classroom pane host with its load-policy leaf, the theme toggle and floating-layer owner, plus the floating-layer-owner wiring the dialog/dropdown/tooltip portals stamp. Also add the workbench-shell locale copy for all 12 locales, port the reference tests for the ported modules, and drop types/workbench-sibling-slices.d.ts now that every declaration has a real implementation. * docs: keep ported comments in English and deployment-neutral * docs: announce 1.0.0 and refresh the feature overview (#1216) * docs: announce 1.0.0 and refresh the feature overview * docs: finalize 1.0.0 README after feature merge * fix(agent): control-plane routes answer 404, not 500, without a database The agent control-plane routes gated only on the runtime flag, so an enabled-but-unconfigured deployment (flag on, DATABASE_URL empty) answered 500 from a store that cannot connect. Gate them on the configured check instead, matching the stage/material routes: the whole surface is cleanly absent until both the flag and the database are present. The status probe keeps reporting both bits. * test: mock both runtime gate exports in the control-plane route suites * fix(agent): abort in-flight TTS on cancel and bound each provider request with a timeout (#1217) The generate_tts / scene-tts path checked the runner's AbortSignal between actions but never created the provider HTTP requests with it, so a session cancel left a hung synthesis fetch in flight until a restart repaired the tool result. Thread the signal end-to-end: TTSModelConfig carries an optional signal, generateTTS combines it with a per-request timeout (TTS_REQUEST_TIMEOUT_MS, default 30s, ported from the reference runtime's TTS bounds) via AbortSignal.any, and every provider fetch (openai, azure, glm, qwen incl. voice-clone + audio download, voxcpm, minimax, doubao, elevenlabs, lemonade) is created with that signal. A timeout now fails the tool call with TTSRequestTimeoutError (a clear retryable error) instead of wedging the session; a caller cancel propagates as the interruption so the runner settles the session as cancelled without a restart. Tests: hung-provider simulation rejects at the timeout with the retryable error; abort mid-flight aborts the captured request signal and surfaces the interrupted shape; removing the signal wiring makes the abort tests fail (red), restoring them turns green. * fix(workbench): PG-mode home listing via owner stages; keep the interrupted terminal course card (#1218) Finding 1: with server persistence on, listStages resolved to the generic GET /api/persistence/documents listing, which the capability model deliberately answers 403 FORBIDDEN_DOCUMENTS for (reads by id, listings owner-only). The home/workspace library now lists through the owner-scoped GET /api/stages surface (same anonymous-owner cookie the workbench uses) when server persistence is enabled; the server-side 403 is untouched. Finding 2: a run interrupted (session_interrupted) and repaired (session_resumed) that ends cancelled before agent_end stranded its pending classroom sightings, so the timeline's terminal card lost the course the answer produced. session_end (cancelled) now flushes the pending sightings into the same course card set agent_end paints, before the stopped caption. * chore(workbench): remove the bookmark concept and the saved-courses drawer (#1219) * chore(classroom): remove the bookmark ('collected') concept entirely The stage-meta viewer port introduced a bookmark surface (stage_bookmarks table, POST /api/bookmarks, the isBookmarked sidecar field, and a readOnly rule that let a saved course stay editable). The product has no such concept, so remove it as a closure: - delete the /api/bookmarks route and the stage_bookmarks table plus its query helpers from the persistence bootstrap - drop isBookmarked from GET /api/stage-meta/[stageId] - simplify the classroom read-only rule to readOnly = !isOwner across the sidecar client, ownership signal, classroom load, stage store and the classroom page - keep publish/unpublish, generation-complete, isOwner and isPublic exactly as they were - update the gate and stage-meta route suites and the README mentions The workspace rail's Bookmark glyphs and comments describe the upstream saved-courses (favorites) section, which is driven by isOwner and renders no collect affordance; they are kept as unrelated homonyms. * chore(workbench): remove the saved-courses drawer UI The first pass removed the bookmark data model but kept the rail's "Saved courses" drawer, judging it a separate surface driven by `isOwner === false`. The home/workspace listing is owner-scoped, so that flag can never occur: `allSaved` is permanently empty and the drawer (plus the collapsed-rail Bookmark mini-button) is a dead affordance. Remove it: the SavedDrawer component and its mount, the savedOpen / savedSection state, the allSaved / matchedSaved derivations, the 'saved' variant of the course-list renderers, the mini Bookmark glyph, the drawer-only CSS, and the drawer's i18n keys from all 12 locales. The courses tab is now exactly one folders tree. The authored/favorites split in workspace-tree.ts goes with it; the tree module no longer reads `isOwner`. The discovery course type keeps the field — the shell still reads it for read-only gating. Upstream has no collect concept; the drawer could only ever render empty here. The reference implementation HAS this drawer (its favorites come from its account system), so this removal is a deliberate upstream product decision, not a fidelity bug. * fix(workbench): restore the attach entry, add the rail settings entry, pin all three entry points (#1221) * fix(workbench): restore the composer attach entry by gating it on the live runtime The AttachButton's rollout probe read a `materialsEnabled` field that this branch's /api/agent/runtime never answers (the materials routes gate on the runtime itself, like the stages), so the gate could never pass and the attach button never rendered — the Pro launch and chat composers showed only the @-mention and enhance glyphs. Substitute the field with the runtime's `enabled` value, which IS the upload action's precondition: POST /api/materials answers 404 whenever it is false, so the render condition now equals the action precondition (no dead button). The button's label (`proMode.attach`) is a user-visible string that becomes visible again; port the reference implementation's own translations verbatim into the 11 locales that still carried the Chinese copy. * feat(workbench): add the settings entry to the rail's bottom-left cluster The reference's rail foot carries a cluster of utilities (its saved-courses drawer, the language switcher, the display toggle). This branch removed the drawer — it could only ever render empty here — and the product decision is to fill that freed spot with the settings entry. Add a settings trigger to the foot cluster (expanded rail, beside the language and display toggles, and on the collapsed strip) and mount the model/provider SettingsDialog in the rail, wired to the trigger. It is the same dialog the classic home opens from its header pill; the workspace had no settings entry of its own, so nothing is duplicated within a surface. * test(workbench): pin the restored upload, attach, and settings entry points Covers the three restored entry points: - the courses-tab upload control: rendered beside the course name filter, wired to the discovery hook's ZIP import trigger, disabled while an import runs, and gated by the same condition as its action (the courses tab); - the composer attach control: an actual render of AttachButton under both probe answers (visible when the runtime says the upload path is live, hidden otherwise), its mounts in the launch and chat composers, the branch's runtime-field substitution in the probe, and the reference's own `proMode.attach` copy in all 12 locales; - the settings entry: the trigger in the rail's foot cluster (expanded and collapsed), beside the language and display toggles, opening the SettingsDialog the rail mounts. * chore(config): the Pro workbench flag implies the MAIC Editor gate (#1223) A workbench build without the editor toggle has no way to edit a course: enabling NEXT_PUBLIC_PRO_WORKBENCH_ENABLED while forgetting NEXT_PUBLIC_MAIC_EDITOR_ENABLED produced exactly that split-brain bundle. The workbench IS Pro mode, so its flag now implies the editor gate; the standalone flag remains for deployments that want the classroom editor without the workbench. Documents both flags in .env.example. * fix(agent): wake SSE tails and the runner on durable deltas (streaming fidelity) (#1222) The Pro workbench chat did not stream: the session/owner SSE routes polled the durable event log on a 5s/30s clock with no wakeup, so message_update deltas (written at 150ms cadence) reached the browser in poll-sized blocks and the thinking strip only mounted after the whole reasoning text had accumulated. Port the reference's LISTEN/NOTIFY delta path: - storage: add in-transaction wake hooks (onSessionEventAppended, onOwnerEventAppended, onCancelRequested) so a host queues pg_notify in the same transaction as the durable append; align readEventsAfterForReplay to rank the bounded page so the first delta after the cursor is always kept (the live tail can never starve). Bump @openmaic/storage to 0.18.0. - app: port the process-wide event-notify bus (dedicated LISTEN client, self-check probe, reconnect backoff; notify through the storage transaction surface), wire the store hooks, subscribe both SSE routes before the initial read with the reference's initializing gate, and give the runner one {kind:'session'} subscription whose wake runs the cancel check and the message drain. Polls stay as the lossy-NOTIFY backstop. - lifecycle: start/stop the bus from instrumentation. Tests: storage hook + compaction contract; route wakeup latency; runner wakeup wiring with a fake agent; bus unit tests; PG contracts proving a real append wakes the routes and a live SSE route forwards a message_update on the wakeup, and that a rolled-back append never wakes. Also fix the pre-existing park-attempt-budget PG test TRUNCATE (missing CASCADE against newer FK tables). * fix(storage): asset writes self-deadlocked against pooled PostgreSQL (#1225) * fix(storage): refuse the non-transactional byte-write deadlock configuration A byte store whose plain write() runs on its own pooled connection cannot be invoked from inside a registry write transaction: after the transaction has claimed the blob-row lock, that write blocks on the lock the transaction just took while the transaction waits on the write - a self-deadlock PostgreSQL cannot detect (one side is idle in transaction). There is no lock-safe ordering for such a writer: bytes must be written after the row claim (writing before it lets the collector delete the bytes while the upsert waits), and any second-connection write after the claim is the deadlock. The configuration is therefore detected and refused: - AssetByteStore gains writesOutsideRegistryDatabase?: true, declaring that the layer's plain byte operations cannot contend for the registry's row locks. - PgAssetStore refuses put()/replace() up front (and defends coordinatedWrite) when the byte store has no writeWith and does not declare the flag, throwing a clear configuration error before any row is claimed. - The collector mirrors the guard on its delete path (deleteWith or a declared out-of-registry layer, else a configuration error). - The object store declares the flag (its out-of-transaction write remains legitimate); the in-registry PostgreSQL byte column provides writeWith / deleteWith instead. - Write transactions (put/replace/remove) set SET LOCAL lock_timeout = 30s so any future lock-contention variant fails loudly instead of hanging. Bumps @openmaic/storage to 0.18.0. * fix(persistence): forward the transactional byte methods through the lazy asset byte-store wrapper The no-bucket case of lazyAssetByteStore returned a bare { write, read, delete } and dropped writeWith/readWith even though the underlying PgAssetByteStore has them. The registry's hasTransactionalWriter duck check then failed and put() fell back to the byte store's own pooled connection, which blocks forever on the blob-row lock the registry transaction just took when the bytes live in the same PostgreSQL - the production self-deadlock. The no-bucket layer is statically PgAssetByteStore, so its transaction-pinned methods are forwarded eagerly (typed against the real signatures via PgForwardedByteStore). The bucket case keeps its lazy-probing semantics: no transactional writer exists there, the signed-URL method stays absent or lazy exactly as documented, and the wrapper now declares writesOutsideRegistryDatabase so the registry may run the plain write inside its transaction. New tests pin the wrapper's transactional capability red-to-green and assert put()/resolve() route byte traffic through the transaction-pinned queryable. * fix(home): cap the generate-prep ingest drain at 3s so Generate never waits the full server budget The classic home flow's Generate click drained in-flight ingests for the full 15s server budget. Cap the wait at GENERATE_DRAIN_CAP_MS (3000ms, documented as a UX bound) and reuse the existing timeout fallback: sources that miss the cap proceed on the legacy byte path and each late-resolving id is released. * chore(storage): bump to 0.19.0 over the concurrently landed 0.18.0 * fix(agent): bound every tool call with a timeout; never resurrect a cancelled session (#1226) * fix(agent): bound every tool call with a global timeout and settle it on cancel A tool await that neither resolves nor rejects wedges the session forever: the lease keeps heartbeating and the driver never reaches its next cancel checkpoint. Race every tool execution (in buildAgent) against a hard budget (OPENMAIC_AGENT_TOOL_TIMEOUT_MS, default 10 min, per-tool overrides for known long runners) and against the caller's AbortSignal, so even a signal-ignoring await cannot keep a cancelled session running. On timeout the call rejects with AgentToolTimeoutError; the agent loop turns the rejection into a structured error tool-result the agent can retry or proceed from, and the abort signal is delivered to the tool's in-flight work through a derived controller. Zombie-tool updates after settlement are dropped. * fix(storage): never re-lease a cancel-requested session; settle it as cancelled on claim The claim scan treated a session with cancel_requested_at set as a normal claim candidate: after a restart it re-leased the same session for attempt N+1 and resumed generating despite the pending cancel. claimNextSession now settles such candidates as cancelled under the claim lock (status cancelled, attempt reset, lease and cancel request cleared, terminal session_end event and owner projection) instead of leasing them, then keeps scanning. Bump @openmaic/storage to 0.18.0. * docs: takeaway-style 1.0.0 announcement with bilingual guide links The 1.0.0 head is now a short takeaway block — badge links to the official user guides (English and Chinese), five one-line highlights, and pointers into Features and the workbench setup section — instead of six dense paragraphs. The detailed provider-neutrality and freshness notes move into the Features workbench section, phrased database- neutrally (the announcement no longer names a specific database). Release date corrected to August 27. * fix(workbench): restore editor chrome, mode transition, streaming, materials, mentions, folders (#1229) * fix(workbench): wire workspace folder routes * fix(editor): restore reference workbench chrome * fix(workbench): persist composer materials and course refs * fix(workbench): preserve live reasoning frames * fix(persistence): back off failed streaming saves * chore(workbench): retire stale slice seams * test(editor): cover element pin layer * chore(storage): bump to 0.21.0 for the user-message ref/material fields * chore(editor): translate ported code comments to English * fix(agent): fence durable tool writes and consume cancel requests atomically (#1230) * fix(agent): enforce provider force-off in agent tools and scrub vendor identity from tool results (#1231) * fix(materials): serialize per-owner quota reservations and make crashed uploads reclaimable (#1232) * fix(editor): resolve dock-bar i18n keys, remove dock height drag, wire element referencing (#1233) * fix(workbench): send the opening session message exactly once with refs intact (#1234) * feat(editor): port timeline TTS preview single-flight and voice-all state latching (#1235) * fix(media): restore the reference classic media chain (#1236) * fix(import): adapt imported PPTX canvas size so decks render without overflow (#1237) * fix(editor): complete element referencing — renderer DOM contract and GenUI picking aligned with the reference (#1238) * test(providers): reconcile the provider-config vendor-token debt count after the main merge The integration line's AK/SK fallback for the managed document provider adds occurrences that main's allowlist snapshot predates. Same mixed-composition debt category the group already documents; no new vendor behavior. * test(providers): reconcile vendor-token debt counts with the integration line The main-merge brought main's neutrality-guard snapshot next to integration features it predates (media-extractor fallback chain, local voice-profile deletion semantics, the enabled-TTS helper). Same debt categories the guard already documents; counts updated to the guard's own tally and two grouped entries added. No new vendor behavior. * fix(agent): carry reasoning through the completions dialect so the thinking strip renders (#1239) * feat(skills): add Feynman and spiral curriculum methods (#1240) * feat(agent): port missing reference tools and skills (parity audit) (#1241) * feat(media): retire asset-registry wiring; media and materials follow the reference byte model (#1242) * fix(classroom): center adapted canvases in the stage and send back navigation home during generation (#1243) * feat(settings): skill management with real list, download, delete, and upload (#1244) * feat(settings): skill management section with real list, detail, and zip download * feat(skills): owner skill delete and upload across storage, API, and settings * fixup! feat(settings): skill management section with real list, detail, and zip download chore: neutralize a reference note in the settings header comment * fix(media): persist origin-independent classroom-media references from the agent runtime (#1245) * feat(editor): float the insert toolbar in the outer frame with collapse (#1246) The insert strip was bounded to the slide card, so it could only ever sit on top of slide content: the card's overflow clipped it and it could not be parked in the padding beside the slide. Move it into the studio frame the element picker's panel already roams (CanvasOverlayPortal + the frame selector), so both canvas overlays share one bounding container and their handles behave the same. While picking, the strip rises over the picker and goes inert, which is the z-order CANVAS_OVERLAY_Z already documents. Add a fold beside the grip: the chevron collapses the strip to that grip row and back, with the buttons unmounted rather than hidden. The fold is session-local state owned by EditShell, next to the drag offset, so a surface swap keeps it; nothing is persisted. Expanding a strip parked at the bottom edge re-clamps through the same bounds rule the keyboard move uses. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * fix(workbench): align the chat timeline's left edge with the composer (#1247) * fix(agent): fence session claims while an ask_user question is outstanding (#1248) * fix(agent): settle-time rescue tracks real delivery instead of a count offset (#1249) * fix(persistence): migrate owner_material to oss_key and drop legacy asset_id (#1250) * docs(readme): surface the 1.0.0 user guide badges at the top (#1253) * fix(workbench): show newly created folders in the sidebar without reload (#1254) * docs(readme): add the release version prefix and drop the opt-in framing * fix(workbench): single-source the chat gutter so timeline and composer share a left edge (#1255) The transcript and the composer each established their own column: their own `px-*` gutter and their own `mx-auto w-full max-w-*` centering wrapper. Equal padding values were never enough, because the two columns are centered inside different containing blocks — the transcript's is a scroll container, whose content box is narrower than the composer footer's by the scrollbar's width: transcript text left = pad + (pane - 2*pad - scrollbar - measure) / 2 composer box left = pad + (pane - 2*pad - measure) / 2 The padding cancels out of the difference and what remains is `-scrollbar/2` at every padding value, so the transcript sat half a scrollbar to the left of the composer and tuning the two paddings against each other could not move it. The column is now established once, by the nearest common ancestor of both (`chatColumn`), and the scroll viewport and the composer footer are siblings inside it that add no horizontal inset of their own. The cap carries the gutter on top of the 760px reading measure, so the text column keeps its width. The handed-over question row drops the padding that indented it past the agent's prose; framed rows keep their own inner padding, which is what a card's border sitting on the column edge means. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * fix(workbench): lock pane-embedded classroom to edit mode (#1256) The workspace right pane painted the full learning chrome — speed control, play button, learner avatars, mic bar — for a course the agent had just created, then flipped to edit once the first scene landed. resolveStageChromeMode treated playback as the DEFAULT branch for a hosted classroom, so every shortfall fell into it: a course whose tab opens at stage_link time has no scenes yet, so currentSceneId is null and isHostedSceneEditable is false. A folded pane parked the playback root behind the fold and cross-faded it out over the pane on unfold, and a failed editor chunk dropped into playback permanently. Lock it at the pane instead of defaulting per entry path: - WorkbenchPanelProvider — the single element that mounts a classroom into the workspace — publishes editPinned (visible && !playback). Every entry path passes through it, so none of them decides. - The hosted resolution can no longer degrade to playback. Start Learning (workbenchLearning, new input, split out from pane visibility) is the one door; everything else resolves between the neutral loading shell and edit. - Stage's chrome dispatch is exhaustive on chromeMode, so the playback root is no longer the else-branch of a condition about the current scene. No flicker: chromeMode is resolved during render, and preloadEditor now answers synchronously (isEditorPreloaded) so a remount with the chunk already registered paints edit on the first frame. A failed import is no longer cached forever, so the lock cannot strand the pane. Standalone classrooms keep their stored mode unchanged. --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> | 1 个月前 | |
fix(audio): resolve CDN-backed narration consistently (#1521) * fix(audio): resolve CDN-backed narration consistently (#1515) * fix(audio): keep export fallback resolution consistent --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> | 12 天前 | |
feat(media): write generated media through the asset pool under server-backed persistence (#1392) * feat(media): store generated media in the asset pool when persistence is server-backed With server-backed persistence the document is durable and shared, but generated media stayed in the producing browser: the document kept its gen_img_* / gen_vid_* placeholder and narration kept a browser-derived audio id. Every new browser that opened such a course re-ran generation for every slide, and it never converged, because the address of the generated bytes was never written back into the document. Under server-backed persistence only, the classic generation chain now stores bytes in the asset pool first and writes the id the pool allocated into the document. - The client bootstrap configures the asset seam alongside the document and runtime seams: an HttpAssetStore over the persistence endpoint carrying the same credentials the document store carries, marked server-backed. The seam preflight now covers all three, so a failure still cannot half-configure persistence. - Image, video and TTS generation commit in one fixed order: provider, pool, document, local cache, task. A reference reaches the document only after put returned an id, so a document can never name bytes that were not stored. A failure before the write-back leaves the placeholder with the provider called exactly once; the retry happens on the next owner load. - The write-back is a per-slot rewrite through mutateDocument, which re-reads the current document under the per-stage lock, so it cannot clobber a newer scene. The open course is refreshed with the same rewrite without being marked dirty. - "Has this already been generated?" is answered by the document (the slide exists and no longer holds the placeholder) instead of by this browser's task table. - The classroom's resume effect fails closed on ownership: only a resolved owner starts generation, so a viewer opening a shared course spends nothing. - The local media and audio tables become a per-tab cache. A failed cache write costs a re-download, never the media. Browser-only mode is unchanged: every new call site sits behind the server-backed gate, the local tables stay authoritative there, and placeholders stay in the document. Rendering and export needed no changes. HttpAssetStore.resolve mints an object URL exactly as the browser store does, and the export byte resolver was already pool-first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(classroom): make the generation owner gate a three-outcome rule and apply it everywhere The gate refused everything but a resolved owner, which read a sidecar that answered "no ownership fact exists for this course" as a reason to block. That is the answer a deployment without the sidecar's server-side prerequisites gives for every course, and the answer a course with no ownership record gives: in both, there is nobody the operator's budget needs protecting from, and refusing strands the course's own author behind a question that can never be answered. Ownership is now four states over the sidecar's three outcomes. A definite answer splits into owner and not-owner. An absent record is its own answer, ownerless, and generation proceeds — the behaviour such a deployment had before the gate existed. Only the absence of an answer, a transport failure or a load that has not asked yet, stays unresolved and fails closed: "we could not ask" must never be read as "nobody owns this". One mapper turns a sidecar result into that state, and one predicate decides on it. The workbench classroom pane runs the same resume effect and had no ownership input at all, so a viewer opening a shared course there could still spend the budget. It now asks the sidecar once per course, in parallel with its load and feeding only the generation gate, so its read-only and edit behaviour is unchanged. The shared progressive-load policy carries the gate for it, with both new inputs required rather than defaulted so a future caller cannot omit them into an open budget. Its stale comment claiming ownership could not be expressed here is corrected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the write-back survive autosave, arrive before the scene does, and never leak Independent reviews of the write-back found three ways a durable document could still end up naming a placeholder, and two ways the gate that protects the operator's budget could be walked around. An autosave round captures the store synchronously and writes that capture, so a round already in flight when a rewrite landed wrote the placeholder straight back over the allocated id, and nothing marked the store dirty again to correct it. The rewrite now marks the units it changed, which leaves a corrective flush queued behind the stale one; re-saving a scene that already holds the id is idempotent, losing the id is not. Media is generated from outlines in parallel with scene content and usually finishes first, so the slide that will carry the placeholder does not exist yet and the write-back has nothing to rewrite. That was the ordinary path, not a tail case, and its result was discarded: the task was marked done, the scene was added afterwards with its placeholder intact, and a second pass in the same run could call the provider again. The allocation is now held under the placeholder — which also answers the skip test, so nothing pays twice — and applied when that scene is committed, before its first save. One complete pass now leaves no placeholder behind. A failed commit used to abandon what it had already allocated. A poster upload that failed threw away a stored video and sent the retry to submit the most expensive job in the system again; a rejected write-back left registry rows that name bytes nothing references, which the byte collector cannot reclaim because it only collects blobs no row names. A poster failure now costs the poster, and a write-back that reached nothing reclaims what it allocated. A partial write is left alone, because the document already names it. The ownership gate is fail-closed again. Treating the sidecar's 404 as permission was wrong: the client cannot tell "this course has no owner" from "this deployment told me nothing", so a visitor who opened a shared course could bill the operator. The root cause was the sidecar itself, which gated on the agent runtime although every persisted course has an owner regardless — the persistence route resolves one for every request. It now gates on server persistence, so the configuration that made 404 the universal answer has real ownership facts to report, and the gate can refuse everything but a named owner. Retry affordances answered to no gate at all. A viewer of a shared course with one failed image was shown a Retry button that called the provider. Both retry entry points and every surface that draws them now read one shared permission, so what is offered and what is allowed are the same value. Also: narration regeneration no longer pretends it can replace bytes behind a live id — the exclusivity proof that would allow it is refused by construction once references leave the browser, so it forks to a fresh id and says so; the "already generated" test lets a finished deck answer from the document alone, since scene order stops identifying an outline once slides are inserted or deleted; stored assets record a specific media type rather than a generic transfer type; the pane no longer asks the sidecar in browser-only mode; and the funnel's docstring now states what the per-stage lock actually guarantees, which is same-browser serialization and not a cross-browser compare-and-swap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): park allocations in the deciding turn, never reclaim on an ambiguous write A delta review of the write-back found the first-pass fix still had a window, and the reclamation it added could delete media the document already names. The allocation was parked after an awaited local cache write. A scene committed in that window reconciled against a registry that did not hold it yet, so the document kept the placeholder — and the entry recorded a moment later then answered the skip test as "already handled", so nothing could correct it. Parking now happens inside the write-back, in the same synchronous turn as the decision that nothing could take the reference; no await separates the live check from the park. Allocations parked by an earlier pass are handed to their slides at the start of the next one, before anything decides what still needs generating, so a held allocation whose scene has since arrived becomes a rewrite rather than an answer. Reclaiming on a rejected write was unsound: a rejection does not prove the server did not apply the write, so deleting the asset could break the scene that now names it. The funnel decides instead, and says so: it reclaims only when no store write was ever issued and nothing took the reference. Anything else is placed if its slide exists and parked if it does not, so the next pass reuses the bytes instead of paying for them again. When a write fails after part of it landed, the live store is brought up to the document before the error is rethrown — otherwise the next ordinary flush would overwrite the half that did land, with the ids deliberately not reclaimed. Parked allocations are now cleared with the course. Classic placeholders are reused across runs, so one surviving an interrupted run would be handed to a different slide of the next deck: the previous picture, on a slide whose provider was never asked. Both classroom surfaces clear the arriving course, the deletion cascade clears the deleted one, and clearing the database clears them all. Two more ways generation could start without asking the gate are closed. An overlapping pass — an outline retry re-enters generation with every outline while the first is still working — re-requested elements whose provider call was already in flight; a task that is not done is an answered request, not an unanswered one. And narration regeneration in the timeline editor called the TTS provider and allocated a pool asset with no ownership check at all; it now reads the same permission, which withholds both the per-line and whole-timeline controls and refuses the call. Finally, a pane opened during the stage-link availability gap recorded the sidecar's 404 for a course that was moments from existing and never asked again, leaving the real owner locked out of generation until it remounted. Ownership is re-fetched once the document becomes available; the gate stays closed until an answer arrives, so asking again can only open it for someone entitled to it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the asset routes reachable, and stale snapshots harmless A full-branch audit found that the deployment this project documents could not store a single generated asset, and that several routes into durable storage could still write a placeholder over a reference that had already landed. The persistence route sent asset requests through the development authenticator, which refuses outright in a production build that has not explicitly opted into it — and the documented server-persistence recipe produces exactly that build. Every store and every read answered 401, so images and video failed on every slide while re-billing the provider on each retry, and a narration failure stopped the deck at its first slide. Assets live in one shared partition by design, so there was never anything per-caller for that authenticator to decide: the route now resolves the asset principal itself, alongside the owner it already resolves for documents. Runtime sessions are genuinely per-learner and keep the development authenticator until real session verification replaces it. And narration that cannot be stored no longer fails its scene: the line stays unvoiced and retryable, which is what an image that cannot be stored does to its slide. Placeholders could also come back from behind. A queued autosave's snapshot, an editor-history entry replayed by an undo, the departing save a course switch flushes — each captures content at its own moment, and any of those moments can predate a write-back. Point fixes at each producer would leave the next producer to rediscover the bug, so the check lives at the write boundary every producer passes through, and the allocation record it consults now outlives the parked queue: a placeholder whose rewrite landed long ago is exactly the case it catches. Two ways generation could be lost or repeated are closed. A pass now claims the elements it will reach and releases them however it ends, so an overlapping pass stands down while an aborted one strands nothing — previously its tasks stayed `pending` and every later pass skipped them with no retry control to recover them. And the media abort controller is aborted before being replaced, so a superseded pass stops calling providers instead of running on for a course the user has left. The remaining two are narrower. The workbench pane asks for ownership only after a document load succeeds, and after every later one, mirroring the page route: the load is what creates the ownership row the first time a course is opened, so asking beforehand asked about a course that did not exist yet and locked its author out for the mount. And the ownership gate on the timeline editor now withholds narration regeneration alone; listening back to existing narration and seeing whether a line has any spend nothing and stay available. Known limitation, unchanged and now stated plainly in the comments that used to point at it as a solution: nothing reclaims an unreferenced pool asset. The registry sweep is written but not wired up, and the byte collector only reclaims blobs no registry row names, so every narration regeneration and every abandoned allocation leaves storage behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): gate asset mutations, and make claims and allocation records survive a handoff Opening the asset routes opened all of them. Reads and allocations are meant to be as open as document reads and creates already are, but no authorization hook was supplied, so the handler's default admitted PUT and DELETE too — and those scope by principal key alone, which is one shared constant. Any caller who learned an id, and a document read hands out every id its slides name, could overwrite or destroy another author's media. Mutations now require the deployment's credential, which in a production build without the development-auth opt-in means they are refused outright; reads and allocations stay open. The route comment says what the posture is and what it is not: the deployment-level fence is the access code, and no per-principal quota is configured. The client's own reclaim is best effort to match — losing an argument about deleting an asset must not cost a task its retry, and the bytes are left for server-side reclamation. The pass claim could not survive the handoff it was written for. A retry aborts the live media pass and starts its replacement in the same synchronous block, long before the aborted pass's cleanup runs, so the replacement saw every element still claimed, collected nothing, and returned — leaving each unreached element at pending with nobody coming back for it and no retry control to recover it, which is the exact failure the claim was introduced to prevent. A claim now carries its pass's signal and is retired the moment that signal aborts, and a pass releases only claims it still owns, so a late unwind cannot take its replacement's work. Claims are also acquired at the single point every request passes through, so a single-task retry participates too — previously a retry awaiting its provider was invisible to a pass starting alongside it and both called it. The allocation record could outlive the bytes it named. It was written before the write-back attempted anything and survived the reclaim that followed a failure, so when the slide finally arrived the write boundary stamped a deleted id into the document — and the placeholder it replaced was gone, which reads as already generated and stops anything from retrying. The record is now written only where the allocation is retained, and forgotten wherever a reclaim removes the bytes, including the narration rollback path. The tests follow. The route test drives the real storage handler against an in-memory registry instead of a stub, so it can see what the resolved principal is then allowed to do; the handoff test performs a real abort mid-pass rather than starting from an already-aborted signal; and the guards that could only assert file layout now assert the property they care about, or have been replaced by behaviour. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make media passes serial per course instead of tracking element ownership Three rounds of per-element claims each produced a new way to lose an element. Whole-pass reservations swallowed a Retry for an element the same pass had already failed, leaving it pending with the affordance gone. Retiring a claim by its signal freed an element whose commit was still uploading, so the replacement pass paid for it twice. A claim held for a failed element stranded its retry. The bookkeeping is the defect: every refinement of "who owns this element right now" answered the question at a moment when the answer was already stale. Passes for one course are now serial. A replacement aborts its predecessor, as before, and then waits for it to settle before collecting. That removes the question entirely: a commit already under way finishes — its bytes stored and its reference written, so the new pass sees a resolved slide and skips it — and an element the aborted pass never reached is still a placeholder and gets collected like any other. The claim set, the reservations, the signal retirement and the identity-checked release are all gone. The task table is consulted for one thing only: an element that is generating right now is a single-element retry running alongside the pass, and taking it too would pay twice. Pending is deliberately not a skip reason — it means a pass once intended to reach an element, which an abandoned pass leaves behind with nobody acting on it, and reading that as answered is what stranded elements before. A retry runs concurrently with a pass, because a pass never revisits an element it has processed, and it re-reads the task after its own await and refuses before touching it: marking first and refusing afterwards destroyed the failed state that draws the affordance. Browser-only mode is back to exactly what it was. The abort is now conditional, the waiting does not apply, and the original status-based skip is restored verbatim. Two baseline lines remain changed in each of the two files, and both are behind a server-backed fork whose else-branch is the original. Two smaller things. The allocation record becomes visible when a write goes on the wire rather than when the round trip ends, and the write boundary reconciles under the document lock rather than before it — a save queued during a write-back was otherwise captured with the placeholder and, for a course the user had left, had no corrective flush to follow. And the comments that said a refused reclaim leaves its bytes for server-side reclamation were wrong: nothing collects them, because the registry entry still names its blob and the sweep that would remove it is not wired up. They now say the bytes leak. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a deferred pass re-earn its right to run, and bound the commit it waits on Serializing passes moved their body out of the block that launched them, and three things followed from that. A pass now wakes when its predecessor settles, which can be after the user has left the course. It enqueued before it looked at its signal, into a task table keyed by element id alone — and placeholder ids are not unique across courses, which is why the classroom clears that table on arrival. So a departing course's pass seeded the arriving course's table with tasks carrying the wrong stage id, and a Retry routes by that id: the reference went into the wrong document. A pass now re-validates after the wait, before touching anything shared. The same lateness broke the skip test. `documentSkipIndex` answers only while the live store is on the pass's stage, and returning nothing put the collection loop on the browser-only rule — a silent demotion from "the document is the authority" to "this browser's task table is", on exactly the path where that table has just been cleared. Every element the predecessor had committed was collected again, paid for again, and its second write-back found no placeholder to rewrite, so its bytes were parked where nothing will ever reference them. In server-backed mode an unreadable document now means the pass stands down. And waiting was unbounded. A commit is uncancellable: the asset client takes no signal, and a document write cannot be half-undone. One stalled upload therefore froze the course's media generation for the session — the replacement never collected, the element sat on a skeleton that draws no Retry, and only a reload recovered. The pass's signal is now threaded into the media proxy fetch, and the commit is bounded by a deadline. The deadline is on the wait, not the work: the commit carries on, and if it lands late the document simply ends up correct, while the element becomes retryable and the queue moves on. The tests that were meant to pin the previous round were not sensitive to it. Two asserted end states where the mechanism only changes ordering, and one of them rigged the document read so the assertion held whether or not the pass had waited; a third covered half of what it claimed. They now observe the ordering directly — nothing is issued while another pass for the course is working; in browser-only mode a second pass reaches its provider immediately — and the reconciliation under the document lock has a test that fails when it moves back outside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * revert(media): drop the commit deadline and the abortable download The deadline bought less than it cost. Abandoning a commit after two minutes makes the element retryable while the real commit is still running, so a Retry starts a second commit for the same placeholder against the first: two provider calls, two allocations, and whichever lands second stamps its result over the other's task by element id. The allocation record is keyed by placeholder, so the loser's cleanup erases the winner's record, and the write boundary then puts the raw placeholder back into the document. That is the overlap serial passes were built to remove, reopened through the one door serialization never covered. So a stalled commit holds the course's media queue until it settles or the page is reloaded, and that is written down rather than papered over. The wait is unbounded on purpose: every ceiling on it turns out to be a way of running two commits for one element. Threading the pass signal into the download was also a mistake, in the other direction. The provider call that produced the URL has already been billed, so cancelling the download throws away work that is paid for — and the shared proxy cache records a cancelled request as a transient failure against that URL, which after three of them blocks it for every consumer in the session. Browser-only mode never asked for this: it had no way to observe an abort there, which is exactly why the bytes were kept. The signal is gone from the download again, and `fetchAsBlob` is byte-for-byte what it was before this branch. The regression guard for the stranded-element rule is restored alongside the timing test that was meant to supersede it. It catches a different rule — a task left pending being read as answered — and nothing else does: making the pass skip pending leaves every other suite green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): stop asking the pool for refs it never issued, bound it, and adopt cached bytes Four things a deployment found once this was running for real. A reference this application mints itself — a generation placeholder, a derived narration key — was never in the pool, because the pool allocates every id it holds. Asking anyway used to be an IndexedDB miss; once the pool is server-backed it is a request that answers 404, one per element per load, forever on a course that still holds placeholders. Every lease and probe now checks first. The check is a negative test on shapes this application owns, not an id validator: the pool's id domain stays unconstrained, and anything that is not one of ours is still asked about. The asset store can bound how much one principal holds, and enforces it inside the write transaction, but nothing ever passed the number. It does now, with a default rather than an opt-in: allocation is reachable by any caller a deployment admits, and with one shared principal an unbounded store is unbounded database growth with no operator-visible brake. Refusing asset mutations to unauthenticated callers was not enough, because every authenticated caller resolves to that same shared principal — so authentication decided nothing, and any signed-in visitor could delete any id they learned. Since this branch began storing media the registry is the only copy a course has. Replacing and deleting are now refused to everyone, and the browser no longer tries: an entry nothing references waits for server-side reclamation instead. What a browser must still do is forget its own record of an allocation that reached nothing, or a later save would stamp an id the document has no reason to trust. And a course generated before any of this holds placeholders in its document with its bytes only in the author's browser. Those bytes are paid for, so the author's next load converts them — stored to the pool and written back through the ordinary commit path, with no provider call — instead of buying them again. A row that records only a hosted URL is treated as absent: that URL is the provider's address, not something a document may hold. One renderer expectation moved with this. An untracked placeholder used to paint as pending on first render because asking the pool left a lease in flight; it settled to disabled a moment later either way, and now says so from the start. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): surface a full store as a refusal and convert legacy narration A quota refusal reached the browser as HTTP 500 with a generic message, which reads as a transient failure: the element kept a Retry that would pay a provider again and be refused again. The store raises the contract's own error and the handler maps it to 507, but the store answering a request is not always built by the same bundle as the handler -- the persistence provider is reached from the route bundle and from instrumentation, which is why its state lives on a Symbol.for global -- and `instanceof` is false across that boundary while the declared code is still right. Classify on the code as well as the class, and make the code a permanent, persisted refusal in the browser: recorded locally so it survives a reload, shown as "storage is full", and refused by the retry entry point so a stale button cannot buy a second generation. Every other storage failure stays retryable. Convert what a pre-server-backed course still holds. Generated media is adopted under either key this application has used for it -- the placeholder, and the allocated id of a course converted once and later rolled back -- instead of only the first. Narration is converted by a load-time pass over the open course's speech actions, since nothing re-enters generation for an action that already has an id: bytes to the pool, id written back through a funnel that mirrors the media one, owner-only and server-backed-only. A line whose bytes are in no browser is left alone rather than re-synthesized. Also: the pool guard is now a positive `ast_` test rather than an enumeration of the shapes we mint (imports never reach the pool, so this is safe in both modes); the slide ref collection is an exported pure function so its four lease sites are covered behaviourally; ASSET_QUOTA_BYTES treats every spelling of zero as opting out and refuses a malformed value at startup instead of falling back; the abort signal is re-checked after the cache read, before an uncancellable commit; and the unused `removeAsset` and pool `replace` surfaces are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * chore(storage): release 0.29.1 The asset HTTP handler now recognises a store refusal by the contract code it declares as well as by its class, so a quota refusal raised in another module realm answers 507 instead of 500. Same contract, stricter recognition, no API change: a patch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a full store recoverable and adoption course-safe Narration adoption read the local audio row by its derived key alone. That key carries no stage id and the table is keyed by id alone, so two courses can mint the same one -- a PPTX import numbers its scenes and actions deterministically, which gives every imported deck's first slide `tts_s1_speech-scene-p1`. Locally a collision only means one course plays another's clip in one browser; adopting it wrote that clip into the shared document permanently, for every device and every visitor. A row that names a course is now adopted only into that course, and a row from before that column existed only when the text it recorded is the text of the action being converted. A full asset store was made permanent last round, which was wrong three times over: it overwrote the refused bytes with an empty blob -- on the conversion path that row is a course's only copy of its own media -- it kept sending the rest of the deck to a provider against a ceiling it already knew was reached, and it left no way back once an operator raised that ceiling. A full store is neither the content's fault nor the configuration's, so it is now its own case: the bytes are kept, the pass stops at the first refusal, and the element shows the reason together with a Retry that re-attempts the upload from those bytes. Nothing retries automatically, so no one is re-billed. The narration write-back now reaches the write boundary every producer of a durable write passes through, not only the dirty mark: adoption never deletes the derived row, so a snapshot that reverts the rewrite is adopted again on the next load and allocates a fresh asset every time. Adoption is also mounted by both classroom surfaces rather than one, takes the course's abort signal, and re-validates that this browser still has the course open before each write. ASSET_QUOTA_BYTES is validated from instrumentation, where the README and the docstring already claimed it was: its only other consumer is lazy and memoised, so a malformed ceiling let the process boot and then failed every persistence request, documents and runtime included. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): remember a full store per course, and never lose retained bytes A stopped pass left the elements it never reached as placeholders with no persisted record -- deliberately, since nothing was attempted for them. But that left the next load with no reason not to try: it called a provider for the next placeholder and was refused at exactly the same point, once per reload, indefinitely. A full store is not a property of any slide. It belongs to the deployment and changes for reasons the document knows nothing about, so it is now remembered once per course in the browser's device KV. A pass that finds the marker stands down before spending anything and leaves every placeholder its "storage is full" state and its Retry; the first upload that succeeds clears it and the next pass runs normally. Narration adoption latched per course so it runs once per load, and the latch outlived the abort that leaving a course performs. On a surface that stays mounted across switches -- the workbench pane is one component for every course it shows -- owner course A, visitor course B, then back to A skipped exactly the clips the abort had cut off, and nothing else converts them. The latch is released with the abort now, and a course adopts one run at a time so a re-entry cannot hand a clip a second allocation while the previous run's uncancellable tail is still settling. A quota-blocked element retried into a network error or a 500 lost the bytes that were kept for it: the retry deleted the row before attempting the upload and wrote no replacement for an error carrying no structured code, so the next retry went back to a provider for media this browser had a moment earlier. The row now survives until an upload succeeds, the failure handler keeps whatever bytes the attempt was given, and the retry asks the question a pass asks -- does this browser already hold bytes for this element -- rather than reading an error code that a second failure has already overwritten. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): adopt real legacy narration, queue re-entries, report attempt outcomes Narration adoption admitted a stage-less row only when the text it recorded matched the action being converted. Both of those columns were added to the local audio table by the very change that moved narration onto allocated ids, so a row still carrying a derived key has neither: the rule refused every real pre-allocation course and passed only on fixtures built from post-allocation rows. What the row cannot say, the key can. A derived key names two clips only when two courses share a scene order and an action id, and an action id repeats only when something other than the generator minted it -- an import numbers them by slide position. So a key built from a generated action id is adopted on that basis, a key an import could have reproduced still needs matching text, and a row that names another course is refused however unique its key looks. Handing a re-entering caller the adoption run already in flight undid the latch release it was paired with: that run is bound to the signal the departure just aborted, so it stops at its next clip while the caller -- which has the course open and a live signal -- is told the work is done, and an effect replayed as mount, cleanup, mount adopts nothing at all. A later caller now waits for the uncancellable tail and scans again, which costs a lookup on a course that has nothing left and finishes the clips the abort cut off on one that does. One attempt at an element now reports both facts its callers need instead of a bare boolean: whether the store refused it for room, and whether bytes actually reached the store. Leaving a course clears the task table, so a retry that landed afterwards read "no failed task" as success and deleted the row holding the only copy of the media. Nothing is inferred from that table any more. Reading the localStorage property can throw where storage is denied by policy, typeof included, so the availability check moved inside the guard: this metadata is best-effort, and a rejection here strands a generation pass that has already enqueued its tasks. A retry is never blocked by the per-course "store is full" marker, but a retry that is refused again re-sets it, and adoption now reads and writes the same marker rather than issuing one refused upload per clip on every load. The two canvas element renderers and both thumbnail renderers show the reason beside the Retry, so a full store does not look like an ordinary failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): probe a full store instead of standing down, and pair notices with a Retry Narration adoption was given both halves of the per-course "the store is full" marker last round: it stood down when the marker was set, and it set the marker when its own upload was refused for room. Those halves are only safe together if something can lift the marker, and for adoption nothing could. It has no affordance of its own, it stood down before reaching its own clear, the media pass returns before its marker gate when there is nothing to generate -- so a narration-only deck, or one whose slides are already satisfied, painted no storage-full element and offered no Retry -- and narration generated rather than adopted allocates directly rather than through the media commit. The course's cached narration was then lost for good, where before it converted on the first load after the ceiling was raised. The gate is a probe now. A marked course attempts exactly one clip per load: refused, it stops and the marker stands, which costs what standing down cost; stored, it lifts the marker and finishes the course. Adoption spends no provider money, so the whole cost of probing a store that is still full is one refused upload. Generated narration lifts the marker too. The three surfaces that gained a failure notice last round drew it for any failure with a reason, including the one refusal that is reachable without server-backed persistence, so a browser-only deck painted something it had not painted before. The notice is drawn beside a Retry and nowhere else, which is what it was added for and what leaves browser-only output unchanged. Both are now asserted through the render harness the surface matrix already had. A caller arriving while a rescan is queued shares it rather than appending another. One rescan converts whatever the run in flight left and every later one would find an allocated id on every action, so a chain bought nothing and turned a single stalled upload into a course that never adopts again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): treat a refusal for room as a fact about one clip, not the deck The asset store checks each write against the headroom it has left, so a store that refuses a long opening clip can still hold every short clip behind it. Narration adoption assumed the opposite: it broke the deck at the first refusal and then re-attempted that same first clip on every later load, because the document names it first. A deck whose longest clip exceeds current headroom therefore never converted the clips that would have fit, with no affordance to recover it -- the state the probe was introduced to remove, reached through a narrower door. An unmarked load now attempts every clip, skipping the ones that do not fit, and remembers the condition only if the load ends with clips it still could not store. A marked load spends its single upload on the smallest clip left rather than the first one named: that is the clip that answers the question the marker asks, because if the smallest does not fit nothing does. The media pass keeps stopping at its first refusal, and for a reason adoption does not share -- every element it attempts costs a provider call. A rescan several callers share took the newest caller's signal, and the newest caller is not necessarily the one still there: a surface that opened a course and closed it again would stop work a surface still showing that course was waiting for, and that surface is latched, so it would never ask again. The shared run now takes a signal that is aborted only once every caller has left. The comment claiming the shared rescan contains a stalled upload was wrong -- the rescan is chained off the run in flight, so a stalled upload leaves every caller pending exactly as a chain would. It claims the bounded queue it actually provides, and the stall is recorded as a limitation. The failed-state containers took their stacking classes unconditionally, so markup differed in browser-only mode even though nothing moved on screen. Those classes are applied only when there is a notice to stack, and the tests assert the exact class attribute rather than a substring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): stop narration adoption writing the media pass's store-full marker The marker means "do not call a provider for this course". A path is entitled to write it only if its own refusal cost a provider call, and narration adoption's refusals cost nothing: it uploads bytes this browser already holds. The store also checks each write against the headroom it has left, so a clip that does not fit says nothing about whether a slide's image would. Adoption was writing it anyway, and one over-long narration clip was therefore enough to stand a course's entire image pass down on every later load -- on a store that had just accepted adoption's other clips. The author could still recover each element by hand, every load, for ever. Three rounds of narrowing this seam produced a finding each time, so it is removed rather than narrowed again. Gone: the marker read, the single-clip probe, the smallest-clip selection, and the up-front read of every row into an array -- which also retires a sampled-then-stale flag and the retention of a whole deck's blobs for the length of a run, and returns the loop to streaming one row at a time. Adoption's rule is now that every load attempts every clip it holds, once; any failure skips that clip and the load continues. The noise the coupling was meant to avoid does not arise, because after the first load the clips still outstanding are exactly the ones that did not fit -- normally none, or one. A successful write still clears the marker, and that is a different kind of statement: a write that went through is a fact this run established, where a refusal is an inference about what some other write would cost. For a course whose media needs nothing, adoption and generated narration are also the only paths that can establish it. The failure module still documented the deck-wide premise this contradicts. It now says what is true: the check is per write, and the media pass stops the deck as a judgement about cost rather than about certainty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bound a full store's cost from the store's own arithmetic Removing the store-full marker from narration adoption removed its bound too, and the code then asserted the bound was unnecessary. It is, on a store with room for most of a deck. On the store the whole mechanism exists for -- the ceiling reached, nothing fitting -- the outstanding set after every load is the entire deck, so a thirty-clip course posted thirty full blobs on every load, indefinitely. Each of those is not a cheap refusal: the bytes are uploaded, the server hashes the whole payload, and only then takes a per-principal lock and sums every entry that principal owns before saying no. The bound needs no flag, no key and nothing carried between loads. The store asks whether `used + addedBytes` exceeds the ceiling, and `used` only grows while a run is uploading, so a clip refused for want of room implies every clip at least that large is refused for the rest of that run. The run keeps the smallest size it has been refused and skips anything no smaller without uploading it; a smaller clip is still attempted, because it may fit. A deck the store refuses entirely now costs one upload per successive size minimum instead of one per clip, and a deck it has room for costs nothing extra, because nothing is refused. Only a refusal for room lowers the bar: a dropped connection says nothing about how much room there is. The deck-wide certainty premise the failure module retracted last round still stood verbatim at the site that implements the stand-down. Both copies now say the same thing: the check is per write, and the pass stops the deck as a judgement about cost rather than about certainty. The comment on adoption's marker clear now names its price. Narration of a few hundred bytes fits in headroom an image does not, so a proven write can let the next pass buy one more image that is refused again -- bounded at one, and the price of the alternative being a course whose media never generates again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(media): state the adoption bound exactly, and stop three comments describing the old rule The comment introducing the in-load bound gave its cost as "at most a handful, and the first load pays the most". Neither clause is a property of the rule. A clip is skipped only when something no larger was already refused, so a fully-refused deck costs one upload per successive size minimum in document order: one when the clips grow, about ln N for an arbitrary order, and one per clip when they only shrink -- a long opener followed by terser lines is exactly that shape. And no load is cheaper than the first, because the bound resets per run and a refused clip stays outstanding. The comment now says that, and points at what would make it exactly one for any ordering: the store returning its remaining headroom in the refusal's existing details channel, which the server leaves empty today. Two other comments still described the previous rule -- "attempts every clip it holds, every load" -- one of them twenty lines above the paragraph that introduces the bound, in the same block. Both now say what the code does. The bound's soundness is worth stating where a maintainer will look for it: quota is charged at full length with no discount for a duplicate, the sum it is checked against joins entries to blobs so the collector cannot lower it, the check takes a per-principal lock before summing, and replace and delete are refused to every browser. Nothing a run can do makes room appear inside it. One test installed a row implementation and replaced it wholesale a few lines later, so the first was dead and the survivor dropped the text the first clip's import-shaped key needs for the ownership rule -- it passed on the coincidence that the fixture's default text is the action's. Merged into one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): keep a Retry from re-buying parked media, and state the store seam once Five findings from an inline review. The standalone classroom route asked the ownership sidecar once per load and recorded only stage ownership when that ask failed. Every non-answer fails closed, so one transient 5xx left the genuine author with no resume, no Retry affordance and no legacy narration converted for the rest of the load, with nothing to change it short of a reload. The failure now records the fail-closed answer explicitly -- an answer an earlier load established must not outlive the failure that replaced it -- and an unresolved answer is asked for again, a few times over a few seconds. A real answer, however unwelcome, is final. A Retry could pay a provider for media the pool already held. When the bytes are stored and only the write-back fails in a way that keeps the allocation, it is parked and no local row exists, because that row is written only after a successful write-back. Retry now reads the parked queue exactly as the pass does and re-attempts the write-back: it re-keys the task done when the document takes it, leaves the entry parked when the slide still does not exist, and stays failed and retryable when the document refuses again. Object URLs a parked allocation owns are revoked when the entry is dropped. The commit path leaves them alone while the entry is parked, because it is then the only thing holding bytes this tab can render, so a course switch or a stage deletion was pinning the whole blob for the life of the tab. An entry a slide has already taken is left alone: the task table is displaying those URLs. The fallback lookup for cached bytes is a stage-scoped scan, and the keyed lookup misses for every row the commit path writes, so a pass was materializing and sorting the course's whole media table once per element. One scan per pass now, built on the first miss. It is sound and not merely cheaper: an element asks only for its own placeholder, and every row a pass writes carries the placeholder of the element that wrote it. "The store accepted a write, so it is not out of room" was enforced at three call sites under slightly different conditions, which made it a convention the next pool write path could silently break. It is stated once, in putAsset, for the course whose bytes it just stored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 16 天前 | |
feat(media): write generated media through the asset pool under server-backed persistence (#1392) * feat(media): store generated media in the asset pool when persistence is server-backed With server-backed persistence the document is durable and shared, but generated media stayed in the producing browser: the document kept its gen_img_* / gen_vid_* placeholder and narration kept a browser-derived audio id. Every new browser that opened such a course re-ran generation for every slide, and it never converged, because the address of the generated bytes was never written back into the document. Under server-backed persistence only, the classic generation chain now stores bytes in the asset pool first and writes the id the pool allocated into the document. - The client bootstrap configures the asset seam alongside the document and runtime seams: an HttpAssetStore over the persistence endpoint carrying the same credentials the document store carries, marked server-backed. The seam preflight now covers all three, so a failure still cannot half-configure persistence. - Image, video and TTS generation commit in one fixed order: provider, pool, document, local cache, task. A reference reaches the document only after put returned an id, so a document can never name bytes that were not stored. A failure before the write-back leaves the placeholder with the provider called exactly once; the retry happens on the next owner load. - The write-back is a per-slot rewrite through mutateDocument, which re-reads the current document under the per-stage lock, so it cannot clobber a newer scene. The open course is refreshed with the same rewrite without being marked dirty. - "Has this already been generated?" is answered by the document (the slide exists and no longer holds the placeholder) instead of by this browser's task table. - The classroom's resume effect fails closed on ownership: only a resolved owner starts generation, so a viewer opening a shared course spends nothing. - The local media and audio tables become a per-tab cache. A failed cache write costs a re-download, never the media. Browser-only mode is unchanged: every new call site sits behind the server-backed gate, the local tables stay authoritative there, and placeholders stay in the document. Rendering and export needed no changes. HttpAssetStore.resolve mints an object URL exactly as the browser store does, and the export byte resolver was already pool-first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(classroom): make the generation owner gate a three-outcome rule and apply it everywhere The gate refused everything but a resolved owner, which read a sidecar that answered "no ownership fact exists for this course" as a reason to block. That is the answer a deployment without the sidecar's server-side prerequisites gives for every course, and the answer a course with no ownership record gives: in both, there is nobody the operator's budget needs protecting from, and refusing strands the course's own author behind a question that can never be answered. Ownership is now four states over the sidecar's three outcomes. A definite answer splits into owner and not-owner. An absent record is its own answer, ownerless, and generation proceeds — the behaviour such a deployment had before the gate existed. Only the absence of an answer, a transport failure or a load that has not asked yet, stays unresolved and fails closed: "we could not ask" must never be read as "nobody owns this". One mapper turns a sidecar result into that state, and one predicate decides on it. The workbench classroom pane runs the same resume effect and had no ownership input at all, so a viewer opening a shared course there could still spend the budget. It now asks the sidecar once per course, in parallel with its load and feeding only the generation gate, so its read-only and edit behaviour is unchanged. The shared progressive-load policy carries the gate for it, with both new inputs required rather than defaulted so a future caller cannot omit them into an open budget. Its stale comment claiming ownership could not be expressed here is corrected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the write-back survive autosave, arrive before the scene does, and never leak Independent reviews of the write-back found three ways a durable document could still end up naming a placeholder, and two ways the gate that protects the operator's budget could be walked around. An autosave round captures the store synchronously and writes that capture, so a round already in flight when a rewrite landed wrote the placeholder straight back over the allocated id, and nothing marked the store dirty again to correct it. The rewrite now marks the units it changed, which leaves a corrective flush queued behind the stale one; re-saving a scene that already holds the id is idempotent, losing the id is not. Media is generated from outlines in parallel with scene content and usually finishes first, so the slide that will carry the placeholder does not exist yet and the write-back has nothing to rewrite. That was the ordinary path, not a tail case, and its result was discarded: the task was marked done, the scene was added afterwards with its placeholder intact, and a second pass in the same run could call the provider again. The allocation is now held under the placeholder — which also answers the skip test, so nothing pays twice — and applied when that scene is committed, before its first save. One complete pass now leaves no placeholder behind. A failed commit used to abandon what it had already allocated. A poster upload that failed threw away a stored video and sent the retry to submit the most expensive job in the system again; a rejected write-back left registry rows that name bytes nothing references, which the byte collector cannot reclaim because it only collects blobs no row names. A poster failure now costs the poster, and a write-back that reached nothing reclaims what it allocated. A partial write is left alone, because the document already names it. The ownership gate is fail-closed again. Treating the sidecar's 404 as permission was wrong: the client cannot tell "this course has no owner" from "this deployment told me nothing", so a visitor who opened a shared course could bill the operator. The root cause was the sidecar itself, which gated on the agent runtime although every persisted course has an owner regardless — the persistence route resolves one for every request. It now gates on server persistence, so the configuration that made 404 the universal answer has real ownership facts to report, and the gate can refuse everything but a named owner. Retry affordances answered to no gate at all. A viewer of a shared course with one failed image was shown a Retry button that called the provider. Both retry entry points and every surface that draws them now read one shared permission, so what is offered and what is allowed are the same value. Also: narration regeneration no longer pretends it can replace bytes behind a live id — the exclusivity proof that would allow it is refused by construction once references leave the browser, so it forks to a fresh id and says so; the "already generated" test lets a finished deck answer from the document alone, since scene order stops identifying an outline once slides are inserted or deleted; stored assets record a specific media type rather than a generic transfer type; the pane no longer asks the sidecar in browser-only mode; and the funnel's docstring now states what the per-stage lock actually guarantees, which is same-browser serialization and not a cross-browser compare-and-swap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): park allocations in the deciding turn, never reclaim on an ambiguous write A delta review of the write-back found the first-pass fix still had a window, and the reclamation it added could delete media the document already names. The allocation was parked after an awaited local cache write. A scene committed in that window reconciled against a registry that did not hold it yet, so the document kept the placeholder — and the entry recorded a moment later then answered the skip test as "already handled", so nothing could correct it. Parking now happens inside the write-back, in the same synchronous turn as the decision that nothing could take the reference; no await separates the live check from the park. Allocations parked by an earlier pass are handed to their slides at the start of the next one, before anything decides what still needs generating, so a held allocation whose scene has since arrived becomes a rewrite rather than an answer. Reclaiming on a rejected write was unsound: a rejection does not prove the server did not apply the write, so deleting the asset could break the scene that now names it. The funnel decides instead, and says so: it reclaims only when no store write was ever issued and nothing took the reference. Anything else is placed if its slide exists and parked if it does not, so the next pass reuses the bytes instead of paying for them again. When a write fails after part of it landed, the live store is brought up to the document before the error is rethrown — otherwise the next ordinary flush would overwrite the half that did land, with the ids deliberately not reclaimed. Parked allocations are now cleared with the course. Classic placeholders are reused across runs, so one surviving an interrupted run would be handed to a different slide of the next deck: the previous picture, on a slide whose provider was never asked. Both classroom surfaces clear the arriving course, the deletion cascade clears the deleted one, and clearing the database clears them all. Two more ways generation could start without asking the gate are closed. An overlapping pass — an outline retry re-enters generation with every outline while the first is still working — re-requested elements whose provider call was already in flight; a task that is not done is an answered request, not an unanswered one. And narration regeneration in the timeline editor called the TTS provider and allocated a pool asset with no ownership check at all; it now reads the same permission, which withholds both the per-line and whole-timeline controls and refuses the call. Finally, a pane opened during the stage-link availability gap recorded the sidecar's 404 for a course that was moments from existing and never asked again, leaving the real owner locked out of generation until it remounted. Ownership is re-fetched once the document becomes available; the gate stays closed until an answer arrives, so asking again can only open it for someone entitled to it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the asset routes reachable, and stale snapshots harmless A full-branch audit found that the deployment this project documents could not store a single generated asset, and that several routes into durable storage could still write a placeholder over a reference that had already landed. The persistence route sent asset requests through the development authenticator, which refuses outright in a production build that has not explicitly opted into it — and the documented server-persistence recipe produces exactly that build. Every store and every read answered 401, so images and video failed on every slide while re-billing the provider on each retry, and a narration failure stopped the deck at its first slide. Assets live in one shared partition by design, so there was never anything per-caller for that authenticator to decide: the route now resolves the asset principal itself, alongside the owner it already resolves for documents. Runtime sessions are genuinely per-learner and keep the development authenticator until real session verification replaces it. And narration that cannot be stored no longer fails its scene: the line stays unvoiced and retryable, which is what an image that cannot be stored does to its slide. Placeholders could also come back from behind. A queued autosave's snapshot, an editor-history entry replayed by an undo, the departing save a course switch flushes — each captures content at its own moment, and any of those moments can predate a write-back. Point fixes at each producer would leave the next producer to rediscover the bug, so the check lives at the write boundary every producer passes through, and the allocation record it consults now outlives the parked queue: a placeholder whose rewrite landed long ago is exactly the case it catches. Two ways generation could be lost or repeated are closed. A pass now claims the elements it will reach and releases them however it ends, so an overlapping pass stands down while an aborted one strands nothing — previously its tasks stayed `pending` and every later pass skipped them with no retry control to recover them. And the media abort controller is aborted before being replaced, so a superseded pass stops calling providers instead of running on for a course the user has left. The remaining two are narrower. The workbench pane asks for ownership only after a document load succeeds, and after every later one, mirroring the page route: the load is what creates the ownership row the first time a course is opened, so asking beforehand asked about a course that did not exist yet and locked its author out for the mount. And the ownership gate on the timeline editor now withholds narration regeneration alone; listening back to existing narration and seeing whether a line has any spend nothing and stay available. Known limitation, unchanged and now stated plainly in the comments that used to point at it as a solution: nothing reclaims an unreferenced pool asset. The registry sweep is written but not wired up, and the byte collector only reclaims blobs no registry row names, so every narration regeneration and every abandoned allocation leaves storage behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): gate asset mutations, and make claims and allocation records survive a handoff Opening the asset routes opened all of them. Reads and allocations are meant to be as open as document reads and creates already are, but no authorization hook was supplied, so the handler's default admitted PUT and DELETE too — and those scope by principal key alone, which is one shared constant. Any caller who learned an id, and a document read hands out every id its slides name, could overwrite or destroy another author's media. Mutations now require the deployment's credential, which in a production build without the development-auth opt-in means they are refused outright; reads and allocations stay open. The route comment says what the posture is and what it is not: the deployment-level fence is the access code, and no per-principal quota is configured. The client's own reclaim is best effort to match — losing an argument about deleting an asset must not cost a task its retry, and the bytes are left for server-side reclamation. The pass claim could not survive the handoff it was written for. A retry aborts the live media pass and starts its replacement in the same synchronous block, long before the aborted pass's cleanup runs, so the replacement saw every element still claimed, collected nothing, and returned — leaving each unreached element at pending with nobody coming back for it and no retry control to recover it, which is the exact failure the claim was introduced to prevent. A claim now carries its pass's signal and is retired the moment that signal aborts, and a pass releases only claims it still owns, so a late unwind cannot take its replacement's work. Claims are also acquired at the single point every request passes through, so a single-task retry participates too — previously a retry awaiting its provider was invisible to a pass starting alongside it and both called it. The allocation record could outlive the bytes it named. It was written before the write-back attempted anything and survived the reclaim that followed a failure, so when the slide finally arrived the write boundary stamped a deleted id into the document — and the placeholder it replaced was gone, which reads as already generated and stops anything from retrying. The record is now written only where the allocation is retained, and forgotten wherever a reclaim removes the bytes, including the narration rollback path. The tests follow. The route test drives the real storage handler against an in-memory registry instead of a stub, so it can see what the resolved principal is then allowed to do; the handoff test performs a real abort mid-pass rather than starting from an already-aborted signal; and the guards that could only assert file layout now assert the property they care about, or have been replaced by behaviour. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make media passes serial per course instead of tracking element ownership Three rounds of per-element claims each produced a new way to lose an element. Whole-pass reservations swallowed a Retry for an element the same pass had already failed, leaving it pending with the affordance gone. Retiring a claim by its signal freed an element whose commit was still uploading, so the replacement pass paid for it twice. A claim held for a failed element stranded its retry. The bookkeeping is the defect: every refinement of "who owns this element right now" answered the question at a moment when the answer was already stale. Passes for one course are now serial. A replacement aborts its predecessor, as before, and then waits for it to settle before collecting. That removes the question entirely: a commit already under way finishes — its bytes stored and its reference written, so the new pass sees a resolved slide and skips it — and an element the aborted pass never reached is still a placeholder and gets collected like any other. The claim set, the reservations, the signal retirement and the identity-checked release are all gone. The task table is consulted for one thing only: an element that is generating right now is a single-element retry running alongside the pass, and taking it too would pay twice. Pending is deliberately not a skip reason — it means a pass once intended to reach an element, which an abandoned pass leaves behind with nobody acting on it, and reading that as answered is what stranded elements before. A retry runs concurrently with a pass, because a pass never revisits an element it has processed, and it re-reads the task after its own await and refuses before touching it: marking first and refusing afterwards destroyed the failed state that draws the affordance. Browser-only mode is back to exactly what it was. The abort is now conditional, the waiting does not apply, and the original status-based skip is restored verbatim. Two baseline lines remain changed in each of the two files, and both are behind a server-backed fork whose else-branch is the original. Two smaller things. The allocation record becomes visible when a write goes on the wire rather than when the round trip ends, and the write boundary reconciles under the document lock rather than before it — a save queued during a write-back was otherwise captured with the placeholder and, for a course the user had left, had no corrective flush to follow. And the comments that said a refused reclaim leaves its bytes for server-side reclamation were wrong: nothing collects them, because the registry entry still names its blob and the sweep that would remove it is not wired up. They now say the bytes leak. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a deferred pass re-earn its right to run, and bound the commit it waits on Serializing passes moved their body out of the block that launched them, and three things followed from that. A pass now wakes when its predecessor settles, which can be after the user has left the course. It enqueued before it looked at its signal, into a task table keyed by element id alone — and placeholder ids are not unique across courses, which is why the classroom clears that table on arrival. So a departing course's pass seeded the arriving course's table with tasks carrying the wrong stage id, and a Retry routes by that id: the reference went into the wrong document. A pass now re-validates after the wait, before touching anything shared. The same lateness broke the skip test. `documentSkipIndex` answers only while the live store is on the pass's stage, and returning nothing put the collection loop on the browser-only rule — a silent demotion from "the document is the authority" to "this browser's task table is", on exactly the path where that table has just been cleared. Every element the predecessor had committed was collected again, paid for again, and its second write-back found no placeholder to rewrite, so its bytes were parked where nothing will ever reference them. In server-backed mode an unreadable document now means the pass stands down. And waiting was unbounded. A commit is uncancellable: the asset client takes no signal, and a document write cannot be half-undone. One stalled upload therefore froze the course's media generation for the session — the replacement never collected, the element sat on a skeleton that draws no Retry, and only a reload recovered. The pass's signal is now threaded into the media proxy fetch, and the commit is bounded by a deadline. The deadline is on the wait, not the work: the commit carries on, and if it lands late the document simply ends up correct, while the element becomes retryable and the queue moves on. The tests that were meant to pin the previous round were not sensitive to it. Two asserted end states where the mechanism only changes ordering, and one of them rigged the document read so the assertion held whether or not the pass had waited; a third covered half of what it claimed. They now observe the ordering directly — nothing is issued while another pass for the course is working; in browser-only mode a second pass reaches its provider immediately — and the reconciliation under the document lock has a test that fails when it moves back outside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * revert(media): drop the commit deadline and the abortable download The deadline bought less than it cost. Abandoning a commit after two minutes makes the element retryable while the real commit is still running, so a Retry starts a second commit for the same placeholder against the first: two provider calls, two allocations, and whichever lands second stamps its result over the other's task by element id. The allocation record is keyed by placeholder, so the loser's cleanup erases the winner's record, and the write boundary then puts the raw placeholder back into the document. That is the overlap serial passes were built to remove, reopened through the one door serialization never covered. So a stalled commit holds the course's media queue until it settles or the page is reloaded, and that is written down rather than papered over. The wait is unbounded on purpose: every ceiling on it turns out to be a way of running two commits for one element. Threading the pass signal into the download was also a mistake, in the other direction. The provider call that produced the URL has already been billed, so cancelling the download throws away work that is paid for — and the shared proxy cache records a cancelled request as a transient failure against that URL, which after three of them blocks it for every consumer in the session. Browser-only mode never asked for this: it had no way to observe an abort there, which is exactly why the bytes were kept. The signal is gone from the download again, and `fetchAsBlob` is byte-for-byte what it was before this branch. The regression guard for the stranded-element rule is restored alongside the timing test that was meant to supersede it. It catches a different rule — a task left pending being read as answered — and nothing else does: making the pass skip pending leaves every other suite green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): stop asking the pool for refs it never issued, bound it, and adopt cached bytes Four things a deployment found once this was running for real. A reference this application mints itself — a generation placeholder, a derived narration key — was never in the pool, because the pool allocates every id it holds. Asking anyway used to be an IndexedDB miss; once the pool is server-backed it is a request that answers 404, one per element per load, forever on a course that still holds placeholders. Every lease and probe now checks first. The check is a negative test on shapes this application owns, not an id validator: the pool's id domain stays unconstrained, and anything that is not one of ours is still asked about. The asset store can bound how much one principal holds, and enforces it inside the write transaction, but nothing ever passed the number. It does now, with a default rather than an opt-in: allocation is reachable by any caller a deployment admits, and with one shared principal an unbounded store is unbounded database growth with no operator-visible brake. Refusing asset mutations to unauthenticated callers was not enough, because every authenticated caller resolves to that same shared principal — so authentication decided nothing, and any signed-in visitor could delete any id they learned. Since this branch began storing media the registry is the only copy a course has. Replacing and deleting are now refused to everyone, and the browser no longer tries: an entry nothing references waits for server-side reclamation instead. What a browser must still do is forget its own record of an allocation that reached nothing, or a later save would stamp an id the document has no reason to trust. And a course generated before any of this holds placeholders in its document with its bytes only in the author's browser. Those bytes are paid for, so the author's next load converts them — stored to the pool and written back through the ordinary commit path, with no provider call — instead of buying them again. A row that records only a hosted URL is treated as absent: that URL is the provider's address, not something a document may hold. One renderer expectation moved with this. An untracked placeholder used to paint as pending on first render because asking the pool left a lease in flight; it settled to disabled a moment later either way, and now says so from the start. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): surface a full store as a refusal and convert legacy narration A quota refusal reached the browser as HTTP 500 with a generic message, which reads as a transient failure: the element kept a Retry that would pay a provider again and be refused again. The store raises the contract's own error and the handler maps it to 507, but the store answering a request is not always built by the same bundle as the handler -- the persistence provider is reached from the route bundle and from instrumentation, which is why its state lives on a Symbol.for global -- and `instanceof` is false across that boundary while the declared code is still right. Classify on the code as well as the class, and make the code a permanent, persisted refusal in the browser: recorded locally so it survives a reload, shown as "storage is full", and refused by the retry entry point so a stale button cannot buy a second generation. Every other storage failure stays retryable. Convert what a pre-server-backed course still holds. Generated media is adopted under either key this application has used for it -- the placeholder, and the allocated id of a course converted once and later rolled back -- instead of only the first. Narration is converted by a load-time pass over the open course's speech actions, since nothing re-enters generation for an action that already has an id: bytes to the pool, id written back through a funnel that mirrors the media one, owner-only and server-backed-only. A line whose bytes are in no browser is left alone rather than re-synthesized. Also: the pool guard is now a positive `ast_` test rather than an enumeration of the shapes we mint (imports never reach the pool, so this is safe in both modes); the slide ref collection is an exported pure function so its four lease sites are covered behaviourally; ASSET_QUOTA_BYTES treats every spelling of zero as opting out and refuses a malformed value at startup instead of falling back; the abort signal is re-checked after the cache read, before an uncancellable commit; and the unused `removeAsset` and pool `replace` surfaces are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * chore(storage): release 0.29.1 The asset HTTP handler now recognises a store refusal by the contract code it declares as well as by its class, so a quota refusal raised in another module realm answers 507 instead of 500. Same contract, stricter recognition, no API change: a patch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a full store recoverable and adoption course-safe Narration adoption read the local audio row by its derived key alone. That key carries no stage id and the table is keyed by id alone, so two courses can mint the same one -- a PPTX import numbers its scenes and actions deterministically, which gives every imported deck's first slide `tts_s1_speech-scene-p1`. Locally a collision only means one course plays another's clip in one browser; adopting it wrote that clip into the shared document permanently, for every device and every visitor. A row that names a course is now adopted only into that course, and a row from before that column existed only when the text it recorded is the text of the action being converted. A full asset store was made permanent last round, which was wrong three times over: it overwrote the refused bytes with an empty blob -- on the conversion path that row is a course's only copy of its own media -- it kept sending the rest of the deck to a provider against a ceiling it already knew was reached, and it left no way back once an operator raised that ceiling. A full store is neither the content's fault nor the configuration's, so it is now its own case: the bytes are kept, the pass stops at the first refusal, and the element shows the reason together with a Retry that re-attempts the upload from those bytes. Nothing retries automatically, so no one is re-billed. The narration write-back now reaches the write boundary every producer of a durable write passes through, not only the dirty mark: adoption never deletes the derived row, so a snapshot that reverts the rewrite is adopted again on the next load and allocates a fresh asset every time. Adoption is also mounted by both classroom surfaces rather than one, takes the course's abort signal, and re-validates that this browser still has the course open before each write. ASSET_QUOTA_BYTES is validated from instrumentation, where the README and the docstring already claimed it was: its only other consumer is lazy and memoised, so a malformed ceiling let the process boot and then failed every persistence request, documents and runtime included. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): remember a full store per course, and never lose retained bytes A stopped pass left the elements it never reached as placeholders with no persisted record -- deliberately, since nothing was attempted for them. But that left the next load with no reason not to try: it called a provider for the next placeholder and was refused at exactly the same point, once per reload, indefinitely. A full store is not a property of any slide. It belongs to the deployment and changes for reasons the document knows nothing about, so it is now remembered once per course in the browser's device KV. A pass that finds the marker stands down before spending anything and leaves every placeholder its "storage is full" state and its Retry; the first upload that succeeds clears it and the next pass runs normally. Narration adoption latched per course so it runs once per load, and the latch outlived the abort that leaving a course performs. On a surface that stays mounted across switches -- the workbench pane is one component for every course it shows -- owner course A, visitor course B, then back to A skipped exactly the clips the abort had cut off, and nothing else converts them. The latch is released with the abort now, and a course adopts one run at a time so a re-entry cannot hand a clip a second allocation while the previous run's uncancellable tail is still settling. A quota-blocked element retried into a network error or a 500 lost the bytes that were kept for it: the retry deleted the row before attempting the upload and wrote no replacement for an error carrying no structured code, so the next retry went back to a provider for media this browser had a moment earlier. The row now survives until an upload succeeds, the failure handler keeps whatever bytes the attempt was given, and the retry asks the question a pass asks -- does this browser already hold bytes for this element -- rather than reading an error code that a second failure has already overwritten. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): adopt real legacy narration, queue re-entries, report attempt outcomes Narration adoption admitted a stage-less row only when the text it recorded matched the action being converted. Both of those columns were added to the local audio table by the very change that moved narration onto allocated ids, so a row still carrying a derived key has neither: the rule refused every real pre-allocation course and passed only on fixtures built from post-allocation rows. What the row cannot say, the key can. A derived key names two clips only when two courses share a scene order and an action id, and an action id repeats only when something other than the generator minted it -- an import numbers them by slide position. So a key built from a generated action id is adopted on that basis, a key an import could have reproduced still needs matching text, and a row that names another course is refused however unique its key looks. Handing a re-entering caller the adoption run already in flight undid the latch release it was paired with: that run is bound to the signal the departure just aborted, so it stops at its next clip while the caller -- which has the course open and a live signal -- is told the work is done, and an effect replayed as mount, cleanup, mount adopts nothing at all. A later caller now waits for the uncancellable tail and scans again, which costs a lookup on a course that has nothing left and finishes the clips the abort cut off on one that does. One attempt at an element now reports both facts its callers need instead of a bare boolean: whether the store refused it for room, and whether bytes actually reached the store. Leaving a course clears the task table, so a retry that landed afterwards read "no failed task" as success and deleted the row holding the only copy of the media. Nothing is inferred from that table any more. Reading the localStorage property can throw where storage is denied by policy, typeof included, so the availability check moved inside the guard: this metadata is best-effort, and a rejection here strands a generation pass that has already enqueued its tasks. A retry is never blocked by the per-course "store is full" marker, but a retry that is refused again re-sets it, and adoption now reads and writes the same marker rather than issuing one refused upload per clip on every load. The two canvas element renderers and both thumbnail renderers show the reason beside the Retry, so a full store does not look like an ordinary failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): probe a full store instead of standing down, and pair notices with a Retry Narration adoption was given both halves of the per-course "the store is full" marker last round: it stood down when the marker was set, and it set the marker when its own upload was refused for room. Those halves are only safe together if something can lift the marker, and for adoption nothing could. It has no affordance of its own, it stood down before reaching its own clear, the media pass returns before its marker gate when there is nothing to generate -- so a narration-only deck, or one whose slides are already satisfied, painted no storage-full element and offered no Retry -- and narration generated rather than adopted allocates directly rather than through the media commit. The course's cached narration was then lost for good, where before it converted on the first load after the ceiling was raised. The gate is a probe now. A marked course attempts exactly one clip per load: refused, it stops and the marker stands, which costs what standing down cost; stored, it lifts the marker and finishes the course. Adoption spends no provider money, so the whole cost of probing a store that is still full is one refused upload. Generated narration lifts the marker too. The three surfaces that gained a failure notice last round drew it for any failure with a reason, including the one refusal that is reachable without server-backed persistence, so a browser-only deck painted something it had not painted before. The notice is drawn beside a Retry and nowhere else, which is what it was added for and what leaves browser-only output unchanged. Both are now asserted through the render harness the surface matrix already had. A caller arriving while a rescan is queued shares it rather than appending another. One rescan converts whatever the run in flight left and every later one would find an allocated id on every action, so a chain bought nothing and turned a single stalled upload into a course that never adopts again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): treat a refusal for room as a fact about one clip, not the deck The asset store checks each write against the headroom it has left, so a store that refuses a long opening clip can still hold every short clip behind it. Narration adoption assumed the opposite: it broke the deck at the first refusal and then re-attempted that same first clip on every later load, because the document names it first. A deck whose longest clip exceeds current headroom therefore never converted the clips that would have fit, with no affordance to recover it -- the state the probe was introduced to remove, reached through a narrower door. An unmarked load now attempts every clip, skipping the ones that do not fit, and remembers the condition only if the load ends with clips it still could not store. A marked load spends its single upload on the smallest clip left rather than the first one named: that is the clip that answers the question the marker asks, because if the smallest does not fit nothing does. The media pass keeps stopping at its first refusal, and for a reason adoption does not share -- every element it attempts costs a provider call. A rescan several callers share took the newest caller's signal, and the newest caller is not necessarily the one still there: a surface that opened a course and closed it again would stop work a surface still showing that course was waiting for, and that surface is latched, so it would never ask again. The shared run now takes a signal that is aborted only once every caller has left. The comment claiming the shared rescan contains a stalled upload was wrong -- the rescan is chained off the run in flight, so a stalled upload leaves every caller pending exactly as a chain would. It claims the bounded queue it actually provides, and the stall is recorded as a limitation. The failed-state containers took their stacking classes unconditionally, so markup differed in browser-only mode even though nothing moved on screen. Those classes are applied only when there is a notice to stack, and the tests assert the exact class attribute rather than a substring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): stop narration adoption writing the media pass's store-full marker The marker means "do not call a provider for this course". A path is entitled to write it only if its own refusal cost a provider call, and narration adoption's refusals cost nothing: it uploads bytes this browser already holds. The store also checks each write against the headroom it has left, so a clip that does not fit says nothing about whether a slide's image would. Adoption was writing it anyway, and one over-long narration clip was therefore enough to stand a course's entire image pass down on every later load -- on a store that had just accepted adoption's other clips. The author could still recover each element by hand, every load, for ever. Three rounds of narrowing this seam produced a finding each time, so it is removed rather than narrowed again. Gone: the marker read, the single-clip probe, the smallest-clip selection, and the up-front read of every row into an array -- which also retires a sampled-then-stale flag and the retention of a whole deck's blobs for the length of a run, and returns the loop to streaming one row at a time. Adoption's rule is now that every load attempts every clip it holds, once; any failure skips that clip and the load continues. The noise the coupling was meant to avoid does not arise, because after the first load the clips still outstanding are exactly the ones that did not fit -- normally none, or one. A successful write still clears the marker, and that is a different kind of statement: a write that went through is a fact this run established, where a refusal is an inference about what some other write would cost. For a course whose media needs nothing, adoption and generated narration are also the only paths that can establish it. The failure module still documented the deck-wide premise this contradicts. It now says what is true: the check is per write, and the media pass stops the deck as a judgement about cost rather than about certainty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bound a full store's cost from the store's own arithmetic Removing the store-full marker from narration adoption removed its bound too, and the code then asserted the bound was unnecessary. It is, on a store with room for most of a deck. On the store the whole mechanism exists for -- the ceiling reached, nothing fitting -- the outstanding set after every load is the entire deck, so a thirty-clip course posted thirty full blobs on every load, indefinitely. Each of those is not a cheap refusal: the bytes are uploaded, the server hashes the whole payload, and only then takes a per-principal lock and sums every entry that principal owns before saying no. The bound needs no flag, no key and nothing carried between loads. The store asks whether `used + addedBytes` exceeds the ceiling, and `used` only grows while a run is uploading, so a clip refused for want of room implies every clip at least that large is refused for the rest of that run. The run keeps the smallest size it has been refused and skips anything no smaller without uploading it; a smaller clip is still attempted, because it may fit. A deck the store refuses entirely now costs one upload per successive size minimum instead of one per clip, and a deck it has room for costs nothing extra, because nothing is refused. Only a refusal for room lowers the bar: a dropped connection says nothing about how much room there is. The deck-wide certainty premise the failure module retracted last round still stood verbatim at the site that implements the stand-down. Both copies now say the same thing: the check is per write, and the pass stops the deck as a judgement about cost rather than about certainty. The comment on adoption's marker clear now names its price. Narration of a few hundred bytes fits in headroom an image does not, so a proven write can let the next pass buy one more image that is refused again -- bounded at one, and the price of the alternative being a course whose media never generates again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(media): state the adoption bound exactly, and stop three comments describing the old rule The comment introducing the in-load bound gave its cost as "at most a handful, and the first load pays the most". Neither clause is a property of the rule. A clip is skipped only when something no larger was already refused, so a fully-refused deck costs one upload per successive size minimum in document order: one when the clips grow, about ln N for an arbitrary order, and one per clip when they only shrink -- a long opener followed by terser lines is exactly that shape. And no load is cheaper than the first, because the bound resets per run and a refused clip stays outstanding. The comment now says that, and points at what would make it exactly one for any ordering: the store returning its remaining headroom in the refusal's existing details channel, which the server leaves empty today. Two other comments still described the previous rule -- "attempts every clip it holds, every load" -- one of them twenty lines above the paragraph that introduces the bound, in the same block. Both now say what the code does. The bound's soundness is worth stating where a maintainer will look for it: quota is charged at full length with no discount for a duplicate, the sum it is checked against joins entries to blobs so the collector cannot lower it, the check takes a per-principal lock before summing, and replace and delete are refused to every browser. Nothing a run can do makes room appear inside it. One test installed a row implementation and replaced it wholesale a few lines later, so the first was dead and the survivor dropped the text the first clip's import-shaped key needs for the ownership rule -- it passed on the coincidence that the fixture's default text is the action's. Merged into one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): keep a Retry from re-buying parked media, and state the store seam once Five findings from an inline review. The standalone classroom route asked the ownership sidecar once per load and recorded only stage ownership when that ask failed. Every non-answer fails closed, so one transient 5xx left the genuine author with no resume, no Retry affordance and no legacy narration converted for the rest of the load, with nothing to change it short of a reload. The failure now records the fail-closed answer explicitly -- an answer an earlier load established must not outlive the failure that replaced it -- and an unresolved answer is asked for again, a few times over a few seconds. A real answer, however unwelcome, is final. A Retry could pay a provider for media the pool already held. When the bytes are stored and only the write-back fails in a way that keeps the allocation, it is parked and no local row exists, because that row is written only after a successful write-back. Retry now reads the parked queue exactly as the pass does and re-attempts the write-back: it re-keys the task done when the document takes it, leaves the entry parked when the slide still does not exist, and stays failed and retryable when the document refuses again. Object URLs a parked allocation owns are revoked when the entry is dropped. The commit path leaves them alone while the entry is parked, because it is then the only thing holding bytes this tab can render, so a course switch or a stage deletion was pinning the whole blob for the life of the tab. An entry a slide has already taken is left alone: the task table is displaying those URLs. The fallback lookup for cached bytes is a stage-scoped scan, and the keyed lookup misses for every row the commit path writes, so a pass was materializing and sorting the course's whole media table once per element. One scan per pass now, built on the first miss. It is sound and not merely cheaper: an element asks only for its own placeholder, and every row a pass writes carries the placeholder of the element that wrote it. "The store accepted a write, so it is not out of room" was enforced at three call sites under slightly different conditions, which made it a convention the next pool write path could silently break. It is stated once, in putAsset, for the course whose bytes it just stored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 16 天前 | |
feat(media): write generated media through the asset pool under server-backed persistence (#1392) * feat(media): store generated media in the asset pool when persistence is server-backed With server-backed persistence the document is durable and shared, but generated media stayed in the producing browser: the document kept its gen_img_* / gen_vid_* placeholder and narration kept a browser-derived audio id. Every new browser that opened such a course re-ran generation for every slide, and it never converged, because the address of the generated bytes was never written back into the document. Under server-backed persistence only, the classic generation chain now stores bytes in the asset pool first and writes the id the pool allocated into the document. - The client bootstrap configures the asset seam alongside the document and runtime seams: an HttpAssetStore over the persistence endpoint carrying the same credentials the document store carries, marked server-backed. The seam preflight now covers all three, so a failure still cannot half-configure persistence. - Image, video and TTS generation commit in one fixed order: provider, pool, document, local cache, task. A reference reaches the document only after put returned an id, so a document can never name bytes that were not stored. A failure before the write-back leaves the placeholder with the provider called exactly once; the retry happens on the next owner load. - The write-back is a per-slot rewrite through mutateDocument, which re-reads the current document under the per-stage lock, so it cannot clobber a newer scene. The open course is refreshed with the same rewrite without being marked dirty. - "Has this already been generated?" is answered by the document (the slide exists and no longer holds the placeholder) instead of by this browser's task table. - The classroom's resume effect fails closed on ownership: only a resolved owner starts generation, so a viewer opening a shared course spends nothing. - The local media and audio tables become a per-tab cache. A failed cache write costs a re-download, never the media. Browser-only mode is unchanged: every new call site sits behind the server-backed gate, the local tables stay authoritative there, and placeholders stay in the document. Rendering and export needed no changes. HttpAssetStore.resolve mints an object URL exactly as the browser store does, and the export byte resolver was already pool-first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(classroom): make the generation owner gate a three-outcome rule and apply it everywhere The gate refused everything but a resolved owner, which read a sidecar that answered "no ownership fact exists for this course" as a reason to block. That is the answer a deployment without the sidecar's server-side prerequisites gives for every course, and the answer a course with no ownership record gives: in both, there is nobody the operator's budget needs protecting from, and refusing strands the course's own author behind a question that can never be answered. Ownership is now four states over the sidecar's three outcomes. A definite answer splits into owner and not-owner. An absent record is its own answer, ownerless, and generation proceeds — the behaviour such a deployment had before the gate existed. Only the absence of an answer, a transport failure or a load that has not asked yet, stays unresolved and fails closed: "we could not ask" must never be read as "nobody owns this". One mapper turns a sidecar result into that state, and one predicate decides on it. The workbench classroom pane runs the same resume effect and had no ownership input at all, so a viewer opening a shared course there could still spend the budget. It now asks the sidecar once per course, in parallel with its load and feeding only the generation gate, so its read-only and edit behaviour is unchanged. The shared progressive-load policy carries the gate for it, with both new inputs required rather than defaulted so a future caller cannot omit them into an open budget. Its stale comment claiming ownership could not be expressed here is corrected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the write-back survive autosave, arrive before the scene does, and never leak Independent reviews of the write-back found three ways a durable document could still end up naming a placeholder, and two ways the gate that protects the operator's budget could be walked around. An autosave round captures the store synchronously and writes that capture, so a round already in flight when a rewrite landed wrote the placeholder straight back over the allocated id, and nothing marked the store dirty again to correct it. The rewrite now marks the units it changed, which leaves a corrective flush queued behind the stale one; re-saving a scene that already holds the id is idempotent, losing the id is not. Media is generated from outlines in parallel with scene content and usually finishes first, so the slide that will carry the placeholder does not exist yet and the write-back has nothing to rewrite. That was the ordinary path, not a tail case, and its result was discarded: the task was marked done, the scene was added afterwards with its placeholder intact, and a second pass in the same run could call the provider again. The allocation is now held under the placeholder — which also answers the skip test, so nothing pays twice — and applied when that scene is committed, before its first save. One complete pass now leaves no placeholder behind. A failed commit used to abandon what it had already allocated. A poster upload that failed threw away a stored video and sent the retry to submit the most expensive job in the system again; a rejected write-back left registry rows that name bytes nothing references, which the byte collector cannot reclaim because it only collects blobs no row names. A poster failure now costs the poster, and a write-back that reached nothing reclaims what it allocated. A partial write is left alone, because the document already names it. The ownership gate is fail-closed again. Treating the sidecar's 404 as permission was wrong: the client cannot tell "this course has no owner" from "this deployment told me nothing", so a visitor who opened a shared course could bill the operator. The root cause was the sidecar itself, which gated on the agent runtime although every persisted course has an owner regardless — the persistence route resolves one for every request. It now gates on server persistence, so the configuration that made 404 the universal answer has real ownership facts to report, and the gate can refuse everything but a named owner. Retry affordances answered to no gate at all. A viewer of a shared course with one failed image was shown a Retry button that called the provider. Both retry entry points and every surface that draws them now read one shared permission, so what is offered and what is allowed are the same value. Also: narration regeneration no longer pretends it can replace bytes behind a live id — the exclusivity proof that would allow it is refused by construction once references leave the browser, so it forks to a fresh id and says so; the "already generated" test lets a finished deck answer from the document alone, since scene order stops identifying an outline once slides are inserted or deleted; stored assets record a specific media type rather than a generic transfer type; the pane no longer asks the sidecar in browser-only mode; and the funnel's docstring now states what the per-stage lock actually guarantees, which is same-browser serialization and not a cross-browser compare-and-swap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): park allocations in the deciding turn, never reclaim on an ambiguous write A delta review of the write-back found the first-pass fix still had a window, and the reclamation it added could delete media the document already names. The allocation was parked after an awaited local cache write. A scene committed in that window reconciled against a registry that did not hold it yet, so the document kept the placeholder — and the entry recorded a moment later then answered the skip test as "already handled", so nothing could correct it. Parking now happens inside the write-back, in the same synchronous turn as the decision that nothing could take the reference; no await separates the live check from the park. Allocations parked by an earlier pass are handed to their slides at the start of the next one, before anything decides what still needs generating, so a held allocation whose scene has since arrived becomes a rewrite rather than an answer. Reclaiming on a rejected write was unsound: a rejection does not prove the server did not apply the write, so deleting the asset could break the scene that now names it. The funnel decides instead, and says so: it reclaims only when no store write was ever issued and nothing took the reference. Anything else is placed if its slide exists and parked if it does not, so the next pass reuses the bytes instead of paying for them again. When a write fails after part of it landed, the live store is brought up to the document before the error is rethrown — otherwise the next ordinary flush would overwrite the half that did land, with the ids deliberately not reclaimed. Parked allocations are now cleared with the course. Classic placeholders are reused across runs, so one surviving an interrupted run would be handed to a different slide of the next deck: the previous picture, on a slide whose provider was never asked. Both classroom surfaces clear the arriving course, the deletion cascade clears the deleted one, and clearing the database clears them all. Two more ways generation could start without asking the gate are closed. An overlapping pass — an outline retry re-enters generation with every outline while the first is still working — re-requested elements whose provider call was already in flight; a task that is not done is an answered request, not an unanswered one. And narration regeneration in the timeline editor called the TTS provider and allocated a pool asset with no ownership check at all; it now reads the same permission, which withholds both the per-line and whole-timeline controls and refuses the call. Finally, a pane opened during the stage-link availability gap recorded the sidecar's 404 for a course that was moments from existing and never asked again, leaving the real owner locked out of generation until it remounted. Ownership is re-fetched once the document becomes available; the gate stays closed until an answer arrives, so asking again can only open it for someone entitled to it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the asset routes reachable, and stale snapshots harmless A full-branch audit found that the deployment this project documents could not store a single generated asset, and that several routes into durable storage could still write a placeholder over a reference that had already landed. The persistence route sent asset requests through the development authenticator, which refuses outright in a production build that has not explicitly opted into it — and the documented server-persistence recipe produces exactly that build. Every store and every read answered 401, so images and video failed on every slide while re-billing the provider on each retry, and a narration failure stopped the deck at its first slide. Assets live in one shared partition by design, so there was never anything per-caller for that authenticator to decide: the route now resolves the asset principal itself, alongside the owner it already resolves for documents. Runtime sessions are genuinely per-learner and keep the development authenticator until real session verification replaces it. And narration that cannot be stored no longer fails its scene: the line stays unvoiced and retryable, which is what an image that cannot be stored does to its slide. Placeholders could also come back from behind. A queued autosave's snapshot, an editor-history entry replayed by an undo, the departing save a course switch flushes — each captures content at its own moment, and any of those moments can predate a write-back. Point fixes at each producer would leave the next producer to rediscover the bug, so the check lives at the write boundary every producer passes through, and the allocation record it consults now outlives the parked queue: a placeholder whose rewrite landed long ago is exactly the case it catches. Two ways generation could be lost or repeated are closed. A pass now claims the elements it will reach and releases them however it ends, so an overlapping pass stands down while an aborted one strands nothing — previously its tasks stayed `pending` and every later pass skipped them with no retry control to recover them. And the media abort controller is aborted before being replaced, so a superseded pass stops calling providers instead of running on for a course the user has left. The remaining two are narrower. The workbench pane asks for ownership only after a document load succeeds, and after every later one, mirroring the page route: the load is what creates the ownership row the first time a course is opened, so asking beforehand asked about a course that did not exist yet and locked its author out for the mount. And the ownership gate on the timeline editor now withholds narration regeneration alone; listening back to existing narration and seeing whether a line has any spend nothing and stay available. Known limitation, unchanged and now stated plainly in the comments that used to point at it as a solution: nothing reclaims an unreferenced pool asset. The registry sweep is written but not wired up, and the byte collector only reclaims blobs no registry row names, so every narration regeneration and every abandoned allocation leaves storage behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): gate asset mutations, and make claims and allocation records survive a handoff Opening the asset routes opened all of them. Reads and allocations are meant to be as open as document reads and creates already are, but no authorization hook was supplied, so the handler's default admitted PUT and DELETE too — and those scope by principal key alone, which is one shared constant. Any caller who learned an id, and a document read hands out every id its slides name, could overwrite or destroy another author's media. Mutations now require the deployment's credential, which in a production build without the development-auth opt-in means they are refused outright; reads and allocations stay open. The route comment says what the posture is and what it is not: the deployment-level fence is the access code, and no per-principal quota is configured. The client's own reclaim is best effort to match — losing an argument about deleting an asset must not cost a task its retry, and the bytes are left for server-side reclamation. The pass claim could not survive the handoff it was written for. A retry aborts the live media pass and starts its replacement in the same synchronous block, long before the aborted pass's cleanup runs, so the replacement saw every element still claimed, collected nothing, and returned — leaving each unreached element at pending with nobody coming back for it and no retry control to recover it, which is the exact failure the claim was introduced to prevent. A claim now carries its pass's signal and is retired the moment that signal aborts, and a pass releases only claims it still owns, so a late unwind cannot take its replacement's work. Claims are also acquired at the single point every request passes through, so a single-task retry participates too — previously a retry awaiting its provider was invisible to a pass starting alongside it and both called it. The allocation record could outlive the bytes it named. It was written before the write-back attempted anything and survived the reclaim that followed a failure, so when the slide finally arrived the write boundary stamped a deleted id into the document — and the placeholder it replaced was gone, which reads as already generated and stops anything from retrying. The record is now written only where the allocation is retained, and forgotten wherever a reclaim removes the bytes, including the narration rollback path. The tests follow. The route test drives the real storage handler against an in-memory registry instead of a stub, so it can see what the resolved principal is then allowed to do; the handoff test performs a real abort mid-pass rather than starting from an already-aborted signal; and the guards that could only assert file layout now assert the property they care about, or have been replaced by behaviour. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make media passes serial per course instead of tracking element ownership Three rounds of per-element claims each produced a new way to lose an element. Whole-pass reservations swallowed a Retry for an element the same pass had already failed, leaving it pending with the affordance gone. Retiring a claim by its signal freed an element whose commit was still uploading, so the replacement pass paid for it twice. A claim held for a failed element stranded its retry. The bookkeeping is the defect: every refinement of "who owns this element right now" answered the question at a moment when the answer was already stale. Passes for one course are now serial. A replacement aborts its predecessor, as before, and then waits for it to settle before collecting. That removes the question entirely: a commit already under way finishes — its bytes stored and its reference written, so the new pass sees a resolved slide and skips it — and an element the aborted pass never reached is still a placeholder and gets collected like any other. The claim set, the reservations, the signal retirement and the identity-checked release are all gone. The task table is consulted for one thing only: an element that is generating right now is a single-element retry running alongside the pass, and taking it too would pay twice. Pending is deliberately not a skip reason — it means a pass once intended to reach an element, which an abandoned pass leaves behind with nobody acting on it, and reading that as answered is what stranded elements before. A retry runs concurrently with a pass, because a pass never revisits an element it has processed, and it re-reads the task after its own await and refuses before touching it: marking first and refusing afterwards destroyed the failed state that draws the affordance. Browser-only mode is back to exactly what it was. The abort is now conditional, the waiting does not apply, and the original status-based skip is restored verbatim. Two baseline lines remain changed in each of the two files, and both are behind a server-backed fork whose else-branch is the original. Two smaller things. The allocation record becomes visible when a write goes on the wire rather than when the round trip ends, and the write boundary reconciles under the document lock rather than before it — a save queued during a write-back was otherwise captured with the placeholder and, for a course the user had left, had no corrective flush to follow. And the comments that said a refused reclaim leaves its bytes for server-side reclamation were wrong: nothing collects them, because the registry entry still names its blob and the sweep that would remove it is not wired up. They now say the bytes leak. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a deferred pass re-earn its right to run, and bound the commit it waits on Serializing passes moved their body out of the block that launched them, and three things followed from that. A pass now wakes when its predecessor settles, which can be after the user has left the course. It enqueued before it looked at its signal, into a task table keyed by element id alone — and placeholder ids are not unique across courses, which is why the classroom clears that table on arrival. So a departing course's pass seeded the arriving course's table with tasks carrying the wrong stage id, and a Retry routes by that id: the reference went into the wrong document. A pass now re-validates after the wait, before touching anything shared. The same lateness broke the skip test. `documentSkipIndex` answers only while the live store is on the pass's stage, and returning nothing put the collection loop on the browser-only rule — a silent demotion from "the document is the authority" to "this browser's task table is", on exactly the path where that table has just been cleared. Every element the predecessor had committed was collected again, paid for again, and its second write-back found no placeholder to rewrite, so its bytes were parked where nothing will ever reference them. In server-backed mode an unreadable document now means the pass stands down. And waiting was unbounded. A commit is uncancellable: the asset client takes no signal, and a document write cannot be half-undone. One stalled upload therefore froze the course's media generation for the session — the replacement never collected, the element sat on a skeleton that draws no Retry, and only a reload recovered. The pass's signal is now threaded into the media proxy fetch, and the commit is bounded by a deadline. The deadline is on the wait, not the work: the commit carries on, and if it lands late the document simply ends up correct, while the element becomes retryable and the queue moves on. The tests that were meant to pin the previous round were not sensitive to it. Two asserted end states where the mechanism only changes ordering, and one of them rigged the document read so the assertion held whether or not the pass had waited; a third covered half of what it claimed. They now observe the ordering directly — nothing is issued while another pass for the course is working; in browser-only mode a second pass reaches its provider immediately — and the reconciliation under the document lock has a test that fails when it moves back outside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * revert(media): drop the commit deadline and the abortable download The deadline bought less than it cost. Abandoning a commit after two minutes makes the element retryable while the real commit is still running, so a Retry starts a second commit for the same placeholder against the first: two provider calls, two allocations, and whichever lands second stamps its result over the other's task by element id. The allocation record is keyed by placeholder, so the loser's cleanup erases the winner's record, and the write boundary then puts the raw placeholder back into the document. That is the overlap serial passes were built to remove, reopened through the one door serialization never covered. So a stalled commit holds the course's media queue until it settles or the page is reloaded, and that is written down rather than papered over. The wait is unbounded on purpose: every ceiling on it turns out to be a way of running two commits for one element. Threading the pass signal into the download was also a mistake, in the other direction. The provider call that produced the URL has already been billed, so cancelling the download throws away work that is paid for — and the shared proxy cache records a cancelled request as a transient failure against that URL, which after three of them blocks it for every consumer in the session. Browser-only mode never asked for this: it had no way to observe an abort there, which is exactly why the bytes were kept. The signal is gone from the download again, and `fetchAsBlob` is byte-for-byte what it was before this branch. The regression guard for the stranded-element rule is restored alongside the timing test that was meant to supersede it. It catches a different rule — a task left pending being read as answered — and nothing else does: making the pass skip pending leaves every other suite green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): stop asking the pool for refs it never issued, bound it, and adopt cached bytes Four things a deployment found once this was running for real. A reference this application mints itself — a generation placeholder, a derived narration key — was never in the pool, because the pool allocates every id it holds. Asking anyway used to be an IndexedDB miss; once the pool is server-backed it is a request that answers 404, one per element per load, forever on a course that still holds placeholders. Every lease and probe now checks first. The check is a negative test on shapes this application owns, not an id validator: the pool's id domain stays unconstrained, and anything that is not one of ours is still asked about. The asset store can bound how much one principal holds, and enforces it inside the write transaction, but nothing ever passed the number. It does now, with a default rather than an opt-in: allocation is reachable by any caller a deployment admits, and with one shared principal an unbounded store is unbounded database growth with no operator-visible brake. Refusing asset mutations to unauthenticated callers was not enough, because every authenticated caller resolves to that same shared principal — so authentication decided nothing, and any signed-in visitor could delete any id they learned. Since this branch began storing media the registry is the only copy a course has. Replacing and deleting are now refused to everyone, and the browser no longer tries: an entry nothing references waits for server-side reclamation instead. What a browser must still do is forget its own record of an allocation that reached nothing, or a later save would stamp an id the document has no reason to trust. And a course generated before any of this holds placeholders in its document with its bytes only in the author's browser. Those bytes are paid for, so the author's next load converts them — stored to the pool and written back through the ordinary commit path, with no provider call — instead of buying them again. A row that records only a hosted URL is treated as absent: that URL is the provider's address, not something a document may hold. One renderer expectation moved with this. An untracked placeholder used to paint as pending on first render because asking the pool left a lease in flight; it settled to disabled a moment later either way, and now says so from the start. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): surface a full store as a refusal and convert legacy narration A quota refusal reached the browser as HTTP 500 with a generic message, which reads as a transient failure: the element kept a Retry that would pay a provider again and be refused again. The store raises the contract's own error and the handler maps it to 507, but the store answering a request is not always built by the same bundle as the handler -- the persistence provider is reached from the route bundle and from instrumentation, which is why its state lives on a Symbol.for global -- and `instanceof` is false across that boundary while the declared code is still right. Classify on the code as well as the class, and make the code a permanent, persisted refusal in the browser: recorded locally so it survives a reload, shown as "storage is full", and refused by the retry entry point so a stale button cannot buy a second generation. Every other storage failure stays retryable. Convert what a pre-server-backed course still holds. Generated media is adopted under either key this application has used for it -- the placeholder, and the allocated id of a course converted once and later rolled back -- instead of only the first. Narration is converted by a load-time pass over the open course's speech actions, since nothing re-enters generation for an action that already has an id: bytes to the pool, id written back through a funnel that mirrors the media one, owner-only and server-backed-only. A line whose bytes are in no browser is left alone rather than re-synthesized. Also: the pool guard is now a positive `ast_` test rather than an enumeration of the shapes we mint (imports never reach the pool, so this is safe in both modes); the slide ref collection is an exported pure function so its four lease sites are covered behaviourally; ASSET_QUOTA_BYTES treats every spelling of zero as opting out and refuses a malformed value at startup instead of falling back; the abort signal is re-checked after the cache read, before an uncancellable commit; and the unused `removeAsset` and pool `replace` surfaces are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * chore(storage): release 0.29.1 The asset HTTP handler now recognises a store refusal by the contract code it declares as well as by its class, so a quota refusal raised in another module realm answers 507 instead of 500. Same contract, stricter recognition, no API change: a patch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a full store recoverable and adoption course-safe Narration adoption read the local audio row by its derived key alone. That key carries no stage id and the table is keyed by id alone, so two courses can mint the same one -- a PPTX import numbers its scenes and actions deterministically, which gives every imported deck's first slide `tts_s1_speech-scene-p1`. Locally a collision only means one course plays another's clip in one browser; adopting it wrote that clip into the shared document permanently, for every device and every visitor. A row that names a course is now adopted only into that course, and a row from before that column existed only when the text it recorded is the text of the action being converted. A full asset store was made permanent last round, which was wrong three times over: it overwrote the refused bytes with an empty blob -- on the conversion path that row is a course's only copy of its own media -- it kept sending the rest of the deck to a provider against a ceiling it already knew was reached, and it left no way back once an operator raised that ceiling. A full store is neither the content's fault nor the configuration's, so it is now its own case: the bytes are kept, the pass stops at the first refusal, and the element shows the reason together with a Retry that re-attempts the upload from those bytes. Nothing retries automatically, so no one is re-billed. The narration write-back now reaches the write boundary every producer of a durable write passes through, not only the dirty mark: adoption never deletes the derived row, so a snapshot that reverts the rewrite is adopted again on the next load and allocates a fresh asset every time. Adoption is also mounted by both classroom surfaces rather than one, takes the course's abort signal, and re-validates that this browser still has the course open before each write. ASSET_QUOTA_BYTES is validated from instrumentation, where the README and the docstring already claimed it was: its only other consumer is lazy and memoised, so a malformed ceiling let the process boot and then failed every persistence request, documents and runtime included. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): remember a full store per course, and never lose retained bytes A stopped pass left the elements it never reached as placeholders with no persisted record -- deliberately, since nothing was attempted for them. But that left the next load with no reason not to try: it called a provider for the next placeholder and was refused at exactly the same point, once per reload, indefinitely. A full store is not a property of any slide. It belongs to the deployment and changes for reasons the document knows nothing about, so it is now remembered once per course in the browser's device KV. A pass that finds the marker stands down before spending anything and leaves every placeholder its "storage is full" state and its Retry; the first upload that succeeds clears it and the next pass runs normally. Narration adoption latched per course so it runs once per load, and the latch outlived the abort that leaving a course performs. On a surface that stays mounted across switches -- the workbench pane is one component for every course it shows -- owner course A, visitor course B, then back to A skipped exactly the clips the abort had cut off, and nothing else converts them. The latch is released with the abort now, and a course adopts one run at a time so a re-entry cannot hand a clip a second allocation while the previous run's uncancellable tail is still settling. A quota-blocked element retried into a network error or a 500 lost the bytes that were kept for it: the retry deleted the row before attempting the upload and wrote no replacement for an error carrying no structured code, so the next retry went back to a provider for media this browser had a moment earlier. The row now survives until an upload succeeds, the failure handler keeps whatever bytes the attempt was given, and the retry asks the question a pass asks -- does this browser already hold bytes for this element -- rather than reading an error code that a second failure has already overwritten. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): adopt real legacy narration, queue re-entries, report attempt outcomes Narration adoption admitted a stage-less row only when the text it recorded matched the action being converted. Both of those columns were added to the local audio table by the very change that moved narration onto allocated ids, so a row still carrying a derived key has neither: the rule refused every real pre-allocation course and passed only on fixtures built from post-allocation rows. What the row cannot say, the key can. A derived key names two clips only when two courses share a scene order and an action id, and an action id repeats only when something other than the generator minted it -- an import numbers them by slide position. So a key built from a generated action id is adopted on that basis, a key an import could have reproduced still needs matching text, and a row that names another course is refused however unique its key looks. Handing a re-entering caller the adoption run already in flight undid the latch release it was paired with: that run is bound to the signal the departure just aborted, so it stops at its next clip while the caller -- which has the course open and a live signal -- is told the work is done, and an effect replayed as mount, cleanup, mount adopts nothing at all. A later caller now waits for the uncancellable tail and scans again, which costs a lookup on a course that has nothing left and finishes the clips the abort cut off on one that does. One attempt at an element now reports both facts its callers need instead of a bare boolean: whether the store refused it for room, and whether bytes actually reached the store. Leaving a course clears the task table, so a retry that landed afterwards read "no failed task" as success and deleted the row holding the only copy of the media. Nothing is inferred from that table any more. Reading the localStorage property can throw where storage is denied by policy, typeof included, so the availability check moved inside the guard: this metadata is best-effort, and a rejection here strands a generation pass that has already enqueued its tasks. A retry is never blocked by the per-course "store is full" marker, but a retry that is refused again re-sets it, and adoption now reads and writes the same marker rather than issuing one refused upload per clip on every load. The two canvas element renderers and both thumbnail renderers show the reason beside the Retry, so a full store does not look like an ordinary failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): probe a full store instead of standing down, and pair notices with a Retry Narration adoption was given both halves of the per-course "the store is full" marker last round: it stood down when the marker was set, and it set the marker when its own upload was refused for room. Those halves are only safe together if something can lift the marker, and for adoption nothing could. It has no affordance of its own, it stood down before reaching its own clear, the media pass returns before its marker gate when there is nothing to generate -- so a narration-only deck, or one whose slides are already satisfied, painted no storage-full element and offered no Retry -- and narration generated rather than adopted allocates directly rather than through the media commit. The course's cached narration was then lost for good, where before it converted on the first load after the ceiling was raised. The gate is a probe now. A marked course attempts exactly one clip per load: refused, it stops and the marker stands, which costs what standing down cost; stored, it lifts the marker and finishes the course. Adoption spends no provider money, so the whole cost of probing a store that is still full is one refused upload. Generated narration lifts the marker too. The three surfaces that gained a failure notice last round drew it for any failure with a reason, including the one refusal that is reachable without server-backed persistence, so a browser-only deck painted something it had not painted before. The notice is drawn beside a Retry and nowhere else, which is what it was added for and what leaves browser-only output unchanged. Both are now asserted through the render harness the surface matrix already had. A caller arriving while a rescan is queued shares it rather than appending another. One rescan converts whatever the run in flight left and every later one would find an allocated id on every action, so a chain bought nothing and turned a single stalled upload into a course that never adopts again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): treat a refusal for room as a fact about one clip, not the deck The asset store checks each write against the headroom it has left, so a store that refuses a long opening clip can still hold every short clip behind it. Narration adoption assumed the opposite: it broke the deck at the first refusal and then re-attempted that same first clip on every later load, because the document names it first. A deck whose longest clip exceeds current headroom therefore never converted the clips that would have fit, with no affordance to recover it -- the state the probe was introduced to remove, reached through a narrower door. An unmarked load now attempts every clip, skipping the ones that do not fit, and remembers the condition only if the load ends with clips it still could not store. A marked load spends its single upload on the smallest clip left rather than the first one named: that is the clip that answers the question the marker asks, because if the smallest does not fit nothing does. The media pass keeps stopping at its first refusal, and for a reason adoption does not share -- every element it attempts costs a provider call. A rescan several callers share took the newest caller's signal, and the newest caller is not necessarily the one still there: a surface that opened a course and closed it again would stop work a surface still showing that course was waiting for, and that surface is latched, so it would never ask again. The shared run now takes a signal that is aborted only once every caller has left. The comment claiming the shared rescan contains a stalled upload was wrong -- the rescan is chained off the run in flight, so a stalled upload leaves every caller pending exactly as a chain would. It claims the bounded queue it actually provides, and the stall is recorded as a limitation. The failed-state containers took their stacking classes unconditionally, so markup differed in browser-only mode even though nothing moved on screen. Those classes are applied only when there is a notice to stack, and the tests assert the exact class attribute rather than a substring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): stop narration adoption writing the media pass's store-full marker The marker means "do not call a provider for this course". A path is entitled to write it only if its own refusal cost a provider call, and narration adoption's refusals cost nothing: it uploads bytes this browser already holds. The store also checks each write against the headroom it has left, so a clip that does not fit says nothing about whether a slide's image would. Adoption was writing it anyway, and one over-long narration clip was therefore enough to stand a course's entire image pass down on every later load -- on a store that had just accepted adoption's other clips. The author could still recover each element by hand, every load, for ever. Three rounds of narrowing this seam produced a finding each time, so it is removed rather than narrowed again. Gone: the marker read, the single-clip probe, the smallest-clip selection, and the up-front read of every row into an array -- which also retires a sampled-then-stale flag and the retention of a whole deck's blobs for the length of a run, and returns the loop to streaming one row at a time. Adoption's rule is now that every load attempts every clip it holds, once; any failure skips that clip and the load continues. The noise the coupling was meant to avoid does not arise, because after the first load the clips still outstanding are exactly the ones that did not fit -- normally none, or one. A successful write still clears the marker, and that is a different kind of statement: a write that went through is a fact this run established, where a refusal is an inference about what some other write would cost. For a course whose media needs nothing, adoption and generated narration are also the only paths that can establish it. The failure module still documented the deck-wide premise this contradicts. It now says what is true: the check is per write, and the media pass stops the deck as a judgement about cost rather than about certainty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bound a full store's cost from the store's own arithmetic Removing the store-full marker from narration adoption removed its bound too, and the code then asserted the bound was unnecessary. It is, on a store with room for most of a deck. On the store the whole mechanism exists for -- the ceiling reached, nothing fitting -- the outstanding set after every load is the entire deck, so a thirty-clip course posted thirty full blobs on every load, indefinitely. Each of those is not a cheap refusal: the bytes are uploaded, the server hashes the whole payload, and only then takes a per-principal lock and sums every entry that principal owns before saying no. The bound needs no flag, no key and nothing carried between loads. The store asks whether `used + addedBytes` exceeds the ceiling, and `used` only grows while a run is uploading, so a clip refused for want of room implies every clip at least that large is refused for the rest of that run. The run keeps the smallest size it has been refused and skips anything no smaller without uploading it; a smaller clip is still attempted, because it may fit. A deck the store refuses entirely now costs one upload per successive size minimum instead of one per clip, and a deck it has room for costs nothing extra, because nothing is refused. Only a refusal for room lowers the bar: a dropped connection says nothing about how much room there is. The deck-wide certainty premise the failure module retracted last round still stood verbatim at the site that implements the stand-down. Both copies now say the same thing: the check is per write, and the pass stops the deck as a judgement about cost rather than about certainty. The comment on adoption's marker clear now names its price. Narration of a few hundred bytes fits in headroom an image does not, so a proven write can let the next pass buy one more image that is refused again -- bounded at one, and the price of the alternative being a course whose media never generates again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(media): state the adoption bound exactly, and stop three comments describing the old rule The comment introducing the in-load bound gave its cost as "at most a handful, and the first load pays the most". Neither clause is a property of the rule. A clip is skipped only when something no larger was already refused, so a fully-refused deck costs one upload per successive size minimum in document order: one when the clips grow, about ln N for an arbitrary order, and one per clip when they only shrink -- a long opener followed by terser lines is exactly that shape. And no load is cheaper than the first, because the bound resets per run and a refused clip stays outstanding. The comment now says that, and points at what would make it exactly one for any ordering: the store returning its remaining headroom in the refusal's existing details channel, which the server leaves empty today. Two other comments still described the previous rule -- "attempts every clip it holds, every load" -- one of them twenty lines above the paragraph that introduces the bound, in the same block. Both now say what the code does. The bound's soundness is worth stating where a maintainer will look for it: quota is charged at full length with no discount for a duplicate, the sum it is checked against joins entries to blobs so the collector cannot lower it, the check takes a per-principal lock before summing, and replace and delete are refused to every browser. Nothing a run can do makes room appear inside it. One test installed a row implementation and replaced it wholesale a few lines later, so the first was dead and the survivor dropped the text the first clip's import-shaped key needs for the ownership rule -- it passed on the coincidence that the fixture's default text is the action's. Merged into one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): keep a Retry from re-buying parked media, and state the store seam once Five findings from an inline review. The standalone classroom route asked the ownership sidecar once per load and recorded only stage ownership when that ask failed. Every non-answer fails closed, so one transient 5xx left the genuine author with no resume, no Retry affordance and no legacy narration converted for the rest of the load, with nothing to change it short of a reload. The failure now records the fail-closed answer explicitly -- an answer an earlier load established must not outlive the failure that replaced it -- and an unresolved answer is asked for again, a few times over a few seconds. A real answer, however unwelcome, is final. A Retry could pay a provider for media the pool already held. When the bytes are stored and only the write-back fails in a way that keeps the allocation, it is parked and no local row exists, because that row is written only after a successful write-back. Retry now reads the parked queue exactly as the pass does and re-attempts the write-back: it re-keys the task done when the document takes it, leaves the entry parked when the slide still does not exist, and stays failed and retryable when the document refuses again. Object URLs a parked allocation owns are revoked when the entry is dropped. The commit path leaves them alone while the entry is parked, because it is then the only thing holding bytes this tab can render, so a course switch or a stage deletion was pinning the whole blob for the life of the tab. An entry a slide has already taken is left alone: the task table is displaying those URLs. The fallback lookup for cached bytes is a stage-scoped scan, and the keyed lookup misses for every row the commit path writes, so a pass was materializing and sorting the course's whole media table once per element. One scan per pass now, built on the first miss. It is sound and not merely cheaper: an element asks only for its own placeholder, and every row a pass writes carries the placeholder of the element that wrote it. "The store accepted a write, so it is not out of room" was enforced at three call sites under slightly different conditions, which made it a convention the next pool write path could silently break. It is stated once, in putAsset, for the course whose bytes it just stored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 16 天前 | |
Update Doubao Seed model catalog (#827) Closes #826 | 2 个月前 | |
feat(media): write generated media through the asset pool under server-backed persistence (#1392) * feat(media): store generated media in the asset pool when persistence is server-backed With server-backed persistence the document is durable and shared, but generated media stayed in the producing browser: the document kept its gen_img_* / gen_vid_* placeholder and narration kept a browser-derived audio id. Every new browser that opened such a course re-ran generation for every slide, and it never converged, because the address of the generated bytes was never written back into the document. Under server-backed persistence only, the classic generation chain now stores bytes in the asset pool first and writes the id the pool allocated into the document. - The client bootstrap configures the asset seam alongside the document and runtime seams: an HttpAssetStore over the persistence endpoint carrying the same credentials the document store carries, marked server-backed. The seam preflight now covers all three, so a failure still cannot half-configure persistence. - Image, video and TTS generation commit in one fixed order: provider, pool, document, local cache, task. A reference reaches the document only after put returned an id, so a document can never name bytes that were not stored. A failure before the write-back leaves the placeholder with the provider called exactly once; the retry happens on the next owner load. - The write-back is a per-slot rewrite through mutateDocument, which re-reads the current document under the per-stage lock, so it cannot clobber a newer scene. The open course is refreshed with the same rewrite without being marked dirty. - "Has this already been generated?" is answered by the document (the slide exists and no longer holds the placeholder) instead of by this browser's task table. - The classroom's resume effect fails closed on ownership: only a resolved owner starts generation, so a viewer opening a shared course spends nothing. - The local media and audio tables become a per-tab cache. A failed cache write costs a re-download, never the media. Browser-only mode is unchanged: every new call site sits behind the server-backed gate, the local tables stay authoritative there, and placeholders stay in the document. Rendering and export needed no changes. HttpAssetStore.resolve mints an object URL exactly as the browser store does, and the export byte resolver was already pool-first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(classroom): make the generation owner gate a three-outcome rule and apply it everywhere The gate refused everything but a resolved owner, which read a sidecar that answered "no ownership fact exists for this course" as a reason to block. That is the answer a deployment without the sidecar's server-side prerequisites gives for every course, and the answer a course with no ownership record gives: in both, there is nobody the operator's budget needs protecting from, and refusing strands the course's own author behind a question that can never be answered. Ownership is now four states over the sidecar's three outcomes. A definite answer splits into owner and not-owner. An absent record is its own answer, ownerless, and generation proceeds — the behaviour such a deployment had before the gate existed. Only the absence of an answer, a transport failure or a load that has not asked yet, stays unresolved and fails closed: "we could not ask" must never be read as "nobody owns this". One mapper turns a sidecar result into that state, and one predicate decides on it. The workbench classroom pane runs the same resume effect and had no ownership input at all, so a viewer opening a shared course there could still spend the budget. It now asks the sidecar once per course, in parallel with its load and feeding only the generation gate, so its read-only and edit behaviour is unchanged. The shared progressive-load policy carries the gate for it, with both new inputs required rather than defaulted so a future caller cannot omit them into an open budget. Its stale comment claiming ownership could not be expressed here is corrected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the write-back survive autosave, arrive before the scene does, and never leak Independent reviews of the write-back found three ways a durable document could still end up naming a placeholder, and two ways the gate that protects the operator's budget could be walked around. An autosave round captures the store synchronously and writes that capture, so a round already in flight when a rewrite landed wrote the placeholder straight back over the allocated id, and nothing marked the store dirty again to correct it. The rewrite now marks the units it changed, which leaves a corrective flush queued behind the stale one; re-saving a scene that already holds the id is idempotent, losing the id is not. Media is generated from outlines in parallel with scene content and usually finishes first, so the slide that will carry the placeholder does not exist yet and the write-back has nothing to rewrite. That was the ordinary path, not a tail case, and its result was discarded: the task was marked done, the scene was added afterwards with its placeholder intact, and a second pass in the same run could call the provider again. The allocation is now held under the placeholder — which also answers the skip test, so nothing pays twice — and applied when that scene is committed, before its first save. One complete pass now leaves no placeholder behind. A failed commit used to abandon what it had already allocated. A poster upload that failed threw away a stored video and sent the retry to submit the most expensive job in the system again; a rejected write-back left registry rows that name bytes nothing references, which the byte collector cannot reclaim because it only collects blobs no row names. A poster failure now costs the poster, and a write-back that reached nothing reclaims what it allocated. A partial write is left alone, because the document already names it. The ownership gate is fail-closed again. Treating the sidecar's 404 as permission was wrong: the client cannot tell "this course has no owner" from "this deployment told me nothing", so a visitor who opened a shared course could bill the operator. The root cause was the sidecar itself, which gated on the agent runtime although every persisted course has an owner regardless — the persistence route resolves one for every request. It now gates on server persistence, so the configuration that made 404 the universal answer has real ownership facts to report, and the gate can refuse everything but a named owner. Retry affordances answered to no gate at all. A viewer of a shared course with one failed image was shown a Retry button that called the provider. Both retry entry points and every surface that draws them now read one shared permission, so what is offered and what is allowed are the same value. Also: narration regeneration no longer pretends it can replace bytes behind a live id — the exclusivity proof that would allow it is refused by construction once references leave the browser, so it forks to a fresh id and says so; the "already generated" test lets a finished deck answer from the document alone, since scene order stops identifying an outline once slides are inserted or deleted; stored assets record a specific media type rather than a generic transfer type; the pane no longer asks the sidecar in browser-only mode; and the funnel's docstring now states what the per-stage lock actually guarantees, which is same-browser serialization and not a cross-browser compare-and-swap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): park allocations in the deciding turn, never reclaim on an ambiguous write A delta review of the write-back found the first-pass fix still had a window, and the reclamation it added could delete media the document already names. The allocation was parked after an awaited local cache write. A scene committed in that window reconciled against a registry that did not hold it yet, so the document kept the placeholder — and the entry recorded a moment later then answered the skip test as "already handled", so nothing could correct it. Parking now happens inside the write-back, in the same synchronous turn as the decision that nothing could take the reference; no await separates the live check from the park. Allocations parked by an earlier pass are handed to their slides at the start of the next one, before anything decides what still needs generating, so a held allocation whose scene has since arrived becomes a rewrite rather than an answer. Reclaiming on a rejected write was unsound: a rejection does not prove the server did not apply the write, so deleting the asset could break the scene that now names it. The funnel decides instead, and says so: it reclaims only when no store write was ever issued and nothing took the reference. Anything else is placed if its slide exists and parked if it does not, so the next pass reuses the bytes instead of paying for them again. When a write fails after part of it landed, the live store is brought up to the document before the error is rethrown — otherwise the next ordinary flush would overwrite the half that did land, with the ids deliberately not reclaimed. Parked allocations are now cleared with the course. Classic placeholders are reused across runs, so one surviving an interrupted run would be handed to a different slide of the next deck: the previous picture, on a slide whose provider was never asked. Both classroom surfaces clear the arriving course, the deletion cascade clears the deleted one, and clearing the database clears them all. Two more ways generation could start without asking the gate are closed. An overlapping pass — an outline retry re-enters generation with every outline while the first is still working — re-requested elements whose provider call was already in flight; a task that is not done is an answered request, not an unanswered one. And narration regeneration in the timeline editor called the TTS provider and allocated a pool asset with no ownership check at all; it now reads the same permission, which withholds both the per-line and whole-timeline controls and refuses the call. Finally, a pane opened during the stage-link availability gap recorded the sidecar's 404 for a course that was moments from existing and never asked again, leaving the real owner locked out of generation until it remounted. Ownership is re-fetched once the document becomes available; the gate stays closed until an answer arrives, so asking again can only open it for someone entitled to it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the asset routes reachable, and stale snapshots harmless A full-branch audit found that the deployment this project documents could not store a single generated asset, and that several routes into durable storage could still write a placeholder over a reference that had already landed. The persistence route sent asset requests through the development authenticator, which refuses outright in a production build that has not explicitly opted into it — and the documented server-persistence recipe produces exactly that build. Every store and every read answered 401, so images and video failed on every slide while re-billing the provider on each retry, and a narration failure stopped the deck at its first slide. Assets live in one shared partition by design, so there was never anything per-caller for that authenticator to decide: the route now resolves the asset principal itself, alongside the owner it already resolves for documents. Runtime sessions are genuinely per-learner and keep the development authenticator until real session verification replaces it. And narration that cannot be stored no longer fails its scene: the line stays unvoiced and retryable, which is what an image that cannot be stored does to its slide. Placeholders could also come back from behind. A queued autosave's snapshot, an editor-history entry replayed by an undo, the departing save a course switch flushes — each captures content at its own moment, and any of those moments can predate a write-back. Point fixes at each producer would leave the next producer to rediscover the bug, so the check lives at the write boundary every producer passes through, and the allocation record it consults now outlives the parked queue: a placeholder whose rewrite landed long ago is exactly the case it catches. Two ways generation could be lost or repeated are closed. A pass now claims the elements it will reach and releases them however it ends, so an overlapping pass stands down while an aborted one strands nothing — previously its tasks stayed `pending` and every later pass skipped them with no retry control to recover them. And the media abort controller is aborted before being replaced, so a superseded pass stops calling providers instead of running on for a course the user has left. The remaining two are narrower. The workbench pane asks for ownership only after a document load succeeds, and after every later one, mirroring the page route: the load is what creates the ownership row the first time a course is opened, so asking beforehand asked about a course that did not exist yet and locked its author out for the mount. And the ownership gate on the timeline editor now withholds narration regeneration alone; listening back to existing narration and seeing whether a line has any spend nothing and stay available. Known limitation, unchanged and now stated plainly in the comments that used to point at it as a solution: nothing reclaims an unreferenced pool asset. The registry sweep is written but not wired up, and the byte collector only reclaims blobs no registry row names, so every narration regeneration and every abandoned allocation leaves storage behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): gate asset mutations, and make claims and allocation records survive a handoff Opening the asset routes opened all of them. Reads and allocations are meant to be as open as document reads and creates already are, but no authorization hook was supplied, so the handler's default admitted PUT and DELETE too — and those scope by principal key alone, which is one shared constant. Any caller who learned an id, and a document read hands out every id its slides name, could overwrite or destroy another author's media. Mutations now require the deployment's credential, which in a production build without the development-auth opt-in means they are refused outright; reads and allocations stay open. The route comment says what the posture is and what it is not: the deployment-level fence is the access code, and no per-principal quota is configured. The client's own reclaim is best effort to match — losing an argument about deleting an asset must not cost a task its retry, and the bytes are left for server-side reclamation. The pass claim could not survive the handoff it was written for. A retry aborts the live media pass and starts its replacement in the same synchronous block, long before the aborted pass's cleanup runs, so the replacement saw every element still claimed, collected nothing, and returned — leaving each unreached element at pending with nobody coming back for it and no retry control to recover it, which is the exact failure the claim was introduced to prevent. A claim now carries its pass's signal and is retired the moment that signal aborts, and a pass releases only claims it still owns, so a late unwind cannot take its replacement's work. Claims are also acquired at the single point every request passes through, so a single-task retry participates too — previously a retry awaiting its provider was invisible to a pass starting alongside it and both called it. The allocation record could outlive the bytes it named. It was written before the write-back attempted anything and survived the reclaim that followed a failure, so when the slide finally arrived the write boundary stamped a deleted id into the document — and the placeholder it replaced was gone, which reads as already generated and stops anything from retrying. The record is now written only where the allocation is retained, and forgotten wherever a reclaim removes the bytes, including the narration rollback path. The tests follow. The route test drives the real storage handler against an in-memory registry instead of a stub, so it can see what the resolved principal is then allowed to do; the handoff test performs a real abort mid-pass rather than starting from an already-aborted signal; and the guards that could only assert file layout now assert the property they care about, or have been replaced by behaviour. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make media passes serial per course instead of tracking element ownership Three rounds of per-element claims each produced a new way to lose an element. Whole-pass reservations swallowed a Retry for an element the same pass had already failed, leaving it pending with the affordance gone. Retiring a claim by its signal freed an element whose commit was still uploading, so the replacement pass paid for it twice. A claim held for a failed element stranded its retry. The bookkeeping is the defect: every refinement of "who owns this element right now" answered the question at a moment when the answer was already stale. Passes for one course are now serial. A replacement aborts its predecessor, as before, and then waits for it to settle before collecting. That removes the question entirely: a commit already under way finishes — its bytes stored and its reference written, so the new pass sees a resolved slide and skips it — and an element the aborted pass never reached is still a placeholder and gets collected like any other. The claim set, the reservations, the signal retirement and the identity-checked release are all gone. The task table is consulted for one thing only: an element that is generating right now is a single-element retry running alongside the pass, and taking it too would pay twice. Pending is deliberately not a skip reason — it means a pass once intended to reach an element, which an abandoned pass leaves behind with nobody acting on it, and reading that as answered is what stranded elements before. A retry runs concurrently with a pass, because a pass never revisits an element it has processed, and it re-reads the task after its own await and refuses before touching it: marking first and refusing afterwards destroyed the failed state that draws the affordance. Browser-only mode is back to exactly what it was. The abort is now conditional, the waiting does not apply, and the original status-based skip is restored verbatim. Two baseline lines remain changed in each of the two files, and both are behind a server-backed fork whose else-branch is the original. Two smaller things. The allocation record becomes visible when a write goes on the wire rather than when the round trip ends, and the write boundary reconciles under the document lock rather than before it — a save queued during a write-back was otherwise captured with the placeholder and, for a course the user had left, had no corrective flush to follow. And the comments that said a refused reclaim leaves its bytes for server-side reclamation were wrong: nothing collects them, because the registry entry still names its blob and the sweep that would remove it is not wired up. They now say the bytes leak. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a deferred pass re-earn its right to run, and bound the commit it waits on Serializing passes moved their body out of the block that launched them, and three things followed from that. A pass now wakes when its predecessor settles, which can be after the user has left the course. It enqueued before it looked at its signal, into a task table keyed by element id alone — and placeholder ids are not unique across courses, which is why the classroom clears that table on arrival. So a departing course's pass seeded the arriving course's table with tasks carrying the wrong stage id, and a Retry routes by that id: the reference went into the wrong document. A pass now re-validates after the wait, before touching anything shared. The same lateness broke the skip test. `documentSkipIndex` answers only while the live store is on the pass's stage, and returning nothing put the collection loop on the browser-only rule — a silent demotion from "the document is the authority" to "this browser's task table is", on exactly the path where that table has just been cleared. Every element the predecessor had committed was collected again, paid for again, and its second write-back found no placeholder to rewrite, so its bytes were parked where nothing will ever reference them. In server-backed mode an unreadable document now means the pass stands down. And waiting was unbounded. A commit is uncancellable: the asset client takes no signal, and a document write cannot be half-undone. One stalled upload therefore froze the course's media generation for the session — the replacement never collected, the element sat on a skeleton that draws no Retry, and only a reload recovered. The pass's signal is now threaded into the media proxy fetch, and the commit is bounded by a deadline. The deadline is on the wait, not the work: the commit carries on, and if it lands late the document simply ends up correct, while the element becomes retryable and the queue moves on. The tests that were meant to pin the previous round were not sensitive to it. Two asserted end states where the mechanism only changes ordering, and one of them rigged the document read so the assertion held whether or not the pass had waited; a third covered half of what it claimed. They now observe the ordering directly — nothing is issued while another pass for the course is working; in browser-only mode a second pass reaches its provider immediately — and the reconciliation under the document lock has a test that fails when it moves back outside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * revert(media): drop the commit deadline and the abortable download The deadline bought less than it cost. Abandoning a commit after two minutes makes the element retryable while the real commit is still running, so a Retry starts a second commit for the same placeholder against the first: two provider calls, two allocations, and whichever lands second stamps its result over the other's task by element id. The allocation record is keyed by placeholder, so the loser's cleanup erases the winner's record, and the write boundary then puts the raw placeholder back into the document. That is the overlap serial passes were built to remove, reopened through the one door serialization never covered. So a stalled commit holds the course's media queue until it settles or the page is reloaded, and that is written down rather than papered over. The wait is unbounded on purpose: every ceiling on it turns out to be a way of running two commits for one element. Threading the pass signal into the download was also a mistake, in the other direction. The provider call that produced the URL has already been billed, so cancelling the download throws away work that is paid for — and the shared proxy cache records a cancelled request as a transient failure against that URL, which after three of them blocks it for every consumer in the session. Browser-only mode never asked for this: it had no way to observe an abort there, which is exactly why the bytes were kept. The signal is gone from the download again, and `fetchAsBlob` is byte-for-byte what it was before this branch. The regression guard for the stranded-element rule is restored alongside the timing test that was meant to supersede it. It catches a different rule — a task left pending being read as answered — and nothing else does: making the pass skip pending leaves every other suite green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): stop asking the pool for refs it never issued, bound it, and adopt cached bytes Four things a deployment found once this was running for real. A reference this application mints itself — a generation placeholder, a derived narration key — was never in the pool, because the pool allocates every id it holds. Asking anyway used to be an IndexedDB miss; once the pool is server-backed it is a request that answers 404, one per element per load, forever on a course that still holds placeholders. Every lease and probe now checks first. The check is a negative test on shapes this application owns, not an id validator: the pool's id domain stays unconstrained, and anything that is not one of ours is still asked about. The asset store can bound how much one principal holds, and enforces it inside the write transaction, but nothing ever passed the number. It does now, with a default rather than an opt-in: allocation is reachable by any caller a deployment admits, and with one shared principal an unbounded store is unbounded database growth with no operator-visible brake. Refusing asset mutations to unauthenticated callers was not enough, because every authenticated caller resolves to that same shared principal — so authentication decided nothing, and any signed-in visitor could delete any id they learned. Since this branch began storing media the registry is the only copy a course has. Replacing and deleting are now refused to everyone, and the browser no longer tries: an entry nothing references waits for server-side reclamation instead. What a browser must still do is forget its own record of an allocation that reached nothing, or a later save would stamp an id the document has no reason to trust. And a course generated before any of this holds placeholders in its document with its bytes only in the author's browser. Those bytes are paid for, so the author's next load converts them — stored to the pool and written back through the ordinary commit path, with no provider call — instead of buying them again. A row that records only a hosted URL is treated as absent: that URL is the provider's address, not something a document may hold. One renderer expectation moved with this. An untracked placeholder used to paint as pending on first render because asking the pool left a lease in flight; it settled to disabled a moment later either way, and now says so from the start. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): surface a full store as a refusal and convert legacy narration A quota refusal reached the browser as HTTP 500 with a generic message, which reads as a transient failure: the element kept a Retry that would pay a provider again and be refused again. The store raises the contract's own error and the handler maps it to 507, but the store answering a request is not always built by the same bundle as the handler -- the persistence provider is reached from the route bundle and from instrumentation, which is why its state lives on a Symbol.for global -- and `instanceof` is false across that boundary while the declared code is still right. Classify on the code as well as the class, and make the code a permanent, persisted refusal in the browser: recorded locally so it survives a reload, shown as "storage is full", and refused by the retry entry point so a stale button cannot buy a second generation. Every other storage failure stays retryable. Convert what a pre-server-backed course still holds. Generated media is adopted under either key this application has used for it -- the placeholder, and the allocated id of a course converted once and later rolled back -- instead of only the first. Narration is converted by a load-time pass over the open course's speech actions, since nothing re-enters generation for an action that already has an id: bytes to the pool, id written back through a funnel that mirrors the media one, owner-only and server-backed-only. A line whose bytes are in no browser is left alone rather than re-synthesized. Also: the pool guard is now a positive `ast_` test rather than an enumeration of the shapes we mint (imports never reach the pool, so this is safe in both modes); the slide ref collection is an exported pure function so its four lease sites are covered behaviourally; ASSET_QUOTA_BYTES treats every spelling of zero as opting out and refuses a malformed value at startup instead of falling back; the abort signal is re-checked after the cache read, before an uncancellable commit; and the unused `removeAsset` and pool `replace` surfaces are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * chore(storage): release 0.29.1 The asset HTTP handler now recognises a store refusal by the contract code it declares as well as by its class, so a quota refusal raised in another module realm answers 507 instead of 500. Same contract, stricter recognition, no API change: a patch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a full store recoverable and adoption course-safe Narration adoption read the local audio row by its derived key alone. That key carries no stage id and the table is keyed by id alone, so two courses can mint the same one -- a PPTX import numbers its scenes and actions deterministically, which gives every imported deck's first slide `tts_s1_speech-scene-p1`. Locally a collision only means one course plays another's clip in one browser; adopting it wrote that clip into the shared document permanently, for every device and every visitor. A row that names a course is now adopted only into that course, and a row from before that column existed only when the text it recorded is the text of the action being converted. A full asset store was made permanent last round, which was wrong three times over: it overwrote the refused bytes with an empty blob -- on the conversion path that row is a course's only copy of its own media -- it kept sending the rest of the deck to a provider against a ceiling it already knew was reached, and it left no way back once an operator raised that ceiling. A full store is neither the content's fault nor the configuration's, so it is now its own case: the bytes are kept, the pass stops at the first refusal, and the element shows the reason together with a Retry that re-attempts the upload from those bytes. Nothing retries automatically, so no one is re-billed. The narration write-back now reaches the write boundary every producer of a durable write passes through, not only the dirty mark: adoption never deletes the derived row, so a snapshot that reverts the rewrite is adopted again on the next load and allocates a fresh asset every time. Adoption is also mounted by both classroom surfaces rather than one, takes the course's abort signal, and re-validates that this browser still has the course open before each write. ASSET_QUOTA_BYTES is validated from instrumentation, where the README and the docstring already claimed it was: its only other consumer is lazy and memoised, so a malformed ceiling let the process boot and then failed every persistence request, documents and runtime included. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): remember a full store per course, and never lose retained bytes A stopped pass left the elements it never reached as placeholders with no persisted record -- deliberately, since nothing was attempted for them. But that left the next load with no reason not to try: it called a provider for the next placeholder and was refused at exactly the same point, once per reload, indefinitely. A full store is not a property of any slide. It belongs to the deployment and changes for reasons the document knows nothing about, so it is now remembered once per course in the browser's device KV. A pass that finds the marker stands down before spending anything and leaves every placeholder its "storage is full" state and its Retry; the first upload that succeeds clears it and the next pass runs normally. Narration adoption latched per course so it runs once per load, and the latch outlived the abort that leaving a course performs. On a surface that stays mounted across switches -- the workbench pane is one component for every course it shows -- owner course A, visitor course B, then back to A skipped exactly the clips the abort had cut off, and nothing else converts them. The latch is released with the abort now, and a course adopts one run at a time so a re-entry cannot hand a clip a second allocation while the previous run's uncancellable tail is still settling. A quota-blocked element retried into a network error or a 500 lost the bytes that were kept for it: the retry deleted the row before attempting the upload and wrote no replacement for an error carrying no structured code, so the next retry went back to a provider for media this browser had a moment earlier. The row now survives until an upload succeeds, the failure handler keeps whatever bytes the attempt was given, and the retry asks the question a pass asks -- does this browser already hold bytes for this element -- rather than reading an error code that a second failure has already overwritten. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): adopt real legacy narration, queue re-entries, report attempt outcomes Narration adoption admitted a stage-less row only when the text it recorded matched the action being converted. Both of those columns were added to the local audio table by the very change that moved narration onto allocated ids, so a row still carrying a derived key has neither: the rule refused every real pre-allocation course and passed only on fixtures built from post-allocation rows. What the row cannot say, the key can. A derived key names two clips only when two courses share a scene order and an action id, and an action id repeats only when something other than the generator minted it -- an import numbers them by slide position. So a key built from a generated action id is adopted on that basis, a key an import could have reproduced still needs matching text, and a row that names another course is refused however unique its key looks. Handing a re-entering caller the adoption run already in flight undid the latch release it was paired with: that run is bound to the signal the departure just aborted, so it stops at its next clip while the caller -- which has the course open and a live signal -- is told the work is done, and an effect replayed as mount, cleanup, mount adopts nothing at all. A later caller now waits for the uncancellable tail and scans again, which costs a lookup on a course that has nothing left and finishes the clips the abort cut off on one that does. One attempt at an element now reports both facts its callers need instead of a bare boolean: whether the store refused it for room, and whether bytes actually reached the store. Leaving a course clears the task table, so a retry that landed afterwards read "no failed task" as success and deleted the row holding the only copy of the media. Nothing is inferred from that table any more. Reading the localStorage property can throw where storage is denied by policy, typeof included, so the availability check moved inside the guard: this metadata is best-effort, and a rejection here strands a generation pass that has already enqueued its tasks. A retry is never blocked by the per-course "store is full" marker, but a retry that is refused again re-sets it, and adoption now reads and writes the same marker rather than issuing one refused upload per clip on every load. The two canvas element renderers and both thumbnail renderers show the reason beside the Retry, so a full store does not look like an ordinary failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): probe a full store instead of standing down, and pair notices with a Retry Narration adoption was given both halves of the per-course "the store is full" marker last round: it stood down when the marker was set, and it set the marker when its own upload was refused for room. Those halves are only safe together if something can lift the marker, and for adoption nothing could. It has no affordance of its own, it stood down before reaching its own clear, the media pass returns before its marker gate when there is nothing to generate -- so a narration-only deck, or one whose slides are already satisfied, painted no storage-full element and offered no Retry -- and narration generated rather than adopted allocates directly rather than through the media commit. The course's cached narration was then lost for good, where before it converted on the first load after the ceiling was raised. The gate is a probe now. A marked course attempts exactly one clip per load: refused, it stops and the marker stands, which costs what standing down cost; stored, it lifts the marker and finishes the course. Adoption spends no provider money, so the whole cost of probing a store that is still full is one refused upload. Generated narration lifts the marker too. The three surfaces that gained a failure notice last round drew it for any failure with a reason, including the one refusal that is reachable without server-backed persistence, so a browser-only deck painted something it had not painted before. The notice is drawn beside a Retry and nowhere else, which is what it was added for and what leaves browser-only output unchanged. Both are now asserted through the render harness the surface matrix already had. A caller arriving while a rescan is queued shares it rather than appending another. One rescan converts whatever the run in flight left and every later one would find an allocated id on every action, so a chain bought nothing and turned a single stalled upload into a course that never adopts again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): treat a refusal for room as a fact about one clip, not the deck The asset store checks each write against the headroom it has left, so a store that refuses a long opening clip can still hold every short clip behind it. Narration adoption assumed the opposite: it broke the deck at the first refusal and then re-attempted that same first clip on every later load, because the document names it first. A deck whose longest clip exceeds current headroom therefore never converted the clips that would have fit, with no affordance to recover it -- the state the probe was introduced to remove, reached through a narrower door. An unmarked load now attempts every clip, skipping the ones that do not fit, and remembers the condition only if the load ends with clips it still could not store. A marked load spends its single upload on the smallest clip left rather than the first one named: that is the clip that answers the question the marker asks, because if the smallest does not fit nothing does. The media pass keeps stopping at its first refusal, and for a reason adoption does not share -- every element it attempts costs a provider call. A rescan several callers share took the newest caller's signal, and the newest caller is not necessarily the one still there: a surface that opened a course and closed it again would stop work a surface still showing that course was waiting for, and that surface is latched, so it would never ask again. The shared run now takes a signal that is aborted only once every caller has left. The comment claiming the shared rescan contains a stalled upload was wrong -- the rescan is chained off the run in flight, so a stalled upload leaves every caller pending exactly as a chain would. It claims the bounded queue it actually provides, and the stall is recorded as a limitation. The failed-state containers took their stacking classes unconditionally, so markup differed in browser-only mode even though nothing moved on screen. Those classes are applied only when there is a notice to stack, and the tests assert the exact class attribute rather than a substring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): stop narration adoption writing the media pass's store-full marker The marker means "do not call a provider for this course". A path is entitled to write it only if its own refusal cost a provider call, and narration adoption's refusals cost nothing: it uploads bytes this browser already holds. The store also checks each write against the headroom it has left, so a clip that does not fit says nothing about whether a slide's image would. Adoption was writing it anyway, and one over-long narration clip was therefore enough to stand a course's entire image pass down on every later load -- on a store that had just accepted adoption's other clips. The author could still recover each element by hand, every load, for ever. Three rounds of narrowing this seam produced a finding each time, so it is removed rather than narrowed again. Gone: the marker read, the single-clip probe, the smallest-clip selection, and the up-front read of every row into an array -- which also retires a sampled-then-stale flag and the retention of a whole deck's blobs for the length of a run, and returns the loop to streaming one row at a time. Adoption's rule is now that every load attempts every clip it holds, once; any failure skips that clip and the load continues. The noise the coupling was meant to avoid does not arise, because after the first load the clips still outstanding are exactly the ones that did not fit -- normally none, or one. A successful write still clears the marker, and that is a different kind of statement: a write that went through is a fact this run established, where a refusal is an inference about what some other write would cost. For a course whose media needs nothing, adoption and generated narration are also the only paths that can establish it. The failure module still documented the deck-wide premise this contradicts. It now says what is true: the check is per write, and the media pass stops the deck as a judgement about cost rather than about certainty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bound a full store's cost from the store's own arithmetic Removing the store-full marker from narration adoption removed its bound too, and the code then asserted the bound was unnecessary. It is, on a store with room for most of a deck. On the store the whole mechanism exists for -- the ceiling reached, nothing fitting -- the outstanding set after every load is the entire deck, so a thirty-clip course posted thirty full blobs on every load, indefinitely. Each of those is not a cheap refusal: the bytes are uploaded, the server hashes the whole payload, and only then takes a per-principal lock and sums every entry that principal owns before saying no. The bound needs no flag, no key and nothing carried between loads. The store asks whether `used + addedBytes` exceeds the ceiling, and `used` only grows while a run is uploading, so a clip refused for want of room implies every clip at least that large is refused for the rest of that run. The run keeps the smallest size it has been refused and skips anything no smaller without uploading it; a smaller clip is still attempted, because it may fit. A deck the store refuses entirely now costs one upload per successive size minimum instead of one per clip, and a deck it has room for costs nothing extra, because nothing is refused. Only a refusal for room lowers the bar: a dropped connection says nothing about how much room there is. The deck-wide certainty premise the failure module retracted last round still stood verbatim at the site that implements the stand-down. Both copies now say the same thing: the check is per write, and the pass stops the deck as a judgement about cost rather than about certainty. The comment on adoption's marker clear now names its price. Narration of a few hundred bytes fits in headroom an image does not, so a proven write can let the next pass buy one more image that is refused again -- bounded at one, and the price of the alternative being a course whose media never generates again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(media): state the adoption bound exactly, and stop three comments describing the old rule The comment introducing the in-load bound gave its cost as "at most a handful, and the first load pays the most". Neither clause is a property of the rule. A clip is skipped only when something no larger was already refused, so a fully-refused deck costs one upload per successive size minimum in document order: one when the clips grow, about ln N for an arbitrary order, and one per clip when they only shrink -- a long opener followed by terser lines is exactly that shape. And no load is cheaper than the first, because the bound resets per run and a refused clip stays outstanding. The comment now says that, and points at what would make it exactly one for any ordering: the store returning its remaining headroom in the refusal's existing details channel, which the server leaves empty today. Two other comments still described the previous rule -- "attempts every clip it holds, every load" -- one of them twenty lines above the paragraph that introduces the bound, in the same block. Both now say what the code does. The bound's soundness is worth stating where a maintainer will look for it: quota is charged at full length with no discount for a duplicate, the sum it is checked against joins entries to blobs so the collector cannot lower it, the check takes a per-principal lock before summing, and replace and delete are refused to every browser. Nothing a run can do makes room appear inside it. One test installed a row implementation and replaced it wholesale a few lines later, so the first was dead and the survivor dropped the text the first clip's import-shaped key needs for the ownership rule -- it passed on the coincidence that the fixture's default text is the action's. Merged into one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): keep a Retry from re-buying parked media, and state the store seam once Five findings from an inline review. The standalone classroom route asked the ownership sidecar once per load and recorded only stage ownership when that ask failed. Every non-answer fails closed, so one transient 5xx left the genuine author with no resume, no Retry affordance and no legacy narration converted for the rest of the load, with nothing to change it short of a reload. The failure now records the fail-closed answer explicitly -- an answer an earlier load established must not outlive the failure that replaced it -- and an unresolved answer is asked for again, a few times over a few seconds. A real answer, however unwelcome, is final. A Retry could pay a provider for media the pool already held. When the bytes are stored and only the write-back fails in a way that keeps the allocation, it is parked and no local row exists, because that row is written only after a successful write-back. Retry now reads the parked queue exactly as the pass does and re-attempts the write-back: it re-keys the task done when the document takes it, leaves the entry parked when the slide still does not exist, and stays failed and retryable when the document refuses again. Object URLs a parked allocation owns are revoked when the entry is dropped. The commit path leaves them alone while the entry is parked, because it is then the only thing holding bytes this tab can render, so a course switch or a stage deletion was pinning the whole blob for the life of the tab. An entry a slide has already taken is left alone: the task table is displaying those URLs. The fallback lookup for cached bytes is a stage-scoped scan, and the keyed lookup misses for every row the commit path writes, so a pass was materializing and sorting the course's whole media table once per element. One scan per pass now, built on the first miss. It is sound and not merely cheaper: an element asks only for its own placeholder, and every row a pass writes carries the placeholder of the element that wrote it. "The store accepted a write, so it is not out of room" was enforced at three call sites under slightly different conditions, which made it a convention the next pool write path could silently break. It is stated once, in putAsset, for the course whose bytes it just stored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 16 天前 | |
feat(dsl): standardize the asset manifest and converge the export paths (#1007 part 3) (#1117) * feat(dsl): standardize the document asset manifest Add asset-manifest.ts to @openmaic/dsl: the canonical AssetManifestEntry shape (ref, kind, and byteSize/mimeType/duration/voice/prompt metadata where available) plus enumerateAssetManifest, the pure document-to-manifest enumeration. An entry's ref is the reference exactly as the document holds it -- the manifest is the id-based reference enumeration with metadata, not a content hash and not a resolution result. The traversal walks the stage whiteboard, each scene's canvas/whiteboards/speech actions, and the stage video-manifest keys in document order, with logical-owner reference counts that match the accounting duplication-safe replacement uses. This settles the media-ref + asset-manifest schema question (#779 open question 4) on the side the asset-pool RFC already implied: the schema is a function of the id semantics decided there. The type lives in the dsl rather than a new @openmaic/exporter package because the enumeration is pure over document types the dsl already owns (Stage/Scene/Slide/Action), so a separate package would add a published artifact and release-workflow surface without adding a capability; the storage contract comment now points at the module. Refs #1007 * refactor(export): drive the classroom ZIP from the asset manifest collectMediaFiles used to scan the whole mediaFiles table for the stage, so any row the document no longer references -- an orphan left by an edit or a superseded regeneration -- rode along into the archive. Both ZIP collectors now take their reference sets from the standardized asset manifest (buildStageAssetManifest wraps the dsl enumeration with the compatibility rows' metadata): only referenced assets are archived, and a referenced asset whose bytes exist only in the pool is still collected via a synthesized record. Byte resolution is unchanged: pool first through resolveStoredBytes / resolveAudioBlob, with the compatibility row kept as the legacy byte fallback and as the metadata source. mediaIndex is now a serialized view of the manifest, and the missing-audio report derives from the manifest's audio entries instead of a second action walk. The audioRef mapping and the legacy audioUrl fetch path (collectLegacyAudioForExport) are untouched. Refs #1007 * refactor(video-export): take the timeline's reference sets from the manifest createVideoTimelineDeps scanned the whole mediaFiles table for the stage and derived its audio id set from its own action walk -- a third, independent answer to "which media does this course use?". Both record loads now key off the standardized asset manifest: media rows are read per manifest ref by compound key instead of by table scan, and the audio id set is the manifest's audio entries. Orphan rows were never reachable through the scene-scoped elementId-to-mediaRef bridge; now they are not even read. The bridge itself is untouched: element ids recur across scenes, so the elementId-to-mediaRef mapping stays scoped per scene, and the legacy audioUrl fallback keeps its own action walk because a URL is not a manifest ref. AssetPlan remains the video IR's view of the same references. Refs #1007 * refactor(export): resolve PPTX media through the shared resolver only Each PPTX element branch carried its own resolution chain: a task-state renderable-URL lookup first, then -- gated on the legacy placeholder predicate -- a stored-bytes override, with the poster block repeating the pattern. One helper now owns resolution for backgrounds, images, video / audio sources, and posters: opaque refs (allocated ids and legacy placeholders alike, no placeholder-pattern gate) resolve pool-first through resolveStoredBytes and embed as data URLs, concrete addresses resolve through the media state machine and keep the caller's fetch path. exportMediaResolution and the resolveStoredMediaBlob wrapper fold into the helper; resolvePptxMediaBinding stays as the state-machine entry the resolution-surface test matrix drives. Refs #1007 * refactor(export): retire the export-side Dexie byte fallbacks Export call sites no longer read bytes off compatibility rows directly. The ZIP collectors and the video timeline's audio load resolve bytes only through the shared resolvers (resolveStoredBytes / resolveAudioBlob), which answer pool-first and keep the compatibility row as their internal legacy fallback level; the row reads that remain at the call sites supply metadata (format/duration/voice/mime/size/prompt) only. The rows themselves stay for legacy and regeneration readers -- what goes is the export paths' own fallback logic. One observable tightening: a failed media row (error set, empty placeholder blob) no longer ships a 0-byte file into the classroom ZIP, and an evicted row no longer ships its empty local blob; referenced-but- byteless assets are simply absent from the archive, as they already were when no row existed. Refs #1007 * test(media): cover the enriched stage asset manifest builder Pins the join between the pure dsl enumeration and the compatibility rows: metadata attaches by ref, rows no document reference names never appear, and a referenced asset with no row keeps a metadata-free entry. Refs #1007 * fix(video-export): widen the deps stage input for the manifest enumeration enumerateAssetManifest reads the stage's whiteboard and videoManifest, so createVideoTimelineDeps declares them on its input instead of the bare id; callers pass only the id today and the optional fields stay absent. Also applies the repo prettier formatting to the files this branch touched. Refs #1007 * fix(dsl): enumerate slide audio elements in the asset manifest Slide audio elements carry their own src, and the manifest skipped them, so a manifest-driven collector could never archive their bytes. The audio slot maps to kind 'audio' alongside narration ids. Refs #1007 * fix(media): harden ref-keyed lookups against prototype-named asset refs AssetRef is an unconstrained string alias, so a media reference can legitimately be "__proto__", "constructor", or any other Object.prototype member. Plain objects keyed by such refs silently drop assignments or answer lookups with the prototype object, which rewrite paths then accept as a mapped id. Convert the remaining ref-keyed lookup tables introduced by the export convergence to prototype-safe structures: the classroom import media/poster alias maps and the legacy-conversion video-manifest reconstruction now use Map / null-prototype containers with explicit membership checks, and every consumed value is validated as a string before it is written into a src / mediaRef / audioId slot. The shared media-task lookup receives the same treatment: one centralized own-property-checked lookupMediaTask now serves the stored-bytes resolver, the PPTX embeddable-src path, the video collection path, and the element/background task resolution, so a prototype-named placeholderRef can no longer hide a re-keyed task from the fallback chain. Adversarial tests drive "__proto__" and "constructor" refs through the import round trip, the PPTX fallback path, the legacy conversion commit path, and the media-task fallback end to end, including a buildPptxBlob regression with a task re-keyed to an allocated id while retaining a prototype-named placeholderRef. * fix(export): use safe archive asset paths * refactor(dsl): centralize slide media slot roles * refactor(export): derive consumer refs from manifest * fix(export): sanitize classroom archive extensions * fix(video-export): preserve narration speech order * fix(export): enforce kind-coherent archive media * fix(export): define media coherence boundary * fix(export): carry task-owned poster binding for PPTX export A video element with no explicit poster falls back to its media task's generated poster URL, but resolveVideoMediaForElement left posterTask undefined for that case, so the PPTX manifest guard saw a foreign URL with no task-ownership exemption and dropped the video element instead of using the established runtime poster fallback. Carry the poster task binding whenever the task poster is the effective poster: the task-owned URL then satisfies the guard's objectUrl exemption end to end. A concrete explicit element poster still stays element-owned and never borrows the binding, and the guard's foreign-ref rejection is preserved (and exported as a directly testable predicate). Coverage: an element with no poster plus a task-provided poster embeds the task poster as the PPTX cover (red at the pre-fix head, green now), and a genuinely unrelated URL with no task ownership is still rejected by the guard. * fix(export): preserve legacy narration source refs in the media index The explicit sourceRef contract was partial: primary audio and generated media entries carried it, but legacy URL narration serialized no source ref. The legacy URL itself is the natural source ref — it is known at fetch time — so wire it through the collected blob into the mediaIndex entry. Import already registers serialized sourceRefs as aliases, so the URL now round-trips as an explicit mapping instead of being reconstructed only from the action's audioRef. Poster siblings are deliberately NOT given their own mediaIndex entry: a sibling poster (media/asset-<n>.poster.<ext>) is a legacy byte copy written from the video record and is not an independently referenced document asset — when the poster is a real document asset it already has its own indexed entry with a sourceRef, and import reconstructs the sibling by path derivation from its parent video entry, reusing the poster's own indexed allocation when one exists. The PR description is narrowed to match; corrected paragraph: "Archive names never interpolate refs — sequential safe paths (media/asset-<n>.<ext>, audio/audio-<n>.<ext>) with the original ref preserved through an explicit sourceRef mapping on every independently indexed media entry: generated media assets, poster assets, primary narration, and legacy URL narration (the legacy URL itself is the entry's sourceRef). Extensions are allowlisted per kind. The one exception is the legacy sibling poster byte copy (media/asset-<n>.poster.<ext>, written next to its video when the video record still carries the pre-pool poster bytes): it is not an independently referenced document asset, so it has no mediaIndex entry or sourceRef of its own — its identity is derivable from its parent video entry (same index), and import reconstructs it by sibling-path derivation from that video entry, reusing the poster's own indexed allocation when one exists." --------- Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 1 个月前 | |
feat(persistence): turn on the server-owned asset lifecycle and release assets on course deletion (#1007 amendment, part 2) (#1473) App wiring for the @openmaic/storage 0.31.0 lifecycle: reference tracking and document references are paired unconditionally and declared at startup, course deletion withdraws references inside the tombstone transaction, ASSET_PENDING_TTL_MS configures the pending window, and the dead client-side reclamation code is removed. | 12 天前 | |
feat(media): write generated media through the asset pool under server-backed persistence (#1392) * feat(media): store generated media in the asset pool when persistence is server-backed With server-backed persistence the document is durable and shared, but generated media stayed in the producing browser: the document kept its gen_img_* / gen_vid_* placeholder and narration kept a browser-derived audio id. Every new browser that opened such a course re-ran generation for every slide, and it never converged, because the address of the generated bytes was never written back into the document. Under server-backed persistence only, the classic generation chain now stores bytes in the asset pool first and writes the id the pool allocated into the document. - The client bootstrap configures the asset seam alongside the document and runtime seams: an HttpAssetStore over the persistence endpoint carrying the same credentials the document store carries, marked server-backed. The seam preflight now covers all three, so a failure still cannot half-configure persistence. - Image, video and TTS generation commit in one fixed order: provider, pool, document, local cache, task. A reference reaches the document only after put returned an id, so a document can never name bytes that were not stored. A failure before the write-back leaves the placeholder with the provider called exactly once; the retry happens on the next owner load. - The write-back is a per-slot rewrite through mutateDocument, which re-reads the current document under the per-stage lock, so it cannot clobber a newer scene. The open course is refreshed with the same rewrite without being marked dirty. - "Has this already been generated?" is answered by the document (the slide exists and no longer holds the placeholder) instead of by this browser's task table. - The classroom's resume effect fails closed on ownership: only a resolved owner starts generation, so a viewer opening a shared course spends nothing. - The local media and audio tables become a per-tab cache. A failed cache write costs a re-download, never the media. Browser-only mode is unchanged: every new call site sits behind the server-backed gate, the local tables stay authoritative there, and placeholders stay in the document. Rendering and export needed no changes. HttpAssetStore.resolve mints an object URL exactly as the browser store does, and the export byte resolver was already pool-first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(classroom): make the generation owner gate a three-outcome rule and apply it everywhere The gate refused everything but a resolved owner, which read a sidecar that answered "no ownership fact exists for this course" as a reason to block. That is the answer a deployment without the sidecar's server-side prerequisites gives for every course, and the answer a course with no ownership record gives: in both, there is nobody the operator's budget needs protecting from, and refusing strands the course's own author behind a question that can never be answered. Ownership is now four states over the sidecar's three outcomes. A definite answer splits into owner and not-owner. An absent record is its own answer, ownerless, and generation proceeds — the behaviour such a deployment had before the gate existed. Only the absence of an answer, a transport failure or a load that has not asked yet, stays unresolved and fails closed: "we could not ask" must never be read as "nobody owns this". One mapper turns a sidecar result into that state, and one predicate decides on it. The workbench classroom pane runs the same resume effect and had no ownership input at all, so a viewer opening a shared course there could still spend the budget. It now asks the sidecar once per course, in parallel with its load and feeding only the generation gate, so its read-only and edit behaviour is unchanged. The shared progressive-load policy carries the gate for it, with both new inputs required rather than defaulted so a future caller cannot omit them into an open budget. Its stale comment claiming ownership could not be expressed here is corrected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the write-back survive autosave, arrive before the scene does, and never leak Independent reviews of the write-back found three ways a durable document could still end up naming a placeholder, and two ways the gate that protects the operator's budget could be walked around. An autosave round captures the store synchronously and writes that capture, so a round already in flight when a rewrite landed wrote the placeholder straight back over the allocated id, and nothing marked the store dirty again to correct it. The rewrite now marks the units it changed, which leaves a corrective flush queued behind the stale one; re-saving a scene that already holds the id is idempotent, losing the id is not. Media is generated from outlines in parallel with scene content and usually finishes first, so the slide that will carry the placeholder does not exist yet and the write-back has nothing to rewrite. That was the ordinary path, not a tail case, and its result was discarded: the task was marked done, the scene was added afterwards with its placeholder intact, and a second pass in the same run could call the provider again. The allocation is now held under the placeholder — which also answers the skip test, so nothing pays twice — and applied when that scene is committed, before its first save. One complete pass now leaves no placeholder behind. A failed commit used to abandon what it had already allocated. A poster upload that failed threw away a stored video and sent the retry to submit the most expensive job in the system again; a rejected write-back left registry rows that name bytes nothing references, which the byte collector cannot reclaim because it only collects blobs no row names. A poster failure now costs the poster, and a write-back that reached nothing reclaims what it allocated. A partial write is left alone, because the document already names it. The ownership gate is fail-closed again. Treating the sidecar's 404 as permission was wrong: the client cannot tell "this course has no owner" from "this deployment told me nothing", so a visitor who opened a shared course could bill the operator. The root cause was the sidecar itself, which gated on the agent runtime although every persisted course has an owner regardless — the persistence route resolves one for every request. It now gates on server persistence, so the configuration that made 404 the universal answer has real ownership facts to report, and the gate can refuse everything but a named owner. Retry affordances answered to no gate at all. A viewer of a shared course with one failed image was shown a Retry button that called the provider. Both retry entry points and every surface that draws them now read one shared permission, so what is offered and what is allowed are the same value. Also: narration regeneration no longer pretends it can replace bytes behind a live id — the exclusivity proof that would allow it is refused by construction once references leave the browser, so it forks to a fresh id and says so; the "already generated" test lets a finished deck answer from the document alone, since scene order stops identifying an outline once slides are inserted or deleted; stored assets record a specific media type rather than a generic transfer type; the pane no longer asks the sidecar in browser-only mode; and the funnel's docstring now states what the per-stage lock actually guarantees, which is same-browser serialization and not a cross-browser compare-and-swap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): park allocations in the deciding turn, never reclaim on an ambiguous write A delta review of the write-back found the first-pass fix still had a window, and the reclamation it added could delete media the document already names. The allocation was parked after an awaited local cache write. A scene committed in that window reconciled against a registry that did not hold it yet, so the document kept the placeholder — and the entry recorded a moment later then answered the skip test as "already handled", so nothing could correct it. Parking now happens inside the write-back, in the same synchronous turn as the decision that nothing could take the reference; no await separates the live check from the park. Allocations parked by an earlier pass are handed to their slides at the start of the next one, before anything decides what still needs generating, so a held allocation whose scene has since arrived becomes a rewrite rather than an answer. Reclaiming on a rejected write was unsound: a rejection does not prove the server did not apply the write, so deleting the asset could break the scene that now names it. The funnel decides instead, and says so: it reclaims only when no store write was ever issued and nothing took the reference. Anything else is placed if its slide exists and parked if it does not, so the next pass reuses the bytes instead of paying for them again. When a write fails after part of it landed, the live store is brought up to the document before the error is rethrown — otherwise the next ordinary flush would overwrite the half that did land, with the ids deliberately not reclaimed. Parked allocations are now cleared with the course. Classic placeholders are reused across runs, so one surviving an interrupted run would be handed to a different slide of the next deck: the previous picture, on a slide whose provider was never asked. Both classroom surfaces clear the arriving course, the deletion cascade clears the deleted one, and clearing the database clears them all. Two more ways generation could start without asking the gate are closed. An overlapping pass — an outline retry re-enters generation with every outline while the first is still working — re-requested elements whose provider call was already in flight; a task that is not done is an answered request, not an unanswered one. And narration regeneration in the timeline editor called the TTS provider and allocated a pool asset with no ownership check at all; it now reads the same permission, which withholds both the per-line and whole-timeline controls and refuses the call. Finally, a pane opened during the stage-link availability gap recorded the sidecar's 404 for a course that was moments from existing and never asked again, leaving the real owner locked out of generation until it remounted. Ownership is re-fetched once the document becomes available; the gate stays closed until an answer arrives, so asking again can only open it for someone entitled to it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the asset routes reachable, and stale snapshots harmless A full-branch audit found that the deployment this project documents could not store a single generated asset, and that several routes into durable storage could still write a placeholder over a reference that had already landed. The persistence route sent asset requests through the development authenticator, which refuses outright in a production build that has not explicitly opted into it — and the documented server-persistence recipe produces exactly that build. Every store and every read answered 401, so images and video failed on every slide while re-billing the provider on each retry, and a narration failure stopped the deck at its first slide. Assets live in one shared partition by design, so there was never anything per-caller for that authenticator to decide: the route now resolves the asset principal itself, alongside the owner it already resolves for documents. Runtime sessions are genuinely per-learner and keep the development authenticator until real session verification replaces it. And narration that cannot be stored no longer fails its scene: the line stays unvoiced and retryable, which is what an image that cannot be stored does to its slide. Placeholders could also come back from behind. A queued autosave's snapshot, an editor-history entry replayed by an undo, the departing save a course switch flushes — each captures content at its own moment, and any of those moments can predate a write-back. Point fixes at each producer would leave the next producer to rediscover the bug, so the check lives at the write boundary every producer passes through, and the allocation record it consults now outlives the parked queue: a placeholder whose rewrite landed long ago is exactly the case it catches. Two ways generation could be lost or repeated are closed. A pass now claims the elements it will reach and releases them however it ends, so an overlapping pass stands down while an aborted one strands nothing — previously its tasks stayed `pending` and every later pass skipped them with no retry control to recover them. And the media abort controller is aborted before being replaced, so a superseded pass stops calling providers instead of running on for a course the user has left. The remaining two are narrower. The workbench pane asks for ownership only after a document load succeeds, and after every later one, mirroring the page route: the load is what creates the ownership row the first time a course is opened, so asking beforehand asked about a course that did not exist yet and locked its author out for the mount. And the ownership gate on the timeline editor now withholds narration regeneration alone; listening back to existing narration and seeing whether a line has any spend nothing and stay available. Known limitation, unchanged and now stated plainly in the comments that used to point at it as a solution: nothing reclaims an unreferenced pool asset. The registry sweep is written but not wired up, and the byte collector only reclaims blobs no registry row names, so every narration regeneration and every abandoned allocation leaves storage behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): gate asset mutations, and make claims and allocation records survive a handoff Opening the asset routes opened all of them. Reads and allocations are meant to be as open as document reads and creates already are, but no authorization hook was supplied, so the handler's default admitted PUT and DELETE too — and those scope by principal key alone, which is one shared constant. Any caller who learned an id, and a document read hands out every id its slides name, could overwrite or destroy another author's media. Mutations now require the deployment's credential, which in a production build without the development-auth opt-in means they are refused outright; reads and allocations stay open. The route comment says what the posture is and what it is not: the deployment-level fence is the access code, and no per-principal quota is configured. The client's own reclaim is best effort to match — losing an argument about deleting an asset must not cost a task its retry, and the bytes are left for server-side reclamation. The pass claim could not survive the handoff it was written for. A retry aborts the live media pass and starts its replacement in the same synchronous block, long before the aborted pass's cleanup runs, so the replacement saw every element still claimed, collected nothing, and returned — leaving each unreached element at pending with nobody coming back for it and no retry control to recover it, which is the exact failure the claim was introduced to prevent. A claim now carries its pass's signal and is retired the moment that signal aborts, and a pass releases only claims it still owns, so a late unwind cannot take its replacement's work. Claims are also acquired at the single point every request passes through, so a single-task retry participates too — previously a retry awaiting its provider was invisible to a pass starting alongside it and both called it. The allocation record could outlive the bytes it named. It was written before the write-back attempted anything and survived the reclaim that followed a failure, so when the slide finally arrived the write boundary stamped a deleted id into the document — and the placeholder it replaced was gone, which reads as already generated and stops anything from retrying. The record is now written only where the allocation is retained, and forgotten wherever a reclaim removes the bytes, including the narration rollback path. The tests follow. The route test drives the real storage handler against an in-memory registry instead of a stub, so it can see what the resolved principal is then allowed to do; the handoff test performs a real abort mid-pass rather than starting from an already-aborted signal; and the guards that could only assert file layout now assert the property they care about, or have been replaced by behaviour. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make media passes serial per course instead of tracking element ownership Three rounds of per-element claims each produced a new way to lose an element. Whole-pass reservations swallowed a Retry for an element the same pass had already failed, leaving it pending with the affordance gone. Retiring a claim by its signal freed an element whose commit was still uploading, so the replacement pass paid for it twice. A claim held for a failed element stranded its retry. The bookkeeping is the defect: every refinement of "who owns this element right now" answered the question at a moment when the answer was already stale. Passes for one course are now serial. A replacement aborts its predecessor, as before, and then waits for it to settle before collecting. That removes the question entirely: a commit already under way finishes — its bytes stored and its reference written, so the new pass sees a resolved slide and skips it — and an element the aborted pass never reached is still a placeholder and gets collected like any other. The claim set, the reservations, the signal retirement and the identity-checked release are all gone. The task table is consulted for one thing only: an element that is generating right now is a single-element retry running alongside the pass, and taking it too would pay twice. Pending is deliberately not a skip reason — it means a pass once intended to reach an element, which an abandoned pass leaves behind with nobody acting on it, and reading that as answered is what stranded elements before. A retry runs concurrently with a pass, because a pass never revisits an element it has processed, and it re-reads the task after its own await and refuses before touching it: marking first and refusing afterwards destroyed the failed state that draws the affordance. Browser-only mode is back to exactly what it was. The abort is now conditional, the waiting does not apply, and the original status-based skip is restored verbatim. Two baseline lines remain changed in each of the two files, and both are behind a server-backed fork whose else-branch is the original. Two smaller things. The allocation record becomes visible when a write goes on the wire rather than when the round trip ends, and the write boundary reconciles under the document lock rather than before it — a save queued during a write-back was otherwise captured with the placeholder and, for a course the user had left, had no corrective flush to follow. And the comments that said a refused reclaim leaves its bytes for server-side reclamation were wrong: nothing collects them, because the registry entry still names its blob and the sweep that would remove it is not wired up. They now say the bytes leak. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a deferred pass re-earn its right to run, and bound the commit it waits on Serializing passes moved their body out of the block that launched them, and three things followed from that. A pass now wakes when its predecessor settles, which can be after the user has left the course. It enqueued before it looked at its signal, into a task table keyed by element id alone — and placeholder ids are not unique across courses, which is why the classroom clears that table on arrival. So a departing course's pass seeded the arriving course's table with tasks carrying the wrong stage id, and a Retry routes by that id: the reference went into the wrong document. A pass now re-validates after the wait, before touching anything shared. The same lateness broke the skip test. `documentSkipIndex` answers only while the live store is on the pass's stage, and returning nothing put the collection loop on the browser-only rule — a silent demotion from "the document is the authority" to "this browser's task table is", on exactly the path where that table has just been cleared. Every element the predecessor had committed was collected again, paid for again, and its second write-back found no placeholder to rewrite, so its bytes were parked where nothing will ever reference them. In server-backed mode an unreadable document now means the pass stands down. And waiting was unbounded. A commit is uncancellable: the asset client takes no signal, and a document write cannot be half-undone. One stalled upload therefore froze the course's media generation for the session — the replacement never collected, the element sat on a skeleton that draws no Retry, and only a reload recovered. The pass's signal is now threaded into the media proxy fetch, and the commit is bounded by a deadline. The deadline is on the wait, not the work: the commit carries on, and if it lands late the document simply ends up correct, while the element becomes retryable and the queue moves on. The tests that were meant to pin the previous round were not sensitive to it. Two asserted end states where the mechanism only changes ordering, and one of them rigged the document read so the assertion held whether or not the pass had waited; a third covered half of what it claimed. They now observe the ordering directly — nothing is issued while another pass for the course is working; in browser-only mode a second pass reaches its provider immediately — and the reconciliation under the document lock has a test that fails when it moves back outside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * revert(media): drop the commit deadline and the abortable download The deadline bought less than it cost. Abandoning a commit after two minutes makes the element retryable while the real commit is still running, so a Retry starts a second commit for the same placeholder against the first: two provider calls, two allocations, and whichever lands second stamps its result over the other's task by element id. The allocation record is keyed by placeholder, so the loser's cleanup erases the winner's record, and the write boundary then puts the raw placeholder back into the document. That is the overlap serial passes were built to remove, reopened through the one door serialization never covered. So a stalled commit holds the course's media queue until it settles or the page is reloaded, and that is written down rather than papered over. The wait is unbounded on purpose: every ceiling on it turns out to be a way of running two commits for one element. Threading the pass signal into the download was also a mistake, in the other direction. The provider call that produced the URL has already been billed, so cancelling the download throws away work that is paid for — and the shared proxy cache records a cancelled request as a transient failure against that URL, which after three of them blocks it for every consumer in the session. Browser-only mode never asked for this: it had no way to observe an abort there, which is exactly why the bytes were kept. The signal is gone from the download again, and `fetchAsBlob` is byte-for-byte what it was before this branch. The regression guard for the stranded-element rule is restored alongside the timing test that was meant to supersede it. It catches a different rule — a task left pending being read as answered — and nothing else does: making the pass skip pending leaves every other suite green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): stop asking the pool for refs it never issued, bound it, and adopt cached bytes Four things a deployment found once this was running for real. A reference this application mints itself — a generation placeholder, a derived narration key — was never in the pool, because the pool allocates every id it holds. Asking anyway used to be an IndexedDB miss; once the pool is server-backed it is a request that answers 404, one per element per load, forever on a course that still holds placeholders. Every lease and probe now checks first. The check is a negative test on shapes this application owns, not an id validator: the pool's id domain stays unconstrained, and anything that is not one of ours is still asked about. The asset store can bound how much one principal holds, and enforces it inside the write transaction, but nothing ever passed the number. It does now, with a default rather than an opt-in: allocation is reachable by any caller a deployment admits, and with one shared principal an unbounded store is unbounded database growth with no operator-visible brake. Refusing asset mutations to unauthenticated callers was not enough, because every authenticated caller resolves to that same shared principal — so authentication decided nothing, and any signed-in visitor could delete any id they learned. Since this branch began storing media the registry is the only copy a course has. Replacing and deleting are now refused to everyone, and the browser no longer tries: an entry nothing references waits for server-side reclamation instead. What a browser must still do is forget its own record of an allocation that reached nothing, or a later save would stamp an id the document has no reason to trust. And a course generated before any of this holds placeholders in its document with its bytes only in the author's browser. Those bytes are paid for, so the author's next load converts them — stored to the pool and written back through the ordinary commit path, with no provider call — instead of buying them again. A row that records only a hosted URL is treated as absent: that URL is the provider's address, not something a document may hold. One renderer expectation moved with this. An untracked placeholder used to paint as pending on first render because asking the pool left a lease in flight; it settled to disabled a moment later either way, and now says so from the start. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): surface a full store as a refusal and convert legacy narration A quota refusal reached the browser as HTTP 500 with a generic message, which reads as a transient failure: the element kept a Retry that would pay a provider again and be refused again. The store raises the contract's own error and the handler maps it to 507, but the store answering a request is not always built by the same bundle as the handler -- the persistence provider is reached from the route bundle and from instrumentation, which is why its state lives on a Symbol.for global -- and `instanceof` is false across that boundary while the declared code is still right. Classify on the code as well as the class, and make the code a permanent, persisted refusal in the browser: recorded locally so it survives a reload, shown as "storage is full", and refused by the retry entry point so a stale button cannot buy a second generation. Every other storage failure stays retryable. Convert what a pre-server-backed course still holds. Generated media is adopted under either key this application has used for it -- the placeholder, and the allocated id of a course converted once and later rolled back -- instead of only the first. Narration is converted by a load-time pass over the open course's speech actions, since nothing re-enters generation for an action that already has an id: bytes to the pool, id written back through a funnel that mirrors the media one, owner-only and server-backed-only. A line whose bytes are in no browser is left alone rather than re-synthesized. Also: the pool guard is now a positive `ast_` test rather than an enumeration of the shapes we mint (imports never reach the pool, so this is safe in both modes); the slide ref collection is an exported pure function so its four lease sites are covered behaviourally; ASSET_QUOTA_BYTES treats every spelling of zero as opting out and refuses a malformed value at startup instead of falling back; the abort signal is re-checked after the cache read, before an uncancellable commit; and the unused `removeAsset` and pool `replace` surfaces are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * chore(storage): release 0.29.1 The asset HTTP handler now recognises a store refusal by the contract code it declares as well as by its class, so a quota refusal raised in another module realm answers 507 instead of 500. Same contract, stricter recognition, no API change: a patch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a full store recoverable and adoption course-safe Narration adoption read the local audio row by its derived key alone. That key carries no stage id and the table is keyed by id alone, so two courses can mint the same one -- a PPTX import numbers its scenes and actions deterministically, which gives every imported deck's first slide `tts_s1_speech-scene-p1`. Locally a collision only means one course plays another's clip in one browser; adopting it wrote that clip into the shared document permanently, for every device and every visitor. A row that names a course is now adopted only into that course, and a row from before that column existed only when the text it recorded is the text of the action being converted. A full asset store was made permanent last round, which was wrong three times over: it overwrote the refused bytes with an empty blob -- on the conversion path that row is a course's only copy of its own media -- it kept sending the rest of the deck to a provider against a ceiling it already knew was reached, and it left no way back once an operator raised that ceiling. A full store is neither the content's fault nor the configuration's, so it is now its own case: the bytes are kept, the pass stops at the first refusal, and the element shows the reason together with a Retry that re-attempts the upload from those bytes. Nothing retries automatically, so no one is re-billed. The narration write-back now reaches the write boundary every producer of a durable write passes through, not only the dirty mark: adoption never deletes the derived row, so a snapshot that reverts the rewrite is adopted again on the next load and allocates a fresh asset every time. Adoption is also mounted by both classroom surfaces rather than one, takes the course's abort signal, and re-validates that this browser still has the course open before each write. ASSET_QUOTA_BYTES is validated from instrumentation, where the README and the docstring already claimed it was: its only other consumer is lazy and memoised, so a malformed ceiling let the process boot and then failed every persistence request, documents and runtime included. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): remember a full store per course, and never lose retained bytes A stopped pass left the elements it never reached as placeholders with no persisted record -- deliberately, since nothing was attempted for them. But that left the next load with no reason not to try: it called a provider for the next placeholder and was refused at exactly the same point, once per reload, indefinitely. A full store is not a property of any slide. It belongs to the deployment and changes for reasons the document knows nothing about, so it is now remembered once per course in the browser's device KV. A pass that finds the marker stands down before spending anything and leaves every placeholder its "storage is full" state and its Retry; the first upload that succeeds clears it and the next pass runs normally. Narration adoption latched per course so it runs once per load, and the latch outlived the abort that leaving a course performs. On a surface that stays mounted across switches -- the workbench pane is one component for every course it shows -- owner course A, visitor course B, then back to A skipped exactly the clips the abort had cut off, and nothing else converts them. The latch is released with the abort now, and a course adopts one run at a time so a re-entry cannot hand a clip a second allocation while the previous run's uncancellable tail is still settling. A quota-blocked element retried into a network error or a 500 lost the bytes that were kept for it: the retry deleted the row before attempting the upload and wrote no replacement for an error carrying no structured code, so the next retry went back to a provider for media this browser had a moment earlier. The row now survives until an upload succeeds, the failure handler keeps whatever bytes the attempt was given, and the retry asks the question a pass asks -- does this browser already hold bytes for this element -- rather than reading an error code that a second failure has already overwritten. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): adopt real legacy narration, queue re-entries, report attempt outcomes Narration adoption admitted a stage-less row only when the text it recorded matched the action being converted. Both of those columns were added to the local audio table by the very change that moved narration onto allocated ids, so a row still carrying a derived key has neither: the rule refused every real pre-allocation course and passed only on fixtures built from post-allocation rows. What the row cannot say, the key can. A derived key names two clips only when two courses share a scene order and an action id, and an action id repeats only when something other than the generator minted it -- an import numbers them by slide position. So a key built from a generated action id is adopted on that basis, a key an import could have reproduced still needs matching text, and a row that names another course is refused however unique its key looks. Handing a re-entering caller the adoption run already in flight undid the latch release it was paired with: that run is bound to the signal the departure just aborted, so it stops at its next clip while the caller -- which has the course open and a live signal -- is told the work is done, and an effect replayed as mount, cleanup, mount adopts nothing at all. A later caller now waits for the uncancellable tail and scans again, which costs a lookup on a course that has nothing left and finishes the clips the abort cut off on one that does. One attempt at an element now reports both facts its callers need instead of a bare boolean: whether the store refused it for room, and whether bytes actually reached the store. Leaving a course clears the task table, so a retry that landed afterwards read "no failed task" as success and deleted the row holding the only copy of the media. Nothing is inferred from that table any more. Reading the localStorage property can throw where storage is denied by policy, typeof included, so the availability check moved inside the guard: this metadata is best-effort, and a rejection here strands a generation pass that has already enqueued its tasks. A retry is never blocked by the per-course "store is full" marker, but a retry that is refused again re-sets it, and adoption now reads and writes the same marker rather than issuing one refused upload per clip on every load. The two canvas element renderers and both thumbnail renderers show the reason beside the Retry, so a full store does not look like an ordinary failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): probe a full store instead of standing down, and pair notices with a Retry Narration adoption was given both halves of the per-course "the store is full" marker last round: it stood down when the marker was set, and it set the marker when its own upload was refused for room. Those halves are only safe together if something can lift the marker, and for adoption nothing could. It has no affordance of its own, it stood down before reaching its own clear, the media pass returns before its marker gate when there is nothing to generate -- so a narration-only deck, or one whose slides are already satisfied, painted no storage-full element and offered no Retry -- and narration generated rather than adopted allocates directly rather than through the media commit. The course's cached narration was then lost for good, where before it converted on the first load after the ceiling was raised. The gate is a probe now. A marked course attempts exactly one clip per load: refused, it stops and the marker stands, which costs what standing down cost; stored, it lifts the marker and finishes the course. Adoption spends no provider money, so the whole cost of probing a store that is still full is one refused upload. Generated narration lifts the marker too. The three surfaces that gained a failure notice last round drew it for any failure with a reason, including the one refusal that is reachable without server-backed persistence, so a browser-only deck painted something it had not painted before. The notice is drawn beside a Retry and nowhere else, which is what it was added for and what leaves browser-only output unchanged. Both are now asserted through the render harness the surface matrix already had. A caller arriving while a rescan is queued shares it rather than appending another. One rescan converts whatever the run in flight left and every later one would find an allocated id on every action, so a chain bought nothing and turned a single stalled upload into a course that never adopts again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): treat a refusal for room as a fact about one clip, not the deck The asset store checks each write against the headroom it has left, so a store that refuses a long opening clip can still hold every short clip behind it. Narration adoption assumed the opposite: it broke the deck at the first refusal and then re-attempted that same first clip on every later load, because the document names it first. A deck whose longest clip exceeds current headroom therefore never converted the clips that would have fit, with no affordance to recover it -- the state the probe was introduced to remove, reached through a narrower door. An unmarked load now attempts every clip, skipping the ones that do not fit, and remembers the condition only if the load ends with clips it still could not store. A marked load spends its single upload on the smallest clip left rather than the first one named: that is the clip that answers the question the marker asks, because if the smallest does not fit nothing does. The media pass keeps stopping at its first refusal, and for a reason adoption does not share -- every element it attempts costs a provider call. A rescan several callers share took the newest caller's signal, and the newest caller is not necessarily the one still there: a surface that opened a course and closed it again would stop work a surface still showing that course was waiting for, and that surface is latched, so it would never ask again. The shared run now takes a signal that is aborted only once every caller has left. The comment claiming the shared rescan contains a stalled upload was wrong -- the rescan is chained off the run in flight, so a stalled upload leaves every caller pending exactly as a chain would. It claims the bounded queue it actually provides, and the stall is recorded as a limitation. The failed-state containers took their stacking classes unconditionally, so markup differed in browser-only mode even though nothing moved on screen. Those classes are applied only when there is a notice to stack, and the tests assert the exact class attribute rather than a substring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): stop narration adoption writing the media pass's store-full marker The marker means "do not call a provider for this course". A path is entitled to write it only if its own refusal cost a provider call, and narration adoption's refusals cost nothing: it uploads bytes this browser already holds. The store also checks each write against the headroom it has left, so a clip that does not fit says nothing about whether a slide's image would. Adoption was writing it anyway, and one over-long narration clip was therefore enough to stand a course's entire image pass down on every later load -- on a store that had just accepted adoption's other clips. The author could still recover each element by hand, every load, for ever. Three rounds of narrowing this seam produced a finding each time, so it is removed rather than narrowed again. Gone: the marker read, the single-clip probe, the smallest-clip selection, and the up-front read of every row into an array -- which also retires a sampled-then-stale flag and the retention of a whole deck's blobs for the length of a run, and returns the loop to streaming one row at a time. Adoption's rule is now that every load attempts every clip it holds, once; any failure skips that clip and the load continues. The noise the coupling was meant to avoid does not arise, because after the first load the clips still outstanding are exactly the ones that did not fit -- normally none, or one. A successful write still clears the marker, and that is a different kind of statement: a write that went through is a fact this run established, where a refusal is an inference about what some other write would cost. For a course whose media needs nothing, adoption and generated narration are also the only paths that can establish it. The failure module still documented the deck-wide premise this contradicts. It now says what is true: the check is per write, and the media pass stops the deck as a judgement about cost rather than about certainty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bound a full store's cost from the store's own arithmetic Removing the store-full marker from narration adoption removed its bound too, and the code then asserted the bound was unnecessary. It is, on a store with room for most of a deck. On the store the whole mechanism exists for -- the ceiling reached, nothing fitting -- the outstanding set after every load is the entire deck, so a thirty-clip course posted thirty full blobs on every load, indefinitely. Each of those is not a cheap refusal: the bytes are uploaded, the server hashes the whole payload, and only then takes a per-principal lock and sums every entry that principal owns before saying no. The bound needs no flag, no key and nothing carried between loads. The store asks whether `used + addedBytes` exceeds the ceiling, and `used` only grows while a run is uploading, so a clip refused for want of room implies every clip at least that large is refused for the rest of that run. The run keeps the smallest size it has been refused and skips anything no smaller without uploading it; a smaller clip is still attempted, because it may fit. A deck the store refuses entirely now costs one upload per successive size minimum instead of one per clip, and a deck it has room for costs nothing extra, because nothing is refused. Only a refusal for room lowers the bar: a dropped connection says nothing about how much room there is. The deck-wide certainty premise the failure module retracted last round still stood verbatim at the site that implements the stand-down. Both copies now say the same thing: the check is per write, and the pass stops the deck as a judgement about cost rather than about certainty. The comment on adoption's marker clear now names its price. Narration of a few hundred bytes fits in headroom an image does not, so a proven write can let the next pass buy one more image that is refused again -- bounded at one, and the price of the alternative being a course whose media never generates again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(media): state the adoption bound exactly, and stop three comments describing the old rule The comment introducing the in-load bound gave its cost as "at most a handful, and the first load pays the most". Neither clause is a property of the rule. A clip is skipped only when something no larger was already refused, so a fully-refused deck costs one upload per successive size minimum in document order: one when the clips grow, about ln N for an arbitrary order, and one per clip when they only shrink -- a long opener followed by terser lines is exactly that shape. And no load is cheaper than the first, because the bound resets per run and a refused clip stays outstanding. The comment now says that, and points at what would make it exactly one for any ordering: the store returning its remaining headroom in the refusal's existing details channel, which the server leaves empty today. Two other comments still described the previous rule -- "attempts every clip it holds, every load" -- one of them twenty lines above the paragraph that introduces the bound, in the same block. Both now say what the code does. The bound's soundness is worth stating where a maintainer will look for it: quota is charged at full length with no discount for a duplicate, the sum it is checked against joins entries to blobs so the collector cannot lower it, the check takes a per-principal lock before summing, and replace and delete are refused to every browser. Nothing a run can do makes room appear inside it. One test installed a row implementation and replaced it wholesale a few lines later, so the first was dead and the survivor dropped the text the first clip's import-shaped key needs for the ownership rule -- it passed on the coincidence that the fixture's default text is the action's. Merged into one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): keep a Retry from re-buying parked media, and state the store seam once Five findings from an inline review. The standalone classroom route asked the ownership sidecar once per load and recorded only stage ownership when that ask failed. Every non-answer fails closed, so one transient 5xx left the genuine author with no resume, no Retry affordance and no legacy narration converted for the rest of the load, with nothing to change it short of a reload. The failure now records the fail-closed answer explicitly -- an answer an earlier load established must not outlive the failure that replaced it -- and an unresolved answer is asked for again, a few times over a few seconds. A real answer, however unwelcome, is final. A Retry could pay a provider for media the pool already held. When the bytes are stored and only the write-back fails in a way that keeps the allocation, it is parked and no local row exists, because that row is written only after a successful write-back. Retry now reads the parked queue exactly as the pass does and re-attempts the write-back: it re-keys the task done when the document takes it, leaves the entry parked when the slide still does not exist, and stays failed and retryable when the document refuses again. Object URLs a parked allocation owns are revoked when the entry is dropped. The commit path leaves them alone while the entry is parked, because it is then the only thing holding bytes this tab can render, so a course switch or a stage deletion was pinning the whole blob for the life of the tab. An entry a slide has already taken is left alone: the task table is displaying those URLs. The fallback lookup for cached bytes is a stage-scoped scan, and the keyed lookup misses for every row the commit path writes, so a pass was materializing and sorting the course's whole media table once per element. One scan per pass now, built on the first miss. It is sound and not merely cheaper: an element asks only for its own placeholder, and every row a pass writes carries the placeholder of the element that wrote it. "The store accepted a write, so it is not out of room" was enforced at three call sites under slightly different conditions, which made it a convention the next pool write path could silently break. It is stated once, in putAsset, for the course whose bytes it just stored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 16 天前 | |
feat(media): allocate generated assets through the registry (#1007 part 2, step b) (#1039) * feat(media): establish shared asset ownership primitives Introduce global browser asset-pool ownership, asset-reference collection, stage reclamation planning, and lease-based URL access. Carry allocated media identity through storage and generation boundaries with regression coverage. * fix(media): enforce safe resolution across every consumer Route image, video, thumbnail, presentation, and video-export consumers through one resolution state machine. Prevent opaque allocated or generated references from reaching render and export sinks, with fallback and ownership tests. * fix(media): protect document-owned assets across mutations Allocate pool bytes before compatibility writes and document commits, then roll back uncommitted generations safely. Preserve document ownership across edits, retries, imports, speech generation, scene changes, and stage deletion. * fix(media): scope retries and tighten ownership guard Scope retries to the target scene and slide and refuse ambiguous shared-reference mutations. Expand retry rendering coverage and keep direct pool URL resolution behind the shared lease owner. * fix(storage): avoid nested lock during stage cleanup Execute prepared reclamation plans against an explicitly deleted document so the compatibility cascade cannot re-enter the per-document lock. Cover deletion of stage-owned media rows even when the document has no references. * refactor(media): confine reclamation to stage deletion Remove inline pool and compatibility-row cleanup from element, speech, scene, and audio replacement flows. Keep whole-stage reclamation behind explicit stage ownership, preserve stage-less legacy audio rows, and document deferred document-truth sweeping. * fix(media): close retry and resolution gaps Restore shared source tasks after successful forks, scope retries across both whiteboard locations, and wait for parallel TTS workers before rollback. Resolve background media through import, rendering, and PPTX export paths while keeping retry controls visible over last-good bytes. * test(media): execute the consumer safety matrix Replace source-substring checks with resolver seam execution across all six UI consumers, stage hydration, and both exporters. Pin each rollback layer independently and harden the ownership guard against aliased pool imports. * chore(packages): publish the additive DSL field Bump the DSL patch version for the optional speech-action field. Keep the transitional reclamation policy app-owned and document the legitimate transaction rollback removals there. * fix(media): make retry rollback task-safe Delay allocation task re-keying until final document reconciliation succeeds, restore shared source tasks on failed forks, and require exact stage-whiteboard targets before falling through from a missed scene. * fix(media): clear private assets and refresh leases Delete the asset-pool database during the confirmed local-data wipe. Notify the app lease layer after same-id replacement so mounted consumers re-resolve current bytes without reaching into storage internals. * test(media): pin closure safety guards Exercise the CSS allocation boundary, unfiltered legacy-row ownership, delete-and-undo byte survival, and real distinct consumer seams. Restore the video-only manifest overwrite condition and share the direct video resolution hook across both element variants. * test(media): narrow rollback element assertion Narrow the reconciled slide element to an image before checking its source so the rollback regression remains type-safe under the full root compiler configuration. * fix(media): close lease refresh races Gate the first batch publication by unique resolved refs, then publish every replacement snapshot without mutating the prior React state object. Serialize invalidation behind pending releases, evict rejected refreshes, and register replacement observation at the pool boundary. * fix(media): reopen pool after clear failures Always evict the singleton once its store has been closed, including blocked and failed database deletion paths. Report blocked deletion as deferred and prove a later write uses a fresh live store. * fix(media): isolate shared retry progress Track forked regeneration under the selected element until it receives a fresh asset identity, leaving the shared source task and bytes untouched. Surface targeted failures through renderer task lookup and skip the redundant fork reconciliation lock. * fix(media): close final asset retry gaps Keep blocked asset clears fail-loud until a successful retry, with actionable settings guidance. Clear failed shared-fork state across durable and live key spaces, and pin renderer lookups, lease publication identity, and committed-ref rollback protection. * fix: preserve actionable retry failures Localize the blocked cache-clear recovery hint across every supported locale and pin the deferred-error mapping. Retain durable fork failure rows while retries run, deleting them only after successful generation so unstructured failures survive reload. * fix(media): hydrate legacy stored video thumbnails Home-page recent-video thumbnails regressed for legacy Dexie mediaFiles rows keyed by gen_vid placeholders: the reworked hydration resolved the row's bytes through the sealed resolver but never surfaced the stored blob (and its poster) as object URLs for the preview card, so the CI recent-video-thumbnail e2e specs found no visible element. Hydration now materializes legacy stored video rows into blob URLs for both the element src and poster while keeping the resolver invariants: opaque refs still never reach a DOM src, and concrete addresses are never blanked. Unit pins cover the seam so the vitest suite catches this class without a browser. * fix(media): resolve sole restored legacy video Classroom playback restored tasks by exact document media references. Legacy gen_vid references can outlive the key used by the one persisted video row, leaving the player on a placeholder even though bytes were restored. Select the sole completed stage video only for legacy sequential refs after exact and reconciled matches. Keep exact failures authoritative, refuse ambiguous candidates, and cover success, ambiguity, and failure precedence in unit tests. * fix(media): scope legacy video recovery Decide restored legacy video recovery once from the complete document and record it through the shared task lookup consumed by playback, editing, and resolved slides. Keep ambiguous documents as placeholders, apply the same decision to thumbnail hydration, and preserve exact failure precedence. * fix(media): exclude claimed video recovery tasks Model restored legacy recovery as a two-pass match across document video elements and task rows. Remove tasks claimed by exact, targeted, or placeholder lookup before applying the sole-candidate fallback, and cover the ownership/cardinality matrix. * fix(media): unify video element resolution Centralize source, task, poster, and legacy recovery decisions for every video consumer. Ensure direct URLs win over opaque refs and play_video waits on element-targeted retry tasks. * refactor(media): route legacy recovery through resolver Let document-aware consumers request legacy video recovery through the unified element binding API, keeping thumbnail hydration on a single decision path. * fix(media): preserve import refs and prefer pool bytes Recognize unambiguous extensionless relative media addresses during classroom import. Resolve allocated export and thumbnail assets from the shared pool before falling back to lagging compatibility rows. * fix(export): preserve concrete video sources Route PPTX video elements through the shared media binding resolver so an unresolved opaque reference cannot replace a playable source. Complete browser and persisted-store cleanup when asset-pool deletion is deferred, while retaining distinct hard-failure behavior. * fix(media): fork retries without exclusive ownership Enumerate logical asset owners across every persisted document before allowing global pool replacement. Thread explicit targets through fresh-id rewrites and cover cross-document aliases plus unreadable ownership. * fix(media): guard global asset reclamation Share a fail-closed persisted-document liveness check between stage deletion and retry replacement. Preserve cross-document pool aliases while deleting stage-owned compatibility rows and cover enumeration failures. * fix(media): cover complete slide asset references Route slide media traversal through a shared mutable slot contract so backgrounds participate in export, thumbnail hydration, collection, and rewrite lifecycles. Snapshot complete surviving-document refs once per reclamation and preserve manifest-only owners while retaining fail-closed behavior. * fix(media): preserve exclusive retry asset ids Allow targeted retries to replace exclusively owned pool assets in place. Keep shared and unprovable ownership paths on fresh allocations, and pin production-shaped retries plus compatibility-row cleanup. * fix(media): revalidate asset bindings at completion Recheck repository-wide ownership before replacing generated media and fork scoped retries when exclusivity changed. Route video export selection through the unified resolver and keep concrete posters independent of task state. * fix(media): count unflushed owners and broadcast replacements The completion-time exclusivity proof read only the persisted document, but slide duplication updates the Zustand aggregate synchronously and schedules persistence behind a debounce. A retry finishing inside that window saw a single persisted owner and replaced the bytes behind a reference the duplicate also held. The proof now also counts owners in the live stage snapshot when that snapshot represents the stage being retried, so an unflushed duplicate forks instead. Same-id replacement notifications were realm-local, so a second tab showing the same classroom kept its lease pinned to the superseded blob URL. The notification now travels over a BroadcastChannel; each receiving realm runs its own observers against its own pool, so a spoofed message can at most force a re-resolve. A missing or failing channel never fails the replacement. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bind replacement listeners and spare shared audio rows A realm that only renders never sends a replacement, so binding the channel from the sender path left passive tabs deaf to peers. Binding now happens where the observer is registered, when the asset-pool module loads, and the receiving realm resolves its own pool lazily so a cleared or unavailable pool degrades to the next resolve instead of throwing. Observer notifications ran under Promise.all, so a rejection surfaced after BrowserAssetStore.replace had already committed and turned a durable success into a reported failure. They are settled individually now; the callback is wrapped because a synchronous throw would otherwise escape before allSettled sees the array. audioFiles rows are keyed globally by audioId, so deleting a stage removed the sole row for an id a surviving document still referenced — playback and both export paths read that table directly and cannot fall back to the preserved pool blob. Rows are now filtered against surviving references, while a failed enumeration still withholds only the irreversible pool removal. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): serve replaced bytes from the pool everywhere Unknown survivor liveness deleted every planned audioFiles row. Those rows are keyed globally by audioId, and playback plus both export paths read the table directly, so losing one is as irreversible for them as removing the pool entry. Unknown liveness now preserves the rows too, leaving bounded garbage for a later pass that can prove exclusivity. The earlier pool-first change covered PPTX, video collection and thumbnail hydration but missed classroom ZIP export, which still serialized the stale compatibility row after a lagged same-id replacement, shipping media the classroom no longer renders. Auditing every direct reader of the media and audio tables surfaced the same gap in playback: speech regeneration also replaces bytes under a stable id and does not roll the pool back when the compatibility write fails, so the player kept serving superseded narration. It now resolves the pool first and falls back to stored rows for legacy and imported audio. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(audio): replace exclusively owned speech clips in place regenerateSpeechAudio accepted the action's audioId but always passed undefined as replaceAssetId, so regenerating an exclusively owned pool-backed clip allocated a new asset every time, rewrote the action and orphaned the previous pool entry and compatibility row until stage reclamation — contradicting the stable-id path generateAndStoreTTS already implements for media. Ownership is now established before synthesis, and the rule itself moved to the shared reference module so media retries, poster replacement and speech regeneration consume one implementation instead of restating it. An exclusively owned clip keeps its id and has its bytes replaced; a shared clip, a legacy id with no pool entry, or unprovable ownership still gets a fresh allocation so other holders keep their audio. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(media): drop the import left behind by the ownership move Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): resolve audio pool-first and fence peer realms Stable-id TTS regeneration commits replaced narration to the pool before the audioFiles mirror write, so a failed mirror leaves the row stale. AudioPlayer already resolved pool-first, but audioObjectUrl, collectAudioFiles and the video timeline dependencies still read Dexie directly and would serve the superseded clip. All allocated-audio readers now share one resolver, with Dexie kept as the fallback for legacy and imported rows. The exclusivity proof modelled unflushed owners in the active realm only, so another tab duplicating the same asset during its save debounce could still be observed as a single owner and have its bytes replaced globally. A peer's pending state cannot be read across realms, so presence is probed instead: any realm holding the stage forces the fork path. A deferred asset-pool deletion no longer reloads the page. The database is still on disk and the guidance asks the user to close the other tab and retry, which the reload discarded. The decision moved into a shared helper so the rule is pinned rather than living inline in the component. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): fail closed when presence cannot be probed The presence helper documented that an unanswerable probe must count as a peer, but every unavailable path — no BroadcastChannel, a constructor that threw, a send that failed, and the window before the pool's asynchronous binding completes — returned false, so the ownership proof cleared a single local owner and replaced globally shared bytes in place. Probing now returns present, absent or unknown, and only a probe that was actually sent and went unanswered is absent; the ownership decision treats unknown exactly like present. The pool declares its binding intent synchronously so a probe issued during the load-time window waits for the bind instead of concluding that presence is unavailable, and releases that gate if the import fails. Coverage reaches the write boundary: with presence unknown, a production-shaped targeted retry forks to a fresh id instead of calling replace, and the original bytes stay intact for a peer's unflushed owner. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(dsl): bump to 0.6.3 after the release dedupe took 0.6.2 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(i18n): add the deferred-clear guidance to fr-FR The locale landed on main after this branch added the key, so the alignment check flagged it as the one missing translation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 1 个月前 | |
feat(dsl): standardize the asset manifest and converge the export paths (#1007 part 3) (#1117) * feat(dsl): standardize the document asset manifest Add asset-manifest.ts to @openmaic/dsl: the canonical AssetManifestEntry shape (ref, kind, and byteSize/mimeType/duration/voice/prompt metadata where available) plus enumerateAssetManifest, the pure document-to-manifest enumeration. An entry's ref is the reference exactly as the document holds it -- the manifest is the id-based reference enumeration with metadata, not a content hash and not a resolution result. The traversal walks the stage whiteboard, each scene's canvas/whiteboards/speech actions, and the stage video-manifest keys in document order, with logical-owner reference counts that match the accounting duplication-safe replacement uses. This settles the media-ref + asset-manifest schema question (#779 open question 4) on the side the asset-pool RFC already implied: the schema is a function of the id semantics decided there. The type lives in the dsl rather than a new @openmaic/exporter package because the enumeration is pure over document types the dsl already owns (Stage/Scene/Slide/Action), so a separate package would add a published artifact and release-workflow surface without adding a capability; the storage contract comment now points at the module. Refs #1007 * refactor(export): drive the classroom ZIP from the asset manifest collectMediaFiles used to scan the whole mediaFiles table for the stage, so any row the document no longer references -- an orphan left by an edit or a superseded regeneration -- rode along into the archive. Both ZIP collectors now take their reference sets from the standardized asset manifest (buildStageAssetManifest wraps the dsl enumeration with the compatibility rows' metadata): only referenced assets are archived, and a referenced asset whose bytes exist only in the pool is still collected via a synthesized record. Byte resolution is unchanged: pool first through resolveStoredBytes / resolveAudioBlob, with the compatibility row kept as the legacy byte fallback and as the metadata source. mediaIndex is now a serialized view of the manifest, and the missing-audio report derives from the manifest's audio entries instead of a second action walk. The audioRef mapping and the legacy audioUrl fetch path (collectLegacyAudioForExport) are untouched. Refs #1007 * refactor(video-export): take the timeline's reference sets from the manifest createVideoTimelineDeps scanned the whole mediaFiles table for the stage and derived its audio id set from its own action walk -- a third, independent answer to "which media does this course use?". Both record loads now key off the standardized asset manifest: media rows are read per manifest ref by compound key instead of by table scan, and the audio id set is the manifest's audio entries. Orphan rows were never reachable through the scene-scoped elementId-to-mediaRef bridge; now they are not even read. The bridge itself is untouched: element ids recur across scenes, so the elementId-to-mediaRef mapping stays scoped per scene, and the legacy audioUrl fallback keeps its own action walk because a URL is not a manifest ref. AssetPlan remains the video IR's view of the same references. Refs #1007 * refactor(export): resolve PPTX media through the shared resolver only Each PPTX element branch carried its own resolution chain: a task-state renderable-URL lookup first, then -- gated on the legacy placeholder predicate -- a stored-bytes override, with the poster block repeating the pattern. One helper now owns resolution for backgrounds, images, video / audio sources, and posters: opaque refs (allocated ids and legacy placeholders alike, no placeholder-pattern gate) resolve pool-first through resolveStoredBytes and embed as data URLs, concrete addresses resolve through the media state machine and keep the caller's fetch path. exportMediaResolution and the resolveStoredMediaBlob wrapper fold into the helper; resolvePptxMediaBinding stays as the state-machine entry the resolution-surface test matrix drives. Refs #1007 * refactor(export): retire the export-side Dexie byte fallbacks Export call sites no longer read bytes off compatibility rows directly. The ZIP collectors and the video timeline's audio load resolve bytes only through the shared resolvers (resolveStoredBytes / resolveAudioBlob), which answer pool-first and keep the compatibility row as their internal legacy fallback level; the row reads that remain at the call sites supply metadata (format/duration/voice/mime/size/prompt) only. The rows themselves stay for legacy and regeneration readers -- what goes is the export paths' own fallback logic. One observable tightening: a failed media row (error set, empty placeholder blob) no longer ships a 0-byte file into the classroom ZIP, and an evicted row no longer ships its empty local blob; referenced-but- byteless assets are simply absent from the archive, as they already were when no row existed. Refs #1007 * test(media): cover the enriched stage asset manifest builder Pins the join between the pure dsl enumeration and the compatibility rows: metadata attaches by ref, rows no document reference names never appear, and a referenced asset with no row keeps a metadata-free entry. Refs #1007 * fix(video-export): widen the deps stage input for the manifest enumeration enumerateAssetManifest reads the stage's whiteboard and videoManifest, so createVideoTimelineDeps declares them on its input instead of the bare id; callers pass only the id today and the optional fields stay absent. Also applies the repo prettier formatting to the files this branch touched. Refs #1007 * fix(dsl): enumerate slide audio elements in the asset manifest Slide audio elements carry their own src, and the manifest skipped them, so a manifest-driven collector could never archive their bytes. The audio slot maps to kind 'audio' alongside narration ids. Refs #1007 * fix(media): harden ref-keyed lookups against prototype-named asset refs AssetRef is an unconstrained string alias, so a media reference can legitimately be "__proto__", "constructor", or any other Object.prototype member. Plain objects keyed by such refs silently drop assignments or answer lookups with the prototype object, which rewrite paths then accept as a mapped id. Convert the remaining ref-keyed lookup tables introduced by the export convergence to prototype-safe structures: the classroom import media/poster alias maps and the legacy-conversion video-manifest reconstruction now use Map / null-prototype containers with explicit membership checks, and every consumed value is validated as a string before it is written into a src / mediaRef / audioId slot. The shared media-task lookup receives the same treatment: one centralized own-property-checked lookupMediaTask now serves the stored-bytes resolver, the PPTX embeddable-src path, the video collection path, and the element/background task resolution, so a prototype-named placeholderRef can no longer hide a re-keyed task from the fallback chain. Adversarial tests drive "__proto__" and "constructor" refs through the import round trip, the PPTX fallback path, the legacy conversion commit path, and the media-task fallback end to end, including a buildPptxBlob regression with a task re-keyed to an allocated id while retaining a prototype-named placeholderRef. * fix(export): use safe archive asset paths * refactor(dsl): centralize slide media slot roles * refactor(export): derive consumer refs from manifest * fix(export): sanitize classroom archive extensions * fix(video-export): preserve narration speech order * fix(export): enforce kind-coherent archive media * fix(export): define media coherence boundary * fix(export): carry task-owned poster binding for PPTX export A video element with no explicit poster falls back to its media task's generated poster URL, but resolveVideoMediaForElement left posterTask undefined for that case, so the PPTX manifest guard saw a foreign URL with no task-ownership exemption and dropped the video element instead of using the established runtime poster fallback. Carry the poster task binding whenever the task poster is the effective poster: the task-owned URL then satisfies the guard's objectUrl exemption end to end. A concrete explicit element poster still stays element-owned and never borrows the binding, and the guard's foreign-ref rejection is preserved (and exported as a directly testable predicate). Coverage: an element with no poster plus a task-provided poster embeds the task poster as the PPTX cover (red at the pre-fix head, green now), and a genuinely unrelated URL with no task ownership is still rejected by the guard. * fix(export): preserve legacy narration source refs in the media index The explicit sourceRef contract was partial: primary audio and generated media entries carried it, but legacy URL narration serialized no source ref. The legacy URL itself is the natural source ref — it is known at fetch time — so wire it through the collected blob into the mediaIndex entry. Import already registers serialized sourceRefs as aliases, so the URL now round-trips as an explicit mapping instead of being reconstructed only from the action's audioRef. Poster siblings are deliberately NOT given their own mediaIndex entry: a sibling poster (media/asset-<n>.poster.<ext>) is a legacy byte copy written from the video record and is not an independently referenced document asset — when the poster is a real document asset it already has its own indexed entry with a sourceRef, and import reconstructs the sibling by path derivation from its parent video entry, reusing the poster's own indexed allocation when one exists. The PR description is narrowed to match; corrected paragraph: "Archive names never interpolate refs — sequential safe paths (media/asset-<n>.<ext>, audio/audio-<n>.<ext>) with the original ref preserved through an explicit sourceRef mapping on every independently indexed media entry: generated media assets, poster assets, primary narration, and legacy URL narration (the legacy URL itself is the entry's sourceRef). Extensions are allowlisted per kind. The one exception is the legacy sibling poster byte copy (media/asset-<n>.poster.<ext>, written next to its video when the video record still carries the pre-pool poster bytes): it is not an independently referenced document asset, so it has no mediaIndex entry or sourceRef of its own — its identity is derivable from its parent video entry (same index), and import reconstructs it by sibling-path derivation from that video entry, reusing the poster's own indexed allocation when one exists." --------- Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 1 个月前 | |
feat(token-plan): add TokenDance one-key preset for every modality (#1525) * feat(token-plan): add TokenDance one-key preset for every modality TokenDance is a model gateway: chat and images are OpenAI-compatible at /gateway/v1, and the same key authenticates vendor-protocol routes on the same host (Ark, MiniMax, Bocha). The preset reuses the existing adapters with those route prefixes as base URLs, so one key lights up LLM, image, video, TTS and web search from Settings -> Token Plan. - providers: add a built-in `tokendance` OpenAI-compatible provider (TOKENDANCE_* env prefix, logo, provider name in all locales) - token-plan: add the TokenDance preset (Seedream image, MiniMax H3 video, MiniMax speech TTS, Bocha web search) - seedream: use a base URL that already ends in a version segment verbatim, so gateway routes like `/ark/v3` do not get `/api/v3` appended - minimax-video: route H3-family models through the v2 task API (content array submit, task-envelope poll); connectivity checks for H3 probe auth on the v2 query route instead of submitting a billable task - README: add a one-key quick example and replace the Gemini-specific model recommendation with a provider-agnostic setup recommendation Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qrrq9CPwb718mpouz8Y2KL * fix(token-plan): accept preset web-search base URLs and report H3 dimensions per ratio - web-search: the client base URL allowlist also accepts the exact base URL a built-in token plan preset writes for that provider, derived from TOKEN_PLAN_PRESETS. Applying a plan whose web-search route is not an official vendor host previously stored a URL that the route rejected with 400. Any other client URL is still rejected. - minimax-video: report H3 v2 clip dimensions for 16:9, 9:16, 4:3 and 1:1 instead of assuming landscape for every non-portrait ratio. - tests: pin the allowlist for every preset, the 1:1 H3 dimensions, and clear TOKENDANCE_* in the provider-config env isolation list. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qrrq9CPwb718mpouz8Y2KL --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> | 11 天前 | |
feat(media): add OpenRouter image and video providers (#1356) * feat(media): add OpenRouter image and video providers OpenMAIC ships six separate video providers (Veo, Kling, Seedance, MiniMax, Grok, HappyHorse) and seven image providers, each needing its own key. OpenRouter fronts those same model families behind one key and one account, so this adds it as a provider on both sides. Both use OpenRouter's dedicated media endpoints, not chat-completions: - Image: POST /images -> { data: [{ b64_json }] } - Video: POST /videos -> 202 { id, status }, poll GET /videos/{id}, then GET /videos/{id}/content for the mp4 bytes The model list is fetched live from GET /images/models and GET /videos/models through /api/openrouter-models rather than pinned in the registry: OpenRouter hosts 48 image and 28 video models today and adds more, so a hardcoded shortlist would decide for the operator which models exist. The registry keeps a three-entry seed as an offline fallback, and the existing custom-model UI still accepts any model id. Both catalogs answer unauthenticated, so the picker fills before a key is pasted; a key is forwarded when present for proxied base URLs. Adapter contracts are covered by stubbed-fetch tests (request shape, empty-response handling, and the video job state machine including terminal failure). No test performs a billable call. Closes #1355 * fix(media): validate the key and tolerate a pasted endpoint URL Three fixes found while configuring the new provider: 1. Both connectivity probes hit the model catalogs, which answer 200 unauthenticated — so "Test Connection" reported success for any string, including an invalid key. Probe GET /key instead: equally cheap, and it actually rejects a bad key. 2. The settings field is labelled "Base URL" but the panel echoes it back as "Request URL", so pasting the full endpoint (https://openrouter.ai/api/v1/images) is the natural mistake. That built /api/v1/images/images and 404'd. Trim a trailing slash and a trailing /images or /videos so both forms work; a proxy path that merely contains the word is left alone. 3. The image and video settings panels read `data.message` on a failed test, but failures answer with `error` (apiError) and only successes carry `message`. Every failing connectivity test — for any provider, not just OpenRouter — rendered "connection failed: undefined" instead of the reason. Pre-existing; surfaced by 1 and 2 above. Closes #1355 * fix(media): make every OpenRouter model selectable, and always select a provider Two gaps found while configuring the new provider. The settings Models list is a read-only catalog for every provider; the actual model picker is the media popover. That picker built its groups from the static registry array, so OpenRouter offered only the three-entry seed while settings listed the full live catalog — the models were visible but not choosable. Feed the same live catalog into the popover, fetched only once the provider is usable so an unconfigured install makes no request. Separately, `imageProviderId`/`videoProviderId` are empty until a provider is chosen (first-run auto-config leaves them blank when the server reports no media provider). Opening the settings panel on an empty id selected nothing: the header rendered the missing name key as "settings.undefined", and Test Connection posted a blank x-image-provider/x-video-provider, so it failed with "No image/video provider configured" whatever key was typed. Fall back to the first catalog entry so the panel always has a selection. Pre-existing and not specific to OpenRouter. Closes #1355 * fix(tts): request a browser-playable format from custom providers `generateOpenAITTS` serves every custom OpenAI-compatible TTS provider but never sent `response_format`, so it inherited whatever each provider defaults to. OpenAI defaults to mp3; OpenRouter's /audio/speech defaults to raw `pcm`. The unknown content type then fell through to the `'mp3'` default below, the client built `data:audio/mp3;base64,…` from headerless PCM samples, and playback failed with "no supported source was found" — while the server logged a clean 200, because the audio really was generated. Name the format instead of inheriting it. Also stop mislabelling an unrecognised body: `pcm`/`l16` now raises a message naming the cause, and `aac`/`opus` are recognised. Two supporting fixes: - /api/openrouter-models normalises its base URL the way the adapters do and falls back to the public catalog when a custom base URL fails, so a typo in a free-text settings field cannot empty the model picker. Also types the headers object so tsc accepts the conditional. - provider-neutrality-guard pins exact per-vendor occurrence counts in lib/server/provider-config.ts. Adding the image and video env entries raises "openrouter" from 2 to 6 (each entry contributes both its key and its value); CI failed without the bump. Closes #1355 * fix(security): never send the operator key to a client-chosen host Review found `/api/openrouter-models` was an SSRF and key-exfiltration path, and the finding is correct. The route took `x-base-url` from the caller at highest precedence while preferring the *server* env key, so any caller could make the server send the operator's OpenRouter credential as an `Authorization: Bearer` header to an arbitrary URL. The route's own comment claimed it followed `/api/verify-image-provider`; that pattern runs `validateUrlForSSRF` on client base URLs, and this route did not. The boundary is now explicit: the server key travels only to the operator's own base URL. A client-supplied URL is SSRF-validated and carries only that caller's own `x-api-key` — the server key is dropped — and the unauthenticated public-catalog fallback never forwards a credential chosen for a different host. Redirects are no longer followed (`redirect: 'manual'`), since a redirect would carry the Authorization header off-host and reopen the same hole, and upstream reads are bounded by a timeout. The per-URL cache is now keyed by destination *and* a hash of the credential, and bounded to 64 entries with oldest-first eviction, so client-supplied URLs cannot grow it without limit and one caller's key-authorised catalog is never served to another. Also from the review: - The image adapter discarded the reported `media_type`. The orchestration layer wraps a bare `base64` as `data:image/png` unconditionally, so jpeg/webp results were mislabelled; the adapter now returns a data URL carrying the real type. - Adapter generation and poll requests set `redirect: 'manual'`, matching the `/key` probe that already did. - `runPolledTask` accepts an `AbortSignal` so the sleep between polls is cancellable; the video adapter passes the caller's signal. Without it a cancelled generation still slept out a full 10s interval. Tests cover the highest-risk paths the review named: which credential reaches which URL, that an SSRF-rejected destination is never contacted, that the fallback is unauthenticated, cache isolation between callers, and MIME preservation. Findings 2 (base-URL normalisation) and 4 (neutrality-guard debt) were already fixed in d553a08, pushed after the review was submitted; CI is green on that commit. Closes #1355 * ci: retry flaky voice clone timeout * fix(vercel): keep OpenRouter catalogs within Hobby function limit * fix(vercel): avoid tracing self-hosted sharp binaries --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> | 10 天前 | |
feat(agent-runtime): workbench generate_image / generate_video write through the asset pool (#1007 part 6) (#1524) The workbench tools store generated bytes in the asset pool under the shared principal and write the allocated id into the document, the same discipline as the classic chain since #1392; the runner's putScene creates the reference rows and commits the allocations. The video completion patch rewrites every placeholder slot and retires anything that would shadow the new id; immediate render is preserved by leasing the id at the render boundary. A store-full refusal fails the tool with a model-readable error and writes nothing. Legacy /api/classroom-media documents keep rendering. Closes #1522. | 11 天前 | |
feat(media): write generated media through the asset pool under server-backed persistence (#1392) * feat(media): store generated media in the asset pool when persistence is server-backed With server-backed persistence the document is durable and shared, but generated media stayed in the producing browser: the document kept its gen_img_* / gen_vid_* placeholder and narration kept a browser-derived audio id. Every new browser that opened such a course re-ran generation for every slide, and it never converged, because the address of the generated bytes was never written back into the document. Under server-backed persistence only, the classic generation chain now stores bytes in the asset pool first and writes the id the pool allocated into the document. - The client bootstrap configures the asset seam alongside the document and runtime seams: an HttpAssetStore over the persistence endpoint carrying the same credentials the document store carries, marked server-backed. The seam preflight now covers all three, so a failure still cannot half-configure persistence. - Image, video and TTS generation commit in one fixed order: provider, pool, document, local cache, task. A reference reaches the document only after put returned an id, so a document can never name bytes that were not stored. A failure before the write-back leaves the placeholder with the provider called exactly once; the retry happens on the next owner load. - The write-back is a per-slot rewrite through mutateDocument, which re-reads the current document under the per-stage lock, so it cannot clobber a newer scene. The open course is refreshed with the same rewrite without being marked dirty. - "Has this already been generated?" is answered by the document (the slide exists and no longer holds the placeholder) instead of by this browser's task table. - The classroom's resume effect fails closed on ownership: only a resolved owner starts generation, so a viewer opening a shared course spends nothing. - The local media and audio tables become a per-tab cache. A failed cache write costs a re-download, never the media. Browser-only mode is unchanged: every new call site sits behind the server-backed gate, the local tables stay authoritative there, and placeholders stay in the document. Rendering and export needed no changes. HttpAssetStore.resolve mints an object URL exactly as the browser store does, and the export byte resolver was already pool-first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(classroom): make the generation owner gate a three-outcome rule and apply it everywhere The gate refused everything but a resolved owner, which read a sidecar that answered "no ownership fact exists for this course" as a reason to block. That is the answer a deployment without the sidecar's server-side prerequisites gives for every course, and the answer a course with no ownership record gives: in both, there is nobody the operator's budget needs protecting from, and refusing strands the course's own author behind a question that can never be answered. Ownership is now four states over the sidecar's three outcomes. A definite answer splits into owner and not-owner. An absent record is its own answer, ownerless, and generation proceeds — the behaviour such a deployment had before the gate existed. Only the absence of an answer, a transport failure or a load that has not asked yet, stays unresolved and fails closed: "we could not ask" must never be read as "nobody owns this". One mapper turns a sidecar result into that state, and one predicate decides on it. The workbench classroom pane runs the same resume effect and had no ownership input at all, so a viewer opening a shared course there could still spend the budget. It now asks the sidecar once per course, in parallel with its load and feeding only the generation gate, so its read-only and edit behaviour is unchanged. The shared progressive-load policy carries the gate for it, with both new inputs required rather than defaulted so a future caller cannot omit them into an open budget. Its stale comment claiming ownership could not be expressed here is corrected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the write-back survive autosave, arrive before the scene does, and never leak Independent reviews of the write-back found three ways a durable document could still end up naming a placeholder, and two ways the gate that protects the operator's budget could be walked around. An autosave round captures the store synchronously and writes that capture, so a round already in flight when a rewrite landed wrote the placeholder straight back over the allocated id, and nothing marked the store dirty again to correct it. The rewrite now marks the units it changed, which leaves a corrective flush queued behind the stale one; re-saving a scene that already holds the id is idempotent, losing the id is not. Media is generated from outlines in parallel with scene content and usually finishes first, so the slide that will carry the placeholder does not exist yet and the write-back has nothing to rewrite. That was the ordinary path, not a tail case, and its result was discarded: the task was marked done, the scene was added afterwards with its placeholder intact, and a second pass in the same run could call the provider again. The allocation is now held under the placeholder — which also answers the skip test, so nothing pays twice — and applied when that scene is committed, before its first save. One complete pass now leaves no placeholder behind. A failed commit used to abandon what it had already allocated. A poster upload that failed threw away a stored video and sent the retry to submit the most expensive job in the system again; a rejected write-back left registry rows that name bytes nothing references, which the byte collector cannot reclaim because it only collects blobs no row names. A poster failure now costs the poster, and a write-back that reached nothing reclaims what it allocated. A partial write is left alone, because the document already names it. The ownership gate is fail-closed again. Treating the sidecar's 404 as permission was wrong: the client cannot tell "this course has no owner" from "this deployment told me nothing", so a visitor who opened a shared course could bill the operator. The root cause was the sidecar itself, which gated on the agent runtime although every persisted course has an owner regardless — the persistence route resolves one for every request. It now gates on server persistence, so the configuration that made 404 the universal answer has real ownership facts to report, and the gate can refuse everything but a named owner. Retry affordances answered to no gate at all. A viewer of a shared course with one failed image was shown a Retry button that called the provider. Both retry entry points and every surface that draws them now read one shared permission, so what is offered and what is allowed are the same value. Also: narration regeneration no longer pretends it can replace bytes behind a live id — the exclusivity proof that would allow it is refused by construction once references leave the browser, so it forks to a fresh id and says so; the "already generated" test lets a finished deck answer from the document alone, since scene order stops identifying an outline once slides are inserted or deleted; stored assets record a specific media type rather than a generic transfer type; the pane no longer asks the sidecar in browser-only mode; and the funnel's docstring now states what the per-stage lock actually guarantees, which is same-browser serialization and not a cross-browser compare-and-swap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): park allocations in the deciding turn, never reclaim on an ambiguous write A delta review of the write-back found the first-pass fix still had a window, and the reclamation it added could delete media the document already names. The allocation was parked after an awaited local cache write. A scene committed in that window reconciled against a registry that did not hold it yet, so the document kept the placeholder — and the entry recorded a moment later then answered the skip test as "already handled", so nothing could correct it. Parking now happens inside the write-back, in the same synchronous turn as the decision that nothing could take the reference; no await separates the live check from the park. Allocations parked by an earlier pass are handed to their slides at the start of the next one, before anything decides what still needs generating, so a held allocation whose scene has since arrived becomes a rewrite rather than an answer. Reclaiming on a rejected write was unsound: a rejection does not prove the server did not apply the write, so deleting the asset could break the scene that now names it. The funnel decides instead, and says so: it reclaims only when no store write was ever issued and nothing took the reference. Anything else is placed if its slide exists and parked if it does not, so the next pass reuses the bytes instead of paying for them again. When a write fails after part of it landed, the live store is brought up to the document before the error is rethrown — otherwise the next ordinary flush would overwrite the half that did land, with the ids deliberately not reclaimed. Parked allocations are now cleared with the course. Classic placeholders are reused across runs, so one surviving an interrupted run would be handed to a different slide of the next deck: the previous picture, on a slide whose provider was never asked. Both classroom surfaces clear the arriving course, the deletion cascade clears the deleted one, and clearing the database clears them all. Two more ways generation could start without asking the gate are closed. An overlapping pass — an outline retry re-enters generation with every outline while the first is still working — re-requested elements whose provider call was already in flight; a task that is not done is an answered request, not an unanswered one. And narration regeneration in the timeline editor called the TTS provider and allocated a pool asset with no ownership check at all; it now reads the same permission, which withholds both the per-line and whole-timeline controls and refuses the call. Finally, a pane opened during the stage-link availability gap recorded the sidecar's 404 for a course that was moments from existing and never asked again, leaving the real owner locked out of generation until it remounted. Ownership is re-fetched once the document becomes available; the gate stays closed until an answer arrives, so asking again can only open it for someone entitled to it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the asset routes reachable, and stale snapshots harmless A full-branch audit found that the deployment this project documents could not store a single generated asset, and that several routes into durable storage could still write a placeholder over a reference that had already landed. The persistence route sent asset requests through the development authenticator, which refuses outright in a production build that has not explicitly opted into it — and the documented server-persistence recipe produces exactly that build. Every store and every read answered 401, so images and video failed on every slide while re-billing the provider on each retry, and a narration failure stopped the deck at its first slide. Assets live in one shared partition by design, so there was never anything per-caller for that authenticator to decide: the route now resolves the asset principal itself, alongside the owner it already resolves for documents. Runtime sessions are genuinely per-learner and keep the development authenticator until real session verification replaces it. And narration that cannot be stored no longer fails its scene: the line stays unvoiced and retryable, which is what an image that cannot be stored does to its slide. Placeholders could also come back from behind. A queued autosave's snapshot, an editor-history entry replayed by an undo, the departing save a course switch flushes — each captures content at its own moment, and any of those moments can predate a write-back. Point fixes at each producer would leave the next producer to rediscover the bug, so the check lives at the write boundary every producer passes through, and the allocation record it consults now outlives the parked queue: a placeholder whose rewrite landed long ago is exactly the case it catches. Two ways generation could be lost or repeated are closed. A pass now claims the elements it will reach and releases them however it ends, so an overlapping pass stands down while an aborted one strands nothing — previously its tasks stayed `pending` and every later pass skipped them with no retry control to recover them. And the media abort controller is aborted before being replaced, so a superseded pass stops calling providers instead of running on for a course the user has left. The remaining two are narrower. The workbench pane asks for ownership only after a document load succeeds, and after every later one, mirroring the page route: the load is what creates the ownership row the first time a course is opened, so asking beforehand asked about a course that did not exist yet and locked its author out for the mount. And the ownership gate on the timeline editor now withholds narration regeneration alone; listening back to existing narration and seeing whether a line has any spend nothing and stay available. Known limitation, unchanged and now stated plainly in the comments that used to point at it as a solution: nothing reclaims an unreferenced pool asset. The registry sweep is written but not wired up, and the byte collector only reclaims blobs no registry row names, so every narration regeneration and every abandoned allocation leaves storage behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): gate asset mutations, and make claims and allocation records survive a handoff Opening the asset routes opened all of them. Reads and allocations are meant to be as open as document reads and creates already are, but no authorization hook was supplied, so the handler's default admitted PUT and DELETE too — and those scope by principal key alone, which is one shared constant. Any caller who learned an id, and a document read hands out every id its slides name, could overwrite or destroy another author's media. Mutations now require the deployment's credential, which in a production build without the development-auth opt-in means they are refused outright; reads and allocations stay open. The route comment says what the posture is and what it is not: the deployment-level fence is the access code, and no per-principal quota is configured. The client's own reclaim is best effort to match — losing an argument about deleting an asset must not cost a task its retry, and the bytes are left for server-side reclamation. The pass claim could not survive the handoff it was written for. A retry aborts the live media pass and starts its replacement in the same synchronous block, long before the aborted pass's cleanup runs, so the replacement saw every element still claimed, collected nothing, and returned — leaving each unreached element at pending with nobody coming back for it and no retry control to recover it, which is the exact failure the claim was introduced to prevent. A claim now carries its pass's signal and is retired the moment that signal aborts, and a pass releases only claims it still owns, so a late unwind cannot take its replacement's work. Claims are also acquired at the single point every request passes through, so a single-task retry participates too — previously a retry awaiting its provider was invisible to a pass starting alongside it and both called it. The allocation record could outlive the bytes it named. It was written before the write-back attempted anything and survived the reclaim that followed a failure, so when the slide finally arrived the write boundary stamped a deleted id into the document — and the placeholder it replaced was gone, which reads as already generated and stops anything from retrying. The record is now written only where the allocation is retained, and forgotten wherever a reclaim removes the bytes, including the narration rollback path. The tests follow. The route test drives the real storage handler against an in-memory registry instead of a stub, so it can see what the resolved principal is then allowed to do; the handoff test performs a real abort mid-pass rather than starting from an already-aborted signal; and the guards that could only assert file layout now assert the property they care about, or have been replaced by behaviour. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make media passes serial per course instead of tracking element ownership Three rounds of per-element claims each produced a new way to lose an element. Whole-pass reservations swallowed a Retry for an element the same pass had already failed, leaving it pending with the affordance gone. Retiring a claim by its signal freed an element whose commit was still uploading, so the replacement pass paid for it twice. A claim held for a failed element stranded its retry. The bookkeeping is the defect: every refinement of "who owns this element right now" answered the question at a moment when the answer was already stale. Passes for one course are now serial. A replacement aborts its predecessor, as before, and then waits for it to settle before collecting. That removes the question entirely: a commit already under way finishes — its bytes stored and its reference written, so the new pass sees a resolved slide and skips it — and an element the aborted pass never reached is still a placeholder and gets collected like any other. The claim set, the reservations, the signal retirement and the identity-checked release are all gone. The task table is consulted for one thing only: an element that is generating right now is a single-element retry running alongside the pass, and taking it too would pay twice. Pending is deliberately not a skip reason — it means a pass once intended to reach an element, which an abandoned pass leaves behind with nobody acting on it, and reading that as answered is what stranded elements before. A retry runs concurrently with a pass, because a pass never revisits an element it has processed, and it re-reads the task after its own await and refuses before touching it: marking first and refusing afterwards destroyed the failed state that draws the affordance. Browser-only mode is back to exactly what it was. The abort is now conditional, the waiting does not apply, and the original status-based skip is restored verbatim. Two baseline lines remain changed in each of the two files, and both are behind a server-backed fork whose else-branch is the original. Two smaller things. The allocation record becomes visible when a write goes on the wire rather than when the round trip ends, and the write boundary reconciles under the document lock rather than before it — a save queued during a write-back was otherwise captured with the placeholder and, for a course the user had left, had no corrective flush to follow. And the comments that said a refused reclaim leaves its bytes for server-side reclamation were wrong: nothing collects them, because the registry entry still names its blob and the sweep that would remove it is not wired up. They now say the bytes leak. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a deferred pass re-earn its right to run, and bound the commit it waits on Serializing passes moved their body out of the block that launched them, and three things followed from that. A pass now wakes when its predecessor settles, which can be after the user has left the course. It enqueued before it looked at its signal, into a task table keyed by element id alone — and placeholder ids are not unique across courses, which is why the classroom clears that table on arrival. So a departing course's pass seeded the arriving course's table with tasks carrying the wrong stage id, and a Retry routes by that id: the reference went into the wrong document. A pass now re-validates after the wait, before touching anything shared. The same lateness broke the skip test. `documentSkipIndex` answers only while the live store is on the pass's stage, and returning nothing put the collection loop on the browser-only rule — a silent demotion from "the document is the authority" to "this browser's task table is", on exactly the path where that table has just been cleared. Every element the predecessor had committed was collected again, paid for again, and its second write-back found no placeholder to rewrite, so its bytes were parked where nothing will ever reference them. In server-backed mode an unreadable document now means the pass stands down. And waiting was unbounded. A commit is uncancellable: the asset client takes no signal, and a document write cannot be half-undone. One stalled upload therefore froze the course's media generation for the session — the replacement never collected, the element sat on a skeleton that draws no Retry, and only a reload recovered. The pass's signal is now threaded into the media proxy fetch, and the commit is bounded by a deadline. The deadline is on the wait, not the work: the commit carries on, and if it lands late the document simply ends up correct, while the element becomes retryable and the queue moves on. The tests that were meant to pin the previous round were not sensitive to it. Two asserted end states where the mechanism only changes ordering, and one of them rigged the document read so the assertion held whether or not the pass had waited; a third covered half of what it claimed. They now observe the ordering directly — nothing is issued while another pass for the course is working; in browser-only mode a second pass reaches its provider immediately — and the reconciliation under the document lock has a test that fails when it moves back outside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * revert(media): drop the commit deadline and the abortable download The deadline bought less than it cost. Abandoning a commit after two minutes makes the element retryable while the real commit is still running, so a Retry starts a second commit for the same placeholder against the first: two provider calls, two allocations, and whichever lands second stamps its result over the other's task by element id. The allocation record is keyed by placeholder, so the loser's cleanup erases the winner's record, and the write boundary then puts the raw placeholder back into the document. That is the overlap serial passes were built to remove, reopened through the one door serialization never covered. So a stalled commit holds the course's media queue until it settles or the page is reloaded, and that is written down rather than papered over. The wait is unbounded on purpose: every ceiling on it turns out to be a way of running two commits for one element. Threading the pass signal into the download was also a mistake, in the other direction. The provider call that produced the URL has already been billed, so cancelling the download throws away work that is paid for — and the shared proxy cache records a cancelled request as a transient failure against that URL, which after three of them blocks it for every consumer in the session. Browser-only mode never asked for this: it had no way to observe an abort there, which is exactly why the bytes were kept. The signal is gone from the download again, and `fetchAsBlob` is byte-for-byte what it was before this branch. The regression guard for the stranded-element rule is restored alongside the timing test that was meant to supersede it. It catches a different rule — a task left pending being read as answered — and nothing else does: making the pass skip pending leaves every other suite green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): stop asking the pool for refs it never issued, bound it, and adopt cached bytes Four things a deployment found once this was running for real. A reference this application mints itself — a generation placeholder, a derived narration key — was never in the pool, because the pool allocates every id it holds. Asking anyway used to be an IndexedDB miss; once the pool is server-backed it is a request that answers 404, one per element per load, forever on a course that still holds placeholders. Every lease and probe now checks first. The check is a negative test on shapes this application owns, not an id validator: the pool's id domain stays unconstrained, and anything that is not one of ours is still asked about. The asset store can bound how much one principal holds, and enforces it inside the write transaction, but nothing ever passed the number. It does now, with a default rather than an opt-in: allocation is reachable by any caller a deployment admits, and with one shared principal an unbounded store is unbounded database growth with no operator-visible brake. Refusing asset mutations to unauthenticated callers was not enough, because every authenticated caller resolves to that same shared principal — so authentication decided nothing, and any signed-in visitor could delete any id they learned. Since this branch began storing media the registry is the only copy a course has. Replacing and deleting are now refused to everyone, and the browser no longer tries: an entry nothing references waits for server-side reclamation instead. What a browser must still do is forget its own record of an allocation that reached nothing, or a later save would stamp an id the document has no reason to trust. And a course generated before any of this holds placeholders in its document with its bytes only in the author's browser. Those bytes are paid for, so the author's next load converts them — stored to the pool and written back through the ordinary commit path, with no provider call — instead of buying them again. A row that records only a hosted URL is treated as absent: that URL is the provider's address, not something a document may hold. One renderer expectation moved with this. An untracked placeholder used to paint as pending on first render because asking the pool left a lease in flight; it settled to disabled a moment later either way, and now says so from the start. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): surface a full store as a refusal and convert legacy narration A quota refusal reached the browser as HTTP 500 with a generic message, which reads as a transient failure: the element kept a Retry that would pay a provider again and be refused again. The store raises the contract's own error and the handler maps it to 507, but the store answering a request is not always built by the same bundle as the handler -- the persistence provider is reached from the route bundle and from instrumentation, which is why its state lives on a Symbol.for global -- and `instanceof` is false across that boundary while the declared code is still right. Classify on the code as well as the class, and make the code a permanent, persisted refusal in the browser: recorded locally so it survives a reload, shown as "storage is full", and refused by the retry entry point so a stale button cannot buy a second generation. Every other storage failure stays retryable. Convert what a pre-server-backed course still holds. Generated media is adopted under either key this application has used for it -- the placeholder, and the allocated id of a course converted once and later rolled back -- instead of only the first. Narration is converted by a load-time pass over the open course's speech actions, since nothing re-enters generation for an action that already has an id: bytes to the pool, id written back through a funnel that mirrors the media one, owner-only and server-backed-only. A line whose bytes are in no browser is left alone rather than re-synthesized. Also: the pool guard is now a positive `ast_` test rather than an enumeration of the shapes we mint (imports never reach the pool, so this is safe in both modes); the slide ref collection is an exported pure function so its four lease sites are covered behaviourally; ASSET_QUOTA_BYTES treats every spelling of zero as opting out and refuses a malformed value at startup instead of falling back; the abort signal is re-checked after the cache read, before an uncancellable commit; and the unused `removeAsset` and pool `replace` surfaces are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * chore(storage): release 0.29.1 The asset HTTP handler now recognises a store refusal by the contract code it declares as well as by its class, so a quota refusal raised in another module realm answers 507 instead of 500. Same contract, stricter recognition, no API change: a patch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a full store recoverable and adoption course-safe Narration adoption read the local audio row by its derived key alone. That key carries no stage id and the table is keyed by id alone, so two courses can mint the same one -- a PPTX import numbers its scenes and actions deterministically, which gives every imported deck's first slide `tts_s1_speech-scene-p1`. Locally a collision only means one course plays another's clip in one browser; adopting it wrote that clip into the shared document permanently, for every device and every visitor. A row that names a course is now adopted only into that course, and a row from before that column existed only when the text it recorded is the text of the action being converted. A full asset store was made permanent last round, which was wrong three times over: it overwrote the refused bytes with an empty blob -- on the conversion path that row is a course's only copy of its own media -- it kept sending the rest of the deck to a provider against a ceiling it already knew was reached, and it left no way back once an operator raised that ceiling. A full store is neither the content's fault nor the configuration's, so it is now its own case: the bytes are kept, the pass stops at the first refusal, and the element shows the reason together with a Retry that re-attempts the upload from those bytes. Nothing retries automatically, so no one is re-billed. The narration write-back now reaches the write boundary every producer of a durable write passes through, not only the dirty mark: adoption never deletes the derived row, so a snapshot that reverts the rewrite is adopted again on the next load and allocates a fresh asset every time. Adoption is also mounted by both classroom surfaces rather than one, takes the course's abort signal, and re-validates that this browser still has the course open before each write. ASSET_QUOTA_BYTES is validated from instrumentation, where the README and the docstring already claimed it was: its only other consumer is lazy and memoised, so a malformed ceiling let the process boot and then failed every persistence request, documents and runtime included. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): remember a full store per course, and never lose retained bytes A stopped pass left the elements it never reached as placeholders with no persisted record -- deliberately, since nothing was attempted for them. But that left the next load with no reason not to try: it called a provider for the next placeholder and was refused at exactly the same point, once per reload, indefinitely. A full store is not a property of any slide. It belongs to the deployment and changes for reasons the document knows nothing about, so it is now remembered once per course in the browser's device KV. A pass that finds the marker stands down before spending anything and leaves every placeholder its "storage is full" state and its Retry; the first upload that succeeds clears it and the next pass runs normally. Narration adoption latched per course so it runs once per load, and the latch outlived the abort that leaving a course performs. On a surface that stays mounted across switches -- the workbench pane is one component for every course it shows -- owner course A, visitor course B, then back to A skipped exactly the clips the abort had cut off, and nothing else converts them. The latch is released with the abort now, and a course adopts one run at a time so a re-entry cannot hand a clip a second allocation while the previous run's uncancellable tail is still settling. A quota-blocked element retried into a network error or a 500 lost the bytes that were kept for it: the retry deleted the row before attempting the upload and wrote no replacement for an error carrying no structured code, so the next retry went back to a provider for media this browser had a moment earlier. The row now survives until an upload succeeds, the failure handler keeps whatever bytes the attempt was given, and the retry asks the question a pass asks -- does this browser already hold bytes for this element -- rather than reading an error code that a second failure has already overwritten. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): adopt real legacy narration, queue re-entries, report attempt outcomes Narration adoption admitted a stage-less row only when the text it recorded matched the action being converted. Both of those columns were added to the local audio table by the very change that moved narration onto allocated ids, so a row still carrying a derived key has neither: the rule refused every real pre-allocation course and passed only on fixtures built from post-allocation rows. What the row cannot say, the key can. A derived key names two clips only when two courses share a scene order and an action id, and an action id repeats only when something other than the generator minted it -- an import numbers them by slide position. So a key built from a generated action id is adopted on that basis, a key an import could have reproduced still needs matching text, and a row that names another course is refused however unique its key looks. Handing a re-entering caller the adoption run already in flight undid the latch release it was paired with: that run is bound to the signal the departure just aborted, so it stops at its next clip while the caller -- which has the course open and a live signal -- is told the work is done, and an effect replayed as mount, cleanup, mount adopts nothing at all. A later caller now waits for the uncancellable tail and scans again, which costs a lookup on a course that has nothing left and finishes the clips the abort cut off on one that does. One attempt at an element now reports both facts its callers need instead of a bare boolean: whether the store refused it for room, and whether bytes actually reached the store. Leaving a course clears the task table, so a retry that landed afterwards read "no failed task" as success and deleted the row holding the only copy of the media. Nothing is inferred from that table any more. Reading the localStorage property can throw where storage is denied by policy, typeof included, so the availability check moved inside the guard: this metadata is best-effort, and a rejection here strands a generation pass that has already enqueued its tasks. A retry is never blocked by the per-course "store is full" marker, but a retry that is refused again re-sets it, and adoption now reads and writes the same marker rather than issuing one refused upload per clip on every load. The two canvas element renderers and both thumbnail renderers show the reason beside the Retry, so a full store does not look like an ordinary failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): probe a full store instead of standing down, and pair notices with a Retry Narration adoption was given both halves of the per-course "the store is full" marker last round: it stood down when the marker was set, and it set the marker when its own upload was refused for room. Those halves are only safe together if something can lift the marker, and for adoption nothing could. It has no affordance of its own, it stood down before reaching its own clear, the media pass returns before its marker gate when there is nothing to generate -- so a narration-only deck, or one whose slides are already satisfied, painted no storage-full element and offered no Retry -- and narration generated rather than adopted allocates directly rather than through the media commit. The course's cached narration was then lost for good, where before it converted on the first load after the ceiling was raised. The gate is a probe now. A marked course attempts exactly one clip per load: refused, it stops and the marker stands, which costs what standing down cost; stored, it lifts the marker and finishes the course. Adoption spends no provider money, so the whole cost of probing a store that is still full is one refused upload. Generated narration lifts the marker too. The three surfaces that gained a failure notice last round drew it for any failure with a reason, including the one refusal that is reachable without server-backed persistence, so a browser-only deck painted something it had not painted before. The notice is drawn beside a Retry and nowhere else, which is what it was added for and what leaves browser-only output unchanged. Both are now asserted through the render harness the surface matrix already had. A caller arriving while a rescan is queued shares it rather than appending another. One rescan converts whatever the run in flight left and every later one would find an allocated id on every action, so a chain bought nothing and turned a single stalled upload into a course that never adopts again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): treat a refusal for room as a fact about one clip, not the deck The asset store checks each write against the headroom it has left, so a store that refuses a long opening clip can still hold every short clip behind it. Narration adoption assumed the opposite: it broke the deck at the first refusal and then re-attempted that same first clip on every later load, because the document names it first. A deck whose longest clip exceeds current headroom therefore never converted the clips that would have fit, with no affordance to recover it -- the state the probe was introduced to remove, reached through a narrower door. An unmarked load now attempts every clip, skipping the ones that do not fit, and remembers the condition only if the load ends with clips it still could not store. A marked load spends its single upload on the smallest clip left rather than the first one named: that is the clip that answers the question the marker asks, because if the smallest does not fit nothing does. The media pass keeps stopping at its first refusal, and for a reason adoption does not share -- every element it attempts costs a provider call. A rescan several callers share took the newest caller's signal, and the newest caller is not necessarily the one still there: a surface that opened a course and closed it again would stop work a surface still showing that course was waiting for, and that surface is latched, so it would never ask again. The shared run now takes a signal that is aborted only once every caller has left. The comment claiming the shared rescan contains a stalled upload was wrong -- the rescan is chained off the run in flight, so a stalled upload leaves every caller pending exactly as a chain would. It claims the bounded queue it actually provides, and the stall is recorded as a limitation. The failed-state containers took their stacking classes unconditionally, so markup differed in browser-only mode even though nothing moved on screen. Those classes are applied only when there is a notice to stack, and the tests assert the exact class attribute rather than a substring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): stop narration adoption writing the media pass's store-full marker The marker means "do not call a provider for this course". A path is entitled to write it only if its own refusal cost a provider call, and narration adoption's refusals cost nothing: it uploads bytes this browser already holds. The store also checks each write against the headroom it has left, so a clip that does not fit says nothing about whether a slide's image would. Adoption was writing it anyway, and one over-long narration clip was therefore enough to stand a course's entire image pass down on every later load -- on a store that had just accepted adoption's other clips. The author could still recover each element by hand, every load, for ever. Three rounds of narrowing this seam produced a finding each time, so it is removed rather than narrowed again. Gone: the marker read, the single-clip probe, the smallest-clip selection, and the up-front read of every row into an array -- which also retires a sampled-then-stale flag and the retention of a whole deck's blobs for the length of a run, and returns the loop to streaming one row at a time. Adoption's rule is now that every load attempts every clip it holds, once; any failure skips that clip and the load continues. The noise the coupling was meant to avoid does not arise, because after the first load the clips still outstanding are exactly the ones that did not fit -- normally none, or one. A successful write still clears the marker, and that is a different kind of statement: a write that went through is a fact this run established, where a refusal is an inference about what some other write would cost. For a course whose media needs nothing, adoption and generated narration are also the only paths that can establish it. The failure module still documented the deck-wide premise this contradicts. It now says what is true: the check is per write, and the media pass stops the deck as a judgement about cost rather than about certainty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bound a full store's cost from the store's own arithmetic Removing the store-full marker from narration adoption removed its bound too, and the code then asserted the bound was unnecessary. It is, on a store with room for most of a deck. On the store the whole mechanism exists for -- the ceiling reached, nothing fitting -- the outstanding set after every load is the entire deck, so a thirty-clip course posted thirty full blobs on every load, indefinitely. Each of those is not a cheap refusal: the bytes are uploaded, the server hashes the whole payload, and only then takes a per-principal lock and sums every entry that principal owns before saying no. The bound needs no flag, no key and nothing carried between loads. The store asks whether `used + addedBytes` exceeds the ceiling, and `used` only grows while a run is uploading, so a clip refused for want of room implies every clip at least that large is refused for the rest of that run. The run keeps the smallest size it has been refused and skips anything no smaller without uploading it; a smaller clip is still attempted, because it may fit. A deck the store refuses entirely now costs one upload per successive size minimum instead of one per clip, and a deck it has room for costs nothing extra, because nothing is refused. Only a refusal for room lowers the bar: a dropped connection says nothing about how much room there is. The deck-wide certainty premise the failure module retracted last round still stood verbatim at the site that implements the stand-down. Both copies now say the same thing: the check is per write, and the pass stops the deck as a judgement about cost rather than about certainty. The comment on adoption's marker clear now names its price. Narration of a few hundred bytes fits in headroom an image does not, so a proven write can let the next pass buy one more image that is refused again -- bounded at one, and the price of the alternative being a course whose media never generates again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(media): state the adoption bound exactly, and stop three comments describing the old rule The comment introducing the in-load bound gave its cost as "at most a handful, and the first load pays the most". Neither clause is a property of the rule. A clip is skipped only when something no larger was already refused, so a fully-refused deck costs one upload per successive size minimum in document order: one when the clips grow, about ln N for an arbitrary order, and one per clip when they only shrink -- a long opener followed by terser lines is exactly that shape. And no load is cheaper than the first, because the bound resets per run and a refused clip stays outstanding. The comment now says that, and points at what would make it exactly one for any ordering: the store returning its remaining headroom in the refusal's existing details channel, which the server leaves empty today. Two other comments still described the previous rule -- "attempts every clip it holds, every load" -- one of them twenty lines above the paragraph that introduces the bound, in the same block. Both now say what the code does. The bound's soundness is worth stating where a maintainer will look for it: quota is charged at full length with no discount for a duplicate, the sum it is checked against joins entries to blobs so the collector cannot lower it, the check takes a per-principal lock before summing, and replace and delete are refused to every browser. Nothing a run can do makes room appear inside it. One test installed a row implementation and replaced it wholesale a few lines later, so the first was dead and the survivor dropped the text the first clip's import-shaped key needs for the ownership rule -- it passed on the coincidence that the fixture's default text is the action's. Merged into one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): keep a Retry from re-buying parked media, and state the store seam once Five findings from an inline review. The standalone classroom route asked the ownership sidecar once per load and recorded only stage ownership when that ask failed. Every non-answer fails closed, so one transient 5xx left the genuine author with no resume, no Retry affordance and no legacy narration converted for the rest of the load, with nothing to change it short of a reload. The failure now records the fail-closed answer explicitly -- an answer an earlier load established must not outlive the failure that replaced it -- and an unresolved answer is asked for again, a few times over a few seconds. A real answer, however unwelcome, is final. A Retry could pay a provider for media the pool already held. When the bytes are stored and only the write-back fails in a way that keeps the allocation, it is parked and no local row exists, because that row is written only after a successful write-back. Retry now reads the parked queue exactly as the pass does and re-attempts the write-back: it re-keys the task done when the document takes it, leaves the entry parked when the slide still does not exist, and stays failed and retryable when the document refuses again. Object URLs a parked allocation owns are revoked when the entry is dropped. The commit path leaves them alone while the entry is parked, because it is then the only thing holding bytes this tab can render, so a course switch or a stage deletion was pinning the whole blob for the life of the tab. An entry a slide has already taken is left alone: the task table is displaying those URLs. The fallback lookup for cached bytes is a stage-scoped scan, and the keyed lookup misses for every row the commit path writes, so a pass was materializing and sorting the course's whole media table once per element. One scan per pass now, built on the first miss. It is sound and not merely cheaper: an element asks only for its own placeholder, and every row a pass writes carries the placeholder of the element that wrote it. "The store accepted a write, so it is not out of room" was enforced at three call sites under slightly different conditions, which made it a convention the next pool write path could silently break. It is stated once, in putAsset, for the course whose bytes it just stored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 16 天前 | |
feat(media): write generated media through the asset pool under server-backed persistence (#1392) * feat(media): store generated media in the asset pool when persistence is server-backed With server-backed persistence the document is durable and shared, but generated media stayed in the producing browser: the document kept its gen_img_* / gen_vid_* placeholder and narration kept a browser-derived audio id. Every new browser that opened such a course re-ran generation for every slide, and it never converged, because the address of the generated bytes was never written back into the document. Under server-backed persistence only, the classic generation chain now stores bytes in the asset pool first and writes the id the pool allocated into the document. - The client bootstrap configures the asset seam alongside the document and runtime seams: an HttpAssetStore over the persistence endpoint carrying the same credentials the document store carries, marked server-backed. The seam preflight now covers all three, so a failure still cannot half-configure persistence. - Image, video and TTS generation commit in one fixed order: provider, pool, document, local cache, task. A reference reaches the document only after put returned an id, so a document can never name bytes that were not stored. A failure before the write-back leaves the placeholder with the provider called exactly once; the retry happens on the next owner load. - The write-back is a per-slot rewrite through mutateDocument, which re-reads the current document under the per-stage lock, so it cannot clobber a newer scene. The open course is refreshed with the same rewrite without being marked dirty. - "Has this already been generated?" is answered by the document (the slide exists and no longer holds the placeholder) instead of by this browser's task table. - The classroom's resume effect fails closed on ownership: only a resolved owner starts generation, so a viewer opening a shared course spends nothing. - The local media and audio tables become a per-tab cache. A failed cache write costs a re-download, never the media. Browser-only mode is unchanged: every new call site sits behind the server-backed gate, the local tables stay authoritative there, and placeholders stay in the document. Rendering and export needed no changes. HttpAssetStore.resolve mints an object URL exactly as the browser store does, and the export byte resolver was already pool-first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(classroom): make the generation owner gate a three-outcome rule and apply it everywhere The gate refused everything but a resolved owner, which read a sidecar that answered "no ownership fact exists for this course" as a reason to block. That is the answer a deployment without the sidecar's server-side prerequisites gives for every course, and the answer a course with no ownership record gives: in both, there is nobody the operator's budget needs protecting from, and refusing strands the course's own author behind a question that can never be answered. Ownership is now four states over the sidecar's three outcomes. A definite answer splits into owner and not-owner. An absent record is its own answer, ownerless, and generation proceeds — the behaviour such a deployment had before the gate existed. Only the absence of an answer, a transport failure or a load that has not asked yet, stays unresolved and fails closed: "we could not ask" must never be read as "nobody owns this". One mapper turns a sidecar result into that state, and one predicate decides on it. The workbench classroom pane runs the same resume effect and had no ownership input at all, so a viewer opening a shared course there could still spend the budget. It now asks the sidecar once per course, in parallel with its load and feeding only the generation gate, so its read-only and edit behaviour is unchanged. The shared progressive-load policy carries the gate for it, with both new inputs required rather than defaulted so a future caller cannot omit them into an open budget. Its stale comment claiming ownership could not be expressed here is corrected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the write-back survive autosave, arrive before the scene does, and never leak Independent reviews of the write-back found three ways a durable document could still end up naming a placeholder, and two ways the gate that protects the operator's budget could be walked around. An autosave round captures the store synchronously and writes that capture, so a round already in flight when a rewrite landed wrote the placeholder straight back over the allocated id, and nothing marked the store dirty again to correct it. The rewrite now marks the units it changed, which leaves a corrective flush queued behind the stale one; re-saving a scene that already holds the id is idempotent, losing the id is not. Media is generated from outlines in parallel with scene content and usually finishes first, so the slide that will carry the placeholder does not exist yet and the write-back has nothing to rewrite. That was the ordinary path, not a tail case, and its result was discarded: the task was marked done, the scene was added afterwards with its placeholder intact, and a second pass in the same run could call the provider again. The allocation is now held under the placeholder — which also answers the skip test, so nothing pays twice — and applied when that scene is committed, before its first save. One complete pass now leaves no placeholder behind. A failed commit used to abandon what it had already allocated. A poster upload that failed threw away a stored video and sent the retry to submit the most expensive job in the system again; a rejected write-back left registry rows that name bytes nothing references, which the byte collector cannot reclaim because it only collects blobs no row names. A poster failure now costs the poster, and a write-back that reached nothing reclaims what it allocated. A partial write is left alone, because the document already names it. The ownership gate is fail-closed again. Treating the sidecar's 404 as permission was wrong: the client cannot tell "this course has no owner" from "this deployment told me nothing", so a visitor who opened a shared course could bill the operator. The root cause was the sidecar itself, which gated on the agent runtime although every persisted course has an owner regardless — the persistence route resolves one for every request. It now gates on server persistence, so the configuration that made 404 the universal answer has real ownership facts to report, and the gate can refuse everything but a named owner. Retry affordances answered to no gate at all. A viewer of a shared course with one failed image was shown a Retry button that called the provider. Both retry entry points and every surface that draws them now read one shared permission, so what is offered and what is allowed are the same value. Also: narration regeneration no longer pretends it can replace bytes behind a live id — the exclusivity proof that would allow it is refused by construction once references leave the browser, so it forks to a fresh id and says so; the "already generated" test lets a finished deck answer from the document alone, since scene order stops identifying an outline once slides are inserted or deleted; stored assets record a specific media type rather than a generic transfer type; the pane no longer asks the sidecar in browser-only mode; and the funnel's docstring now states what the per-stage lock actually guarantees, which is same-browser serialization and not a cross-browser compare-and-swap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): park allocations in the deciding turn, never reclaim on an ambiguous write A delta review of the write-back found the first-pass fix still had a window, and the reclamation it added could delete media the document already names. The allocation was parked after an awaited local cache write. A scene committed in that window reconciled against a registry that did not hold it yet, so the document kept the placeholder — and the entry recorded a moment later then answered the skip test as "already handled", so nothing could correct it. Parking now happens inside the write-back, in the same synchronous turn as the decision that nothing could take the reference; no await separates the live check from the park. Allocations parked by an earlier pass are handed to their slides at the start of the next one, before anything decides what still needs generating, so a held allocation whose scene has since arrived becomes a rewrite rather than an answer. Reclaiming on a rejected write was unsound: a rejection does not prove the server did not apply the write, so deleting the asset could break the scene that now names it. The funnel decides instead, and says so: it reclaims only when no store write was ever issued and nothing took the reference. Anything else is placed if its slide exists and parked if it does not, so the next pass reuses the bytes instead of paying for them again. When a write fails after part of it landed, the live store is brought up to the document before the error is rethrown — otherwise the next ordinary flush would overwrite the half that did land, with the ids deliberately not reclaimed. Parked allocations are now cleared with the course. Classic placeholders are reused across runs, so one surviving an interrupted run would be handed to a different slide of the next deck: the previous picture, on a slide whose provider was never asked. Both classroom surfaces clear the arriving course, the deletion cascade clears the deleted one, and clearing the database clears them all. Two more ways generation could start without asking the gate are closed. An overlapping pass — an outline retry re-enters generation with every outline while the first is still working — re-requested elements whose provider call was already in flight; a task that is not done is an answered request, not an unanswered one. And narration regeneration in the timeline editor called the TTS provider and allocated a pool asset with no ownership check at all; it now reads the same permission, which withholds both the per-line and whole-timeline controls and refuses the call. Finally, a pane opened during the stage-link availability gap recorded the sidecar's 404 for a course that was moments from existing and never asked again, leaving the real owner locked out of generation until it remounted. Ownership is re-fetched once the document becomes available; the gate stays closed until an answer arrives, so asking again can only open it for someone entitled to it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the asset routes reachable, and stale snapshots harmless A full-branch audit found that the deployment this project documents could not store a single generated asset, and that several routes into durable storage could still write a placeholder over a reference that had already landed. The persistence route sent asset requests through the development authenticator, which refuses outright in a production build that has not explicitly opted into it — and the documented server-persistence recipe produces exactly that build. Every store and every read answered 401, so images and video failed on every slide while re-billing the provider on each retry, and a narration failure stopped the deck at its first slide. Assets live in one shared partition by design, so there was never anything per-caller for that authenticator to decide: the route now resolves the asset principal itself, alongside the owner it already resolves for documents. Runtime sessions are genuinely per-learner and keep the development authenticator until real session verification replaces it. And narration that cannot be stored no longer fails its scene: the line stays unvoiced and retryable, which is what an image that cannot be stored does to its slide. Placeholders could also come back from behind. A queued autosave's snapshot, an editor-history entry replayed by an undo, the departing save a course switch flushes — each captures content at its own moment, and any of those moments can predate a write-back. Point fixes at each producer would leave the next producer to rediscover the bug, so the check lives at the write boundary every producer passes through, and the allocation record it consults now outlives the parked queue: a placeholder whose rewrite landed long ago is exactly the case it catches. Two ways generation could be lost or repeated are closed. A pass now claims the elements it will reach and releases them however it ends, so an overlapping pass stands down while an aborted one strands nothing — previously its tasks stayed `pending` and every later pass skipped them with no retry control to recover them. And the media abort controller is aborted before being replaced, so a superseded pass stops calling providers instead of running on for a course the user has left. The remaining two are narrower. The workbench pane asks for ownership only after a document load succeeds, and after every later one, mirroring the page route: the load is what creates the ownership row the first time a course is opened, so asking beforehand asked about a course that did not exist yet and locked its author out for the mount. And the ownership gate on the timeline editor now withholds narration regeneration alone; listening back to existing narration and seeing whether a line has any spend nothing and stay available. Known limitation, unchanged and now stated plainly in the comments that used to point at it as a solution: nothing reclaims an unreferenced pool asset. The registry sweep is written but not wired up, and the byte collector only reclaims blobs no registry row names, so every narration regeneration and every abandoned allocation leaves storage behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): gate asset mutations, and make claims and allocation records survive a handoff Opening the asset routes opened all of them. Reads and allocations are meant to be as open as document reads and creates already are, but no authorization hook was supplied, so the handler's default admitted PUT and DELETE too — and those scope by principal key alone, which is one shared constant. Any caller who learned an id, and a document read hands out every id its slides name, could overwrite or destroy another author's media. Mutations now require the deployment's credential, which in a production build without the development-auth opt-in means they are refused outright; reads and allocations stay open. The route comment says what the posture is and what it is not: the deployment-level fence is the access code, and no per-principal quota is configured. The client's own reclaim is best effort to match — losing an argument about deleting an asset must not cost a task its retry, and the bytes are left for server-side reclamation. The pass claim could not survive the handoff it was written for. A retry aborts the live media pass and starts its replacement in the same synchronous block, long before the aborted pass's cleanup runs, so the replacement saw every element still claimed, collected nothing, and returned — leaving each unreached element at pending with nobody coming back for it and no retry control to recover it, which is the exact failure the claim was introduced to prevent. A claim now carries its pass's signal and is retired the moment that signal aborts, and a pass releases only claims it still owns, so a late unwind cannot take its replacement's work. Claims are also acquired at the single point every request passes through, so a single-task retry participates too — previously a retry awaiting its provider was invisible to a pass starting alongside it and both called it. The allocation record could outlive the bytes it named. It was written before the write-back attempted anything and survived the reclaim that followed a failure, so when the slide finally arrived the write boundary stamped a deleted id into the document — and the placeholder it replaced was gone, which reads as already generated and stops anything from retrying. The record is now written only where the allocation is retained, and forgotten wherever a reclaim removes the bytes, including the narration rollback path. The tests follow. The route test drives the real storage handler against an in-memory registry instead of a stub, so it can see what the resolved principal is then allowed to do; the handoff test performs a real abort mid-pass rather than starting from an already-aborted signal; and the guards that could only assert file layout now assert the property they care about, or have been replaced by behaviour. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make media passes serial per course instead of tracking element ownership Three rounds of per-element claims each produced a new way to lose an element. Whole-pass reservations swallowed a Retry for an element the same pass had already failed, leaving it pending with the affordance gone. Retiring a claim by its signal freed an element whose commit was still uploading, so the replacement pass paid for it twice. A claim held for a failed element stranded its retry. The bookkeeping is the defect: every refinement of "who owns this element right now" answered the question at a moment when the answer was already stale. Passes for one course are now serial. A replacement aborts its predecessor, as before, and then waits for it to settle before collecting. That removes the question entirely: a commit already under way finishes — its bytes stored and its reference written, so the new pass sees a resolved slide and skips it — and an element the aborted pass never reached is still a placeholder and gets collected like any other. The claim set, the reservations, the signal retirement and the identity-checked release are all gone. The task table is consulted for one thing only: an element that is generating right now is a single-element retry running alongside the pass, and taking it too would pay twice. Pending is deliberately not a skip reason — it means a pass once intended to reach an element, which an abandoned pass leaves behind with nobody acting on it, and reading that as answered is what stranded elements before. A retry runs concurrently with a pass, because a pass never revisits an element it has processed, and it re-reads the task after its own await and refuses before touching it: marking first and refusing afterwards destroyed the failed state that draws the affordance. Browser-only mode is back to exactly what it was. The abort is now conditional, the waiting does not apply, and the original status-based skip is restored verbatim. Two baseline lines remain changed in each of the two files, and both are behind a server-backed fork whose else-branch is the original. Two smaller things. The allocation record becomes visible when a write goes on the wire rather than when the round trip ends, and the write boundary reconciles under the document lock rather than before it — a save queued during a write-back was otherwise captured with the placeholder and, for a course the user had left, had no corrective flush to follow. And the comments that said a refused reclaim leaves its bytes for server-side reclamation were wrong: nothing collects them, because the registry entry still names its blob and the sweep that would remove it is not wired up. They now say the bytes leak. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a deferred pass re-earn its right to run, and bound the commit it waits on Serializing passes moved their body out of the block that launched them, and three things followed from that. A pass now wakes when its predecessor settles, which can be after the user has left the course. It enqueued before it looked at its signal, into a task table keyed by element id alone — and placeholder ids are not unique across courses, which is why the classroom clears that table on arrival. So a departing course's pass seeded the arriving course's table with tasks carrying the wrong stage id, and a Retry routes by that id: the reference went into the wrong document. A pass now re-validates after the wait, before touching anything shared. The same lateness broke the skip test. `documentSkipIndex` answers only while the live store is on the pass's stage, and returning nothing put the collection loop on the browser-only rule — a silent demotion from "the document is the authority" to "this browser's task table is", on exactly the path where that table has just been cleared. Every element the predecessor had committed was collected again, paid for again, and its second write-back found no placeholder to rewrite, so its bytes were parked where nothing will ever reference them. In server-backed mode an unreadable document now means the pass stands down. And waiting was unbounded. A commit is uncancellable: the asset client takes no signal, and a document write cannot be half-undone. One stalled upload therefore froze the course's media generation for the session — the replacement never collected, the element sat on a skeleton that draws no Retry, and only a reload recovered. The pass's signal is now threaded into the media proxy fetch, and the commit is bounded by a deadline. The deadline is on the wait, not the work: the commit carries on, and if it lands late the document simply ends up correct, while the element becomes retryable and the queue moves on. The tests that were meant to pin the previous round were not sensitive to it. Two asserted end states where the mechanism only changes ordering, and one of them rigged the document read so the assertion held whether or not the pass had waited; a third covered half of what it claimed. They now observe the ordering directly — nothing is issued while another pass for the course is working; in browser-only mode a second pass reaches its provider immediately — and the reconciliation under the document lock has a test that fails when it moves back outside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * revert(media): drop the commit deadline and the abortable download The deadline bought less than it cost. Abandoning a commit after two minutes makes the element retryable while the real commit is still running, so a Retry starts a second commit for the same placeholder against the first: two provider calls, two allocations, and whichever lands second stamps its result over the other's task by element id. The allocation record is keyed by placeholder, so the loser's cleanup erases the winner's record, and the write boundary then puts the raw placeholder back into the document. That is the overlap serial passes were built to remove, reopened through the one door serialization never covered. So a stalled commit holds the course's media queue until it settles or the page is reloaded, and that is written down rather than papered over. The wait is unbounded on purpose: every ceiling on it turns out to be a way of running two commits for one element. Threading the pass signal into the download was also a mistake, in the other direction. The provider call that produced the URL has already been billed, so cancelling the download throws away work that is paid for — and the shared proxy cache records a cancelled request as a transient failure against that URL, which after three of them blocks it for every consumer in the session. Browser-only mode never asked for this: it had no way to observe an abort there, which is exactly why the bytes were kept. The signal is gone from the download again, and `fetchAsBlob` is byte-for-byte what it was before this branch. The regression guard for the stranded-element rule is restored alongside the timing test that was meant to supersede it. It catches a different rule — a task left pending being read as answered — and nothing else does: making the pass skip pending leaves every other suite green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): stop asking the pool for refs it never issued, bound it, and adopt cached bytes Four things a deployment found once this was running for real. A reference this application mints itself — a generation placeholder, a derived narration key — was never in the pool, because the pool allocates every id it holds. Asking anyway used to be an IndexedDB miss; once the pool is server-backed it is a request that answers 404, one per element per load, forever on a course that still holds placeholders. Every lease and probe now checks first. The check is a negative test on shapes this application owns, not an id validator: the pool's id domain stays unconstrained, and anything that is not one of ours is still asked about. The asset store can bound how much one principal holds, and enforces it inside the write transaction, but nothing ever passed the number. It does now, with a default rather than an opt-in: allocation is reachable by any caller a deployment admits, and with one shared principal an unbounded store is unbounded database growth with no operator-visible brake. Refusing asset mutations to unauthenticated callers was not enough, because every authenticated caller resolves to that same shared principal — so authentication decided nothing, and any signed-in visitor could delete any id they learned. Since this branch began storing media the registry is the only copy a course has. Replacing and deleting are now refused to everyone, and the browser no longer tries: an entry nothing references waits for server-side reclamation instead. What a browser must still do is forget its own record of an allocation that reached nothing, or a later save would stamp an id the document has no reason to trust. And a course generated before any of this holds placeholders in its document with its bytes only in the author's browser. Those bytes are paid for, so the author's next load converts them — stored to the pool and written back through the ordinary commit path, with no provider call — instead of buying them again. A row that records only a hosted URL is treated as absent: that URL is the provider's address, not something a document may hold. One renderer expectation moved with this. An untracked placeholder used to paint as pending on first render because asking the pool left a lease in flight; it settled to disabled a moment later either way, and now says so from the start. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): surface a full store as a refusal and convert legacy narration A quota refusal reached the browser as HTTP 500 with a generic message, which reads as a transient failure: the element kept a Retry that would pay a provider again and be refused again. The store raises the contract's own error and the handler maps it to 507, but the store answering a request is not always built by the same bundle as the handler -- the persistence provider is reached from the route bundle and from instrumentation, which is why its state lives on a Symbol.for global -- and `instanceof` is false across that boundary while the declared code is still right. Classify on the code as well as the class, and make the code a permanent, persisted refusal in the browser: recorded locally so it survives a reload, shown as "storage is full", and refused by the retry entry point so a stale button cannot buy a second generation. Every other storage failure stays retryable. Convert what a pre-server-backed course still holds. Generated media is adopted under either key this application has used for it -- the placeholder, and the allocated id of a course converted once and later rolled back -- instead of only the first. Narration is converted by a load-time pass over the open course's speech actions, since nothing re-enters generation for an action that already has an id: bytes to the pool, id written back through a funnel that mirrors the media one, owner-only and server-backed-only. A line whose bytes are in no browser is left alone rather than re-synthesized. Also: the pool guard is now a positive `ast_` test rather than an enumeration of the shapes we mint (imports never reach the pool, so this is safe in both modes); the slide ref collection is an exported pure function so its four lease sites are covered behaviourally; ASSET_QUOTA_BYTES treats every spelling of zero as opting out and refuses a malformed value at startup instead of falling back; the abort signal is re-checked after the cache read, before an uncancellable commit; and the unused `removeAsset` and pool `replace` surfaces are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * chore(storage): release 0.29.1 The asset HTTP handler now recognises a store refusal by the contract code it declares as well as by its class, so a quota refusal raised in another module realm answers 507 instead of 500. Same contract, stricter recognition, no API change: a patch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a full store recoverable and adoption course-safe Narration adoption read the local audio row by its derived key alone. That key carries no stage id and the table is keyed by id alone, so two courses can mint the same one -- a PPTX import numbers its scenes and actions deterministically, which gives every imported deck's first slide `tts_s1_speech-scene-p1`. Locally a collision only means one course plays another's clip in one browser; adopting it wrote that clip into the shared document permanently, for every device and every visitor. A row that names a course is now adopted only into that course, and a row from before that column existed only when the text it recorded is the text of the action being converted. A full asset store was made permanent last round, which was wrong three times over: it overwrote the refused bytes with an empty blob -- on the conversion path that row is a course's only copy of its own media -- it kept sending the rest of the deck to a provider against a ceiling it already knew was reached, and it left no way back once an operator raised that ceiling. A full store is neither the content's fault nor the configuration's, so it is now its own case: the bytes are kept, the pass stops at the first refusal, and the element shows the reason together with a Retry that re-attempts the upload from those bytes. Nothing retries automatically, so no one is re-billed. The narration write-back now reaches the write boundary every producer of a durable write passes through, not only the dirty mark: adoption never deletes the derived row, so a snapshot that reverts the rewrite is adopted again on the next load and allocates a fresh asset every time. Adoption is also mounted by both classroom surfaces rather than one, takes the course's abort signal, and re-validates that this browser still has the course open before each write. ASSET_QUOTA_BYTES is validated from instrumentation, where the README and the docstring already claimed it was: its only other consumer is lazy and memoised, so a malformed ceiling let the process boot and then failed every persistence request, documents and runtime included. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): remember a full store per course, and never lose retained bytes A stopped pass left the elements it never reached as placeholders with no persisted record -- deliberately, since nothing was attempted for them. But that left the next load with no reason not to try: it called a provider for the next placeholder and was refused at exactly the same point, once per reload, indefinitely. A full store is not a property of any slide. It belongs to the deployment and changes for reasons the document knows nothing about, so it is now remembered once per course in the browser's device KV. A pass that finds the marker stands down before spending anything and leaves every placeholder its "storage is full" state and its Retry; the first upload that succeeds clears it and the next pass runs normally. Narration adoption latched per course so it runs once per load, and the latch outlived the abort that leaving a course performs. On a surface that stays mounted across switches -- the workbench pane is one component for every course it shows -- owner course A, visitor course B, then back to A skipped exactly the clips the abort had cut off, and nothing else converts them. The latch is released with the abort now, and a course adopts one run at a time so a re-entry cannot hand a clip a second allocation while the previous run's uncancellable tail is still settling. A quota-blocked element retried into a network error or a 500 lost the bytes that were kept for it: the retry deleted the row before attempting the upload and wrote no replacement for an error carrying no structured code, so the next retry went back to a provider for media this browser had a moment earlier. The row now survives until an upload succeeds, the failure handler keeps whatever bytes the attempt was given, and the retry asks the question a pass asks -- does this browser already hold bytes for this element -- rather than reading an error code that a second failure has already overwritten. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): adopt real legacy narration, queue re-entries, report attempt outcomes Narration adoption admitted a stage-less row only when the text it recorded matched the action being converted. Both of those columns were added to the local audio table by the very change that moved narration onto allocated ids, so a row still carrying a derived key has neither: the rule refused every real pre-allocation course and passed only on fixtures built from post-allocation rows. What the row cannot say, the key can. A derived key names two clips only when two courses share a scene order and an action id, and an action id repeats only when something other than the generator minted it -- an import numbers them by slide position. So a key built from a generated action id is adopted on that basis, a key an import could have reproduced still needs matching text, and a row that names another course is refused however unique its key looks. Handing a re-entering caller the adoption run already in flight undid the latch release it was paired with: that run is bound to the signal the departure just aborted, so it stops at its next clip while the caller -- which has the course open and a live signal -- is told the work is done, and an effect replayed as mount, cleanup, mount adopts nothing at all. A later caller now waits for the uncancellable tail and scans again, which costs a lookup on a course that has nothing left and finishes the clips the abort cut off on one that does. One attempt at an element now reports both facts its callers need instead of a bare boolean: whether the store refused it for room, and whether bytes actually reached the store. Leaving a course clears the task table, so a retry that landed afterwards read "no failed task" as success and deleted the row holding the only copy of the media. Nothing is inferred from that table any more. Reading the localStorage property can throw where storage is denied by policy, typeof included, so the availability check moved inside the guard: this metadata is best-effort, and a rejection here strands a generation pass that has already enqueued its tasks. A retry is never blocked by the per-course "store is full" marker, but a retry that is refused again re-sets it, and adoption now reads and writes the same marker rather than issuing one refused upload per clip on every load. The two canvas element renderers and both thumbnail renderers show the reason beside the Retry, so a full store does not look like an ordinary failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): probe a full store instead of standing down, and pair notices with a Retry Narration adoption was given both halves of the per-course "the store is full" marker last round: it stood down when the marker was set, and it set the marker when its own upload was refused for room. Those halves are only safe together if something can lift the marker, and for adoption nothing could. It has no affordance of its own, it stood down before reaching its own clear, the media pass returns before its marker gate when there is nothing to generate -- so a narration-only deck, or one whose slides are already satisfied, painted no storage-full element and offered no Retry -- and narration generated rather than adopted allocates directly rather than through the media commit. The course's cached narration was then lost for good, where before it converted on the first load after the ceiling was raised. The gate is a probe now. A marked course attempts exactly one clip per load: refused, it stops and the marker stands, which costs what standing down cost; stored, it lifts the marker and finishes the course. Adoption spends no provider money, so the whole cost of probing a store that is still full is one refused upload. Generated narration lifts the marker too. The three surfaces that gained a failure notice last round drew it for any failure with a reason, including the one refusal that is reachable without server-backed persistence, so a browser-only deck painted something it had not painted before. The notice is drawn beside a Retry and nowhere else, which is what it was added for and what leaves browser-only output unchanged. Both are now asserted through the render harness the surface matrix already had. A caller arriving while a rescan is queued shares it rather than appending another. One rescan converts whatever the run in flight left and every later one would find an allocated id on every action, so a chain bought nothing and turned a single stalled upload into a course that never adopts again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): treat a refusal for room as a fact about one clip, not the deck The asset store checks each write against the headroom it has left, so a store that refuses a long opening clip can still hold every short clip behind it. Narration adoption assumed the opposite: it broke the deck at the first refusal and then re-attempted that same first clip on every later load, because the document names it first. A deck whose longest clip exceeds current headroom therefore never converted the clips that would have fit, with no affordance to recover it -- the state the probe was introduced to remove, reached through a narrower door. An unmarked load now attempts every clip, skipping the ones that do not fit, and remembers the condition only if the load ends with clips it still could not store. A marked load spends its single upload on the smallest clip left rather than the first one named: that is the clip that answers the question the marker asks, because if the smallest does not fit nothing does. The media pass keeps stopping at its first refusal, and for a reason adoption does not share -- every element it attempts costs a provider call. A rescan several callers share took the newest caller's signal, and the newest caller is not necessarily the one still there: a surface that opened a course and closed it again would stop work a surface still showing that course was waiting for, and that surface is latched, so it would never ask again. The shared run now takes a signal that is aborted only once every caller has left. The comment claiming the shared rescan contains a stalled upload was wrong -- the rescan is chained off the run in flight, so a stalled upload leaves every caller pending exactly as a chain would. It claims the bounded queue it actually provides, and the stall is recorded as a limitation. The failed-state containers took their stacking classes unconditionally, so markup differed in browser-only mode even though nothing moved on screen. Those classes are applied only when there is a notice to stack, and the tests assert the exact class attribute rather than a substring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): stop narration adoption writing the media pass's store-full marker The marker means "do not call a provider for this course". A path is entitled to write it only if its own refusal cost a provider call, and narration adoption's refusals cost nothing: it uploads bytes this browser already holds. The store also checks each write against the headroom it has left, so a clip that does not fit says nothing about whether a slide's image would. Adoption was writing it anyway, and one over-long narration clip was therefore enough to stand a course's entire image pass down on every later load -- on a store that had just accepted adoption's other clips. The author could still recover each element by hand, every load, for ever. Three rounds of narrowing this seam produced a finding each time, so it is removed rather than narrowed again. Gone: the marker read, the single-clip probe, the smallest-clip selection, and the up-front read of every row into an array -- which also retires a sampled-then-stale flag and the retention of a whole deck's blobs for the length of a run, and returns the loop to streaming one row at a time. Adoption's rule is now that every load attempts every clip it holds, once; any failure skips that clip and the load continues. The noise the coupling was meant to avoid does not arise, because after the first load the clips still outstanding are exactly the ones that did not fit -- normally none, or one. A successful write still clears the marker, and that is a different kind of statement: a write that went through is a fact this run established, where a refusal is an inference about what some other write would cost. For a course whose media needs nothing, adoption and generated narration are also the only paths that can establish it. The failure module still documented the deck-wide premise this contradicts. It now says what is true: the check is per write, and the media pass stops the deck as a judgement about cost rather than about certainty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bound a full store's cost from the store's own arithmetic Removing the store-full marker from narration adoption removed its bound too, and the code then asserted the bound was unnecessary. It is, on a store with room for most of a deck. On the store the whole mechanism exists for -- the ceiling reached, nothing fitting -- the outstanding set after every load is the entire deck, so a thirty-clip course posted thirty full blobs on every load, indefinitely. Each of those is not a cheap refusal: the bytes are uploaded, the server hashes the whole payload, and only then takes a per-principal lock and sums every entry that principal owns before saying no. The bound needs no flag, no key and nothing carried between loads. The store asks whether `used + addedBytes` exceeds the ceiling, and `used` only grows while a run is uploading, so a clip refused for want of room implies every clip at least that large is refused for the rest of that run. The run keeps the smallest size it has been refused and skips anything no smaller without uploading it; a smaller clip is still attempted, because it may fit. A deck the store refuses entirely now costs one upload per successive size minimum instead of one per clip, and a deck it has room for costs nothing extra, because nothing is refused. Only a refusal for room lowers the bar: a dropped connection says nothing about how much room there is. The deck-wide certainty premise the failure module retracted last round still stood verbatim at the site that implements the stand-down. Both copies now say the same thing: the check is per write, and the pass stops the deck as a judgement about cost rather than about certainty. The comment on adoption's marker clear now names its price. Narration of a few hundred bytes fits in headroom an image does not, so a proven write can let the next pass buy one more image that is refused again -- bounded at one, and the price of the alternative being a course whose media never generates again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(media): state the adoption bound exactly, and stop three comments describing the old rule The comment introducing the in-load bound gave its cost as "at most a handful, and the first load pays the most". Neither clause is a property of the rule. A clip is skipped only when something no larger was already refused, so a fully-refused deck costs one upload per successive size minimum in document order: one when the clips grow, about ln N for an arbitrary order, and one per clip when they only shrink -- a long opener followed by terser lines is exactly that shape. And no load is cheaper than the first, because the bound resets per run and a refused clip stays outstanding. The comment now says that, and points at what would make it exactly one for any ordering: the store returning its remaining headroom in the refusal's existing details channel, which the server leaves empty today. Two other comments still described the previous rule -- "attempts every clip it holds, every load" -- one of them twenty lines above the paragraph that introduces the bound, in the same block. Both now say what the code does. The bound's soundness is worth stating where a maintainer will look for it: quota is charged at full length with no discount for a duplicate, the sum it is checked against joins entries to blobs so the collector cannot lower it, the check takes a per-principal lock before summing, and replace and delete are refused to every browser. Nothing a run can do makes room appear inside it. One test installed a row implementation and replaced it wholesale a few lines later, so the first was dead and the survivor dropped the text the first clip's import-shaped key needs for the ownership rule -- it passed on the coincidence that the fixture's default text is the action's. Merged into one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): keep a Retry from re-buying parked media, and state the store seam once Five findings from an inline review. The standalone classroom route asked the ownership sidecar once per load and recorded only stage ownership when that ask failed. Every non-answer fails closed, so one transient 5xx left the genuine author with no resume, no Retry affordance and no legacy narration converted for the rest of the load, with nothing to change it short of a reload. The failure now records the fail-closed answer explicitly -- an answer an earlier load established must not outlive the failure that replaced it -- and an unresolved answer is asked for again, a few times over a few seconds. A real answer, however unwelcome, is final. A Retry could pay a provider for media the pool already held. When the bytes are stored and only the write-back fails in a way that keeps the allocation, it is parked and no local row exists, because that row is written only after a successful write-back. Retry now reads the parked queue exactly as the pass does and re-attempts the write-back: it re-keys the task done when the document takes it, leaves the entry parked when the slide still does not exist, and stays failed and retryable when the document refuses again. Object URLs a parked allocation owns are revoked when the entry is dropped. The commit path leaves them alone while the entry is parked, because it is then the only thing holding bytes this tab can render, so a course switch or a stage deletion was pinning the whole blob for the life of the tab. An entry a slide has already taken is left alone: the task table is displaying those URLs. The fallback lookup for cached bytes is a stage-scoped scan, and the keyed lookup misses for every row the commit path writes, so a pass was materializing and sorting the course's whole media table once per element. One scan per pass now, built on the first miss. It is sound and not merely cheaper: an element asks only for its own placeholder, and every row a pass writes carries the placeholder of the element that wrote it. "The store accepted a write, so it is not out of room" was enforced at three call sites under slightly different conditions, which made it a convention the next pool write path could silently break. It is stated once, in putAsset, for the course whose bytes it just stored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> | 16 天前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 5 天前 | ||
| 12 天前 | ||
| 1 个月前 | ||
| 16 天前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 16 天前 | ||
| 7 天前 | ||
| 11 天前 | ||
| 16 天前 | ||
| 7 天前 | ||
| 12 天前 | ||
| 16 天前 | ||
| 16 天前 | ||
| 5 天前 | ||
| 5 天前 | ||
| 1 个月前 | ||
| 5 天前 | ||
| 11 天前 | ||
| 1 个月前 | ||
| 11 天前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 2 个月前 | ||
| 16 天前 | ||
| 1 个月前 | ||
| 10 天前 | ||
| 16 天前 | ||
| 16 天前 | ||
| 2 个月前 | ||
| 2 个月前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 12 天前 | ||
| 16 天前 | ||
| 16 天前 | ||
| 16 天前 | ||
| 2 个月前 | ||
| 16 天前 | ||
| 1 个月前 | ||
| 12 天前 | ||
| 16 天前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 11 天前 | ||
| 10 天前 | ||
| 11 天前 | ||
| 16 天前 | ||
| 16 天前 |