| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
Krea 2 regional prompting (#9407) * feat(krea2): add regional prompting * fix(krea2): harden regional prompting * Address findings * test(krea2): cover regional prompting gaps in attention, denoise and canvas graph Adds the coverage the regional prompting change was missing: - attention: the processor's mask/sequence-length shape guard, and that a processor built without regional state ignores the shared mask - regional prompting: that only even main blocks actually apply the mask during a real transformer forward (the existing test only checked the processor map), and that an unsupported mask rank is rejected - denoise: that each CFG pass installs a mask sized for its own text sequence when the positive and negative prompts tokenize to different lengths - canvas graph: a composed suite running the real addRegions, addKrea2LoRAs and regional-guidance validators, covering validator filtering, the IP adapter collector teardown, the LoRA-before-regions ordering the regional encoders depend on, and per-region enhancers Each test was mutation-checked against a deliberately broken implementation to confirm it fails for the right reason. --------- Co-authored-by: Alexander Eichhorn <alex@eichhorn.dev> | 1 个月前 | |
Krea 2 regional prompting (#9407) * feat(krea2): add regional prompting * fix(krea2): harden regional prompting * Address findings * test(krea2): cover regional prompting gaps in attention, denoise and canvas graph Adds the coverage the regional prompting change was missing: - attention: the processor's mask/sequence-length shape guard, and that a processor built without regional state ignores the shared mask - regional prompting: that only even main blocks actually apply the mask during a real transformer forward (the existing test only checked the processor map), and that an unsupported mask rank is rejected - denoise: that each CFG pass installs a mask sized for its own text sequence when the positive and negative prompts tokenize to different lengths - canvas graph: a composed suite running the real addRegions, addKrea2LoRAs and regional-guidance validators, covering validator filtering, the IP adapter collector teardown, the LoRA-before-regions ordering the regional encoders depend on, and per-region enhancers Each test was mutation-checked against a deliberately broken implementation to confirm it fails for the right reason. --------- Co-authored-by: Alexander Eichhorn <alex@eichhorn.dev> | 1 个月前 | |
feat(Model Support): add Krea-2-Turbo/Raw model + LoRA support (WIP) (#9304) * feat(krea2): add Krea-2-Turbo model + LoRA support (WIP) Integrate Krea-2-Turbo (krea/Krea-2-Turbo) text-to-image per NEW_MODEL_INTEGRATION.md: Krea2Transformer2DModel (single-stream MMDiT) + Qwen3-VL text encoder (12-layer hidden-state tap, 4D prompt_embeds) + reused Qwen-Image VAE + FlowMatchEulerDiscrete scheduler. Backend: - taxonomy: BaseModelType.Krea2, ModelType/ModelFormat.Qwen3VLEncoder, Krea2VariantType (Turbo = "krea2_turbo" to avoid Z-Image collision) - config probes: Main_Diffusers/Checkpoint_Krea2, Qwen3VLEncoder, LoRA_LyCORIS_Krea2 (text_fusion/time_mod_proj signature; excluded from the Qwen-Image probe to avoid double-match) - loaders for the diffusers pipeline + standalone Qwen3-VL encoder, with runtime workarounds for the HF model's version mismatches (AutoTokenizer, extra_special_tokens={}, rope_parameters->rope_scaling) - native sampling (pack/unpack, position_ids, linear-mu shift) and hand-written Euler denoise loop; reuses qwen_image l2i/i2l - invocations: model_loader, text_encoder, denoise, lora_loader, plus two ecosystem enhancers (conditioning rebalance, seed variance) - LoRA conversion for diffusers PEFT (lora_transformer- prefix) Frontend: - 'krea-2' base + qwen3_vl_encoder type/format across model maps, buildKrea2Graph, addKrea2LoRAs, graph-builder denoise/base lists, optimal dimension 1024, regenerated schema.ts Fixes: - estimate transformer working memory in krea2_denoise so the cache reserves activation headroom and offloads more model under partial loading; fixes fp8 + LoRA OOM at 1024 (model was placed before LoRA patches were applied, leaving no room for their activations) WIP: requires diffusers main (>=0.39 dev) for Krea2Transformer2DModel; pyproject.toml temporarily pins diffusers to git main. * fix(krea2): support single-file VAE/encoder mix-and-match end-to-end Allow non-diffusers Krea-2 transformers (GGUF/fp8) to run with standalone single-file VAE + Qwen3-VL encoder, fixing several blockers found in testing. - buildKrea2Graph: drop the hard "requires Diffusers-format" assert; instead require both a VAE and a Qwen3-VL encoder to be selected when the transformer is not diffusers (mirrors readiness.ts). - Qwen3-VL encoder remap: handle both single-file key conventions — implicit (model.layers.*) and explicit (model.language_model.*). The old blind model.* -> language_model.* turned the bf16 file's keys into language_model.language_model.* (398 meta tensors -> "Cannot copy out of meta tensor" crash). Both files now load 0 missing / 0 unexpected / 0 meta. - Qwen3-VL tokenizer/config: broaden the offline-cache fallback from OSError to Exception so a partial HF cache (config present, vocab missing) re-fetches instead of dying with TypeError. - Qwen3-VL encoder fp8: keep an fp8 source checkpoint fp8-resident with per-layer upcast (storage float8_e4m3fn, compute bf16) instead of dequantizing to bf16. Halves resident VRAM (~8.9GB -> ~4.4GB), avoiding partial-load thrashing alongside a large transformer. Auto-enabled for fp8 sources on CUDA; bf16 files stay bf16. - Qwen-Image VAE: a native-layout qwen_image_vae single file is classified with the Anima base and loaded as AutoencoderKLWan, but the qwen l2i/i2l nodes need AutoencoderKLQwenImage. Add backend/krea2/vae_compat.py::as_qwen_image_vae to reinterpret a Wan VAE as AutoencoderKLQwenImage (state dicts are identical, 194/194 keys); both qwen VAE nodes use it. Idempotent for real QwenImage VAEs. * fix(krea2): re-apply Wan→QwenImage VAE adapter after upstream merge An upstream merge reintroduced the AutoencoderKLQwenImage isinstance asserts in the qwen VAE nodes (without the import → F821) and dropped the adapter in the i2l path. A native-layout qwen_image_vae single file is classified with the Anima base and loaded as AutoencoderKLWan, so the asserts fail at runtime. - qwen_image_latents_to_image: drop the reintroduced pre-device assert (the as_qwen_image_vae adapter inside model_on_device already handles the class). - qwen_image_image_to_latents: restore the as_qwen_image_vae import + adapter call, remove both asserts. - estimate_vae_working_memory_qwen_image only reads tensor shape + element size, so it runs correctly on either VAE class before the adapter. * fix(graph): make isMainModelWithoutUnet a type guard incl. krea2_model_loader krea2_model_loader was added to MainModelLoaderNodes but not to isMainModelWithoutUnet, and the guard wasn't a type predicate — so it never narrowed modelLoader. OutputFields of the loader union collapses to the common 'vae' field, making g.addEdge(modelLoader, 'unet', ...) in addInpaint/addOutpaint fail to type-check ('unet' not assignable to 'vae'). Redefine the guard as a type predicate keyed on the inverse (only main_model_loader/sdxl_model_loader expose a unet), so every transformer-based loader is treated as unet-less automatically and the negated branch narrows to the unet-bearing loaders. * test(krea2): add Krea2VariantType type-test to satisfy knip zKrea2VariantType was only used within common.ts (in zAnyModelVariant) and never referenced externally, so knip flagged it as an unused export. Every sibling variant enum avoids this by being asserted in common.test-d.ts; add the missing Krea2VariantType assertion, which both uses the export and verifies the manual zod enum matches the generated S['Krea2VariantType']. * build: pin diffusers to 0.39.0 (first stable release with Krea-2) diffusers 0.39.0 is the first stable release containing Krea2Pipeline / Krea2Transformer2DModel (plus the Qwen-Image VAE and Qwen3-VL text encoder Krea-2 relies on). Replace the temporary git-main dependency with the pinned release and update the lockfile's diffusers entry (version, sdist, wheel, specifier) to the official PyPI 0.39.0 artifacts. * feat(krea2): metadata recall, starter models, and config-probe tests Close the remaining gaps against docs/new-model-integration for Krea-2. Metadata recall (§7): buildKrea2Graph now records the standalone VAE, Qwen3-VL encoder, and both conditioning enhancers (seed-variance + rebalance) to image metadata, and parsing.tsx adds the matching recall handlers (base-guarded to 'krea-2'), so a Krea-2 image's VAE/encoder/enhancer settings restore on recall — important for reproducing single-file/GGUF generations. Starter models: add Krea-2 Raw (Base variant, full pipeline), Krea-2 Turbo GGUF Q4_K_M / Q8_0 (vantagewithai/Krea-2-Turbo-GGUF) with Qwen-Image VAE + Qwen3-VL encoder dependencies, and a standalone Qwen3-VL 4B encoder (Qwen/Qwen3-VL-4B- Instruct). The VAE dependency reuses the existing diffusers qwen_image_vae starter; krea2_turbo gains its explicit Turbo variant. Tests: add config-probe unit tests for Krea-2 variant detection (name heuristic, _has_krea2_keys, GGUF/checkpoint/diffusers variant, default settings) and for the single-file Qwen3-VL encoder probe (visual-tower vs. text-only Qwen3). 31 tests. * Chore Ruff * fix(krea2): validate denoise inputs and VAE compatibility * test(krea2): add loader, denoise, graph, listener, recall and starter-model coverage Backend: - test_krea2_state_dict_utils.py: cover the pure loader transforms (prefix strip, native<->diffusers conversion, scaled-fp8 dequant, Qwen3-VL key remap) and _reject_incomplete_load parametrized over the single-file/GGUF/encoder call sites (rejects meta-tensor partial loads, names the missing params) - test_krea2_denoise.py: _prepare_cfg_scale (broadcast/length/type), the per-step cfg list vs. img2img-clip regression, _validate_inputs happy path, and _get_noise determinism/shape - test_starter_models.py: Krea-2 bundle registration, diffusers+GGUF+standalone membership, and GGUF entries declaring their VAE + Qwen3-VL dependencies Frontend: - modelSelected.test.ts: Krea-2 standalone-component defaulting (auto-select on GGUF, anima-VAE fallback, no-overwrite, diffusers clears overrides, clear on switch away) - parsing.test.tsx: Krea2 VAE/encoder + enhancer recall gating (parses only for krea-2, never clobbers otherwise) - buildKrea2Graph.test.ts: CFG negative-conditioning gating, enhancer node insertion/chaining, non-diffusers standalone-model assertion, metadata - ImageMetadataActions.test.tsx: require all eight Krea2 recall handlers Fix: exclude krea-2 from the generic VAEModel metadata handler (it has a dedicated Krea2VAEModel handler), matching the existing z-image/flux2 exclusions. * fix(ui): wire Krea metadata recall actions * fix(krea2): address adversarial review findings * fix Krea-2 review findings * fix(krea2): use per-prompt position ids for the CFG uncond pass The rotary position ids (text tokens + image grid) were built once from the positive prompt's length and reused for the negative pass. When the negative prompt tokenizes to a different length than the positive one, the rotary embedding ends up a different length than the uncond query sequence and the transformer crashes in apply_rotary_emb ("tensor a (N) must match tensor b (M)"). Build a separate neg_position_ids from the negative prompt's length and pass it to the uncond transformer call. txt2img with CFG off (distilled Turbo) is unaffected — only the cond pass runs there. Adds a regression test that drives differing positive/negative prompt lengths and asserts len(position_ids) == text_len + image_tokens for each pass. * fix(krea2): resolve deep-review findings across encoder, loaders, LoRA and tokenization HIGH: - qwen3_encoder: _has_qwen_vl_visual_tower now also matches the nested model.visual.* layout (mirroring _is_qwen3_vl_encoder_state_dict), so a single-file Qwen3-VL 4B encoder no longer matches BOTH the text-only Qwen3 and the Qwen3-VL configs. The nondeterministic tie-break could register it as the wrong type and hide it from Krea-2's encoder dropdown, hard-blocking the single-file/GGUF install path. MEDIUM: - krea2_text_encoder: tokenize (prefix+prompt) and the assistant-turn suffix separately and concatenate, so prompts over the token budget keep the suffix (append-after-truncate) instead of having it silently cut off. LOW: - main: _get_krea2_variant_from_name lets "turbo" win and only matches "raw"/"base" as whole tokens, so Turbo files like "krea2_turbo_baseline_q4.gguf" are not read as Base. - krea2_lora_conversion_utils: raise a descriptive ValueError (not a bare KeyError) when a PEFT layer has lora_A without a matching lora_B. - factory: read config.json as UTF-8 so a non-ASCII config is not mis-treated as unrecognized (and the model dir wrongly rejected) under a cp1252 locale. - krea2 loader: _reject_incomplete_load also inspects named_buffers(), so a checkpoint missing a persistent buffer fails at load time rather than mid-inference. Tests: Qwen3-VL dual-match rejection, long-prompt suffix preservation, addKrea2LoRAs reroute, rebalance gains/validation, seed-variance determinism/out-of-place, variant filename heuristic, incomplete-LoRA error, UTF-8 config dir, meta-buffer rejection. * fix(krea2): reshape native final-layer modulation + honor scheduler shift config - loader: last.modulation.lin (native/GGUF) is reshaped to (2, hidden) to match diffusers Krea2FinalLayer.scale_shift_table, not just renamed. assign=True would otherwise install a flat 1-D parameter (which the meta-only completeness guard cannot catch), failing at inference on the primary GGUF/native path. Verified the final table is (2, hidden) and the per-block tables are (6, hidden) against the installed Krea2Transformer2DModel. - denoise: the resolution-aware timestep shift (mu) now reads base_shift/max_shift/ base_image_seq_len/max_image_seq_len from the loaded scheduler's config, falling back to the Krea-2 defaults, so a Raw checkpoint shipping a customized scheduler_config.json is sampled with its own shift parameters. Adds a converter test asserting last.modulation.lin -> final_layer.scale_shift_table is reshaped to (2, hidden). * fix(krea2): resolve remaining review findings * test(krea2): cover converter tensor shapes and scheduler-config mu path Add the two regression guards the loader/denoise fixes were missing: - Validate _convert_krea2_native_to_diffusers against the real Krea2Transformer2DModel (built on the meta device from KREA2_TRANSFORMER_CONFIG). Every scale_shift_table is sized from the actual module dims and asserted after conversion, pinning final_layer.scale_shift_table to (2, hidden) and each per-block table to (6, hidden). The stub-based boundary tests could not catch a wrong-shaped converted tensor, since load_state_dict(assign=True) installs any shape and _reject_incomplete_load only checks the meta device, not shapes. - Exercise the resolution-aware mu branch in Krea2Denoise._run_diffusion (shift=None, undistilled/Base config) so it is no longer dead: assert the mu passed to set_timesteps is derived from the loaded scheduler's base_shift/max_shift/base_image_seq_len/max_image_seq_len, and falls back to the Krea-2 defaults for absent keys. * fix(krea2): resolve final adversarial review findings * fix(krea2): calibrate seed variance to embedding std; fix randomize slider The Seed Variance enhancer added noise at an absolute magnitude (strength=20), so its effect depended on the embedding scale. Conditioning Rebalance multiplies the embeddings by up to ~20x, so with rebalance off the same noise overwhelmed the signal and prompt following collapsed (reported by lstein). Calibrate the noise to the embedding's standard deviation instead — the same approach the Z-Image Seed Variance enhancer already uses — so a given strength behaves consistently regardless of embedding scale. strength is now a std multiplier in [0, 2] (default 0.1); 0 or randomize_percent 0 is a no-op. Also fix the Randomize Percent slider: with sliderMin=1 and a coarse step of 5 the grid was anchored at 1 (1, 6, 11, 21, 26, ...) and never hit round tens. Anchor at 0 with a coarse step of 10, and relax the backend floor to ge=0.0. * Chore knip * fix(krea2): sync metadata recall ranges; accept diffusion_model LoRA layout Follow-up to the seed-variance recalibration: the metadata recall parsers still used the old ranges, so recalling an image dispatched state the backend rejects. - Krea2SeedVarianceStrength recall now parses 0..2 (the std-multiplier range), not 0..100 — recalling the old absolute value 20 no longer produces invalid state that buildKrea2Graph forwards to a failing generation. - Krea2SeedVarianceRandomizePercent recall now allows 0 (the disabled value), matching the slider, param state, and invocation. - LoRA_LyCORIS_Krea2_Config accepts a transformer-only LoRA using the diffusion_model.transformer_blocks.* layout under an explicit Krea-2 override; the converter already handles the diffusion_model. prefix. Adds range boundary tests for both recall parsers and diffusion_model.* LoRA accept/reject tests. * fix(krea2): reject orphan LoRA halves, invalid rebalance weights, incompatible VAE Three install/queue-time guards so malformed inputs are rejected up front instead of failing mid-generation: - LoRA identification now requires every lora_A/B (or lora_down/up) weight to have its partner half. A valid layer plus a dangling half previously installed and then crashed during LoRA conversion; both the explicit-override and the automatic-detection paths now validate completeness. - Krea-2 Conditioning Rebalance weights are validated as exactly 12 finite numbers before generation: in readiness (blocks the queue), in metadata recall (rejects instead of dispatching invalid state), and in the input field (isInvalid). Mirrors Krea2ConditioningRebalanceInvocation._parse_weights. - as_qwen_image_vae now requires the Qwen-Image geometry (16 latent channels, 8x spatial, no patchification) and rejects Wan 2.2's 48-channel / patchified VAE before encode/decode, rather than failing on 16-vs-48 normalization. Adds LoRA orphan-pair tests, rebalance-weight validator + recall tests, and Wan VAE geometry accept/reject tests. * Chore ruff + pnpm fix * fix(krea2): tighten LoRA validation + fix DoRA/alias conversion bugs Addresses five install/convert-time issues so malformed Krea-2 LoRAs and rebalance weights are rejected up front (or converted correctly): - Explicit Krea-2 override now rejects an orphaned lora_A/B (or lora_down/up) half anywhere in the state dict, not just under the approved prefixes — a transformer_blocks pair plus a dangling text_fusion half previously installed and then crashed during conversion. - Krea-2 LoRA detection now requires a complete weight pair; a file with only dora_scale (no A/B weights) is rejected instead of failing later on load. - Rebalance weights are restricted to decimal/scientific notation, rejecting the hex/binary/octal literals (0x10, 0b10, 0o10) that JS Number() accepts but the backend's Python float() rejects. - The converter now recognizes the standard PEFT/Diffusers DoRA magnitude key lora_magnitude_vector.weight, mapping it to dora_scale so a valid DoRA adapter loads as a DoRALayer instead of being split into a bogus layer. - Conflicting transformer./diffusion_model. aliases that normalize to the same target layer now raise explicitly instead of silently overwriting one. Adds tests for each case. * test(krea2): cover adversarial validation cases * feat(krea2): support native (ComfyUI) Krea-2 LoRAs; fix Anima misdetection Native Krea-2 LoRAs (e.g. sliders) name modules differently from InvokeAI's diffusers Krea2Transformer2DModel: diffusion_model.blocks.N with attn.wq/wk/wv/ wo/gate, mlp.{down,gate,up}, and a txtfusion stage. These were misidentified as Anima (whose strict detector matched the bare blocks.N.mlp.*) and, even when forced to Krea-2, could not be applied because the converter only understood the diffusers PEFT layout. - Add a verified 1:1 native->diffusers key remap in the Krea-2 LoRA converter (blocks->transformer_blocks, attn.wq/wk/wv->to_q/to_k/to_v, attn.wo->to_out.0, attn.gate->to_gate, mlp->ff, txtfusion->text_fusion). Every native module maps onto a real Linear in the diffusers model (checked against all 512 keys of a real slider LoRA). DoRA magnitude survives the remap. - Extend Krea-2 LoRA detection (config + converter) to recognize the native signature (txtfusion, or the gated attention attn.wq + attn.gate). - Tighten the Anima strict detector to require the Anima-specific mlp.layer_N / mlp_layerN naming instead of a bare mlp, so a native Krea-2 LoRA is no longer false-matched as Anima. No Anima/Wan regressions. Adds native remap, DoRA-through-remap, diffusers-untouched, and native identification tests. * fix(krea2): complete native LoRA normalization * fix(krea2): use memory-efficient attention to fit VRAM (was OOM/hang) Krea-2's transformer uses grouped-query attention (48 query / 12 KV heads) and its stock processor calls scaled_dot_product_attention with enable_gqa=True. PyTorch only supports enable_gqa on the math SDPA backend, which materializes the full O(seq^2) score matrix: ~6.75 GB per attention at 1280x720 (3600 tokens) and ~40 GB at 2560x1440. On builds without flash attention (e.g. Windows) there is no fused fallback, so generation either OOMs or the model cache offloads the transformer to RAM and the forward pass crawls. - Add Krea2MemoryEfficientAttnProcessor: expands the KV heads (repeat_interleave) so enable_gqa is not needed, and runs under the memory-efficient SDPA kernel (O(seq) memory, supports the padding mask). Numerically equivalent to the stock processor; measured ~6.75 GB -> ~1.41 GB per block at 3600 tokens. Installed on the transformer in krea2_denoise before the denoise loop. - Recalibrate _estimate_working_memory: with O(seq) attention the activation footprint is small and ~linear, so the previous ~2.6 MiB/token (O(seq^2)) figure no longer applies. The new estimate reserves realistic headroom (~8.5 GB at 2560x1440 instead of an impossible ~36 GB), so the idle Qwen3-VL encoder is evicted and the fp8 transformer stays resident on a 24 GB card. Adds processor equivalence tests (GQA and non-GQA) and a working-memory bound regression test. * fix(krea2): reject mixed-layout key collisions instead of silently overwriting Both Krea-2 key normalizers (native->diffusers transformer keys, ComfyUI single-file Qwen3-VL encoder keys) mapped each source key to one target key and wrote it straight into the output dict. A malformed mixed-layout checkpoint that carries both a native key and its already-normalized alias (e.g. blocks.0.attn.wq.weight and transformer_blocks.0.attn.to_q.weight, or a bare layers.1.weight and its model.-prefixed twin) collapses both onto one target key, and the surviving tensor depended on dict iteration order. Route every write through a shared _put_unique_key helper that raises an actionable RuntimeError naming both colliding source keys, so such a checkpoint fails at load time instead of silently dropping a tensor. Add order-independent collision regression tests for both normalizers. --------- Co-authored-by: Jonathan <34005131+JPPhoto@users.noreply.github.com> Co-authored-by: JPPhoto <jpollack@jpollackphoto.com> Co-authored-by: Lincoln Stein <lincoln.stein@gmail.com> | 2 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 1 个月前 | ||
| 1 个月前 | ||
| 2 个月前 |