Ssimple-zhengrefactor: rename Spark3 model to Spark2_5
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
refactor: rename Spark3 model to Spark2_5 | 1 个月前 | |
cmake : introduce semantic versioning (#26839) * cmake : introduce semantic versioning (wip) This commit introduces semantic versioning to llama.cpp. * squash! cmake : introduce semantic versioning (wip) * cmake : update test-cmake README notes [no ci] * include libmtmd in output so show its semversioned * ci : add make-release workflow * ci : fix build number check in build-cmake-pkg.yml * examples : remove trailing whitespace * ci : abort if upstream ggml version does not exist * ci : extract step contents into scripts * ci : add GGML_NATIVE=OFF to ubuntu job * examples : remove CI build information from test-cmake [no ci] This commit removes the nightly/release information that I added previously to keep this focused only on using building and installing llama.cpp with cmake and being able to quickly verify changes or troubleshoot issues. * ci : merge scripts into single script * remove -dev-build_number support This commit removes the incremental build number (versioning) support that I added. This was incorrect and we should only use the semver for the version. Releases will be tag a nightly build and package maintainers/managers that build from source can use the tag and it is therefor important that the correct version is reported. So a nightly-build will report the semver without the build number. The build number and commit as availble via cmake and test-cmake has been updated to include an example of using them: console $ ./build.sh [test-cmake] version: 0.1.0, build: 10360 (08c69e381) ... Refs: https://github.com/ggml-org/llama.cpp/pull/26839#discussion_r3755836969 * docs: add initial release.md documentation * cmake : clean-up and add LLAMA_BUILD_IS_DEV option * ci : remove version input from make-release job * ci : add LLAMA_BUILD_IS_DEV=OFF to build-cmake-pkg.yml Refs: https://github.com/danbev/llama.cpp/actions/runs/31576801921/job/94050639145 * docs : update release notes with LLAMA_BUILD_IS_DEV info [no ci] * ci : add TODO to winget workflow [no ci] --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> | 1 个月前 | |
llama : check LoRA tensor data is within file bounds (#27056) * llama : check LoRA tensor data is within file bounds * Update src/llama-adapter.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> | 1 个月前 | |
llama : re-enable manual LoRA adapter free (#19983) * Re-enable manual LoRA adapter free * Remove stale "all adapters must be loaded before context creation" stale comments | 6 个月前 | |
refactor: rename Spark3 model to Spark2_5 | 1 个月前 | |
refactor: rename Spark3 model to Spark2_5 | 1 个月前 | |
llama-batch: fix allowed decreasing pos in a seq (#25449) | 2 个月前 | |
llama-batch: add n_keep_tail in split_equal for recurrent models (#25278) | 2 个月前 | |
chat : add Granite 4.1 chat template (#23518) | 4 个月前 | |
chat : add Granite 4.1 chat template (#23518) | 4 个月前 | |
model : BailingMoE3 Support (#26608) * Adding support for bailingmoe3 * Adds speculative decoding support * Make BailingMoE3 safe gate metadata optional * bailingmoe3: apply trained SwiGLU clamps * common: fix Bailing V3 tool argument parsing * llama-model-saver, instantiate float vector metadata writer * bailingmoe3: support Q-LoRA (Ling-3.0-tiny) Ling-3.0-flash sets q_lora_rank: None and projects Q directly, so the current implementation loads a single ATTN_Q tensor. Ling-3.0-tiny sets q_lora_rank: 256 and routes Q through a LoRA bottleneck instead: q_a_proj -> q_a_layernorm -> q_b_proj Conversion therefore failed with: ValueError: Can not map tensor 'model.layers.3.attention.q_a_layernorm.weight' Add the missing path, mirroring the existing deepseek2 MLA implementation: * constants.py - add ATTN_Q_A / ATTN_Q_B / ATTN_Q_A_NORM to BAILINGMOE3 * tensor_mapping.py - map model.layers.{bid}.attention.q_{a,b}_proj and q_a_layernorm * conversion - emit attention.q_lora_rank when the config has it * bailingmoe3.cpp - read n_lora_q; create the Q-LoRA tensors and build Q through the bottleneck when q_lora_rank > 0 Everything is gated on q_lora_rank > 0. Ling-3.0-flash's config has no q_lora_rank, the converter only emits the key when present, hparams.n_lora_q defaults to 0, and get_key(..., required=false) leaves the target untouched when the key is absent - so flash keeps taking the existing direct-Q branch. The LoRA path produces the same shape as the direct projection, so the nope/rope split, RoPE application and wk_b absorption downstream are unchanged. * small mtp change * bailingmoe3: support separate MTP GGUF and Q-LoRA MTP * gguf: remove duplicate add_kda_gate_lower_bound definition --------- Co-authored-by: bloomer <bloomer@booper.brushtail.me> Co-authored-by: Dyluhn <dylanranejohnston1@gmail.com> | 1 个月前 | |
llama: refactor fused ops (#24646) | 2 个月前 | |
cparams : rename LLAMA_MAX_PARALLEL_SEQUENCES to LLAMA_MAX_SEQ (#14188) ggml-ci | 1 年前 | |
llama : support multi-output backend sampling (#25532) * Enable backend sampling with token speculation * Clamp the mask sum before converting it into the sampled index * Add a numeric context parameter declaring the maximum outputs one sequence * More fixes * Don't reuse memory for output views. * Match dist between CPU and GPU * Fix CPU and backend sampling mismatches * Simpify some of the changes * Fix tests on Vulkan * More test fixes * Rebase changes * Rebase and address review comments * Address review comments * Address review comments * Update src/llama-sampler.cpp Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> | 1 个月前 | |
mtmd: support Qwen3-TTS (note: breaking change to llama-tts binary) (#26254) * convert text model * main model load ok * convert encoder ok * speaker encoder loading ok * speaker enc graph * adapt vocab for backbone (with some tricks) * add suppress_tokens * poc new mtmd gen api * convert code_predictor to gguf * load gen_code model ok * add clip_encode * wire up * code gen cgraph init version Co-authored-by: Pascal <admin@serveurperso.com> * code2wav convert to gguf * code2wav graph ok * wire up in/out * (wip) subgraph * wire up * wip, correct code2wav * demo (to be removed) * code2wav preserve kv between calls * demo voice clone * llama: add llama_model_get_tok_embd * mtmd_helper_gen_audio API * fix clamp cold prefix Co-authored-by: Pascal <admin@serveurperso.com> * fuse snake op Co-authored-by: Pascal <admin@serveurperso.com> * demo: use proper sampling * update dev docs * polymorphism helper * revamp llama-tts binary * update docs * fix compile * fix lint * nits * add guide + docs * more timings info * clean up code comments * security fixes * update docs * use ggml_build_forward_select, clean up comments * fix ci * use ISO 639-1 language code * rename CODE2WAV --> GEN_WAV, update docs * clean up * clean up tts.cpp * add seq_id * add step_prompt() * mtmd_helper_model_can_chat * clean up comments --------- Co-authored-by: Pascal <admin@serveurperso.com> | 2 个月前 | |
grammar : degrade max repetition >= 2000 to unbounded (#26613) | 2 个月前 | |
common/grammar : replace problematic backtracking regex [\s\S]* (#18342) * grammar : add support for std::regex_search() with trigger patterns * common : update hermes2 pro trigger to search instead of match * common : use regex_search with anchoring for partial matching * common : adjust regex partial tests to use new pattern * grammar : check pattern directly instead of adding a type * common : adjust existing patterns to match new semantics | 8 个月前 | |
model: add Kimi-K3 text model (#26185) * model: add Kimi-K3 text model Hybrid KDA (linear) + MLA (full) attention as in Kimi-Linear-48B, plus five things that architecture does not have: 1. cross-layer residual attention (attn_res_block_size) 2. latent MoE (routed experts run at n_expert_latent) 3. situ activation (replaces SwiGLU everywhere) 4. MLA output gate (sigmoid gate before o_proj) 5. full-rank KDA gate (single ssm_g instead of ssm_g_a/ssm_g_b) K3's text_config reports KimiLinearForCausalLM - the older 48B architecture - so get_model_architecture routes on the top-level name instead. The KDA decay gate has two forms, selected by linear_attn_config's gate_lower_bound. It is not a clamp: when set it swaps the activation entirely (fla/ops/kda/gate.py), from -exp(A_log)*softplus(x) to lower_bound*sigmoid(exp(A_log)*x). K3 sets it to -5.0; kimi-linear leaves it unset, so that path is unchanged. Cross-layer residuals reuse ggml_dsv4_hc_pre for the weighted sum. That op is CPU + CUDA only, so Metal/Vulkan will fall back per-node until those kernels exist. The routed experts ship as compressed-tensors "mxfp4-pack-quantized". That is bit-compatible with ggml's MXFP4 - same E2M1 code assignment, same E8M0 scale byte, only the nibble positions within a block differ - so they are repacked rather than dequantized, losslessly and without a ~5.5 TB bf16 round-trip. The repack is built lazily because gguf_writer holds every added tensor until the final write. DeepSeek-V4 was already doing the identical bit-shuffling, so it now shares the helper. Verified against Moonshot's own code path (transformers + fla's Triton KDA kernels) on a tiny model exercising every K3-specific feature. Final-position logits vs the fp32 reference: 6.7e-05 rel / corr 1.00000000 for both the chunked and the recurrent delta-net path. MXFP4 blocks dequantize to the source weights with 0.0e+00 error. Assisted-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * model: fix ty errors in the Kimi-K3 converter - _res_parts buffers (kind, tensor) pairs, not bare tensors - get_tensors must return an Iterator, matching ModelBase - LazyBase's func takes one argument, so pass the expert loaders through args instead of the closure - borrowing KimiLinearModel.set_vocab from an unrelated TextModel is deliberate and safe, but not expressible in the signature No behaviour change: the MXFP4 repack still dequantizes to the source weights with 0.0e+00 error and end-to-end logits are unchanged (8.386e-03 rel, corr 0.99996630). Assisted-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Update conversion/kimi_k3.py Co-authored-by: Boris Dvorkin <b_dvorkin@niuitmo.ru> * Increase LLAMA_MAX_EXPERTS from 512 to 1024 * tests : support for Kimi K3 in archs test * chat : add Kimi K3 chat format (reasoning, content, typed tool calls) K3's assistant output is an XTML-ish tagged format built by the template's open_tag/close_tag macros. Two properties break generic parsing: 1. The generation prompt ends with open_tag('think'), so the completion starts inside the think section with no opening marker in the output (thinking_forced_open). 2. Only <|open|>/<|close|>/<|sep|>/<|end_of_msg|> are special tokens; tag names ("think", "response", "message") are ordinary text tokens. Adds common_chat_params_init_kimi_k3 (PEG_NATIVE) with detection on the marker trio, reasoning extraction, response unwrapping, and tool-call parsing of the tools/call/argument tag structure with argument types taken from the tool schema. Includes the K3 chat template fixture and 9 test-chat cases derived from real generations of the full 2.8T model. Verified end-to-end against Kimi-K3-Q2_K (GrEarl/Kimi-K3-GGUF) on 8x B200: content, reasoning_content, streaming deltas, and tool_calls all correct; finish_reason stop/tool_calls as appropriate. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chat : add message_delimiters for Kimi K3 Per-role message-start markers for token-level span splitting. User and assistant messages carry only the role attribute, so their full opener (through <|sep|>) is used; system and tool messages continue with more attributes (type=/tool=/index=), so those delimiters stop after the role's closing quote. Verified against the K3 tiktoken vocabulary that the closing quote is always a standalone token across all attribute variants, so the token-level prefix match stays exact. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: apply nits from @ngxson and text fixes from @danielhanchen * tests : added missing hyperparameters and tensors for Kimi K3 in test-llama-archs * chore : move overly verbose header file comments to Kimi K3 source file * tests : re-enabled KIMI_K3 in test-llama-archs for WebGPU backend * model-saver : emit kda_gate_lower_bound for Kimi K3 Quick fix. The Kimi K3 loader reads kda_gate_lower_bound and gates a graph branch on it (it scales the KDA gate when the bound is above -INFINITY), but the model saver never wrote the key, so a save->load roundtrip silently dropped it back to the -INFINITY default and changed the model's output. The real K3 config sets gate_lower_bound = -5.0. I propose to emit it from the saver, and set it to -5.0 in the test-llama-archs K3 case so the roundtrip check exercises it (the roundtrip fails without the saver line). * Refactor conditional for model architecture check * tests : re-enabled (again) KIMI_K3 and MINIMAX_M3 in test-llama-archs for WebGPU backend * fix code comments * add template on conversion * move repack_mxfp4_blocks to model base * nits * add_value_length * optimize res_stack construction * nits --------- Co-authored-by: Boris Dvorkin <b_dvorkin@niuitmo.ru> Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com> Co-authored-by: Deepankar Singh <singh.deepankar39@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Caleb DeLeeuw <caleb.deleeuw@gmail.com> Co-authored-by: Xuan Son Nguyen <son@huggingface.co> | 1 个月前 | |
model: add Kimi-K3 text model (#26185) * model: add Kimi-K3 text model Hybrid KDA (linear) + MLA (full) attention as in Kimi-Linear-48B, plus five things that architecture does not have: 1. cross-layer residual attention (attn_res_block_size) 2. latent MoE (routed experts run at n_expert_latent) 3. situ activation (replaces SwiGLU everywhere) 4. MLA output gate (sigmoid gate before o_proj) 5. full-rank KDA gate (single ssm_g instead of ssm_g_a/ssm_g_b) K3's text_config reports KimiLinearForCausalLM - the older 48B architecture - so get_model_architecture routes on the top-level name instead. The KDA decay gate has two forms, selected by linear_attn_config's gate_lower_bound. It is not a clamp: when set it swaps the activation entirely (fla/ops/kda/gate.py), from -exp(A_log)*softplus(x) to lower_bound*sigmoid(exp(A_log)*x). K3 sets it to -5.0; kimi-linear leaves it unset, so that path is unchanged. Cross-layer residuals reuse ggml_dsv4_hc_pre for the weighted sum. That op is CPU + CUDA only, so Metal/Vulkan will fall back per-node until those kernels exist. The routed experts ship as compressed-tensors "mxfp4-pack-quantized". That is bit-compatible with ggml's MXFP4 - same E2M1 code assignment, same E8M0 scale byte, only the nibble positions within a block differ - so they are repacked rather than dequantized, losslessly and without a ~5.5 TB bf16 round-trip. The repack is built lazily because gguf_writer holds every added tensor until the final write. DeepSeek-V4 was already doing the identical bit-shuffling, so it now shares the helper. Verified against Moonshot's own code path (transformers + fla's Triton KDA kernels) on a tiny model exercising every K3-specific feature. Final-position logits vs the fp32 reference: 6.7e-05 rel / corr 1.00000000 for both the chunked and the recurrent delta-net path. MXFP4 blocks dequantize to the source weights with 0.0e+00 error. Assisted-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * model: fix ty errors in the Kimi-K3 converter - _res_parts buffers (kind, tensor) pairs, not bare tensors - get_tensors must return an Iterator, matching ModelBase - LazyBase's func takes one argument, so pass the expert loaders through args instead of the closure - borrowing KimiLinearModel.set_vocab from an unrelated TextModel is deliberate and safe, but not expressible in the signature No behaviour change: the MXFP4 repack still dequantizes to the source weights with 0.0e+00 error and end-to-end logits are unchanged (8.386e-03 rel, corr 0.99996630). Assisted-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Update conversion/kimi_k3.py Co-authored-by: Boris Dvorkin <b_dvorkin@niuitmo.ru> * Increase LLAMA_MAX_EXPERTS from 512 to 1024 * tests : support for Kimi K3 in archs test * chat : add Kimi K3 chat format (reasoning, content, typed tool calls) K3's assistant output is an XTML-ish tagged format built by the template's open_tag/close_tag macros. Two properties break generic parsing: 1. The generation prompt ends with open_tag('think'), so the completion starts inside the think section with no opening marker in the output (thinking_forced_open). 2. Only <|open|>/<|close|>/<|sep|>/<|end_of_msg|> are special tokens; tag names ("think", "response", "message") are ordinary text tokens. Adds common_chat_params_init_kimi_k3 (PEG_NATIVE) with detection on the marker trio, reasoning extraction, response unwrapping, and tool-call parsing of the tools/call/argument tag structure with argument types taken from the tool schema. Includes the K3 chat template fixture and 9 test-chat cases derived from real generations of the full 2.8T model. Verified end-to-end against Kimi-K3-Q2_K (GrEarl/Kimi-K3-GGUF) on 8x B200: content, reasoning_content, streaming deltas, and tool_calls all correct; finish_reason stop/tool_calls as appropriate. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chat : add message_delimiters for Kimi K3 Per-role message-start markers for token-level span splitting. User and assistant messages carry only the role attribute, so their full opener (through <|sep|>) is used; system and tool messages continue with more attributes (type=/tool=/index=), so those delimiters stop after the role's closing quote. Verified against the K3 tiktoken vocabulary that the closing quote is always a standalone token across all attribute variants, so the token-level prefix match stays exact. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: apply nits from @ngxson and text fixes from @danielhanchen * tests : added missing hyperparameters and tensors for Kimi K3 in test-llama-archs * chore : move overly verbose header file comments to Kimi K3 source file * tests : re-enabled KIMI_K3 in test-llama-archs for WebGPU backend * model-saver : emit kda_gate_lower_bound for Kimi K3 Quick fix. The Kimi K3 loader reads kda_gate_lower_bound and gates a graph branch on it (it scales the KDA gate when the bound is above -INFINITY), but the model saver never wrote the key, so a save->load roundtrip silently dropped it back to the -INFINITY default and changed the model's output. The real K3 config sets gate_lower_bound = -5.0. I propose to emit it from the saver, and set it to -5.0 in the test-llama-archs K3 case so the roundtrip check exercises it (the roundtrip fails without the saver line). * Refactor conditional for model architecture check * tests : re-enabled (again) KIMI_K3 and MINIMAX_M3 in test-llama-archs for WebGPU backend * fix code comments * add template on conversion * move repack_mxfp4_blocks to model base * nits * add_value_length * optimize res_stack construction * nits --------- Co-authored-by: Boris Dvorkin <b_dvorkin@niuitmo.ru> Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com> Co-authored-by: Deepankar Singh <singh.deepankar39@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Caleb DeLeeuw <caleb.deleeuw@gmail.com> Co-authored-by: Xuan Son Nguyen <son@huggingface.co> | 1 个月前 | |
model : add support for MiniMaxText01ForCausalLM and MiniMaxM1ForCausalLM (#27018) * llama : support for MiniMax-Text-01 model * chore : renames to match the other MiniMax models * model : add logits mask as MiniMax-Text-01 embeddings tensor has zero-valued embeddings for tokens >= 200032 that produce zero logits disrupting the token sampling process * llama : replace hardcoded conditions with hparams.is_recr() * model : used build_rs() for recurrent state management * chore : code cleanup * model : optimized MiniMax-Text-01 by removing the state tranpose operations * chore : removed unnecessary ggml_cont() in MiniMax-Text-01 implementation * llama : add generic logits mask graph input * model : permuted diag_decay dimensions to avoid doing it inside MiniMax-Text-01 graph * chore : code cleanup * chore : code cleanup * model : use token positions when calculating MiniMax-Text-01 decay tensors * convert : add support for MiniMaxM1ForCausalLM as it seems to be the same as MiniMaxText01ForCausalLM * chat : add jinja template for MiniMax-M1 Co-authored-by: QscQ <qscqesze@gmail.com> * chore : code cleanup * tests : MINIMAX_01-related fixes * chore : silence Python lint errors * vocab : remove unnecessary vocab type * convert : update MiniMaxText01Model conversion to use yield when modifying tensors * convert : suppress tokens with zero-valued embeddings during MiniMax-Text-01 conversion * llama : removed logits mask - no longer necessary as token suppression is used instead * model : use common functions to make MiniMax-Text-01 implementation more concise Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> * model : use common functions to make MiniMax-Text-01 implementation more concise Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> * convert : override non-working built-in chat template during conversion * tests : skip arch MINIMAX_01 tests for WebGPU backend (it breaks again) --------- Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com> Co-authored-by: QscQ <qscqesze@gmail.com> Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> | 1 个月前 | |
model : BailingMoE3 Support (#26608) * Adding support for bailingmoe3 * Adds speculative decoding support * Make BailingMoE3 safe gate metadata optional * bailingmoe3: apply trained SwiGLU clamps * common: fix Bailing V3 tool argument parsing * llama-model-saver, instantiate float vector metadata writer * bailingmoe3: support Q-LoRA (Ling-3.0-tiny) Ling-3.0-flash sets q_lora_rank: None and projects Q directly, so the current implementation loads a single ATTN_Q tensor. Ling-3.0-tiny sets q_lora_rank: 256 and routes Q through a LoRA bottleneck instead: q_a_proj -> q_a_layernorm -> q_b_proj Conversion therefore failed with: ValueError: Can not map tensor 'model.layers.3.attention.q_a_layernorm.weight' Add the missing path, mirroring the existing deepseek2 MLA implementation: * constants.py - add ATTN_Q_A / ATTN_Q_B / ATTN_Q_A_NORM to BAILINGMOE3 * tensor_mapping.py - map model.layers.{bid}.attention.q_{a,b}_proj and q_a_layernorm * conversion - emit attention.q_lora_rank when the config has it * bailingmoe3.cpp - read n_lora_q; create the Q-LoRA tensors and build Q through the bottleneck when q_lora_rank > 0 Everything is gated on q_lora_rank > 0. Ling-3.0-flash's config has no q_lora_rank, the converter only emits the key when present, hparams.n_lora_q defaults to 0, and get_key(..., required=false) leaves the target untouched when the key is absent - so flash keeps taking the existing direct-Q branch. The LoRA path produces the same shape as the direct projection, so the nope/rope split, RoPE application and wk_b absorption downstream are unchanged. * small mtp change * bailingmoe3: support separate MTP GGUF and Q-LoRA MTP * gguf: remove duplicate add_kda_gate_lower_bound definition --------- Co-authored-by: bloomer <bloomer@booper.brushtail.me> Co-authored-by: Dyluhn <dylanranejohnston1@gmail.com> | 1 个月前 | |
llama : correct platform-independent loading of BOOL metadata (#21428) * model-loader : fix GGUF bool array conversion * model-loader : fix remaining GGUF bool pointer uses | 5 个月前 | |
llama: refactor fused ops (#24646) | 2 个月前 | |
server : avoid checkpoint data host copies (#22558) * server : avoid checkpoint data host copies * llama : refactor llama_io_read_i | 5 个月前 | |
llama : add option to save memory in device buffers (#22679) * llama : add option to save memory in device buffers * tests : extend llama-save-load-state | 5 个月前 | |
llama : allocate indexer cache only in "full" indexer layers (#26474) Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com> | 2 个月前 | |
llama : allocate indexer cache only in "full" indexer layers (#26474) Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com> | 2 个月前 | |
DeepseekV4 MTP + DSpark (#25784) | 2 个月前 | |
DeepseekV4 MTP + DSpark (#25784) | 2 个月前 | |
llama-batch: add n_keep_tail in split_equal for recurrent models (#25278) | 2 个月前 | |
DeepSeek V4 (#24162) * convert: add dsv4 conversion * add basic setup * add llm_graph_input_dsv4 * add save-load state * add sinkhorn eps - correction by @fairydreaming * add rope fix * cleanup dead code * fix bugs * support pro model: added by @fairydreaming * remove redundant V cache * Chat template * remove debugging leftovers * Add mechanism for inlining templates based on architecture * s/deepseek-v4-flash/deepseek4/g * s/deepseek-v4-flash/deepseek4/g continued * enable graph reuse * enable FA * fix test llama archs * rename * compatibility with antirez ds4 GGUFs * simplified set_gguf_parameters() by calling super class method, replaced moe.score_func with expert_gating_func. * reserve worst-case kv-cache * revert max split inputs * address review comments * add padding to enable FA * pad only the final value of plan.n_kv to 256 * remove built-in cpp chat template * cont: remove cpp built-in template * rm outdated test * replace ggml_view_3d() with ggml_reshape_3d() Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> * only support n_seq=1 for now * remove unused var * cont: remove unused var * use scale bias * use correct ptr for can_reuse * remove gen-chat-inline-templates.py * simplify graph reuse * cont: cleanup * remove unused inputs * enable partial checkpointing * add correct shape for kq_mask + set llama_model_n_swa to 0 for dsv4 * precompute source_idx + add comment about dummy write * support multi-seq * remove restored_trim_pos * use split_equal when possible * fix indent * address review comments * use LLM_KV * fix ci --------- Co-authored-by: Piotr Wilkin <piotr.wilkin@syndatis.com> Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com> Co-authored-by: Xuan Son Nguyen <son@huggingface.co> Co-authored-by: fairydreaming <166155368+fairydreaming@users.noreply.github.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> | 3 个月前 | |
model: M3: Move MSA into a new memory implementation (#26338) * Move MSA logic from llama-kv-cache into llama-kv-cache-msa * cont : minor * cont : ws fix --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> | 2 个月前 | |
model: M3: Move MSA into a new memory implementation (#26338) * Move MSA logic from llama-kv-cache into llama-kv-cache-msa * cont : minor * cont : ws fix --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> | 2 个月前 | |
model : Granite-Switch Architecture (#25107) * granite-switch: add llama.cpp backend (POC, CPU) New "granite-switch" architecture: a dense, all-attention Granite-4.1 model with N embedded LoRA adapters selected per-token by control tokens. - gguf-py schema (arch, KV keys, stacked LoRA tensor names) + writer helpers - conversion/granite.py: GraniteSwitchModel converter (stacks N adapters + zero base slot into per-projection A/B tensors; emits switch metadata) - C++ arch registration (llama-arch.{h,cpp}, llama-model.{h,cpp}) - src/models/granite_switch.cpp: load + per-token switched-LoRA graph via ggml_mul_mat_id over stacked tensors; sticky per-token index + control-token substitution in llm_graph_input_switch::set_input - llm_graph_input_switch in src/models/models.h Runs end-to-end on CPU: convert 3b checkpoint (842 tensors, stacked dim 13) and generate on both base and control-token paths. Sticky switch state is single-sequence (POC); full multi-sequence machinery is a follow-up. * granite-switch: add Mac (Metal) build + mid-sequence switch demo script Self-contained script to build llama.cpp on Apple Silicon (Metal), convert the composed 3b checkpoint, and run the crisp mid-sequence adapter-switch demos verified on Vela: - answerability: <|answerability|> mid-seq -> "unanswerable" - query_rewrite: <|query_rewrite|> mid-seq -> {"rewritten_question": ...} Each demo runs the same prompt twice, differing only by a control token placed before the assistant turn, so the per-token switch is visible. * granite-switch mac demo: add -no-cnv so each run is one-shot The composed model ships a chat template, so llama-completion auto-enables interactive conversation mode and halts at a > prompt after generating, stalling the script. -no-cnv disables conversation mode: generate once from the raw prompt and exit (also prints special tokens, making the switch visible). * granite-switch: replace global sticky index with in-graph router attention The POC computed the per-token adapter index on the CPU and carried it across ubatches in ONE global mutable int32_t poc_sticky_index, reset only when a ubatch contained sequence position 0. That global had two problems: 1. Concurrency: with multiple sequences in a batch it was last-writer- wins — one sequence's adapter leaked into the others. 2. Multi-turn: an interactive ollama run chat continues one KV cache, so turn 2 never saw position 0 and the index never reset — the adapter stayed stuck on across turns. Port the vLLM/HF backend mechanism faithfully: a single-head causal "router" attention recovers the adapter index in-graph. Per token, only dim 0 carries signal — Q[0]=1, K[0]=+gain for a control token / -gain otherwise, V[0]=adapter slot / 0 — and the causal softmax over the single visible control token recovers that adapter's slot (readback = clamp(round(V[0]), 0, n_adapters)). gain=15 matches config.py and is F16-safe (no F32 cache). The router's K/V live in the model KV cache at an extra layer R == hparams.router_layer (== n_layer). We bump n_layer_all to n_real+1 so the cache allocator gives the router its own per-sequence slot, and set n_layer_nextn=1 so n_layer() stays n_real — the decoder loop and tensor loading are untouched and never reference layer R. The router K is exempted from the k-shift RoPE loop (its dim-0 value is a literal magnitude, not a rotation). Because the selection now lives in the per-sequence KV cache, CONCURRENT requests are isolated for free (problem 1 fixed; verified by scratch/concurrent_switch_test.cpp). set_input becomes stateless pure per-token maps; the global is gone. Single-switch contract / known limitation, identical to vLLM & HF: the gain is flat (no recency), so within one sequence there is no mechanism to revert to base mid-sequence — once an adapter fires it stays on until that sequence ends (problem 2 is therefore NOT fixed by a faithful copy; vLLM/HF avoid it only because each served request is a fresh sequence). A client continuing one KV cache across turns must start a fresh sequence per turn, or opt into a recency-biased router (a deliberate divergence, not done here). Documented in granite_switch.cpp and asserted by scratch/multiturn_leak_test.cpp. Verified (CPU): both demos unchanged (answerability -> "unanswerable", query_rewrite -> rewritten query); concurrent two-sequence isolation passes; multi-turn carry-over matches the vLLM/HF contract. * granite-switch: drop scratch tests and mac demo for upstream PR Remove the local-only development artifacts that should not ship in the upstream PR: - granite-switch-mac-demo.sh (local Metal build + demo driver) - scratch/concurrent_switch_test.cpp - scratch/multiturn_leak_test.cpp Also drop the now-dangling reference to the scratch tests from the granite_switch.cpp header comment. Leaves only the core architecture support (conversion, gguf constants, llama-arch/model/kv-cache, and the granite_switch graph). * granite-switch: trim comments to match native llama.cpp style * granite-switch: trim conversion comments to match native style * granite-switch: drop unused adapter_ranks metadata * granite-switch: rename arch to graniteswitch and drop obid alias * granite-switch: fix non-ASCII comments and document router gain assumption * granite-switch: drop section comments from constants.py to match native style * granite-switch: add functional tensor block comments matching Granite4 Vision style * granite-switch: clarify n_expert_used comment State the actual constraint: mul_mat_id needs n_expert_used == 1, and since the GGUF carries expert_count = 0 the generic loader's n_expert == 0 => n_expert_used == 0 assertion has already passed by the time load_arch_hparams runs, so it is forced to 1 here. * granite-switch: note n_layer_nextn reuse has no MTP The router carving reuses n_layer_nextn, normally the MTP/next-token count. Clarify in the comment that it is borrowed here purely as the trailing-layers lever and that there is no MTP head, to spare readers the double-take. * granite-switch: rename source file and apply review nits * granite-switch: don't force LoRA tensors to F16, follow --outtype instead * granite-switch: drop redundant _permute_qk wrapper, call LlamaModel.permute directly * granite-switch: read router gain from GGUF (control_token_gain) instead of hardcoding 15.0 * granite-switch: derive n_slots() * granite-switch: move llm_graph_input_switch into granite-switch.cpp * granite-switch: cut AI-style narration comments * granite-switch: collapse multi-line comments * granite-switch: rename control_token_* maps to adapter_token_* * granite-switch: cut noise comments * granite-switch: rename embedded LoRA tensors to <base>.lora_a/lora_b * granite-switch: GGML_ASSERT token input to avoid UB on embeddings * granite-switch: TODO for raw embedding input support * granite-switch: collapse LoRA tensor constants to .lora_a/.lora_b suffix * granite-switch: drop n_expert_used hack, guard mul_mat_id buft probe * granite-switch: stop forcing dense expert counts, read from config * granite-switch: renamed control_token_gain metadata key to router_gain * granite-switch: trim header comments to match native style * granite-switch: collapse LoRA tensors to base name + suffix * granite-switch: inline suffix checks in tensor op resolution * granite-switch: drop switch-lora struct comment * granite-switch: guard router layer index and inline n_slots * granite-switch: group adapter metadata under {arch}.adapters.* namespace * granite-switch: add hparams.has_rope(il) for KV-shift rope skipping * granite-switch: skip arch in test-llama-archs (adapter fixture missing, TODO) * granite-switch: Keys.Adapters namespace + simplify n_slots * granite-switch: validate substitute token ids against n_vocab * granite-switch: bound adapter count and lora rank from GGUF * granite-switch: reject MTP context type when router_layer is set * granite-switch: throw on bad adapter metadata instead of GGML_ASSERT * granite-switch: use ASCII +/- in router K signal comment * granite-switch: document n_layer_nextn repurpose and its leak points * granite-switch: gate lora_a/lora_b op mapping on router_layer * granite-switch: label all three preview model sizes | 1 个月前 | |
model: M3: Move MSA into a new memory implementation (#26338) * Move MSA logic from llama-kv-cache into llama-kv-cache-msa * cont : minor * cont : ws fix --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> | 2 个月前 | |
kv-cache : avoid kv cells copies (#24277) | 3 个月前 | |
llama-batch: add n_keep_tail in split_equal for recurrent models (#25278) | 2 个月前 | |
llama + spec: MTP Support (#22673) * spec: support MTP * fix batch size * rename files * cont : simplify (#7) * MTP: clean-up (#9) * MTP: clean-up * review: use llama_context_type instead of llama_graph_type * review: remove llama_model_has_mtp * review: fix convert issues * convert: fix pycheck * review: formatting * use mtp- for identifying mtp models * convert: fix mtp conversion * mtp -> draft-mtp * remove unused llama_arch * add need_embd in speculative * llama: allow partial seq_rm for GDN models for speculative decoding Currently speculative checkpoint needs to restart from a checkpoint after some draft tokens are not accepted, this leads to some wastage in running the target again. This PR adds the ability to rollback upto draft_max by storing the GDN intermediates. * fix pending state * vulkan: add GDN partial rollback * meta: extend check to axis 1 * metal: add GDN partial rollback Extend the gated delta net kernel to store intermediate states for partial rollback support on the Metal backend. - Add K (snapshot slot count) as a function constant - Read input state from slot 0 of the 3D state tensor - Write intermediate states to different slots during token loop - For K=1, maintain backward-compatible single-slot behavior Ref: https://github.com/ggml-org/llama.cpp/commit/8c05923630110223669f069af2000e9cf10c02bc Assisted-by: llama.cpp:local pi * delta_net_base: use ggml_pad instead of new_tensor * review: add need_rs_seq * review: rename part_bounded to n_rs * review: deslop comments * review: rename, add asserts * server : adjust checkpoint logic (#11) * server : adjust checkpoint logic * cont : rm asserts * server-context: fix early exit * spec : fix compatibility with n-gram and add TODOs (#13) * metal : cleanup * llama : fix faulty bitwise check in recurrent memory * server : disable RS-based MTP in combination with other spec types * spec : add TODOs * cont : fix comment * cont : update comment * common : fix logic for ngram + mtp compat * llama-memory: enable checkpointing with partial rollback * cont: add test-case for loading into a dirty ctx * llama-memory-recurrent: clear rs_idx in clear * download: fix mtp path * llama-arch: fix enorm op * docs: update docs * conversion: fix type annotations --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> | 4 个月前 | |
llama-batch: add n_keep_tail in split_equal for recurrent models (#25278) | 2 个月前 | |
llama + spec: MTP Support (#22673) * spec: support MTP * fix batch size * rename files * cont : simplify (#7) * MTP: clean-up (#9) * MTP: clean-up * review: use llama_context_type instead of llama_graph_type * review: remove llama_model_has_mtp * review: fix convert issues * convert: fix pycheck * review: formatting * use mtp- for identifying mtp models * convert: fix mtp conversion * mtp -> draft-mtp * remove unused llama_arch * add need_embd in speculative * llama: allow partial seq_rm for GDN models for speculative decoding Currently speculative checkpoint needs to restart from a checkpoint after some draft tokens are not accepted, this leads to some wastage in running the target again. This PR adds the ability to rollback upto draft_max by storing the GDN intermediates. * fix pending state * vulkan: add GDN partial rollback * meta: extend check to axis 1 * metal: add GDN partial rollback Extend the gated delta net kernel to store intermediate states for partial rollback support on the Metal backend. - Add K (snapshot slot count) as a function constant - Read input state from slot 0 of the 3D state tensor - Write intermediate states to different slots during token loop - For K=1, maintain backward-compatible single-slot behavior Ref: https://github.com/ggml-org/llama.cpp/commit/8c05923630110223669f069af2000e9cf10c02bc Assisted-by: llama.cpp:local pi * delta_net_base: use ggml_pad instead of new_tensor * review: add need_rs_seq * review: rename part_bounded to n_rs * review: deslop comments * review: rename, add asserts * server : adjust checkpoint logic (#11) * server : adjust checkpoint logic * cont : rm asserts * server-context: fix early exit * spec : fix compatibility with n-gram and add TODOs (#13) * metal : cleanup * llama : fix faulty bitwise check in recurrent memory * server : disable RS-based MTP in combination with other spec types * spec : add TODOs * cont : fix comment * cont : update comment * common : fix logic for ngram + mtp compat * llama-memory: enable checkpointing with partial rollback * cont: add test-case for loading into a dirty ctx * llama-memory-recurrent: clear rs_idx in clear * download: fix mtp path * llama-arch: fix enorm op * docs: update docs * conversion: fix type annotations --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> | 4 个月前 | |
llama: various bug fixes (#26051) | 2 个月前 | |
llama : MTP clean-up (#23269) * llama : disable equal splits for recurrent memory with partial rollback * spec : re-enable p-min with MTP drafts * spec : re-enable ngram spec in combination with RS rollback * spec : fix ngram-map-* params * spec : fix acceptance logic in combined ngram + draft configs * graph : fix reuse for combined token + embd batches * spec : log parameters for each speculative implementation - add LOG_INF in each constructor with implementation type and parameters - extract device string logic into common_speculative_get_devices_str() - move 'adding speculative implementation' log from init into constructors Assisted-by: llama.cpp:local pi * spec : extend --spec-default with ngram-map-k4v Assisted-by: llama.cpp:local pi * minor : fix n_embd log * args : update draft.n_max == 3 + regen docs * spec : relax ngram-mod rejection thold to 0.25 @ 5 low * logs : improve * docs : update speculative decoding CLI argument documentation - Add missing draft model CPU scheduling and tensor override parameters - Update --spec-type to include all available types (excluding draft-eagle3 WIP) - Fix default values to match implementation (n_max=3, n_min=0, p_min=0.0) - Remove deprecated options (spec-draft-ctx-size, spec-draft-replace) - Add environment variables for new parameters Assisted-by: llama.cpp:local pi * arg : step-back on adding k4v to the default spec config * cont : fix name | 4 个月前 | |
memory : correctly handle failure in apply() (#14438) ggml-ci | 1 年前 | |
llama : add Gemma4 MTP (#23398) | 3 个月前 | |
Update llama-mmap to use ftello/fseeko (#22497) * Update llama-mmap to work with 32-bit wasm and >2GB models * Update to gguf.cpp style | 5 个月前 | |
llama: fix llama-model-saver (#20503) * llama : add fd-based model loading via llama_model_load_from_fd * llama : address review feedback for fd-based model loading * llama : use FILE pointer instead of fd in public API * llama : use FILE pointer consistently, address review feedback * fixup * fix tensor names * fix llama-model-saver * roundtrip tests * fixup * refactor tests * fix prints * fix model saving * fix CI, disable Chameleon * print seed --------- Co-authored-by: Siddhesh2377 <siddheshsonar2377@gmail.com> | 6 个月前 | |
quant : Optimise memory usage by evicting weights after processing each layer (#22877) * Evict weights from memory after processing each layer * Revert changes * Move unmap to libllama * Unmap weights offloaded to backend * Change member's constness * Remove unmap weights offloaded to backend | 1 个月前 | |
quant : Optimise memory usage by evicting weights after processing each layer (#22877) * Evict weights from memory after processing each layer * Revert changes * Move unmap to libllama * Unmap weights offloaded to backend * Change member's constness * Remove unmap weights offloaded to backend | 1 个月前 | |
refactor: rename Spark3 model to Spark2_5 | 1 个月前 | |
llama: fix llama-model-saver (#20503) * llama : add fd-based model loading via llama_model_load_from_fd * llama : address review feedback for fd-based model loading * llama : use FILE pointer instead of fd in public API * llama : use FILE pointer consistently, address review feedback * fixup * fix tensor names * fix llama-model-saver * roundtrip tests * fixup * refactor tests * fix prints * fix model saving * fix CI, disable Chameleon * print seed --------- Co-authored-by: Siddhesh2377 <siddheshsonar2377@gmail.com> | 6 个月前 | |
refactor: rename Spark3 model to Spark2_5 | 1 个月前 | |
model : BailingMoE3 Support (#26608) * Adding support for bailingmoe3 * Adds speculative decoding support * Make BailingMoE3 safe gate metadata optional * bailingmoe3: apply trained SwiGLU clamps * common: fix Bailing V3 tool argument parsing * llama-model-saver, instantiate float vector metadata writer * bailingmoe3: support Q-LoRA (Ling-3.0-tiny) Ling-3.0-flash sets q_lora_rank: None and projects Q directly, so the current implementation loads a single ATTN_Q tensor. Ling-3.0-tiny sets q_lora_rank: 256 and routes Q through a LoRA bottleneck instead: q_a_proj -> q_a_layernorm -> q_b_proj Conversion therefore failed with: ValueError: Can not map tensor 'model.layers.3.attention.q_a_layernorm.weight' Add the missing path, mirroring the existing deepseek2 MLA implementation: * constants.py - add ATTN_Q_A / ATTN_Q_B / ATTN_Q_A_NORM to BAILINGMOE3 * tensor_mapping.py - map model.layers.{bid}.attention.q_{a,b}_proj and q_a_layernorm * conversion - emit attention.q_lora_rank when the config has it * bailingmoe3.cpp - read n_lora_q; create the Q-LoRA tensors and build Q through the bottleneck when q_lora_rank > 0 Everything is gated on q_lora_rank > 0. Ling-3.0-flash's config has no q_lora_rank, the converter only emits the key when present, hparams.n_lora_q defaults to 0, and get_key(..., required=false) leaves the target untouched when the key is absent - so flash keeps taking the existing direct-Q branch. The LoRA path produces the same shape as the direct projection, so the nope/rope split, RoPE application and wk_b absorption downstream are unchanged. * small mtp change * bailingmoe3: support separate MTP GGUF and Q-LoRA MTP * gguf: remove duplicate add_kda_gate_lower_bound definition --------- Co-authored-by: bloomer <bloomer@booper.brushtail.me> Co-authored-by: Dyluhn <dylanranejohnston1@gmail.com> | 1 个月前 | |
quant : Optimise memory usage by evicting weights after processing each layer (#22877) * Evict weights from memory after processing each layer * Revert changes * Move unmap to libllama * Unmap weights offloaded to backend * Change member's constness * Remove unmap weights offloaded to backend | 1 个月前 | |
llama : refactor src/llama.cpp (#10902) * llama : scatter llama.cpp into multiple modules (wip) * llama : control-vector -> adapter * llama : arch * llama : mmap ggml-ci * ci : remove BUILD_SHARED_LIBS=OFF ggml-ci * llama : arch (cont) ggml-ci * llama : chat ggml-ci * llama : model ggml-ci * llama : hparams ggml-ci * llama : adapter ggml-ci * examples : fix ggml-ci * rebase ggml-ci * minor * llama : kv cache ggml-ci * llama : impl ggml-ci * llama : batch ggml-ci * cont ggml-ci * llama : context ggml-ci * minor * llama : context (cont) ggml-ci * llama : model loader ggml-ci * common : update lora ggml-ci * llama : quant ggml-ci * llama : quant (cont) ggml-ci * minor [no ci] | 1 年前 | |
llama : support multi-output backend sampling (#25532) * Enable backend sampling with token speculation * Clamp the mask sum before converting it into the sampled index * Add a numeric context parameter declaring the maximum outputs one sequence * More fixes * Don't reuse memory for output views. * Match dist between CPU and GPU * Fix CPU and backend sampling mismatches * Simpify some of the changes * Fix tests on Vulkan * More test fixes * Rebase changes * Rebase and address review comments * Address review comments * Address review comments * Update src/llama-sampler.cpp Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> | 1 个月前 | |
llama : support multi-output backend sampling (#25532) * Enable backend sampling with token speculation * Clamp the mask sum before converting it into the sampled index * Add a numeric context parameter declaring the maximum outputs one sequence * More fixes * Don't reuse memory for output views. * Match dist between CPU and GPU * Fix CPU and backend sampling mismatches * Simpify some of the changes * Fix tests on Vulkan * More test fixes * Rebase changes * Rebase and address review comments * Address review comments * Address review comments * Update src/llama-sampler.cpp Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> | 1 个月前 | |
refactor: rename Spark3 model to Spark2_5 | 1 个月前 | |
refactor: rename Spark3 model to Spark2_5 | 1 个月前 | |
llama : allow virtual igpu devices (#26953) * llama : allow virtual igpu devices * cont : better comment | 1 个月前 | |
server : better security control for public deployments (#9776) * server : more explicit endpoint access settings * protect /props endpoint * fix tests * update server docs * fix typo * fix tests | 1 年前 | |
llama : reduce compile time and binary size (#9712) * llama : speed up compile time * fix build * fix build (2) | 1 年前 | |
unicode : include '~' in collapsed symbol class (#26972) The collapsed \p{S} class was missing '~', which split " ~" into separate pre-tokens and prevented the Ġ~ BPE merge used by DeepSeek V4. This caused re-tokenized prompts to diverge from sampled tokens and broke KV cache reuse. Assisted-by: Codex | 1 个月前 | |
vocab: fix Gemma4 tokenizer (#21343) * seems to work * fix case with new line Co-authored-by: sayap <sokann@gmail.com> * gemma 4: fix pre tok regex --------- Co-authored-by: Xuan Son Nguyen <son@huggingface.co> Co-authored-by: sayap <sokann@gmail.com> | 6 个月前 |