ik_llama.cpp

mirror of https://github.com/ikawrakow/ik_llama.cpp.git synced 2026-06-28 04:30:15 -05:00

Author	SHA1	Message	Date
Kawrakow	0c45696db4	Minor logging cleanup (#1873 )	2026-05-24 07:29:32 +03:00
Kawrakow	809a63bbb7	Fix MLA models with ngl < n_layer (#1870 ) * Fix split mode graph with ngl < n_layer (MLA models) * It is actually not related to split mode graph	2026-05-24 07:29:17 +03:00
Kawrakow	a6bb509305	Fix split mode graph with ngl < n_layer (#1869 )	2026-05-23 12:58:09 +03:00
Kawrakow	3f45ba9387	MTP tweaks 3 (#1862 )	2026-05-23 07:23:20 +03:00
Samuel Oliveira Alves	19e09e81d4	Change MTP graph input preparation with additional parameters and validation checks (#1866 )	2026-05-23 07:22:04 +03:00
Kawrakow	b3d39cff8b	Fix split mode graph for Qwen35-MoE + MTP (#1861 )	2026-05-22 09:23:53 +03:00
thad0ctor	b26521b9ef	Fix raw-vs-local device id confusion under -dev/-devd subsets (#1826 ) llm_load_tensors stores `default_layer_device[i]` as a local index into `model.devices` (consistent with `device_mem[]`, `model.splits[]`, and all graph-building consumers), but the four `llama_default_buffer_type_offload(model, default_layer_device[i])` callsites passed it through as if it were a raw post-CVD device id. Under `-dev`/`-devd` subsets where `model.devices != {0..N-1}`, this selected the wrong buffer type. Wrap with `model.devices[...]` to match the existing `model.devices[main_gpu]` pattern on the adjacent lines. llama_init_from_model has the same bug for `main_gpu`: every consumer (auto-fit override at line 3428, MTP clamp, the `model.devices[main_gpu]` translations at lines 3678/3682, and graph-building `splits[main_gpu]`) treats it as a local index, but the five single-GPU backend init paths (CUDA, Vulkan, SYCL, Kompute, CANN) pass `model->main_gpu` straight to the backend init, which expects a raw device id. e.g. `-dev CUDA1` with default `--main-gpu 0` and `split_mode=NONE` called `ggml_backend_cuda_init(0)` instead of `cuda_init(1)`. Compute `main_gpu_id` once and use it for all five paths.	2026-05-22 08:32:52 +03:00
Kawrakow	48a55f74e4	Disable split mode graph for Qwen35-MoE when MTP is enabled (#1858 )	2026-05-21 16:29:35 +03:00
Samuel Oliveira Alves	7b73f45541	Add adaptive sampling clone and free functions to manage memory (#1851 )	2026-05-21 08:11:17 +03:00
David Young	aefb8bdd99	MLA TP -khad: ggml_dequant_hadamard fused op + wv_b/wk_b_pp Hadamard fold (#1852 ) * ggml: ggml_dequant_hadamard fused op for MLA -khad path Adds a new ggml op that fuses (ggml_cast -> F32) + (ggml_hadamard) into a single kernel. Reads a quantized (or F16/F32) source and produces a per- Hadamard-block F32 chunk with the inverse transform applied, without materializing a full-size F32 intermediate buffer. Motivation: the MLA pp_opt path in build_deepseek2.cpp un-encodes the H-applied cache_nope view at every PP call. Today that runs as a cast (quant -> F32) followed by a separate ggml_hadamard kernel, costing two full-size F32 passes per layer per rank per call. Fusing them halves the bandwidth on the un-encode and removes one kernel launch. CUDA kernels in dequant_hadamard.cu lift the Walsh-Hadamard butterfly from hadamard.cu and dequant helpers from dequantize.cuh: * qr=1 layout (q8_0): consecutive dequant pair, stage 1 fused with load * qr=2 layout (q4_0 / q4_1 / q5_0 / q5_1 / q6_0 / iq4_nl): dequant pair at stride qk/2, explicit stage 1 after sync * F16 has a dedicated kernel * F32 source falls back to the standalone Hadamard op CPU impl in iqk_cpu_ops.cpp composes the existing type_traits.to_float dequant with fast_ht for graph completeness. nh in {64, 128, 256, 512}. * MLA-TP: Hadamard pretransform of wv_b/wk_b_pp for -khad Fold the 64-block orthonormal Hadamard into wv_b and wk_b_pp once at context init so the pp_opt mul_mats consume the K cache in its on-disk encoded basis. The per-PP-call cache_nope un-Hadamard is then skipped (rope half still un-applied — it goes to FA via concat, no wk_b multiply). Math is identity by H^T H = I: mul_mat(H@wv_b, H@cache) = wv_b^T @ cache. For mla=2/3 absorb, composes correctly with the existing post-FA ggml_hadamard(kqv_compressed, 64). All-or-nothing across layers under a castable type-allowlist (excludes 1-3 bpw IQ types whose requant blows up beyond PPL noise). Models with ineligible weights fall back to the runtime un-Hadamard path unchanged. Composes with the fused ggml_dequant_hadamard op (prior commit): with the fold active only the rope half still runs the runtime transform, via the fused kernel. * MLA-TP: fix TG with -khad after wv_b/wk_b_pp fold The absorb branch of build_deepseek2_tp_attention applies ggml_hadamard to kqv_compressed after FA, then multiplies by wv_b. Pre-fold this was needed because wv_b was un-encoded; with the wv_b fold (prior commit) the mul_mat already expects H-encoded kqv_compressed: mul_mat(H @ wv_b, kqv_encoded) = wv_b^T @ H @ H @ kqv_unencoded = wv_b^T @ kqv_unencoded (H @ H = I) Skip the post-FA hadamard when model.khad_pretransformed is set so the two H applications cancel instead of double-applying. Affects the absorb branch: TG (n_tokens=1), short-context PP (n_kv < 1024), and models without wk_b_pp. Long-context PP goes through the pp_opt branch and is unrelated/unchanged. Reported by @ikawrakow on PR 1852. Verified across mla={1,2,3} x khad={on,off} x -ctk={q8_0,q4_0} on GLM-4.7-Flash IQ5_K and the unsloth IQ4_XS variant ik used to reproduce. * ggml_hadamard: accept F16 and quant sources; drop GGML_OP_DEQUANT_HADAMARD Per @ikawrakow review on PR 1852: subsume the per-source-type dispatch into the existing GGML_OP_HADAMARD instead of carrying a separate enum entry, op constructor, and standalone files. ggml_hadamard's API is unchanged from the call-site perspective. The constructor's F32-only assertion is dropped; ggml_cuda_op_hadamard and iqk_hadamard now dispatch internally: - F32 source: existing F32 butterfly (unchanged) - F16 source: dedicated kernel - q8_0 / q4_0 / q4_1 / q5_0 / q5_1 / q6_0 / iq4_nl: fused dequant + butterfly kernel (lifted from the deleted dequant_hadamard.cu) - CPU side composes traits.to_float with fast_ht Net diff: -80 lines. Removes dequant_hadamard.{cu,cuh}, the enum entry, op table rows, ggml_dequant_hadamard constructor, dispatch cases, and the DEQUANT_HADAMARD supports_op block. Verified clean build + TG smoke (mla=3 +khad q8 on GLM-4.7-Flash-IQ4_XS, same coherent output as prior commit on feat/dequant-hadamard).	2026-05-21 07:29:15 +03:00
Samuel Oliveira Alves	11a1fea9e2	Move embedding management to speculative (#1825 ) * refactor speculative decoding with companion context and draft result structures * feat: add common speculative feature handling in server context * refactor: move embedings outside server * feat: harden draft input hidden state in llama context * remove unused functions * refactor: streamline speculative feature handling and remove unused code * remove redundant code * remove more unused variables * refactor: implement speculative feature handling	2026-05-20 17:42:48 +03:00
David Young	dd67a9fb24	MLA TP prompt processing optimisation (#1841 ) * MLA TP prompt processing optimisation Adds a per-rank prompt-processing path to build_deepseek2_tp_attention that materialises K/V from the compressed latent cache and runs a standard flash_attn instead of the FlashMLA-3 absorb kernel the TP attention currently uses for all batch sizes. Affects MLA archs under -sm graph (DEEPSEEK2, GLM_DSA, MISTRAL4). Gated on n_tokens >= 128 (set by caller) AND n_kv >= 1024. Below either threshold the absorb path runs unchanged. Token generation takes the absorb path; only prompt processing at non-trivial context materialises. A second piece pre-computes wk_b in a pp_opt-favouring orientation (wk_b_pp: [kv_lora_rank, qk_nope, n_head]) at llm_prepare_mla time, so the per-PP-call materialise can mul_mat against the latent cache directly without an F16 cast + permute + ggml_cont on wk_b each call. Path A (wkv_b in GGUF) and Path B (only wk_b/wv_b in GGUF) both populate wk_b_pp through the standard per-rank replica setup. Measured on 8x RTX 3090, -sm graph -mla 2 -fa on: DSV2.5 IQ2_XS c=8k ub=2048 PP +51% to +60% GLM-4.7-Flash IQ4_XS c=32k ub=2048 PP -6% (PP@0) to +77% (PP@30720) GLM-5.1 IQ1_S q4_0 c=16k ub=2048 PP +5% to +9% PPL parity within +/-0.2 noise (DSV2.5 bit-identical 5.3917, GLM-4.7 8.83 vs 8.96, GLM-5.1 6.96 vs 7.00). Token-generation throughput unchanged within noise. Compute buffer at init: DSV2.5 -54 MiB total (allocator noise) GLM-4.7-Flash +1042 MiB total (~+173 MiB per non-output device) GLM-5.1 0 (MoE intermediates dominate) * MLA TP: respect mla=1 vs mla=3 distinction, rename attn_k_b_pp -> attn_kv_b ikawrakow/ik_llama.cpp#1841 review feedback: the pp_opt path lost the intended trade-off where mla=1 forgoes pp_opt to save VRAM and mla=3 pays the wk_b_pp tensor cost for faster long-context PP. - llm_prepare_mla second pass: gate wk_b_pp synthesis on mla > 1. Models that ship wk_b in their GGUF (mainline format) no longer allocate the pp_opt-favoring K weight under mla=1. - llm_prepare_mla first pass (wk_b synthesis from wkv_b): keep unconditional under -sm graph. The wk_b_pp materialization here shares the wk_b_f32 intermediate with the wk_b synthesis above, and isolating just the wk_b_pp branch leaves the synthesized wk_b in a state that makes the absorb path produce inf on some quant combos (DSV2.5 IQ2_XS). Trade: the synthesized-wkv_b path still pays the wk_b_pp allocation under mla=1, but the bigger compute-buffer saving (no pp_opt branch at runtime) still applies. - build_deepseek2 outer pp_opt: include cparams.mla_attn > 1 in the pp_opt definition itself, so mla=1 is bypassed throughout (TP and non-TP attention paths). - build_deepseek2 tp pp_opt: require wk_b_pp present. Drop the dead runtime wk_b transpose fallback (unreachable now that wk_b_pp is guaranteed when tp_pp_opt fires). - llama_kv_cache_init: have_wkv_b probe now treats wk_b_pp (attn_kv_b) as equivalent to wkv_b for the purposes of allowing mla>1 to stay put. Without this, -sm graph models that have wk_b/wv_b separately in the GGUF (no combined wkv_b) would silently downgrade to mla=1. - Rename the synthesized tensor "attn_k_b_pp.weight" -> "attn_kv_b.weight" to match the mainline naming ik uses. GLM-5.1 in particular benefits: its mla=3 PP improvement over mla=1 is negligible on this arch (~0.4% in our sweeps), so users save the runtime cost by sticking to mla=1.	2026-05-20 17:03:05 +03:00
Kawrakow	6bb3ee3a32	Enable split mode graph for MLA models and partial offload (#1835 )	2026-05-20 07:13:55 +03:00
Kawrakow	997c587a6c	Fix #1837 (#1838 )	2026-05-19 17:56:21 +03:00
David Young	c07a052315	MLA tensor parallelism under -sm graph (DEEPSEEK2/GLM_DSA/MISTRAL4) (#1821 ) * MLA tensor parallelism under -sm graph (DEEPSEEK2/GLM_DSA/MISTRAL4) Extends -sm graph (split-mode graph) to MLA-style attention across the DEEPSEEK2, GLM_DSA, and MISTRAL4 architectures. Previously these archs fell back to -sm layer regardless of the user's flag. Implementation: - Per-rank attention build in build_deepseek2_tp_attention with view-sliced FlashAttention, split-buffer output projection, and ggml_reduce across devices - wk_b / wv_b absorbed weights replicated per device via materialize() in llm_prepare_mla (these can't live in a split buffer) - KV cache replication path (replicated_k_l) for graph-mode TP - distribute_mla_tensors_for_split_mode_graph routes attention/norm tensors into ctx_split; expert tensors stay per-layer - Implements ggml_backend_cuda_split_buffer_get_tensor for the replicated / row-split / col-split inverse paths - Early-reject guard in src/llama.cpp that auto-downgrades -sm graph to -sm layer (with a warning) when incompatible loader flags are set: -ncmoe, -cmoe, -ot, -rtr, -muge New CLI flag: - -gap \| --graph-attn-precision <f16\|f32> (default f16) See the PR description for the full validation matrix (3 archs x 2/4/8 GPU counts), perf numbers, VRAM accounting, and known limitations. * Some tweaks * materialize lambda: per-head split for graph-mode tp_replicate 7dd19e19 changed wk_b/wv_b distribution from mirror to per-head split (split_dim=2) via prepare_split_tensors. That path only fires when wk_b/wv_b are loaded from GGUF. Models that store only wkv_b in GGUF derive wk_b/wv_b at load via llm_prepare_mla, going through the materialize lambda, which was untouched and still produced mirror replicas (split_dim=-1, full n_head per device). build_deepseek2_tp_attention now does mul_mat(wk_b_local, q_nope_perm) without the prior view_3d slice, so a mirror replica passes an n_head tensor where the kernel expects n_head_local. Result: silent SIGSEGV right after model load. Mirror logic in materialize is replaced with the same per-head split as prepare_split_tensors: head_offsets derived from wo split, each rank gets a tensor with ne[2]=n_head_local, data copied from the appropriate source byte slice. Singular `computed` tensor keeps full metadata for tensors_by_name lookups. Tested: 8x3090, -sm graph -mla 3 -fa on now boots cleanly and sweep-benches without crash. Log confirms new path: "Computed blk.X.attn_k_b.weight ... split across N devices on dim=2". * cleanup: indent fix + remove dead view_3d slicing and debug printf - build_deepseek2.cpp: re-indent the self_attention block in build_deepseek2_layer_attention (lines 253-670). Block was at column 0 inside a function body; now at the expected 4/8-space indent. - build_deepseek2.cpp: drop the commented-out view_3d slicing and debug printfs left over after 7dd19e19's switch to direct mul_mat on per-rank wk_b_local / wv_b_local. Update the stale 'wk_b is replicated (split_dim=-1)' comment to match the new split_dim=2 reality. - ggml-cuda.cu: remove the leftover debug printf in ggml_backend_cuda_split_buffer_get_tensor. No behavior change. Verified with a clean rebuild and DSV2.5 + GLM-4.7-Flash sweep-bench runs. * llm_load_tensors: gate incompatible-flag warning to MLA archs The -ncmoe / -rtr / -muge / -ot warning under -sm graph currently fires for all archs that support graph mode. That's an over-reach: the incompatibility is specific to the MLA TP paths (DEEPSEEK2, GLM_DSA, MISTRAL4) — Gemma4 graph mode existed pre-PR and works with those flags. Gate the warning to MLA archs only. Also refreshes two stale comments left over from the wk_b/wv_b mirror -> per-head-split rewrite: - src/llama.cpp llm_prepare_mla: "Replicate wk_b/wv_b ..." now reads "Per-head split wk_b/wv_b ..." to match what the materialize lambda actually does post-823a39e2. - src/llama-load-tensors.cpp distribute_mla_tensors_for_split_mode_graph: drop the wkv_b row-split mention (wkv_b is no longer created under graph mode after 7dd19e19) and correct the wk_b/wv_b distribution description (per-head split, not per-device replicated). --------- Co-authored-by: Kawrakow <iwankawrakow@gmail.com>	2026-05-19 08:36:17 +03:00
Kawrakow	a407b9ca3d	Fix Qwen3.6-MoE low MTP acceptance rate (#1815 ) * Fix Qwen3.6-MoE low MTP acceptance rate * Fix Gemma4 MTP	2026-05-18 07:26:17 +03:00
Kawrakow	0ab9bdf793	Fix Qwen3.5/3.6 MTP and -muge (#1816 )	2026-05-17 17:14:47 +03:00
Kawrakow	1f8c603d9c	Quantize: add extra output tensor for MTP (#1810 ) * Quantize: add extra output tensor for MTP * Consistently use --mtp-requantize-output-tensor	2026-05-17 13:59:56 +03:00
Kawrakow	3e573cfea6	MTP: option to use re-quantized output tensor for better TG performance (#1809 ) * Option to use re-quantized output tensor for MTP * Remove quantize extra output option * Handle interleaved types	2026-05-16 14:40:18 +03:00
Samuel Oliveira Alves	0fcffdb64d	feat: map Gemma 4 tensor and support with imatrix (#1796 )	2026-05-14 09:01:24 +03:00
Kawrakow	949bb8f1d6	More MTP tweaks (#1792 )	2026-05-13 17:55:43 +03:00
Kawrakow	397150caa2	MTP: faster recurrent state restore (#1791 ) * MTP: store ready per step convolution states * Cleanup	2026-05-13 11:00:24 +03:00
firecoperana	cdc288bc97	server: reset cache tokens after pp stops (#1787 ) Co-authored-by: firecoperana <firecoperana>	2026-05-13 09:05:32 +03:00
Kawrakow	cec1a6c1f5	MTP: Reuse graphs (again) (#1780 )	2026-05-12 07:36:12 +03:00
Samuel Oliveira Alves	be8435793e	Pre-allocate buffers for hybrid model checkpoints (#1774 ) * hybrid-spec: improve recurrent checkpoint handling in speculative decoding * change per-step save to support scheduling and asynchronous tensor operations * remove redudant backend tensor fallback * improve recurrent tensor handling for split graph	2026-05-12 07:21:25 +03:00
Kawrakow	eb570eb966	MTP: Avoid per step SSM copy (#1778 ) * Avoid copying the per-step SSM state (CUDA) * Avoid copying the per-step SSM state (CPU) * Allocate only what is necessary for per-step SSM state * Cleanup	2026-05-11 18:15:55 +03:00
Kawrakow	94940cd882	MTP: ebable per step recurrent state for split mode graph (#1773 )	2026-05-11 12:40:04 +03:00
Kawrakow	4bbdb8ed0b	Faster per step recurrent state restore when using MTP (#1767 )	2026-05-10 07:51:06 +03:00
Samuel Oliveira Alves	c2b8bca807	Add MTP Support for Gemma 4 (#1744 ) * gemma-mtp: build the arch to load the MTP model * gemma-mtp: fix mtp kv state * gemma-mtp: refactor some functions and create gguf * gemma-mtp: make usable for embeddings models variant * gemma-mtp: fix qwen mtp load in graph split * gemma-mtp: refactor tensor creation and adjust output tensor handling * Gemma 4 MTP: improve tensor handling, and adjust split mode logic	2026-05-10 07:44:20 +03:00
Kawrakow	2f0b47c19d	Use async copies to save/restore recurrent state (#1759 )	2026-05-09 08:31:56 +03:00
joelfarthing	9a26522af2	qwen35moe : support MTP tail layer (#1745 ) Co-authored-by: Joel Farthing <262452229+joelfarthing@users.noreply.github.com>	2026-05-07 15:46:41 +03:00
Kawrakow	e722f0bb73	MTP tweaks (#1741 )	2026-05-06 08:35:11 +03:00
Kawrakow	8b56d813a9	MTP improvements (#1736 ) * MTP improvements * Cleanup	2026-05-05 08:05:24 +03:00
dmaivel	38c200373f	Fix graph reuse regression with MTP checkpoints (#1728 )	2026-05-03 12:04:39 +03:00
Kawrakow	2dd3818083	MTP: better graph reuse (#1713 )	2026-05-03 08:15:21 +03:00
dmaivel	1beaaa002d	speculative: enable MTP per-step checkpoints with CPU recurrent layers (#1724 )	2026-05-03 08:14:56 +03:00
firecoperana	b8eb8ccbb5	server: revert llama_decode stop (#1722 ) Co-authored-by: firecoperana <firecoperana>	2026-05-02 18:19:59 +03:00
Kawrakow	453a027c17	Revert "server: defer recurrent-state reset to graph build (addresses #1696 r…" (#1704 ) This reverts commit fb05c2e9a2ded4d42861bc000ac01778cd1ba4c9.	2026-04-28 12:33:34 +02:00
markaalonzo	fb05c2e9a2	server: defer recurrent-state reset to graph build (addresses #1696 review) (#1696 ) ikawrakow asked the seq_rm zeroing be hoisted out of the per-cell loop and preferably done as part of the compute graph. This rewrites the fix: - llama_kv_cache gets a pending_recurrent_reset bitmap, sized to qnext_state_slots and indexed by slot index. - llama_kv_cache_seq_rm only marks the bitmap; no CPU-side tensor writes, no work inside the per-cell loop. - delta_net::build_layer_attn_linear ORs the bitmap into the existing reset_state flag, so the in-graph ggml_scale(state, 0.0f) path that already handled batch.pos[0] == 0 now also handles slot reset. - build_qwen3next and build_qwen35 clear the bitmap after ggml_build_forward_expand, once per graph build. - llama_kv_cache_clear also drops the bitmap, since the full-buffer clear it already does makes any pending in-graph reset redundant. Previous behavior is preserved (every recurrent layer column for the released slot is zeroed before the next request reuses it), but the work moves from the CPU into the compute graph, runs once per graph instead of once per cell, and reuses the existing reset path referenced in the review.	2026-04-28 08:00:01 +02:00
Samuel Oliveira Alves	67e6346225	Support for Qwen 3.5 MTP (dense models only) (#1698 ) * qwen-mtp: add dense mtp for one draft * add support for smaller qwen mtp commit * qwen-mtp: fix graph for qwen dense variants * Squashed commit of the following: commit a92a154b38c7fddc84460f8852c900f8d6ce907e Author: SamuelOliveirads <samueloliveira32df@gmail.com> Date: Mon Apr 20 13:30:21 2026 -0300 recurrent model: refactor api commit dfac8f19f6edc0014b4116041b89c1e0dfb173c7 Author: SamuelOliveirads <samueloliveira32df@gmail.com> Date: Mon Apr 20 12:22:29 2026 -0300 recurrent model: implement recurrent kernel checkpoint commit 9c44b117f93e9060030907e1250106358c9ccf47 Author: SamuelOliveirads <samueloliveira32df@gmail.com> Date: Sat Apr 18 11:52:39 2026 -0300 speculative: fix sampler for checkpoints commit e7006393bca20adcd86e2d77021e0b41d7bd9db1 Author: SamuelOliveirads <samueloliveira32df@gmail.com> Date: Fri Apr 17 14:08:25 2026 -0300 server: refactor checkpoint state logic commit 57eabf04df5185cab19539185b8f4b85e578905b Merge: dc4797b7 64234e3c Author: SamuelOliveirads <samueloliveira32df@gmail.com> Date: Fri Apr 17 13:53:41 2026 -0300 Merge branch 'main' into fix/hybrid-cache-speculative commit dc4797b72363482bb35750bb5edc87068116dc0f Author: SamuelOliveirads <samueloliveira32df@gmail.com> Date: Fri Apr 17 13:12:40 2026 -0300 reset ngram mod state for rejected tokens commit 8ff2d943a31b3d54440698db41e62ea121d661be Author: SamuelOliveirads <samueloliveira32df@gmail.com> Date: Fri Apr 17 13:08:04 2026 -0300 server: snapshot recurrent state in tensor commit d93dfb5e6b78a822e7331b33feedbdc47eb5ec79 Author: SamuelOliveirads <samueloliveira32df@gmail.com> Date: Thu Apr 16 22:36:37 2026 -0300 fix: save/restore sampler state during speculative checkpoint When speculative decoding rejects draft tokens and restores the recurrent state checkpoint, the sampler (RNG, grammar, prev tokens) must also be restored to maintain consistency. Without this, the sampler state reflects the rejected draft tokens, leading to potential divergence. Uses common_sampler_clone() to snapshot the sampler before the speculative batch decode, and restores it on rejection. commit d670cf85cd59f23a339faf6fd773889063176801 Author: SamuelOliveirads <samueloliveira32df@gmail.com> Date: Thu Apr 16 21:53:52 2026 -0300 server: spec checkpoints for recurrent models * server: fix leak context between requests * qwen3: allow mtp to run with split graph * qwen3 mtp: selects rows before the ffn	2026-04-28 07:47:50 +02:00
Samuel Oliveira Alves	ea94afe777	Speculative checkpoints for recurrent models (#1669 ) * server: spec checkpoints for recurrent models * fix: save/restore sampler state during speculative checkpoint When speculative decoding rejects draft tokens and restores the recurrent state checkpoint, the sampler (RNG, grammar, prev tokens) must also be restored to maintain consistency. Without this, the sampler state reflects the rejected draft tokens, leading to potential divergence. Uses common_sampler_clone() to snapshot the sampler before the speculative batch decode, and restores it on rejection. * server: snapshot recurrent state in tensor * reset ngram mod state for rejected tokens * server: refactor checkpoint state logic * speculative: fix sampler for checkpoints * recurrent model: implement recurrent kernel checkpoint * recurrent model: refactor api * spec: free rbudget before overwriting	2026-04-24 09:59:30 +02:00
markaalonzo	48819dadaf	server: fix ret=-3 on hybrid/recurrent prompt cache, and clear sticky stop flag (#1673 ) Two related issues that manifest as 'llama_decode ret=-3' on hybrid architectures (e.g. Qwen3.5/3.6 MoE, Qwen3-Next), matching the symptom reported in #1576. 1) server_context::apply_checkpoint() was written around transformer KV semantics (pos_min / pos_max per-token window). For hybrid and pure recurrent models the per-token pos_min threshold does not apply: the recurrent state is a single snapshot, and the server-side checkpoint is a whole-prefix record. The old selector 'cur.pos_min < pos_min_thold' can succeed on a checkpoint whose pos_max is past the current n_past, and — more commonly — fall through to do_reset = true, which zeros slot.n_past / slot.n_past_prompt. Zeroing in-place while the recurrent state in the context is still populated makes the next decode batch disagree with the live state, returning ret=-3. This change gates the checkpoint path on llama_model_has_recurrent(llama_get_model(slot.ctx)): - selector uses pos_max <= slot.n_past && pos_max < pos_next (whole-prefix match, leaves at least one token to decode); - on miss, slot state is preserved rather than zeroed, letting update_slots() continue from the already-valid n_past_prompt; - the erase loop drops any checkpoint whose pos_max > pos_next, matching the rewind semantics for recurrent state. Transformer behavior is unchanged. 2) stop_internal_decode is a file-static global in src/llama.cpp, set by llama_decode_stop() (called on client disconnect) and polled inside the decode loop to bail out with ret=-3. The flag is only cleared on one conditional path in server_slot::release(), so a stop signal that arrives after the interrupted llama_decode() has already returned bleeds into the NEXT decode call and causes an immediate ret=-3 with no work performed. Clear it at the top of the public llama_decode() entry so the signal is scoped to the in-flight decode it was meant for. Build-verified: llama-server with GGML_CUDA=ON, -DCMAKE_CUDA_ARCHITECTURES=86 (sm_86), IQK flash-attn + matmul enabled. No new APIs introduced — llama_model_has_recurrent is already public and already used elsewhere in server-context.cpp. Closes #1576	2026-04-23 09:19:17 +02:00
Kawrakow	e5355e9895	Quantization options (#1677 )	2026-04-23 09:05:39 +02:00
firecoperana	e0596bf614	Autoparser - complete refactoring of parser architecture (#1376 ) * Autoparser - complete refactoring of parser architecture Autoparser: add optional argument reshuffle capability Autoparser: True streaming (#20177) * Relax atomicity constraint for nicer, more pleasent, True Streaming parsing * Whitespace * Remove redundant atomics Revert to OAI-compatible args (#20213) * Revert to OAI-compatible args * Apply workaround::func_args_not_string Fix structured outputs (#20223) * Fix structured outputs * Update common/chat-auto-parser-generator.cpp Co-authored-by: Aldehir Rojas <hello@alde.dev> --------- Co-authored-by: Aldehir Rojas <hello@alde.dev> Fix compile bug (#20203) * Fix compile bug * Update common/chat-auto-parser-helpers.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> # Conflicts: # common/chat-auto-parser-helpers.cpp common : gracefully handle incomplete output (#20191) * common : handle incomplete UTF-8 at end of input in PEG parser * cont : if reached end prematurely, emit needs_more_input to propagate partial output * cont: refactor peg parse context to add lenient flag * cont : remove partial flag, keep lenient flag PEG parser for LFM2 (#20251) * PEG parser for LFM2 * Simplify using python_value() common: map developer role to system (#20215) * Map developer role to system * Simplify common: consolidate PEG string parsers (#20263) * common : consolidate PEG string parsers * cont : fix json_string_content() examples : fix empty items in json_schema_to_grammar.py [no ci] (#19968) * Fix logic for retrieving schema items in `json_schema_to_grammar.py` If `schema['items']` is `{}` and `prefixItems not in schema', as `{}` is Falsy, the original code here will raise an error. I think if `schema['items']` is `{}`, them items should just be `{}` * Apply suggestion from @CISC Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Add tests for arrays with empty items Add two unit tests to `tests/test-json-schema-to-grammar.cpp` that validate handling of arrays when 'items' is an empty schema and when 'prefixItems' is present alongside an empty 'items'. Both tests expect the same generated grammar, ensuring the JSON Schema->grammar conversion treats an empty 'items' schema (and the presence of 'prefixItems') correctly and covering this edge case. --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> Reduce level of content parser warning message to avoid log spam on non-debug verbosity (#20347) do not return if template parse failed add arg to enable parallel tool call common : fix incorrect uses of stoul (#20313) # Conflicts: # common/arg.cpp # src/llama-grammar.cpp examples : fix empty items in json_schema_to_grammar.py [no ci] (#19968) * Fix logic for retrieving schema items in `json_schema_to_grammar.py` If `schema['items']` is `{}` and `prefixItems not in schema', as `{}` is Falsy, the original code here will raise an error. I think if `schema['items']` is `{}`, them items should just be `{}` * Apply suggestion from @CISC Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Add tests for arrays with empty items Add two unit tests to `tests/test-json-schema-to-grammar.cpp` that validate handling of arrays when 'items' is an empty schema and when 'prefixItems' is present alongside an empty 'items'. Both tests expect the same generated grammar, ensuring the JSON Schema->grammar conversion treats an empty 'items' schema (and the presence of 'prefixItems') correctly and covering this edge case. --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> Add support for MiroThinker with new jinja template common/parser: handle reasoning budget (#20297) * v1 * Finished! * Handlie cli * Reasoning sampler * Apply suggestions from code review Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Less explosive terminology :) * Add utf-8 case and tests * common : migrate reasoning budget sampler to common * cont : clean up * cont : expose state and allow passing as initial state * cont : remove unused imports * cont : update state machine doc string --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> Co-authored-by: Alde Rojas <hello@alde.dev> common/parser: use nlohmann::ordered_json to preserve parameter order (#20385) common/parser: add GigaChatV3/3.1 models support (#19931) Co-authored-by: Mishusha <pmv26021975@gmail.com> common/parser: gracefully handle undetected tool parser, print error message. (#20286) fix: prevent nullptr dereference (#20552) common : fix iterator::end() dereference (#20445) # Conflicts: # common/regex-partial.cpp jinja : add capability check for object args (#20612) common/parser: add `--skip-chat-parsing` to force a pure content parser. (#20289) * Add `--force-pure-content` to force a pure content parser. * Update common/arg.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> common : rework gpt-oss parser (#20393) * common : rework gpt-oss parser * cont : fix gpt-oss tests * cont : add structured output test * cont : rename final to final_msg common : fix gpt-oss content removal (#20745) common/parser: add proper reasoning tag prefill reading (#20424) * Implement proper prefill extraction * Refactor cli parameters, update docs, move reasoning budget sampler part to common/reasoning-budget.cpp * Update tools/server/server-task.cpp * refactor: move grammars to variant, remove grammar_external, handle exception internally * Make code less C++y Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> chat : handle tool calls with no required args in TAG_WITH_TAGGED format (#20764) * chat : handle tool calls with no required args in TAG_WITH_TAGGED format * Update tests/test-chat.cpp [no ci] Co-authored-by: Aldehir Rojas <hello@alde.dev> --------- Co-authored-by: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com> Co-authored-by: Aldehir Rojas <hello@alde.dev> common/parser : fix out_of_range crash in throw path (#20424 regression) (#20777) * chat : fix out_of_range crash in throw path (#20424 regression) #20424 introduced effective_input = generation_prompt + input, but the throw path uses input.substr(result.end) where result.end is a position within effective_input. Every thinking model with a non-empty generation_prompt crashes with std::out_of_range instead of the intended error message. Test crashes on unpatched master, passes with fix: cmake -B build -DLLAMA_BUILD_TESTS=ON -DLLAMA_BUILD_TOOLS=OFF cmake --build build --target test-chat ./build/bin/test-chat * Update test-chat.cpp * Update test-chat.cpp * Update test-chat.cpp --------- Co-authored-by: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com> jinja : fix heap OOB read in value equality comparison (#20782) Address GHSA-q9j6-4hhc-rq9p and GHSA-2q4c-9gq5-5vfp. The three-iterator overload of std::equal in value_array_t::equivalent() and value_object_t::equivalent() reads past the end of the shorter container when comparing arrays or objects of different lengths. Use the four-iterator overload (C++14) which checks both range lengths. Found-by: Pwno common : fix typo in debug log ('extracft' -> 'extract') (#20807) common/parser: fix nasty bug causing subtle corruption of generation prompt (#20825) jinja : refactor token advancement (#20864) * refactor token advancement * exercise sub-expressions common/autoparser : detect reasoning markers when enable_thinking changes system prompt (#20859) common : replace wrap_for_generation with a prefix convenience function and fix gpt-oss (#20912) jinja: fix macro with kwargs (#20960) * jinja: fix macro with kwargs * Apply suggestions from code review Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * fix newline problem --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> common : inhibit lazy grammar sampler while reasoning is active (#20970) * common : inhibit grammar while reasoning budget is active * cont : update force_pos in accept * cont : fix tests * cont : tweak should apply logic * cont : return early not using grammar sampler * Add tests * cont : prevent backend sampling when reasoning budget enabled * cont : fix typo --------- Co-authored-by: Piotr Wilkin <piotr.wilkin@syndatis.com> # Conflicts: # common/reasoning-budget.h # common/sampling.cpp # tools/cli/cli.cpp # tools/server/server-common.cpp # tools/server/server-task.cpp common/parser: fix reasoning whitespace bugs + extra parser tests (#21085) * fix whitespace reasoning issues + add reconstruction tests * Proper fix * fix Nemotron autoparser test expectations to include newline in marker common : add reasoning_format = none support to gpt-oss (#21094) common/json-schema: fix: handle non-capturing groups (?:...) in JSON schema pattern converter (#21124) The regex-to-grammar converter in _visit_pattern() crashes with SIGSEGV when a JSON schema "pattern" field contains a non-capturing group (?:...). Root cause: when the parser sees '(' followed by '?', it pushes a warning but does not advance past '?:'. The recursive transform() call then interprets '?' as a quantifier and calls seq.back() on an empty vector, causing undefined behavior. This commonly occurs when serving OpenAI-compatible tool calls from clients that include complex regex patterns in their JSON schemas (e.g., date validation patterns like ^(?:(?:\d\d[2468][048]\|...)-02-29\|...)$). The fix: - Skip '?:' after '(' to treat non-capturing groups as regular groups - For unsupported syntax (?=, ?!, etc.), skip to matching ')' safely, handling escaped characters to avoid miscounting parenthesis depth - Adjust the ')' unbalanced-parentheses check using direct char comparisons instead of substr - Add test cases for non-capturing groups (C++ only, as the JS/Python implementations do not yet support this syntax) common/parser: fix handling of tool definition with missing properties key (#21128) jinja : handle empty expressions correctly (#20913) * Reject empty computed member expressions before returning slices[0] from parse_member_expression_arguments(). * Treat empty computed member expressions with Jinja2 undefined semantics Treat empty computed member expressions like `a[]` as undefined instead of raising a parser error, to match Jinja2 behavior. - return a noop expression for empty computed member arguments - return undefined when a computed member key evaluates to undefined - add Jinja tests covering `a[]\|default('fallback')` and `a[] is undefined` * Handle undefined computed member properties Move undefined-property handling to the common member access path, and add a test covering `a[undefined] is undefined`. * Use default undefined value in member access Initialize val and then return it when property is undefined. Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * empty statement parses to blank_expression instead of noop_statement --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> common : gpt-oss handle builtin and unsolicited tool calls (#21213) fix: tool call parsing for LFM2 and LFM2.5 models (#21242) * fix: tool call parsing for LFM2 and LFM2.5 models' * refactor: add test / break out lfm2 and lfm2.5 parsing logic # Conflicts: # common/chat.cpp Relax prefill parser to allow space. (#21240) * Relax prefill parser to allow space. * Move changes from prefix() to parser generation * Only allow spaces if we're not having a pure content parser next common : add commentary rules for gpt-oss-20b (#21286) add reasoning budget model, mtmd: fix gguf conversion for audio/vision mmproj (#21309) * fix gguf conversion for audio/vision mmproj * fix test # Conflicts: # convert_hf_to_gguf.py # examples/eval-callback/eval-callback.cpp # examples/mtmd/CMakeLists.txt # examples/mtmd/clip-impl.h # examples/mtmd/mtmd.cpp # gguf-py/gguf/constants.py # gguf-py/gguf/gguf_writer.py # gguf-py/gguf/tensor_mapping.py # src/CMakeLists.txt # src/llama-arch.cpp # src/llama-arch.h # src/llama-model.cpp # src/llama-model.h # src/llama-vocab.cpp # src/models/models.h # tests/test-llama-archs.cpp # tools/mtmd/clip-graph.h # tools/mtmd/clip-model.h # tools/mtmd/clip.cpp # tools/mtmd/models/models.h fix: gemma 4 template (#21326) chat : avoid including json in chat.h (#21306) jinja: coerce input for string-specific filters (#21370) common : fix tool call type detection for nullable and enum schemas (#21327) * common : fix tool call type detection for nullable and enum schemas * common, tests : fix grammar delegation for nullable/enum schemas and add tests Fix enum type inference to scan all enum values (not just index 0) so schemas like {"enum": [0, "celsius"]} correctly detect string type. Fix schema_delegates in peg-parser to handle nullable type arrays (["string", "null"]) and typeless enum schemas in raw mode, allowing the tagged parser to use raw text instead of JSON-formatted strings. Add test cases for Qwen3-Coder (TAG_WITH_TAGGED format): - nullable string ["string", "null"] - nullable string with null first ["null", "string"] - nullable integer ["integer", "null"] - enum without explicit type key common/parser: fix call ID detection (Mistral parser mostly) + atomicity for tag-json parsers (#21230) * Fix call ID detection (Mistral parser mostly) + atomicity for tag-json parsers * Rename * Update common/chat-auto-parser-generator.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> common : add gemma 4 specialized parser (#21418) * common : add gemma4 dedicated parser * cont : add '<\|tool_response>' as eog * cont : emit JSON from Gemma4 tool call AST * cont : more fixes * cont : refactor convert function * cont : refine rules and mapping * cont : add more tests * cont : clean up * cont : remove autoparser gemma4 implementation * cont : more cleanup * cont : rename gemma4.jinja to match the others * cont : add custom template to support interleaved thinking * cont : preserve reasoning in model turns * cont : fix initializer error * cont : fix unused vars * cont : fix accidental static * cont : fix specialized_template signature * fix extra semicolon * remove debug line and extra space [no ci] fix reasoning budget parser: fix MiniMax handling (#21573) jinja : support ensure_ascii=true, string repetition and int/float self-filtering (#21623) * feat: jinja engine improvements for reka-edge Port three Jinja engine improvements needed for the reka-edge model: 1. Python-style string repetition ("ab" * 3 → "ababab") 2. ensure_ascii=true support for tojson filter (escapes non-ASCII to \uXXXX) 3. int() builtin on value_int_t (identity, needed for Reka Edge template) * fix: escape invalid utf8 bytes when ensure_ascii=true The json_ensure_ascii_preserving_format function does not correctly handle an edge case where if UTF-8 parsing fails, it adds the non-ascii character back to the output as a raw byte. This commit fixes that by adding the unicode standard replacement character \\ufffd to the output instead. This is the standard behavior for various programming languages like Python, Rust, Go, etc. * chore: address PR comments 1. Add todo comment for supporting string repetition for array/tuples 2. Add support for float identity operation 3. Move invalid ascii test case to test_fuzzing * chore: accept suggestion for common/jinja/value.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> common : simplify autoparser tagged parser rules (#21216) * common : simplify autoparser tagged parser rules * cont : remove upper limit on optional args * cont : revert changes to parsing at the end * cont : undo arbitrary ordering of optional args * cont : fix uninitialized required parameters * revert to simplify merge * re-apply patches * restore flexible optional arg ordering tests common : fix ambiguous grammar rule in gemma4 (#21661) * common : fix ambiguous grammar rule in gemma4 * cont : fix missing comma... common : enable reasoning budget sampler for gemma4 (#21697) * fix: enable reasoning budget sampler for gemma4 Add thinking_start_tag and thinking_end_tag to common_chat_params_init_gemma4(). Without these, the reasoning budget sampler never activates for gemma4. Make the newline after "thought" optional in the PEG parser to handle budget=0 (sampler forces end tag before the newline). Add test case for empty thinking block. Fixes #21487 * use p.space() instead of p.optional(p.literal("\n")) in gemma4 thought parser common : better align to the updated official gemma4 template (#21704) fix: Fix broken structured output when using $refs in json_schema (#21699) chat: dedicated DeepSeek v3.2 parser + "official" template (#21785) Hide render_message_to_json warning common/gemma4 : handle parsing edge cases (#21760) common: skip reasoning budget sampler when no budget is requested (#21870) * common: skip reasoning budget sampler when no budget is requested After I added thinking_start_tag / thinking_end_tag for gemma4 in #21697, the reasoning budget sampler gets unconditionally created even when no budget is configured (the default -1). The same applies to kimi_k2, lfm2, lfm2_5, and ministral_3 which also set these tags. The budget gets converted to INT_MAX, so the sampler never actually forces any tokens but still runs per-token checks (start tag matching in IDLE state, token-to-piece conversion + UTF-8 checks in COUNTING state). More importantly, the mere existence of the sampler (non-null rbudget) disables backend sampling. Backend sampling lets the GPU select tokens directly, avoiding a full logits transfer from GPU to CPU every token. This could explain the 30% speed regression reported in #21784 (98 t/s to 70 t/s on Vulkan). So I added a reasoning_budget_tokens >= 0 check to the sampler creation condition. When the budget is unlimited, the sampler is not created, backend sampling stays enabled, and no per-token overhead is added. When a budget is explicitly set (0, 128, 1024, etc.), the sampler is created and works as before. * common: preserve rbudget when grammar is lazy Following up on the review feedback on #21870: keep the reasoning budget sampler when grammar_lazy is true, so the thinking-block grammar suppression from #20970 still works when tools are in use. This way, we only skip the sampler when both no budget is set AND grammar is not lazy. autoparser: support case of JSON_NATIVE with per-call markers (test case: Reka-Edge) (#21892) * fix grammar * fix add sampled token --------- Co-authored-by: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com> Co-authored-by: firecoperana <firecoperana>	2026-04-22 10:04:13 +02:00
Kawrakow	00ba208a5c	Fix Gemma4 partial offload (#1657 ) * Fix Gemma4 partial offload * Also here	2026-04-19 14:25:05 +02:00
Kawrakow	eaf83865a1	Vision support for Gemma4 (#1635 ) * WIP: Gemma4 vision Crashes on the GPU because of rms_norm requiring ne0 to be multiple of warp_size. Runs on the CPU, but produces garbage. * Remove unnecessary assert in CUDA rms_norm * GLU was not advertised as supported on CUDA * Still not working * This seems to work	2026-04-16 17:26:31 +02:00
Kawrakow	539d1cf989	Disallow speculation for hybrid/recurrent models (#1645 )	2026-04-16 17:21:44 +02:00
dmaivel	4f4bcfbe67	Add --defer-experts flag to defer expert mmap residency on Linux (#1634 ) * Add --defer-experts flag to defer expert mmap residency on Linux * Disable warmup when defer-experts is enabled	2026-04-16 08:54:44 +02:00
Nexes the Elder	67268c8fde	Fix mixed KV cache: type_v_first used instead of type_v_last for last layers (#1626 ) In llama_kv_cache_init() call, params.type_v_first was incorrectly passed twice instead of params.type_v_last. This caused V cache in the last N layers to use type_v_first instead of type_v_last. Fix: Replace second params.type_v_first with params.type_v_last.	2026-04-13 07:23:13 +02:00
Kawrakow	08ae48c667	Better routing for Gemma4-MoE (#1615 )	2026-04-11 15:19:02 +02:00

1 2 3 4 5 ...

373 Commits