ik_llama.cpp/examples at 48819dadaf655b8697574c9fe1b0c9d03e16feb9 - ik_llama.cpp - Jared's Git Server

jdelony/ik_llama.cpp

mirror of https://github.com/ikawrakow/ik_llama.cpp.git synced 2026-06-28 04:30:15 -05:00

History

markaalonzo 48819dadaf

server: fix ret=-3 on hybrid/recurrent prompt cache, and clear sticky stop flag (#1673 )

Two related issues that manifest as 'llama_decode ret=-3' on hybrid
architectures (e.g. Qwen3.5/3.6 MoE, Qwen3-Next), matching the symptom
reported in #1576.

1) server_context::apply_checkpoint() was written around transformer KV
   semantics (pos_min / pos_max per-token window). For hybrid and pure
   recurrent models the per-token pos_min threshold does not apply: the
   recurrent state is a single snapshot, and the server-side checkpoint
   is a whole-prefix record. The old selector 'cur.pos_min < pos_min_thold'
   can succeed on a checkpoint whose pos_max is past the current n_past,
   and — more commonly — fall through to do_reset = true, which zeros
   slot.n_past / slot.n_past_prompt. Zeroing in-place while the recurrent
   state in the context is still populated makes the next decode batch
   disagree with the live state, returning ret=-3.

   This change gates the checkpoint path on
   llama_model_has_recurrent(llama_get_model(slot.ctx)):
   - selector uses pos_max <= slot.n_past && pos_max < pos_next
     (whole-prefix match, leaves at least one token to decode);
   - on miss, slot state is preserved rather than zeroed, letting
     update_slots() continue from the already-valid n_past_prompt;
   - the erase loop drops any checkpoint whose pos_max > pos_next,
     matching the rewind semantics for recurrent state.

   Transformer behavior is unchanged.

2) stop_internal_decode is a file-static global in src/llama.cpp, set by
   llama_decode_stop() (called on client disconnect) and polled inside
   the decode loop to bail out with ret=-3. The flag is only cleared on
   one conditional path in server_slot::release(), so a stop signal that
   arrives after the interrupted llama_decode() has already returned
   bleeds into the NEXT decode call and causes an immediate ret=-3 with
   no work performed. Clear it at the top of the public llama_decode()
   entry so the signal is scoped to the in-flight decode it was meant
   for.

Build-verified: llama-server with GGML_CUDA=ON, -DCMAKE_CUDA_ARCHITECTURES=86
(sm_86), IQK flash-attn + matmul enabled. No new APIs introduced —
llama_model_has_recurrent is already public and already used elsewhere in
server-context.cpp.

Closes #1576

2026-04-23 09:19:17 +02:00

..

Merge mainline - Aug 12 2024 (#17 )

2024-08-12 15:14:32 +02:00

spec : add self speculative decoding, ngram and refactor (#1261 )

2026-02-13 19:04:55 +01:00

spec : add self speculative decoding, ngram and refactor (#1261 )

2026-02-13 19:04:55 +01:00

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

Quantization options (#1677 )

2026-04-23 09:05:39 +02:00

convert-llama2c-to-ggml

spec : add self speculative decoding, ngram and refactor (#1261 )

2026-02-13 19:04:55 +01:00

cvector-generator

spec : add self speculative decoding, ngram and refactor (#1261 )

2026-02-13 19:04:55 +01:00

deprecation-warning

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

spec : add self speculative decoding, ngram and refactor (#1261 )

2026-02-13 19:04:55 +01:00

spec : add self speculative decoding, ngram and refactor (#1261 )

2026-02-13 19:04:55 +01:00

Merge vulkan code from mainline up to commit of 6/28/2025 (#563 )

2025-07-02 08:49:42 +02:00

common : introduce composable PEG parser combinators for chat parsing and new jinja template engine (#1369 )

2026-03-09 11:03:33 +01:00

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

gguf-split: fix the split output files naming (#1336 )

2026-03-02 08:43:47 +01:00

spec : add self speculative decoding, ngram and refactor (#1261 )

2026-02-13 19:04:55 +01:00

Enable imatrix calculation for models with fused ffn_up/gate_exps tensors (#1418 )

2026-03-13 17:57:38 +01:00

spec : add self speculative decoding, ngram and refactor (#1261 )

2026-02-13 19:04:55 +01:00

build: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )

2024-06-13 00:41:52 +01:00

Add --defer-experts flag to defer expert mmap residency on Linux (#1634 )

2026-04-16 08:54:44 +02:00

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

add dry sampler (#513 )

2025-06-19 10:24:53 +03:00

spec : add self speculative decoding, ngram and refactor (#1261 )

2026-02-13 19:04:55 +01:00

spec : add self speculative decoding, ngram and refactor (#1261 )

2026-02-13 19:04:55 +01:00

common : introduce composable PEG parser combinators for chat parsing and new jinja template engine (#1369 )

2026-03-09 11:03:33 +01:00

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

Vision support for Gemma4 (#1635 )

2026-04-16 17:26:31 +02:00

spec : add self speculative decoding, ngram and refactor (#1261 )

2026-02-13 19:04:55 +01:00

Autoparser - complete refactoring of parser architecture (#1376 )

2026-04-22 10:04:13 +02:00

spec : add self speculative decoding, ngram and refactor (#1261 )

2026-02-13 19:04:55 +01:00

spec : add self speculative decoding, ngram and refactor (#1261 )

2026-02-13 19:04:55 +01:00

Quantization options (#1677 )

2026-04-23 09:05:39 +02:00

Quantization options (#1677 )

2026-04-23 09:05:39 +02:00

spec : add self speculative decoding, ngram and refactor (#1261 )

2026-02-13 19:04:55 +01:00

Refactor chat and server file (#1062 )

2025-12-15 08:27:20 +01:00

save-load-state

server: enable checkpoint for recurrent models (#1310 )

2026-02-26 06:51:18 +01:00

server: fix ret=-3 on hybrid/recurrent prompt cache, and clear sticky stop flag (#1673 )

2026-04-23 09:19:17 +02:00

spec : add self speculative decoding, ngram and refactor (#1261 )

2026-02-13 19:04:55 +01:00

Autoparser - complete refactoring of parser architecture (#1376 )

2026-04-22 10:04:13 +02:00

Cleaner log for adjusted splits (#1494 )

2026-03-24 07:49:40 +01:00

Merge mainline - Aug 12 2024 (#17 )

2024-08-12 15:14:32 +02:00

spec : add self speculative decoding, ngram and refactor (#1261 )

2026-02-13 19:04:55 +01:00

base-translate.sh

build: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )

2024-06-13 00:41:52 +01:00

chat-13B.bat

Create chat-13B.bat (#592 )

2023-03-29 20:21:09 +03:00

chat-13B.sh

build: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )

2024-06-13 00:41:52 +01:00

chat-persistent.sh

build: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )

2024-06-13 00:41:52 +01:00

chat-vicuna.sh

build: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )

2024-06-13 00:41:52 +01:00

chat.sh

build: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )

2024-06-13 00:41:52 +01:00

CMakeLists.txt

Port mdmd from mainline + Qwen2/2.5-VL support (#798 )

2025-09-27 08:45:29 +02:00

convert_legacy_llama.py

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

json_schema_pydantic_example.py

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

json_schema_to_grammar.py

Autoparser - complete refactoring of parser architecture (#1376 )

2026-04-22 10:04:13 +02:00

llama.vim

llama.vim : added api key support (#5090 )

2024-01-23 08:51:27 +02:00

llm.vim

llm.vim : stop generation at multiple linebreaks, bind to <F2> (#2879 )

2023-08-30 09:50:55 +03:00

Miku.sh

build: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )

2024-06-13 00:41:52 +01:00

pydantic_models_to_grammar_examples.py

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

pydantic_models_to_grammar.py

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

reason-act.sh

build: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )

2024-06-13 00:41:52 +01:00

regex_to_grammar.py

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

server_embd.py

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

server-llama2-13B.sh

build: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )

2024-06-13 00:41:52 +01:00

ts-type-to-grammar.sh

JSON schema conversion: ⚡️ faster repetitions, min/maxLength for strings, cap number length (#6555 )

2024-04-12 19:43:38 +01:00