Skip to content

fix: report fit math when expert cache pack allocation fails - #7

Open
thecodacus wants to merge 1 commit into
perffrom
fable5/moe-cache-diagnostics
Open

fix: report fit math when expert cache pack allocation fails#7
thecodacus wants to merge 1 commit into
perffrom
fable5/moe-cache-diagnostics

Conversation

@thecodacus

Copy link
Copy Markdown
Owner

When the pack does not fit, the warning now includes the numbers needed to choose a working value:

init_moe_expert_cache: pack allocation failed - expert cache disabled (120 slots need 23811 MiB, device has 8542 MiB free; ~198 MiB/slot -> at most 43 slots fit, and KV/compute buffers still allocate after this)

Motivated by real confusion in deployment: an oversized moe-cache-slots silently falls back to baseline, which reads as "the flag does nothing".

The old warning gave no way to pick a working slot count; users had to
derive per-slot cost from cudaMalloc errors. Now reports requested MiB,
free MiB, per-slot MiB, and the max slot count that could fit.
BAIS1C pushed a commit to BAIS1C/llama.cpp that referenced this pull request Aug 3, 2026
* spec: support MTP

* fix batch size

* rename files

* cont : simplify (thecodacus#7)

* MTP: clean-up (ggml-org#9)

* MTP: clean-up

* review: use llama_context_type instead of llama_graph_type

* review: remove llama_model_has_mtp

* review: fix convert issues

* convert: fix pycheck

* review: formatting

* use `mtp-` for identifying mtp models

* convert: fix mtp conversion

* mtp -> draft-mtp

* remove unused llama_arch

* add need_embd in speculative

* llama: allow partial seq_rm for GDN models for speculative decoding

Currently speculative checkpoint needs to restart from a checkpoint
after some draft tokens are not accepted, this leads to some wastage in
running the target again. This PR adds the ability to rollback upto
`draft_max` by storing the GDN intermediates.

* fix pending state

* vulkan: add GDN partial rollback

* meta: extend check to axis 1

* metal: add GDN partial rollback

Extend the gated delta net kernel to store intermediate states for
partial rollback support on the Metal backend.

- Add K (snapshot slot count) as a function constant
- Read input state from slot 0 of the 3D state tensor
- Write intermediate states to different slots during token loop
- For K=1, maintain backward-compatible single-slot behavior

Ref: ggml-org@8c05923

Assisted-by: llama.cpp:local pi

* delta_net_base: use ggml_pad instead of new_tensor

* review: add need_rs_seq

* review: rename part_bounded to n_rs

* review: deslop comments

* review: rename, add asserts

* server : adjust checkpoint logic (ggml-org#11)

* server : adjust checkpoint logic

* cont : rm asserts

* server-context: fix early exit

* spec : fix compatibility with n-gram and add TODOs (ggml-org#13)

* metal : cleanup

* llama : fix faulty bitwise check in recurrent memory

* server : disable RS-based MTP in combination with other spec types

* spec : add TODOs

* cont : fix comment

* cont : update comment

* common : fix logic for ngram + mtp compat

* llama-memory: enable checkpointing with partial rollback

* cont: add test-case for loading into a dirty ctx

* llama-memory-recurrent: clear rs_idx in clear

* download: fix mtp path

* llama-arch: fix enorm op

* docs: update docs

* conversion: fix type annotations

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant