Skip to content

tests : update speculative params - #26925

Merged
ggerganov merged 1 commit into
masterfrom
gg/tests-server-fix-spec
Aug 12, 2026
Merged

tests : update speculative params#26925
ggerganov merged 1 commit into
masterfrom
gg/tests-server-fix-spec

Conversation

@ggerganov

@ggerganov ggerganov commented Aug 11, 2026

Copy link
Copy Markdown
Member

Overview

cont #25532

Adjust the speculative server test to be more robust to batch variances.

Fixes failures such as:

https://github.com/ggml-org/llama.cpp/actions/runs/31395690987/job/93477913727#step:5:4125

Successful run with this PR:

https://github.com/ggml-org/llama.cpp/actions/runs/31522118519

Requirements

@ggerganov
ggerganov requested a review from a team as a code owner August 11, 2026 19:32
@ggerganov
ggerganov requested review from gaugarg-nv and removed request for a team August 11, 2026 19:32
@ggerganov
ggerganov merged commit a4a4c51 into master Aug 12, 2026
14 checks passed
@ggerganov
ggerganov deleted the gg/tests-server-fix-spec branch August 12, 2026 05:08
gabe-l-hart added a commit to gabe-l-hart/llama.cpp that referenced this pull request Aug 12, 2026
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* origin/master: (383 commits)
  cmake :  introduce semantic versioning  (ggml-org#26839)
  gguf : harden loader against malformed tensor dims and metadata types (ggml-org#25596)
  kleidiai: Add runtime feature detection mechanism for aarch64/kleidiai (ggml-org#26076)
  model : disallow integer dflash sliding_window_pattern (ggml-org#26900)
  sync : ggml
  cmake : add config version support (ggml/1582)
  server : support slot save/restore with media inputs (ggml-org#26640)
  ui: add read_media tool (ggml-org#25877)
  opencl: default FA c8 cluster width to 16 on X1E (ggml-org#26433)
  tests : update speculative params (ggml-org#26925)
  vulkan: add TQ2_0 (ternary) support (ggml-org#25850)
  wavtokenizer-dec : bound posnet/convnext block_count against n_layer_all (ggml-org#26892)
  convert : handle per_layer_config in Gemma4 (transformers 5.15) (ggml-org#26882)
  opencl: use flat mv q5_k when weight exceeds image1d_buffer_t limit (ggml-org#26880)
  chat : fix muse-glimmer detection of tool calls after EOM (ggml-org#26879)
  ci : add missing release check (ggml-org#26923)
  CUDA: only disable CUDA graphs when mul_mat_id actually needs a stream sync (ggml-org#26802)
  cuda : add warp-per-row wkv7 kernel for single-token decode (ggml-org#26111)
  spec : update speculative-simple (ggml-org#26904)
  chat : tighten bare function parsing for Qwen models (ggml-org#26793)
  ...
huaxel pushed a commit to huaxel/CachyLLama that referenced this pull request Aug 12, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants