Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 5 additions & 6 deletions .agents/NOW.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,8 +9,7 @@ benchmark record. Budget: 100 lines.

## Live claims

Working head: `row/backend-rocm-w0` (#41). Prior: benchmark checkpoint
`bench/qwen35-upstream-rebenchmark-20260805` on `upstream/main` @ `59674cf1d`.
Work: exact-chunks on main `1ce0d662b`; sm_120 measured at `3d2581551`.

| Claim / track | State | Next command or step |
|---|---|---|
Expand All @@ -20,7 +19,7 @@ Working head: `row/backend-rocm-w0` (#41). Prior: benchmark checkpoint
| MiniMax-H3 lane | **bf16 shards STREAM both towers (DiT + encoder); Q4_K_M enc cond cos 0.9975, 3.5° med, DIFFUSE** | render A/B on saved embeds |
| Kimi-Linear-48B | **ROW 7 fold LANDS (#122 §21): engine==CLI 128/128; golden 122/128; SACRED green; v13 tokens ABI** | ACTIVE: 19.0 tok/s vs vLLM ~21 (~0.90×) |
| 35B fresh grid | **BOUND** @`1ea26427`: 0.93-1.03x, c16 0.93x. INTAKE + Option A both NEGATIVE | Lever left: prefill glue (#61) |
| Qwen3.5-4B revalidation | 0.9971x @`59674cf1` (#35); TTFT/PSS pass, TPOT/ITL open | `docs/bench-evidence/` |
| Qwen3.5-4B sm_120 | Exact chunks ON: rebased-main reprofile 3.072x kernel / +2.272% run; sealed-vLLM throughput 1.021x PASS. Latency/VRAM OPEN | Spike residual 1.609x conv gap |
| RPi5 A76 CPU | **R5 asm GREEN; llama NOT MET**: 0.461x pf, 0.653x dec, RSS -24% | W6: BF16 GEMM |
| MXFP4 parity | c1 1.020, c2-c8 0.962-0.969. **#82 CLOSED: ptxas-lineage REFUTED (A/B ties our+vLLM PTX all ptxas/JIT; +10us=engine context, not codegen)** | TERMINAL: at parity |
| ROW-SERVE-ASYNC-DENSE-MIRROR | **LANDED+dgx-VERIFIED** (`f9c969ae`): async mirror on classic dense Qwen3; SACRED 184/184 | Residual: sibling scope one-liner |
Expand All @@ -30,7 +29,7 @@ Working head: `row/backend-rocm-w0` (#41). Prior: benchmark checkpoint
| `BACKEND-ROCM` | **(b) fix in; #140 gfx1201 hipBLAS + Gemma-4 MoE landed (contributor, authorship-preserved); W0 green 4 archs** | compile + M2 ([spec](specs/rocm-unified-memory-b.md)) |
| TP spike #287 (PR #143) | **LANDED** ([spec](specs/tensor-parallelism-spike.md)); DSpark rider grounded | dispatch TP-W1 (CPU-able) |
| Release | SPIKE; 30/30 | #129 |
| Surface coverage (`ARCH-ONE-SURFACE`) | ROW 8 + #139 IN; **ROW 6 IN REVIEW (#137): embeddings LIVE — `LlamaModel` arch, PoolingRunner in the step, `vllm_embed` v15, `/v1/embeddings`, fold gate 4/4-231, 9 kills** | Merge #137; real-ckpt oracle cosine residual |
| Surface coverage (`ARCH-ONE-SURFACE`) | ROW 8 + #139 IN; **ROW 6 LANDED (#137): embeddings LIVE — `LlamaModel` arch, PoolingRunner in the step, `vllm_embed` v15, `/v1/embeddings`, fold gate 4/4-231, 9 kills** | Real-checkpoint oracle cosine residual |

In-flight (default-OFF, not pushed): `laguna-fp4proj-prod`, laguna
bf16/legacy/pipeline-gemv, `ds4-hc-expand-fuse`.
Expand All @@ -49,8 +48,8 @@ both gate models, reproduced 2–3x on an idle box. See [gates.md](gates.md) and
0. **`ROAD-V1-MEM`** KV auto-sizing spike LANDED (`specs/kv-sizing.md`, `READY`).
1. **Spike the Parakeet encoder row** (vLLM carries it inside
`nano_nemotron_vl.py`; the transducer half is NOT in vLLM: separate call).
2. **Qwen3.5-4B serving follow-up:** bind the default-ON async-serving path
against the same oracle before attributing the remaining TPOT gap.
2. **Qwen3.5-4B sm_120:** rebased branch is GREEN and reprofiled. Spike the
residual 1.609x conv gap; latency/VRAM and gate models stay open.
2. **Merge the invocation-parity prevention** (CI guard + AGENTS.md checklist);
CUDA build-verify the byte-exact `kGemvHeuristicAlgos` refactor on dgx.
3. **Same-tool re-verify deepseek_v4's bf16 resident tower** (the one other
Expand Down
75 changes: 73 additions & 2 deletions .agents/benchmark-record.md
Original file line number Diff line number Diff line change
Expand Up @@ -15124,8 +15124,6 @@ llama.cpp's Vulkan" when it is really "our CPU tier vs llama.cpp's Vulkan". The
comparison becomes meaningful when native coverage closes — the progress metric is
`vt::GetReferenceTierHits()` reaching 0, and the ops that matter for this model are
the RoPE table build, the sampler tail, and the remaining norm/glue set.
>>>>>>> 814230a0 (bench(vulkan): VK-E unblocked with identical weights; ours quoted as NO RATIO)

#### CORRECTION (2026-08-07, same session): the vllm.cpp Vulkan arm is GPU-BOUND, not CPU-bound

The entry above attributes vllm.cpp's slowness to "our HOST FALLBACK wearing a
Expand Down Expand Up @@ -15983,3 +15981,76 @@ loads, or caching the K/V slice across the query heads that share a KV head —
Recorded because the hypothesis was specific and the refutation is reusable: this
is the second kernel this session where a barrier-count argument looked compelling
and measured flat (the subgroup GEMV was the first).

## 2026-08-07 — sm_120 exact GDN causal-conv chunks: 3.069x kernel, +2.152% enclosing

**Disposition:** ACCEPTED and reproduced on clean current-main transplant
`upstream/main` `f91a5917a`. Exact chunks default ON; latency, VRAM and
27B/35B gates stay open.

The metadata builder enumerates exact `(sequence, 8-token chunk)` work, uploads
its two i32 descriptors once per step and shares them across GDN layers. The
register kernel consumes one descriptor per `grid.y`.
`VT_CONV_EXACT_CHUNKS=0` is the same-binary whole-sequence rollback;
`VT_CONV_REG=0` selects tiled/scalar.

**Correctness.** RED compile evidence:
`/tmp/vllm-agent-runs/gdn-exact-red-escalated.json`. Focused host metadata and
flags, affected Qwen fixtures, full CUDA GDN and cached Qwen3.5-4B 3/3·1672
were green. Three production ON/OFF pairs had identical token files for all
128 requests × 128 outputs.

**Same-binary profile.** Manifest
`/tmp/vllm-agent-runs/qwen35-conv-exact-ab-profile.json`; rollback/default
traces `/tmp/qwen35-conv-exact-{off,on}.nsys-rep`. Rollback: 1728 calls,
720.047171 ms, 416.694 us mean. Exact: 1728 calls, 234.607112 ms, 135.768 us.
That is **3.069x**, saving 485.440 ms. Exact `grid.y=279/280/282` work counts
replace dominant `(64,28..32,1)` whole-sequence grids, confirming the proposed
mechanism. Pinned-vLLM same-tool total is 145.421 ms; residual **1.613x**.

**Enclosing A/B.** Manifest
`/tmp/vllm-agent-runs/qwen35-conv-exact-local-ab.json`; evidence root
`/tmp/qwen35-conv-exact-local-ab-20260807`. Three alternating pairs under one
GPU lock and a 25 GiB user-systemd scope:

| Axis | rollback | exact default | ratio |
|---|---:|---:|---:|
| total | 6641.800 tok/s | 6784.743 tok/s | 1.02152x |
| output | 734.433 tok/s | 750.237 tok/s | 1.02152x |
| TTFT | 1048.927 ms | 1018.040 ms | 0.97055x |
| TPOT / ITL | 35.420 ms | 34.740 ms | 0.98080x |
| E2E | 5547.687 ms | 5430.180 ms | 0.97882x |

Against sealed vLLM, exact local is **1.021246x** throughput,
**1.085812x** TTFT and **1.024597x** TPOT. Exact peak VRAM
13044/13058/13058 MiB, mean 13053.3, versus old local 13054 and vLLM 12820.

**VOID oracle attempts.** `qwen35-conv-exact-default-full-compare.json` exposed
the missing live-driver link path; the harness now adds `/run/opengl-driver/lib`
to `LIBRARY_PATH`. `qwen35-conv-exact-default-full-compare-rerun.json` reached
13/18 legs before a transient Torch bytecode read invalidated Triton AOT cache
keys. Neither attempt supersedes the sealed denominator. Full evidence:
`docs/bench-evidence/qwen35-4b-sm120-main-20260807.md`.

**Clean-transplant reproduction (`f91a5917a`).** Contained CPU/CUDA rebuild;
focused CPU 6/6, full CUDA GDN 66/66·4300, cached 4B 3/3·1672. Same-binary
graph-node traces `/tmp/qwen35-conv-exact-transplant-{off,on}.nsys-rep` and
token files reproduce the mechanism with byte identity: rollback 1728 calls /
720.216507 ms / 416.792 us, exact 1728 / 234.379395 ms / 135.636 us =
**3.072866x**. Profiled enclosing totals are 6587.66→6727.35 tok/s
(**1.021205x**), TTFT 1058.73→1025.46 ms, TPOT 35.70→35.04 ms and E2E
5592.69→5475.91 ms. The transplanted result is therefore reproduced, not merely
carried from its old branch. Against the sealed vLLM conv trace, residual is
**1.611730x**.

**Post-rebase reproduction (`3d2581551` on `upstream/main` `48a54141f`).** The
contained rebuild and all three gates remain green: focused 6/6, CUDA GDN
66/66·4300, cached 4B 3/3·1672. Fresh graph-node traces and token files under
`/tmp/qwen35-conv-exact-rebase-3d2581551-{off,on}.*` are byte-identical.
Rollback is 1728 calls / 718.704016 ms / 415.917 us; exact is 1728 /
233.954533 ms / 135.390 us = **3.07198x**. Profiled enclosing totals are
6589.65→6739.34 tok/s (**1.02272x**), TTFT 1057.63→1022.70 ms, TPOT
35.70→34.99 ms and E2E 5590.92→5466.20 ms. The result therefore survives the
27-commit main advance; against the sealed vLLM conv trace the residual is
**1.60881x**. Trace SHA-256: rollback `6a5dde18e...f97c47`, exact
`f47fb9cc...7aecf9`; both token files `83fcdc45...453545`.
2 changes: 1 addition & 1 deletion .agents/coordination.md
Original file line number Diff line number Diff line change
Expand Up @@ -1619,7 +1619,7 @@ items a-runner/b stay with the async/GDN `runner.cpp` owners.
| `CLAIM-POOLING` | `ENG-POOLER-SEQ` (INVENTORIED-implicit→ACTIVE, W1→**W2**), `ENG-POOLING-RUNNER` (**NEW row, ACTIVE, W3**), `SERVE-POOLING-ENDPOINTS` (INVENTORIED→SPIKE) | Claude Code (opus-4-8) | isolated worktree `.claude/worktrees/claim-pooling-w2w3` (CPU build `build-cpu` `-DVLLM_CPP_CUDA=OFF -DVLLM_CPP_SERVER=ON` Release + CPU run; NO dgx/GPU — the pooling reductions + activations + runner are host arithmetic) | branch `claim-pooling-w2w3`, base `main` `edf68c91` (confirmed via `git rev-parse HEAD`) | Pooling task class HIGH-priority feature-gap #2. W0 spike + W1 CPU pooler OP (prior pass); **W2 pooler HEADS composite + `SequencePooler`/`DispatchPooler` + `PoolerConfig`/`PoolingParams` and W3 pooling RUNNER path (this pass).** Owns ONLY: NEW `include/vllm/model_executor/layers/pooler/{pooling_metadata,methods,activations,common,pooling_params,pooler_config,heads,poolers,dispatch_pooler}.h` + `src/vllm/model_executor/layers/pooler/{methods,activations,heads,poolers,dispatch_pooler}.cpp` + NEW `include/vllm/v1/worker/gpu/pool/pooling_runner.h` + `src/vllm/v1/worker/gpu/pool/pooling_runner.cpp`; NEW `tests/vllm/model_executor/layers/pooler/{test_pooler,test_pooler_heads}.cpp` + `tests/vllm/v1/worker/gpu/pool/test_pooling_runner.cpp`; `CMakeLists.txt` (6 source lines) + `tests/CMakeLists.txt` (3 tests); NEW `.agents/specs/pooling-task-class.md`; the `ENG-POOLER-SEQ` + NEW `ENG-POOLING-RUNNER` engine-matrix rows + `SERVE-POOLING-ENDPOINTS` note + engine Serving/Total rollup (Serving 21→22/ACTIVE 6→7, Total 130→131/ACTIVE 47→48) + `scripts/check-agent-record.py` ENGINE 130→131; the record surfaces (this row, `roadmap_v1.md` gap #2, `docs/STATUS.md`, `docs/BENCHMARKS.md`, `feature-matrix.md` MODEL-POOLING note, `parity-ledger.md`, `state.md`). **NON-COLLISION:** additive NEW files only — the sole edits to existing compiled headers are ADDITIVE (methods.h defaulted virtuals, pooling_metadata.h new fields); ZERO edits to any existing production forward/runner path; NO pooling MODEL row created (concrete embedding model + real-oracle cosine gate is the named W3-model residual), so README/Metal/model-matrix rows untouched. | `ACTIVE` | 2026-07-29 — **W2 + W3 LANDED + CPU-GATED (foreground, NOT pushed).** `test_pooler_heads` 27/27 (240 asserts, Embedding/Classifier heads + SequencePooler + DispatchPooler incl. mixed embed+classify batch + ctor validation) and `test_pooling_runner` 5/5 (14 asserts, runner path + STRUCTURAL cosine-parity gate vs double-precision LAST+normalize ref) — plus W1 `test_pooler` 17/17 unchanged. RED-first proven: disable matryoshka slice + logit_mean → 8 cases/50 asserts fail (heads); CLS-instead-of-LAST drops cosine <0.5 + disable normalize → 2 unit-L2 asserts fail (runner). Clean CPU `-Wall -Wextra -Werror` 0-warn full-library build. **HONEST RESIDUAL:** the cosine gate is STRUCTURAL (synthetic weights) — the real-model `vllm.LLM(task="embed").encode` oracle cosine gate needs a registered concrete embedding model forward (W3-model, no number fabricated). Residuals (spec §Work breakdown): concrete pooling MODEL + real-oracle cosine gate (W3-model), endpoints /v1/embeddings+score+rerank+classify (W4), tokwise AllPool/StepPool (W5). Prior 2026-07-28 — W0 spike + W1 pooler OP LANDED + CPU-GATED: `test_pooler` 17/17 (50 asserts) vs double-precision refs, RED-first proven. |
| `CLAIM-DSV4-GGUF-LOADER` | `QUANT-GGUF-IQ2_XXS` (INVENTORIED→ACTIVE), `QUANT-GGUF-Q2_K` (INVENTORIED→ACTIVE); cross-refs `MODEL-TEXT-deepseek-v4-deepseek-v4-for-causal-lm` (stays `SPIKE`, owned by `CLAIM-DEEPSEEK-V4-IMPL`) | Claude Code (opus-4-8) | isolated worktree `.claude/worktrees/gguf-iquant-dsv4` (CPU-only `build-cpu` `-DVLLM_CPP_CUDA=OFF`; NO GPU, NO 90 GB download — the dequant unit gate uses known packed bytes; the GGUF header was HTTP-range-read, no download) | branch `feat/gguf-iquant-dsv4`, base `main` `4d1be010` (confirmed via `git rev-parse HEAD`) | GGUF IQ2_XXS + Q2_K dequant, the DeepSeek-V4-Flash single-Spark GGUF quant-path brick (W1). Owns ONLY the GGUF/quant PATH (NOT the forward — the forward TUs stay owned by `CLAIM-DEEPSEEK-V4-*`): the two `dequantize_row_*` decoders + grid/sign tables in `src/vt/cpu/cpu_quant_dequant.cpp`; the `kQ2_K`/`kIQ2_XXS` vt block dtype registration in `src/vt/dtype.{h,cpp}` + `src/vt/ops.cpp`; the id-16 reader trait in `gguf_reader.cpp` + ids-10/16 dispatch in `gguf_dequant.cpp`; `tests/vllm/test_gguf_dequant.cpp` + `tests/vt/test_ops_quant_traits.cpp`; the two `QUANT-GGUF-*` rows; NEW `.agents/specs/gguf-iquant-dsv4.md`; the V4 GGUF-loadable note on the model-matrix V4 row (row stays SPIKE); the record surfaces. **NON-COLLISION:** additive within the existing GGUF dequant switch + vt block table — does NOT touch any DeepSeek-V4 forward TU (`deepseek_v4.{cpp,h}`/`_dsa`/`_weights`), README, or Metal; the k-quant/NVFP4 decoders are byte-unchanged. | `ACTIVE` | 2026-07-29 — **W1 LANDED + CPU-GATED (foreground, NOT pushed).** IQ2_XXS (id 16, codebook `iq2xxs_grid`+signs+4-bit scale) + Q2_K (id 10, nibble sub-scale/min) ported 1:1 from llama.cpp `ggml-quants.c` `237ad9b96`; both DEQUANT-ONLY (no vec_dot ⇒ route to expand-bf16). `test_gguf_dequant` **15/15·480** (hand-derived literals: IQ2_XXS grid[1] byte0=0x2b→5.375, ksigns[1] flips j=0,7→±3.0, db 0.125/0.375; Q2_K 5.75/-0.25/2.5/0.25) + `test_ops_quant_traits` **9/9·5643** (dequant-only contract). All 7 changed TUs clean under full `-Werror`; the `voxtral.cpp` GCC-13 `-Werror=array-bounds` FP PROVEN pre-existing (fails at base with this diff's `dtype.h` reverted), neutralized only to link the test binaries. **W2 (V4-GGUF loader) DERIVED not landed:** HTTP-range-read the real `UD-IQ2_XXS` header — `general.architecture=deepseek4`, `general.file_type=19` (=IQ2_XXS), `split.tensors.count=1328`, full `deepseek4.*` config-KV schema; the tensor NAME manifest is beyond the CDN range cap + uncached ⇒ the V4 registry GGUF reject STAYS. Residuals: V4 forward (W3-W8, multi-Spark) + the V4-GGUF name map (W2, manifest-blocked) + a vec_dot perf leaf. |

| `CLAIM-EMBEDDINGS-ONE-SURFACE` | `ENG-POOLING-RUNNER` (live engine-step invocation), `SERVE-POOLING-ENDPOINTS` (SPIKE→ACTIVE, `/v1/embeddings`), `MODEL-EMBED-llama-llama-for-causal-lm` (INVENTORIED→ACTIVE; PARTIAL on merge — only the `LlamaModel` membership registered); cross-refs `ENG-POOLER-SEQ` (stays `CLAIM-POOLING`, ops untouched) | Claude Code (fable-5) helper, task #285 | isolated worktree `/home/mudler/_git/vllm.cpp-embeddings-one-surface` (CPU-only; lean per-target builds under disk pressure) | branch `row/EMBEDDINGS-ONE-SURFACE`, base `main` `b44ad337`, DRAFT PR #137 (the reservation) | ARCH-ONE-SURFACE fold ROW 6: embeddings/pooling through the ONE surface. Owns: NEW `src/vllm/model_executor/models/llama_embedding_registry.cpp` + `LoadLlamaModelEmbeddingWeights` (llama_weights.cpp) + `Qwen3DenseModel::ForwardHidden` (qwen3.{h,cpp} additive tail); the ADDITIVE task-gated pooling plumb (`LoadedModel::pooler()`, runner `pooling_runner_`+`pool_tokens`, `Request/EngineCoreRequest::pooling_params`, `ModelRunnerOutput::pooler_output`, scheduler pooling stop, `EngineCoreOutput/RequestOutput::pooling_output`, `LLMEngine::add_pooling_request/embed`, `ResolveAsyncEnabled(is_pooling_model)`); `vllm_embed`/`vllm_embedding_result_free` ABI v15 (vllm.h + vllm_c.cpp incl. the refuse-both-directions guards + the v13 `vllm_complete_tokens` missing-guard fix); `handle_embeddings` + `set_embedder` + task-conditional route (api_server.{h,cpp}) + server main pooling dispatch; NEW fixture `tests/vllm/models/fixtures/llama_embed_e2e` + `scripts/mm/llama_embed_fixture_gen.py` + `tests/vllm/models/test_llama_embedding_fold.cpp`; test/guard updates (test_capi v15 section + floor pin >= 15, test_dlopen symbols, c_header_compile.c, test_api_server embeddings section, test_model_registry/gguf arch pins, check-supported-models ARCH_TOKEN_RE); allowlist row removal + FEATURES/STATUS/BENCHMARKS rows + matrices + specs. **NON-COLLISION:** every engine hook is task-gated on `is_pooling_model`/`pooling_params` (nullopt/false = byte-identical text path); no SACRED path rewritten; no example added. | `ACTIVE` | 2026-08-08 — CPU-LANDED on the branch: fold gate `test_llama_embedding_fold` 4/4-231 (engine path == direct registry path + f64 LAST+normalize ref + chunked is_valid arm), `test_capi` 48/48-462, `test_dlopen` 30/30, server suite 50/50, registry 24/24-820, engine suites green (scheduler 423, llm_engine 204, engine_core 44, output_processor 77, qwen3_forward 1557, async_llm 342, llama_forward 509); 9 mutation kills (floor pin, refuse both directions, route gating both ways, engine-step invocation, scheduler stop, registry info pin, async-off wire). RESIDUAL: real embedding checkpoint + `LLM(task="embed")` oracle cosine. |
| `CLAIM-EMBEDDINGS-ONE-SURFACE` | `ENG-POOLING-RUNNER` (live engine-step invocation), `SERVE-POOLING-ENDPOINTS` (SPIKE→ACTIVE, `/v1/embeddings`), the `LlamaModel` embedding membership (INVENTORIED→PARTIAL on merge — only that membership registered); cross-refs `ENG-POOLER-SEQ` (stays `CLAIM-POOLING`, ops untouched) | Claude Code (fable-5) helper, task #285 | isolated worktree `/home/mudler/_git/vllm.cpp-embeddings-one-surface` (CPU-only; lean per-target builds under disk pressure) | branch `row/EMBEDDINGS-ONE-SURFACE`, base `main` `b44ad337`, PR #137 MERGED | ARCH-ONE-SURFACE fold ROW 6: embeddings/pooling through the ONE surface. Owns: NEW `src/vllm/model_executor/models/llama_embedding_registry.cpp` + `LoadLlamaModelEmbeddingWeights` (llama_weights.cpp) + `Qwen3DenseModel::ForwardHidden` (qwen3.{h,cpp} additive tail); the ADDITIVE task-gated pooling plumb (`LoadedModel::pooler()`, runner `pooling_runner_`+`pool_tokens`, `Request/EngineCoreRequest::pooling_params`, `ModelRunnerOutput::pooler_output`, scheduler pooling stop, `EngineCoreOutput/RequestOutput::pooling_output`, `LLMEngine::add_pooling_request/embed`, `ResolveAsyncEnabled(is_pooling_model)`); `vllm_embed`/`vllm_embedding_result_free` ABI v15 (vllm.h + vllm_c.cpp incl. the refuse-both-directions guards + the v13 `vllm_complete_tokens` missing-guard fix); `handle_embeddings` + `set_embedder` + task-conditional route (api_server.{h,cpp}) + server main pooling dispatch; NEW fixture `tests/vllm/models/fixtures/llama_embed_e2e` + `scripts/mm/llama_embed_fixture_gen.py` + `tests/vllm/models/test_llama_embedding_fold.cpp`; test/guard updates (test_capi v15 section + floor pin >= 15, test_dlopen symbols, c_header_compile.c, test_api_server embeddings section, test_model_registry/gguf arch pins, check-supported-models ARCH_TOKEN_RE); allowlist row removal + FEATURES/STATUS/BENCHMARKS rows + matrices + specs. **NON-COLLISION:** every engine hook is task-gated on `is_pooling_model`/`pooling_params` (nullopt/false = byte-identical text path); no SACRED path rewritten; no example added. | `DONE` | 2026-08-08 — MERGED in PR #137: fold gate `test_llama_embedding_fold` 4/4-231 (engine path == direct registry path + f64 LAST+normalize ref + chunked is_valid arm), `test_capi` 48/48-462, `test_dlopen` 30/30, server suite 50/50, registry 24/24-820, engine suites green (scheduler 423, llm_engine 204, engine_core 44, output_processor 77, qwen3_forward 1557, async_llm 342, llama_forward 509); 9 mutation kills. RESIDUAL moved to the PARTIAL model row: real embedding checkpoint + `LLM(task="embed")` oracle cosine. |


<!-- claim-view:begin -->
Expand Down
Loading
Loading