diff --git a/.agents/parity-ledger.md b/.agents/parity-ledger.md index e8c5f98fb..8a3d17dfe 100644 --- a/.agents/parity-ledger.md +++ b/.agents/parity-ledger.md @@ -426,7 +426,7 @@ Columns: | 2026-07-14 (`SERVE-GATE-ONLINE` third async-credit execution + direct-script bootstrap repair; `FAILED / VOID`) | Executes clean `b8681ac` through all six explicit ON/OFF timing legs under one uncontended lock. The first required Torch trace then fails before profiler startup because an absolute script invocation cannot import repository-local `tools`; neither trace or completion marker exists, so all timings remain diagnostic-only. The repair bootstraps the repository root only for direct script execution and adds an outside-repository absolute-path regression. Production inference is unchanged; live surfaces replace the preceding async narratives with this checkpoint. | Root `~/work/vllm-async-credit/b8681ac80b3f84af71955cf3a20cece2a118ea1f`; series/raw-set/log-set/corpus-set SHA `e8c7a4b7…86b0` / `65bff32f…6e51` / `79fd3836…f9c8` / `9fb8027a…d30`. Anchors: repaired [profiler](../tools/bench/profile_vllm_online_gate.py#L21), [absolute-path contract](../tests/tools/test_online_gate_trace.py#L29), [async spike](specs/async-serving.md), [online gate](specs/cuda-online-serving-gate.md), and [scoreboard](../docs/BENCHMARKS.md). | Six × **6/6** requests pass. Provisional ON/OFF medians are **160.798982 / 160.582996 total tok/s = 1.001345×**, TPOT **106.279834 / 107.333011 ms = 1.009909×**, and TTFT **812.002454 / 696.263685 ms = 0.857465×**; total CV ≤0.137978%. The whole series is **VOID** and earns no credit because both traces/metadata/summaries/marker are absent. Cleanup returns GPU/lock/port idle; focused contracts pass **6/6**. Binding remains **55/124**, W3 stays unowned `READY`, and a fresh commit/root must repeat every timing and trace arm. | | 2026-07-14 (`SERVE-GATE-ONLINE` fourth async-credit execution; six timings + ON trace, `FAILED / VOID`) | Executes clean `9b1774c` through all six timing legs and a complete explicit-ON Torch trace under one lock. The driver then applies the accepted 48-prompt H1d `--model-key 27` annotation-count contract to the six-prompt c2 diagnostic; the valid trace has 1,539 rather than 1,588 generation annotations, so summary fails closed before OFF. Shape-neutral re-read succeeds; the corrected recipe omits the model key only for this c2 diagnostic. Production inference and H1d validation are unchanged; live surfaces replace prior async narratives with this checkpoint. | Root `~/work/vllm-async-credit/9b1774c014880a0039545ea1be0fa01426cbd900`; series/raw-set/log-set/trace-set SHA `0d204e91…0310` / `9ce64024…4bc` / `6956cb19…7cf0` / `acb402a9…d419`; ON selected trace SHA `ad071f36…bf1`. Anchors: [profiler](../tools/bench/profile_vllm_online_gate.py#L21), shape-neutral [summarizer](../tools/bench/summarize_torch_kernels.py#L23), [async spike](specs/async-serving.md), [online gate](specs/cuda-online-serving-gate.md), and [scoreboard](../docs/BENCHMARKS.md). | Six × **6/6** timing requests pass. Provisional ON/OFF medians are **160.287860 / 160.213485 total tok/s = 1.000464×**, TPOT **106.648618 / 107.594484 ms = 1.008869×**, and TTFT **809.941298 / 697.928448 ms = 0.861703×**; total CV ≤0.249954%. Shape-neutral ON aggregation is **1,803,708 kernels / 171.130066 s**. The whole series remains **VOID** because OFF/summary/manifest/marker are absent. Cleanup returns GPU/lock/port idle; binding remains **55/124**, W3 stays unowned `READY`, and a fresh no-model-key c2 series must repeat every arm. | | 2026-07-14 (`SERVE-GATE-ONLINE` fifth async-credit execution + durable finalizer; complete diagnostic, neutral for speed) | Executes clean `3812d8` through all six explicit async ON/OFF timing legs and both shape-neutral Torch traces under one uncontended lock. Replaces the fragile inline derived step with a standard-library fail-closed finalizer that validates six raw legs, requested/resolved mode metadata, trace contracts and per-kernel totals, records output non-invariance, hashes every immutable artifact and writes the completion marker last. Production inference and local async defaults are unchanged; live surfaces collapse all precursor chronology to this accepted result. | Root `~/work/vllm-async-credit/3812d8d2b4a68d2e501007d01fe10cdf17751d02`; finalizer/summary/manifest/marker/artifact SHA `b3082a6e…1633` / `35b7344a…c323` / `e757b4ad…86c6` / `aa1e410b…369c` / `ead68397…8e56`; selected trace SHA ON/OFF `57413dd1…1cba` / `89bb9900…3c4`. Anchors: [profiler](../tools/bench/profile_vllm_online_gate.py#L21), [finalizer](../tools/bench/finalize_async_credit.py#L1), [ported contracts](../tests/tools/test_async_credit_summary.py#L1), [async spike](specs/async-serving.md), [online gate](specs/cuda-online-serving-gate.md), and [scoreboard](../docs/BENCHMARKS.md). | Six × **6/6** timing requests pass. ON/OFF medians are **160.347697 / 160.003134 tok/s = 1.002153×** total, **106.642353 / 107.739836 ms = 1.010291×** TPOT, and **807.657803 / 696.329685 ms = 0.862159×** TTFT. Shape-neutral traces are **1,798,044 / 170.819890 s** ON and **1,810,902 / 170.478267 s** OFF: **1.002004×** GPU time, so no 1.04× speed credit. ON digests are stable; OFF/pairs vary diagnostically with batch shape. Cleanup passes; focused **26/26**, all tool tests **55/55**, record checker green. Binding stays **55/124**; W3 remains unowned `READY`, and speed work moves to exact low-batch RMSNorm/generated-partition + FP4-tactic mapping. | -| 2026-07-14 (`SERVE-GATE-ONLINE` exact-c2 trace contract; local capture pending) | Reconstructs the accepted c2 oracle raw trace into an exact ordered executed-path contract, then parameterizes only the trace-build CUDA observer/server/driver for an explicitly requested B=S=2 graph while retaining B=16 as the accepted default. The c2 driver keeps the model gate, frozen plans, three local Nsight sessions × four ranges, six-prompt closed-loop corpus, fresh paired oracle trace, cache/lifecycle evidence and one lock. It deliberately emits no accepted low-batch status until a dedicated finalizer lands; production builds and inference dispatch are unchanged. | Accepted oracle root `~/work/vllm-async-credit/3812d8d2b4a68d2e501007d01fe10cdf17751d02`, selected trace SHA `57413dd1…1cba`; local anchors [controller](../include/vt/cuda/cuda_profiler_control.h#L13), [CUDA observer](../src/vt/cuda/cuda_backend.cu#L226), [server flag](../examples/server/main.cpp#L90), [driver](../scripts/dgx-online-serving.sh#L14), [validator](../tools/bench/online_gate.py#L1761), [requested-batch contract](../tests/tools/test_online_gate_client.py#L860), [online-gate spike](specs/cuda-online-serving-gate.md), and [scoreboard](../docs/BENCHMARKS.md). | Oracle evidence is exact: **1,524 clean B=2 windows / 1,160 kernels each**, identical ordered-name/signature SHA `858915dd…fad0` / `b5c6fcac…dd7b`, **177 generated RMSNorm/quant calls / 0.442805 ms** and FP4 tactics **128 Stream-K 128x64x256 + 80 static-persistent 128x32x256** per window. CPU **106/106**, tool **57/57**, focused **35/35**, and record/mutation/doc **18/18** pass; no GPU command ran. Local c2 capture/final status is **PENDING**, so binding remains **55/124**, no residual/speed credit/implementation leaf is claimed, and exact-grid/35B performance stay blocked. | +| 2026-07-14 (`SERVE-GATE-ONLINE` exact-c2 trace contract; local capture pending) | Reconstructs the accepted c2 oracle raw trace into an exact ordered executed-path contract, then parameterizes only the trace-build CUDA observer/server/driver for an explicitly requested B=S=2 graph while retaining B=16 as the accepted default. The c2 driver keeps the model gate, frozen plans, three local Nsight sessions × four ranges, six-prompt closed-loop corpus, fresh paired oracle trace, cache/lifecycle evidence and one lock. It deliberately emits no accepted low-batch status until a dedicated finalizer lands; production builds and inference dispatch are unchanged. | Accepted oracle root `~/work/vllm-async-credit/3812d8d2b4a68d2e501007d01fe10cdf17751d02`, selected trace SHA `57413dd1…1cba`; local anchors [controller](../include/vt/cuda/cuda_profiler_control.h#L13), [CUDA observer](../src/vt/cuda/cuda_backend.cu#L226), [server flag](../src/vllm/entrypoints/openai/server_main.cpp#L692), [driver](../scripts/dgx-online-serving.sh#L14), [validator](../tools/bench/online_gate.py#L1761), [requested-batch contract](../tests/tools/test_online_gate_client.py#L860), [online-gate spike](specs/cuda-online-serving-gate.md), and [scoreboard](../docs/BENCHMARKS.md). | Oracle evidence is exact: **1,524 clean B=2 windows / 1,160 kernels each**, identical ordered-name/signature SHA `858915dd…fad0` / `b5c6fcac…dd7b`, **177 generated RMSNorm/quant calls / 0.442805 ms** and FP4 tactics **128 Stream-K 128x64x256 + 80 static-persistent 128x32x256** per window. CPU **106/106**, tool **57/57**, focused **35/35**, and record/mutation/doc **18/18** pass; no GPU command ran. Local c2 capture/final status is **PENDING**, so binding remains **55/124**, no residual/speed credit/implementation leaf is claimed, and exact-grid/35B performance stay blocked. | | 2026-07-14 (`SERVE-GATE-ONLINE` first exact local-c2 execution + batch-keyed validator repair; `FAILED / VOID`) | Executes clean `ad8b58f` through the real 27B model gate and the first of three local B=2 Nsight sessions under one lock. The four-replay controller, six-prompt client, frozen plan map and graceful target lifecycle pass; range validation then applies c16's 1,107-kernel contract to the valid 1,011-kernel B=2 graph and fails closed before sessions 2/3 or the oracle. The repair retains c16 and keys an independently tested 27B/B=2 contract through validator, summarizer and driver. Production inference is unchanged. | Root `~/work/vllm.cpp-executed-path-c2/ad8b58f8708ce9bdf32aa9043611b3f6049be7fd`; run/execution/model-gate/control/raw-report-set/evidence-set SHA `9f285fd6…0aec` / `2a3d326f…6b56` / `bc9dc95b…da17` / `ee1589df…c719` / `f3aa3ca9…64c6` / `0da532ac…ac35`. Anchors: batch contracts and [validator](../tools/bench/online_gate.py#L99), [driver](../scripts/dgx-online-serving.sh#L666), [ported requested-batch cases](../tests/tools/test_online_gate_client.py#L1083), [online gate](specs/cuda-online-serving-gate.md), and [scoreboard](../docs/BENCHMARKS.md). | Model gate passes **1/1**; local client/probe pass **6/6 + 2/2** and profile control records B=2, four replays, 64 frozen plans, zero tuning. Read-only reconstruction proves all four reports lossless and identical at **1,011 kernels + 7 memcpy + 1 memset**, multiset SHA `6b75bcff…1ce3`; mean traced kernel time is 109.722456 ms, FP4 is the oracle-matched **128+80** split, and RMSNorm-family structure is **177 calls / 2.237944 ms**. The series is **VOID**: no sessions 2/3, fresh oracle, final status, or binding ratio. Focused **35/35**, all tool **57/57**, policy **18/18**, and record/doc checks pass; cleanup returns GPU/lock/port idle. Binding remains **55/124** and full repaired retry is required. | | 2026-07-14 (`SERVE-GATE-ONLINE` exact-c2 complete raw capture + durable-finalizer implementation; raw complete / finalizer `PENDING`) | Executes clean `179a0fc` through the whole repaired paired trace series under one uncontended lock, then adds a standard-library fail-closed finalizer and exact-B=2 ported contracts. The finalizer reuses the batch-keyed raw validator, proves local range invariance, accepts only the observed steady-oracle topology plus bounded drains and launch-signature allowlist, resolves complete family/tactic counts, hashes the immutable artifact set and writes the status marker last. Production inference and accepted c16 behavior are unchanged; live status surfaces retain only this current snapshot. | Raw root `~/work/vllm.cpp-executed-path-c2/179a0fc2afc1c33b63d14de8e50d3fde976c7356`; fresh oracle trace SHA `2b3bf412…785c`; local node SHA `44fcf31f…b93d`; anchors: batch-aware [validator](../tools/bench/online_gate.py#L99), [c2 finalizer](../tools/bench/finalize_low_batch_trace.py#L1), [ported finalizer contracts](../tests/tools/test_low_batch_trace_summary.py#L1), [online-gate spike](specs/cuda-online-serving-gate.md), and [scoreboard](../docs/BENCHMARKS.md). A full read-only finalizer preflight against the raw root passes; committed durable summary/manifest/marker hashes remain pending. | Model gate passes. All **12/12** local ranges are lossless/invariant at **1,011+7+1**; the oracle has **1,522×1,160** invariant steady B=2 windows plus two bounded drains. Diagnostic median local/oracle kernel time is **111.076528 / 105.520831 ms = 1.052650×**. BF16 GEMMs lead at **193/51.662672 ms vs 97/48.798042 ms** (+96 launches/+2.864630 ms), followed by equal-count RMSNorm partitions at **2.249728 vs 0.439491 ms** (+1.810237 ms); FP4 tactics match **128+80** and FP4 time is non-positive. Cross-profiler timing is non-binding; durable final status, speed credit, exact-grid rerun and 35B performance remain **PENDING**. Binding stays **55/124**; cleanup returns GPU/lock/port idle. | | 2026-07-14 (`SERVE-GATE-ONLINE` exact-c2 durable finalization; `COMPLETE DIAGNOSTIC`) | Runs the exact pushed `fe28003` finalizer bytes once against immutable raw `179a0fc`, revalidates the full raw chain, and writes the completion marker last. This changes evidence lifecycle only: production inference, c16 validation and the binding performance grid are unchanged. Live surfaces replace the pending checkpoint with this durable snapshot while detailed chronology remains here and in state. | Root `~/work/vllm.cpp-executed-path-c2/179a0fc2afc1c33b63d14de8e50d3fde976c7356`; summary / manifest / status / artifact-set / finalizer SHA `0ef6a124…0273` / `2556cfd0…2f21` / `9e0143fa…7b57` / `cc248ad2…823a` / `45dbf28a…3311`; run-log SHA `362f0f1e…1cef`. Anchors: committed [finalizer](../tools/bench/finalize_low_batch_trace.py#L1), [ported contracts](../tests/tools/test_low_batch_trace_summary.py#L1), [serving-gate spec](specs/cuda-online-serving-gate.md), and [scoreboard](../docs/BENCHMARKS.md). | Status is **`complete-diagnostic`**. It retains **12/12** invariant local **1,011+7+1** ranges, **1,522×1,160** steady oracle B=2 windows plus two drains, and matching **128+80** FP4 tactics. Diagnostic local/oracle medians remain **111.076528 / 105.520831 ms = 1.052650×**; BF16 GEMMs lead at +96 launches/+2.864630 ms, RMSNorm follows at +1.810237 ms, and FP4 is non-positive. Cross-profiler timing earns no speed credit; binding remains **55/124**. No GPU command ran, GPU/lock/port are idle, and the next legal checkpoint is the BF16-GEMM whole-chain spike before code. | diff --git a/docs/BENCHMARKS.md b/docs/BENCHMARKS.md index 14fe1e7c9..9ecbd9a8c 100644 --- a/docs/BENCHMARKS.md +++ b/docs/BENCHMARKS.md @@ -380,7 +380,7 @@ built on it rather than keeping the flattering one. | Laguna NVFP4 decode | `flock $HOME/gpu.lock ./build-cuda/examples/laguna-gen --model ~/laguna-xs-nvfp4 --gpu` (that directory holds the S-2.1 checkpoint); `drop_caches` first, create the CUDA context before loading weights | | DeepSeek-V4-Flash decode | `deepseek-v4-gen --gpu --kv-cache` on `ds4flash.gguf`, captured under tmux | | Metal vs MLX-LM | Paired A/B harness, interleaved runs, cold legs discarded | -| Vulkan vs llama.cpp Vulkan | Not yet runnable (no model runs on Vulkan). Planned harness in `.agents/specs/vulkan-full-support.md` §5.2 | +| Vulkan vs llama.cpp Vulkan | Same GGUF both arms: ours `-DVLLM_CPP_VULKAN=ON`, llama.cpp `-DGGML_VULKAN=ON` at `237ad9b96` via `llama-bench`; clean legs only, one `flock $HOME/gpu.lock`. GEMV sweep: `benchmarks/vulkan_gemv_ab.cpp` | Build flags, environment variables, and the full gate list are in [BUILD.md](BUILD.md) and [ENVIRONMENT.md](ENVIRONMENT.md). diff --git a/docs/ENVIRONMENT.md b/docs/ENVIRONMENT.md index a73d23e0b..c3071bb90 100644 --- a/docs/ENVIRONMENT.md +++ b/docs/ENVIRONMENT.md @@ -155,6 +155,7 @@ on CUDA/CPU builds beyond the documented behavior. | `VT_GEMMA4_RESIDENT_EXPERTS` | unset | `=1` preloads the Gemma-4 MoE experts resident on the GPU(s) after the first use instead of streaming them per step (discrete-ROCm optimization). No-op (with a stderr note) on a binary built without `-DVLLM_CPP_HIP` | | `VT_GEMMA4_RESIDENT_GPUS` | `2` | Number of GPUs across which resident Gemma-4 experts are spread; clamped to the ROCm device count. Read only when `VT_GEMMA4_RESIDENT_EXPERTS=1` | | `VT_GEMMA4_RESIDENT_MAX_LAYERS` | (all) | Caps how many MoE layers get resident-preloaded, to fit a smaller VRAM budget. Read only when `VT_GEMMA4_RESIDENT_EXPERTS=1` | +| `VLLM_CPP_HTTP_FIXED_POOL` | `1` (fixed) | `=0` reverts the HTTP worker pool to the legacy dynamic mode. Production uses the capacity-derived fixed pool; the opt-out exists for same-binary A/B attribution | | `VT_ROCM_ATTN_CPU_REF` | unset | `=1` routes ROCm paged attention through the CPU reference kernel instead of the HIP kernel — a correctness A/B for the ROCm attention bring-up | | `VT_DEBUG_SAMPLED` | unset | `=1` prints the per-step sampled token id(s) to stderr (sampling-loop debug). Read-only; does not change output. Read once per token, so it does not stall the hot loop | diff --git a/docs/STATUS.md b/docs/STATUS.md index 951ef08c5..c0f6914ee 100644 --- a/docs/STATUS.md +++ b/docs/STATUS.md @@ -127,7 +127,7 @@ token-for-token correctness against the pinned oracle. | Plugin system (out-of-core registration) | Spiked; first CPU brick landed, not yet wired into any production path | The extensibility-first discovery layer. W0 spike over vLLM's plugin surface (general / platform / io_processor / endpoint groups, the `register_model` an out-of-tree plugin calls, the invocation seams) is committed (`.agents/specs/plugin-system.md`, `ENG-PLUGIN-SYSTEM` ACTIVE, `CLAIM-PLUGIN-SYSTEM`). W1 landed `vllm::plugins::LoadGeneralPlugins()` + the out-of-core general-plugin registration seam (`RegisterGeneralPlugin` / `REGISTER_VLLM_GENERAL_PLUGIN`) over the existing `REGISTER_VLLM_MODEL`-style registries (the in-tree factory `MODEL-FACTORY-registry` is record-repaired `DONE` 2026-08-05: 28 self-registering TUs, dgx debt paid by the 2026-07-23 seven-gate run): a 1:1 mirror of `load_general_plugins` (load-once idempotence, the `VLLM_PLUGINS` allowlist, per-plugin failure isolation). Proven by an out-of-core toy-model plugin that registers a toy architecture through the public `RegisterModel` seam — unit-gated RED-first (`test_plugin_system` 1 case / 29 assertions: the toy arch resolves ONLY after LoadGeneralPlugins runs it, and not under `VLLM_PLUGINS=""`). Python entry points have no C++20 analogue, so discovery is the project's static-init/`dlopen` registration idiom (recorded porting-inventory §9). NOT yet wired: real shared-object `dlopen` + the C-ABI `vllm_plugin_register` entry (W2), the engine/CLI `--load-plugins` wiring that calls LoadGeneralPlugins from the construction paths (W3), the platform/quant plugin kinds (W4), and the io_processor/stat_logger/endpoint groups (W5) are named residuals. See docs/BENCHMARKS.md | | Offline Batch API (JSONL file runner) | Spiked; first CPU brick landed, not yet exposed as a CLI | The offline OpenAI Batch API: read a JSONL of OpenAI-format requests, run each through the engine, write a JSONL of responses. W0 spike over vLLM's `run_batch.py` (schema, endpoint dispatch, run loop, file I/O) is committed (`.agents/specs/batch-api.md`, `SERVE-BATCH-API` ACTIVE, `CLAIM-BATCH-API`). W1 landed `RunBatch` (`RunLine`/`RunLines`/`Run`) + `RunBatchFile` — a pure orchestrator over the existing `OpenAIServingChat::create_chat_completion` (NO reimplemented generation), 1:1 with vLLM's endpoint_registry url→handler map: `/v1/chat/completions` dispatch, the `BatchResponseData`/`BatchRequestOutput` schema (`vllm-` ids, custom_id echoed), the `run_request` AllResponse/ErrorResponse/stream branches, and the unsupported-endpoint/url error rows. Unit-gated RED-first (`test_openai_run_batch` 7 cases / 80 assertions over the synthetic serving engine: ordered rows + custom_id echo + per-line BatchRequestOutput schema round-trip, a malformed line isolated into an error row so the batch continues, dispatch + 404 error rows; dropping the custom_id echo fails 9 assertions). Recorded deviation: a malformed line is isolated (batch continues) where upstream aborts the job. NOT yet exposed: the `vllm run-batch` CLI + `BatchFrontendArgs` (W2), embeddings/score/rerank dispatch (W3, rides pooling endpoints), audio transcription/translation + media fetch (W4), and http(s)/data-URL file I/O + metrics + overlapped `AsyncLLM` submission (W5) are named residuals. See docs/BENCHMARKS.md | | Tokenizers | Supported | Byte-level BPE (Qwen/Llama-3/OPT/GPT-2/DeepSeek/OLMo-2) and SentencePiece BPE (Mistral/Gemma), plus GGUF vocab; added-token `lstrip`/`rstrip` whitespace semantics (e.g. Phi-4-mini's special tokens); byte-exact vs the vLLM oracle | -| Multimodal: image to text | Correctness-complete; vision-forward speed beats vLLM; OpenAI-server content-part parse + processor routing + engine mm-request plumbing landed (CPU); end-to-end serving pending | Strict token-exact 32/32 vs vLLM 0.25.0 on Qwen3-VL-4B and Qwen3.6-27B (`Qwen3_5ForConditionalGeneration`); C++ image processor + vision tower + MRoPE/DeepStack backbone with on-GPU greedy sampling; the 27B decode step is now graph-capturable (routes through the production captured decode, token-exact held). The vision tower now defaults to the flash-tiled non-causal attention (`vt::AttentionDenseFlash`, byte-identical to the previous warp kernel — token-identical, goldens unchanged): the per-image tower forward is ~142 ms vs vLLM 0.25.0's ~250 ms eager encode = 0.57x (faster). Attribution-first profiling found the tower attention is serial-latency-bound (not K/V-bandwidth-bound), so flash tiling is only a 1.04x same-binary A/B over the warp kernel here — the tower already beats vLLM; batched serving speed still pending. OpenAI-server multimodal wiring — CPU bricks 1+2 landed (`CLAIM-MM-SERVING-W1`/`W2`, .agents/specs/mm-serving.md): the chat request parses the OpenAI multimodal content-part array (`image_url`/`input_audio`/`audio_url`), decodes the base64/`data:` payloads, routes them through the existing single-sequence processors to a placeholder-expanded prompt + mm-feature handles, AND the engine now carries them — additive `LLMEngine`/`AsyncLLM` `add_request(MultiModalInputs)`/`generate` overloads (via `InputProcessor::process_inputs_mm`, mm_features onto `EngineCoreRequest`/`Request`), chat-template placeholder-string helpers, and a serving_chat `MultiModalChatFn` seam (default unset ⇒ text path byte-identical). The W3 `MultiModalChatFn` seam BODY now lands (`CLAIM-MM-SERVING-E2E`): `MakeQwen3VLImageChatFn` (chat_mm.{h,cpp}) turns an image chat request into the placeholder-EXPANDED engine input — marker-inject at the mm part position, render the chat template, tokenize (the single `<|image_pad|>` maps to one image_token_id via `EncodeWithSpecialTokens`), then `RouteImageRgb` EXPANDS to 196 image tokens + mm_features — and is wired into `examples/server/main.cpp` (guarded on `preprocessor_config.json`; text-only models leave the seam unset ⇒ byte-identical). Unit-gated `test_chat_mm` 8/8 (the W3 seam-body test drives the real tokenizer + chat template → 196 image tokens; RED line = the text-only path renders 0) + `test_input_processor` 10/10 + `test_openai_serving` (the production seam is invoked on an image request + routed to the engine mm generate overload; text-only never touches it; streaming+mm rejected), CPU-only, no model weights; bare-string chat requests byte-identical (proven). **ENGINE MM-FORWARD LANDED (`CLAIM-ENGINE-MM-FORWARD`, 2026-07-28): multimodal now runs through the engine's REGISTERED forward.** The architectural block is resolved: `ModelForwardInput` gains an additive default-nullopt `mm` field (merged inputs_embeds + 3-D MRoPE positions + DeepStack; nullopt for text ⇒ the shared runner path is byte-identical by construction), `Qwen3VLForConditionalGeneration` is now `REGISTER_VLLM_MODEL`-registered, and the registered forward folds the M2c decode into `ModelRegistry::Forward` via the shared `Qwen3VLForwardStepLastLogits` (`Qwen3VLGenerateGreedyViaRegistry` drives every step through the registry). GPU token-exact gate `test_qwen3vl_registry_e2e` on dgx.casa GB10: an image→text run THROUGH `ModelRegistry::Forward` emits the M2c golden tokens 32/32 STRICT (RED-first — unregistered ⇒ resolve throws); the M2c standalone cross-check holds 32/32 (the refactor is byte-neutral). Text inertness (the shared-path RED line): `test_runner` 16/16 + `test_scheduler` 36/36 + `test_model_registry` 24/24 + `test_chat_mm` 8/8 + `test_openai_serving` 41/41 all green. **Residual:** the FULL in-runner scheduler-fed tower run (the batched engine loop building the mm field from staged encoder outputs) + the real server `/v1/chat/completions` GPU end-to-end + video/multi-image/audio through the registered path — recipe in `.agents/specs/mm-serving.md`. **GEMMA-4 IMAGE through the registered forward LANDED (`CLAIM-GEMMA4-MM-E2E`, 2026-07-29):** `Gemma4ForConditionalGeneration` is now `supports_multimodal=true` with an mm branch routing `ModelForwardInput.mm` → `Gemma4Model::ForwardMm` (the SigLIP2 projector output masked-scattered into the `` rows + PLE image-rows→0; 1-D positions, no MRoPE/DeepStack); the driver `Gemma4GenerateGreedyViaRegistry` runs image→text through `ModelRegistry::Forward`. dgx sm_121a gate `test_gemma4_registry_e2e`: **16/18 content tokens BIT-EXACT** vs the STRICT `gemma4_e4b_image` golden (full sentence), the single divergence a terminal-punctuation bf16 near-tie (top1-top2 margin ~0.10-0.12 logit, invariant to vision-input precision ⇒ backbone bf16-accumulation, not the fold); text SACRED 32/32 UNCHANGED (inertness). Residuals: STRICT-18/18 (bit-match vLLM prefill bf16), Gemma-4 AUDIO→text e2e (mel A1 + merge; the G3 audio tower is per-stage proven), speed. **C2 (2026-07-31, `CLAIM-C2-VISION-QKV-BIAS`): both vision towers' attention QKV now fold to ONE `vt::MatmulBT` + a fused merged-`[3H]` BIAS epilogue + a contiguous `vt::QkvSplit` via the shared `models::FusedMergedQkvBiasSplit` helper** — qwen3_vl_vision (weight already resident-fused `[3H,H]`+`[3H]` bias) and gemma4_vision (loader-concatenated `qkv_proj [3H,H]`, no bias, per-slice clamp epilogue kept). Bit-exact (A/B unit `test_ops_qkv_merged_bias` byte-identical, RED-first); qwen3_vl gated e2e by `test_qwen3vl_e2e` image→text token-exact, gemma4_vision CPU-A/B + CUDA build-verified; existing tower gates held. See docs/BENCHMARKS.md | +| Multimodal: image to text | Correctness-complete; vision-forward speed beats vLLM; OpenAI-server content-part parse + processor routing + engine mm-request plumbing landed (CPU); end-to-end serving pending | Strict token-exact 32/32 vs vLLM 0.25.0 on Qwen3-VL-4B and Qwen3.6-27B (`Qwen3_5ForConditionalGeneration`); C++ image processor + vision tower + MRoPE/DeepStack backbone with on-GPU greedy sampling; the 27B decode step is now graph-capturable (routes through the production captured decode, token-exact held). The vision tower now defaults to the flash-tiled non-causal attention (`vt::AttentionDenseFlash`, byte-identical to the previous warp kernel — token-identical, goldens unchanged): the per-image tower forward is ~142 ms vs vLLM 0.25.0's ~250 ms eager encode = 0.57x (faster). Attribution-first profiling found the tower attention is serial-latency-bound (not K/V-bandwidth-bound), so flash tiling is only a 1.04x same-binary A/B over the warp kernel here — the tower already beats vLLM; batched serving speed still pending. OpenAI-server multimodal wiring — CPU bricks 1+2 landed (`CLAIM-MM-SERVING-W1`/`W2`, .agents/specs/mm-serving.md): the chat request parses the OpenAI multimodal content-part array (`image_url`/`input_audio`/`audio_url`), decodes the base64/`data:` payloads, routes them through the existing single-sequence processors to a placeholder-expanded prompt + mm-feature handles, AND the engine now carries them — additive `LLMEngine`/`AsyncLLM` `add_request(MultiModalInputs)`/`generate` overloads (via `InputProcessor::process_inputs_mm`, mm_features onto `EngineCoreRequest`/`Request`), chat-template placeholder-string helpers, and a serving_chat `MultiModalChatFn` seam (default unset ⇒ text path byte-identical). The W3 `MultiModalChatFn` seam BODY now lands (`CLAIM-MM-SERVING-E2E`): `MakeQwen3VLImageChatFn` (chat_mm.{h,cpp}) turns an image chat request into the placeholder-EXPANDED engine input — marker-inject at the mm part position, render the chat template, tokenize (the single `<|image_pad|>` maps to one image_token_id via `EncodeWithSpecialTokens`), then `RouteImageRgb` expands to 196 image tokens + mm_features — wired into `src/vllm/entrypoints/openai/server_main.cpp` (guarded on `preprocessor_config.json`; text-only models leave it unset ⇒ byte-identical). Unit-gated `test_chat_mm` 8/8 (the W3 seam test drives the real tokenizer + chat template → 196 image tokens; RED line = the text-only path renders 0) + `test_input_processor` 10/10 + `test_openai_serving` (the production seam is invoked on an image request + routed to the engine mm generate overload; text-only never touches it; streaming+mm rejected), CPU-only, no model weights; bare-string chat requests byte-identical (proven). **ENGINE MM-FORWARD LANDED (`CLAIM-ENGINE-MM-FORWARD`, 2026-07-28): multimodal now runs through the engine's REGISTERED forward.** The architectural block is resolved: `ModelForwardInput` gains an additive default-nullopt `mm` field (merged inputs_embeds + 3-D MRoPE positions + DeepStack; nullopt for text ⇒ the shared runner path is byte-identical by construction), `Qwen3VLForConditionalGeneration` is now `REGISTER_VLLM_MODEL`-registered, and the registered forward folds the M2c decode into `ModelRegistry::Forward` via the shared `Qwen3VLForwardStepLastLogits` (`Qwen3VLGenerateGreedyViaRegistry` drives every step through the registry). GPU token-exact gate `test_qwen3vl_registry_e2e` on dgx.casa GB10: an image→text run THROUGH `ModelRegistry::Forward` emits the M2c golden tokens 32/32 STRICT (RED-first — unregistered ⇒ resolve throws); the M2c standalone cross-check holds 32/32 (the refactor is byte-neutral). Text inertness (the shared-path RED line): `test_runner` 16/16 + `test_scheduler` 36/36 + `test_model_registry` 24/24 + `test_chat_mm` 8/8 + `test_openai_serving` 41/41 all green. **Residual:** the FULL in-runner scheduler-fed tower run (the batched engine loop building the mm field from staged encoder outputs) + the real server `/v1/chat/completions` GPU end-to-end + video/multi-image/audio through the registered path — recipe in `.agents/specs/mm-serving.md`. **GEMMA-4 IMAGE through the registered forward LANDED (`CLAIM-GEMMA4-MM-E2E`, 2026-07-29):** `Gemma4ForConditionalGeneration` is now `supports_multimodal=true` with an mm branch routing `ModelForwardInput.mm` → `Gemma4Model::ForwardMm` (the SigLIP2 projector output masked-scattered into the `` rows + PLE image-rows→0; 1-D positions, no MRoPE/DeepStack); the driver `Gemma4GenerateGreedyViaRegistry` runs image→text through `ModelRegistry::Forward`. dgx sm_121a gate `test_gemma4_registry_e2e`: **16/18 content tokens BIT-EXACT** vs the STRICT `gemma4_e4b_image` golden (full sentence), the single divergence a terminal-punctuation bf16 near-tie (top1-top2 margin ~0.10-0.12 logit, invariant to vision-input precision ⇒ backbone bf16-accumulation, not the fold); text SACRED 32/32 UNCHANGED (inertness). Residuals: STRICT-18/18 (bit-match vLLM prefill bf16), Gemma-4 AUDIO→text e2e (mel A1 + merge; the G3 audio tower is per-stage proven), speed. **C2 (2026-07-31, `CLAIM-C2-VISION-QKV-BIAS`): both vision towers' attention QKV now fold to ONE `vt::MatmulBT` + a fused merged-`[3H]` BIAS epilogue + a contiguous `vt::QkvSplit` via the shared `models::FusedMergedQkvBiasSplit` helper** — qwen3_vl_vision (weight already resident-fused `[3H,H]`+`[3H]` bias) and gemma4_vision (loader-concatenated `qkv_proj [3H,H]`, no bias, per-slice clamp epilogue kept). Bit-exact (A/B unit `test_ops_qkv_merged_bias` byte-identical, RED-first); qwen3_vl gated e2e by `test_qwen3vl_e2e` image→text token-exact, gemma4_vision CPU-A/B + CUDA build-verified; existing tower gates held. See docs/BENCHMARKS.md | | Multimodal: video to text | Correctness-complete, speed-pending; OpenAI-server content-part parse + engine mm-request plumbing landed (CPU), end-to-end serving pending | End-to-end on Qwen3-VL-4B (near-tie-robust) and Qwen3.6-27B (strict 32/32); reuses the image tower and temporal MRoPE plus video preprocessing (frame sampling, temporal grid) | | Multimodal: audio to text | Decode beats vLLM; encoder TTFT measured and improved but not yet at parity; OpenAI-server content-part parse + audio processor routing + engine mm-request plumbing landed (CPU), end-to-end serving pending | End-to-end on Voxtral-Mini-3B (near-tie-robust, decoder token-exact 48/48) plus a Whisper-class encoder tower. Decode is graph-captured and runs vLLM's FA2 varlen split-KV kernel (the KV block-size is rounded up to a multiple of 16 to meet its precondition): audio TPOT 39.5 ms/token BEATS vLLM 0.25.0's 40.8 ms (0.97x, non-overlapping bands). The Whisper encoder attention now uses a flash-tiled non-causal head-dim-64 kernel (`vt::AttentionDenseFlash`): a block of query-warps shares each streamed K/V tile out of shared memory (FlashAttention K/V tiling), with the per-warp online-softmax math copied verbatim from the warp kernel so the output is bit-identical (token-identical). Same-binary A/B: the encoder self-attention drops from 35.11 to 19.29 ms/layer (1.82x, non-overlapping) and the encoder forward from ~1.83 s to ~1.37 s (1.33x), token-identical output (16/16, STRICT golden unchanged, sanitizer 0). The encoder weights are now device-resident: each of the 487 encoder weight tensors is converted to bf16 and uploaded to the GPU once (mirroring the decoder's residency), reused across forwards instead of being re-marshalled every call; same-binary A/B (`VT_WHISPER_ENC_REMARSHAL`) drops the encoder forward a further ~1.37 s to ~0.73 s (1.89x, non-overlapping), removing ~648 ms of per-call host weight marshalling (nsys: 974 fewer Host-to-Device copies, ~2.5 GB less traffic), byte-identical (16/16, STRICT golden unchanged, sanitizer 0). But encoder TTFT (~0.73 s) is still far above vLLM's 43 ms (~17x, was ~32x): the residual is now GPU-compute-bound (the scalar warp-per-query attention plus the conv GEMMs), so closing it needs a tensor-core MMA head-dim-64 non-causal flash attention. Correctness held under the ratified distributional near-tie gate (teacher-force PASS, 0 divergent, strict prefix exact vs vLLM; STRICT golden unchanged; 16/16). Remaining: a tensor-core MMA encoder-attention kernel and dropping the conv host round-trip (a device im2col kernel), and there is no batched c2+ or `audio_url` serving ingestion. See docs/BENCHMARKS.md | diff --git a/docs/USAGE.md b/docs/USAGE.md index 8f35730d3..2fde77e5f 100644 --- a/docs/USAGE.md +++ b/docs/USAGE.md @@ -39,6 +39,7 @@ build/examples/vllm-cli \ | `--seed S` | (unset) | RNG seed (enables seeded sampling) | | `--stream` | off | Stream token deltas to stdout | | `--speculative-config ''` | (unset) | Speculative decoding, same JSON as vLLM's flag. See [docs/SPECULATIVE-DECODING.md](SPECULATIVE-DECODING.md) | +| `--repeat N` | `1` | Load once, then run N blocking completions. Use it to read a warm decode tok/s without paying model load each time. Not supported with `--stream`, which falls back to 1 | | `-h`, `--help` | | Print usage and exit | Two more example binaries ship alongside it: @@ -152,6 +153,12 @@ to one built without video support. See | `--reasoning-parser ` | `none` | Reasoning parser (`think_auto`, `deepseek_r1`, `deepseek_v3`, `holo2`, `mistral`, `minimax_m2`, `minimax_m2_append_think`, `step3`, `olmo3`). `auto` detects, `none` disables | | `--kv-transfer-config ''` | (unset) | External KV connector, same JSON as vLLM's flag. See [docs/KV-OFFLOAD.md](KV-OFFLOAD.md) | | `--speculative-config ''` | (unset) | Speculative decoding (`mtp`, `dflash`, `ngram`), same JSON as vLLM's flag. See [docs/SPECULATIVE-DECODING.md](SPECULATIVE-DECODING.md) | +| `--enable-log-requests` / `--disable-log-requests` | on | Log each incoming request. Mirrors vLLM's flag of the same name | +| `--enable-log-outputs` | off | Also log the generated output, not just the request | +| `--max-log-len N` | `256` | Truncate logged prompts and outputs to N characters | +| `--enable-metrics` / `--disable-metrics` | on | Serve the metrics endpoint | +| `--enable-thinking` / `--no-enable-thinking` | off | Set the `enable_thinking` chat-template variable for templates that gate a reasoning block on it (Gemma-4 and friends). Our spelling of vLLM's `--default-chat-template-kwargs enable_thinking` | +| `--verbose`, `-v` | off | Verbose server logging | | `-h`, `--help` | | Print usage and exit | For a production deployment, use [LocalAI](https://localai.io), which can embed diff --git a/tests/scripts/test_check_public_doc_tables.py b/tests/scripts/test_check_public_doc_tables.py index 9507e905d..970c12157 100644 --- a/tests/scripts/test_check_public_doc_tables.py +++ b/tests/scripts/test_check_public_doc_tables.py @@ -359,9 +359,6 @@ def test_heading_inside_a_fence_is_not_a_section(self) -> None: self.assertTrue(all("not a heading" not in t for t, _ in sections)) -if __name__ == "__main__": - unittest.main() - # docs/STATUS.md is guarded by a RATCHET, not a budget: it may only shrink. The # mutations therefore prove both directions -- growth fails, shrinkage passes -- @@ -473,3 +470,5 @@ def test_the_ratchet_carries_no_hidden_headroom(self) -> None: f"measured {measured[key]}; lower it in the change that shrank " "the page", ) +if __name__ == "__main__": + unittest.main()