Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .agents/NOW.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ Working head: `row/backend-rocm-w0` (#41). Prior: benchmark checkpoint
| f32-out GEMV audit | Only laguna + ds4 bf16 tower affected; gate models unaffected | Re-verify ds4 tower same-tool |
| Invocation-parity prevention | CI guard + checklist landing | Merge; build-verify `kGemvHeuristicAlgos` on dgx |
| MiniMax-H3 lane | **fl2va COHERENT; ref2va NVFP4 grid DIAGNOSED (#95): NO loader bug** | weights/islands/RoPE all quant-noise-close to coherent GGUF; residual = community-NVFP4 quant fidelity §8.12 |
| Kimi-Linear-48B (KDA+NoPE-MLA+MoE) | **e2e RUNS** (bf16-resident §13): 13/13·656. Token gate **NEAR-TIE 106/128** | device GDN/MLA islands; 1.59 tok/s; default OFF |
| Kimi-Linear-48B (KDA+NoPE-MLA+MoE) | bf16 knobs **106→120/128**, NOT STRICT (§14, `row/KIMI-LINEAR-STRICT-SPEED`); default OFF | residual = device islands; 1.30 tok/s |
| 35B fresh grid | **BOUND** @`1ea26427`: 0.93-1.03x, c16 0.93x. INTAKE + Option A both NEGATIVE | Lever left: prefill glue (#61) |
| Qwen3.5-4B revalidation | 0.9971x @`59674cf1` (#35); TTFT/PSS pass, TPOT/ITL open | `docs/bench-evidence/` |
| MXFP4 parity | c1 1.020, c2-c8 0.962-0.969. **#82 CLOSED: ptxas-lineage REFUTED (A/B ties our+vLLM PTX all ptxas/JIT; +10us=engine context, not codegen)** | TERMINAL: at parity |
Expand Down
19 changes: 19 additions & 0 deletions .agents/benchmark-record.md
Original file line number Diff line number Diff line change
Expand Up @@ -14535,3 +14535,22 @@ Ran #94's prescribed next diagnostic — a layer-by-layer activation diff of the
**Render A/B re-confirmed in the same byte-inert build (256×256/22f/12steps, `pe.f32`, `VT_H3_ACT_DUMP` unset).** NVFP4 t2va = pale multicolour PATCH GRID every frame (0/10/21, per-patch uniform = degenerate latent); FL2VA-GGUF t2va = a COHERENT photorealistic orange cat on a windowsill. The diagnostic hooks are byte-inert (reproduce #94's outputs exactly).

**CONCLUSION.** The mission's "2nd load-path defect" does not exist — our loader materializes the community NVFP4 file's weights faithfully (byte-verified). The residual is the **community NVFP4 checkpoint's quantization fidelity** (metadata `converted_by: "Star Ultimate Model Converter Pro"` — same dubious-converter lineage as the #94 nibble bug; corr to the coherent Q3_K only 0.85-0.94, well below the >0.99 a clean 4-bit quant gives) interacting with the DiT's massive-activation sensitivity. Definitively separating "poor community quant" from "inherent t2va-OOD sensitivity of these fl2va/ref2va finetunes" needs a clean reference — a bf16 ground truth (132 GiB host-f32 = OOM on one GB10) or a same-finetune REF2VA-GGUF control (disk-blocked, 23 GiB free) — both currently blocked; the path forward is an official modelopt-NVFP4 checkpoint, not a loader change. The #94 nibble fix stands as the objectively-correct dequant. The fp4-resident Marlin arm's separate grid stays a distinct, wiring-gated-only residual (untouched, per the mission). Artifacts `~/h3fp4/diff/{nvfp4_t2va,gguf_t2va}.mp4` + `{gguf,nvfp4}{2,3,4}.txt` fingerprints.

---

## 2026-08-07 — Kimi-Linear-48B W7-speed STRICT lever: bf16 regime 106→120/128, plateaus (device islands the residual)

Full 48.9B model on GB10 (sm_121a clean CUDA build, cutlass-4.5.0, 14 GDN AOT symbols; CPU gate `test_kimi_linear_forward` 13/13·656 in the CUDA binary), §12 8-prompt×16-token gate vs the STRICT deterministic golden. Three env-gated numeric knobs added to `kimi_linear_device.cpp` (default OFF, byte-identical). Both flock locks, reclaim-wait ≥90 GiB between every reload, min-avail ≥115 GiB, no reboot.

| Config | env | TOKEN MATCH /128 | tok/s | note |
|---|---|---|---|---|
| control | (none) | 106 | 1.31 | reproduces §13 baseline exactly (build integrity) |
| residual only | `VT_KIMI_BF16_RESIDUAL` | 106 | 1.61 | net-zero: shuffles flips, BREAKS p3 into `163586×` repeat |
| islands only | `VT_KIMI_BF16_ISLANDS` | 106 | 1.30 | fixes p2, destabilizes p3 — net-zero |
| **residual + islands** | both | **120** | 1.30 | **BEST**: p0–p6 all 16/16; only p7 pos-8 flips (`18705`→`58084`) |
| + island-output bf16 | (islands, out-round) | 90 | 1.32 | REGRESSION — reverted |
| + f32 accumulation | `…VT_KIMI_ISLAND_F32ACC` | 91–106 | 0.84 | NEGATIVE — kept as documented A/B |

FLIP LEDGER (deterministic golden arbitrates): the two levers interact — island bf16-input rounding fixes p2 but repeats p3; the bf16 residual stream (vLLM `fused_add_rms_norm` order) re-stabilizes p3. At 120/128 the SOLE divergence is p7 position 8 (golden deterministically `18705`, ours `58084`), a single near-tie that cascades to 8/16 on p7. Further precision-"matching" (output bf16, f32 accum) is a coin-flip that regresses — it is not vLLM's actual GDN-Triton/FA2 kernel arithmetic. Host-precision-matching PLATEAUS at 120/128.

VERDICT: NO arm STRICT (K=3-deterministic golden → STRICT required, not distributional); default STAYS OFF (parity-enablers). NAMED residual (= also the speed lever): the device islands — a NEW per-channel-decay GDN kernel (`g[T,H,D]`; `vt::GdnDecode`/`GdnPrefill` carry only per-head `g[T,Hv]`, ops.h:1797/1846 — NOT a drop-in) + paged `mla::ForwardMlaAttentionBlock` (FA2). Speed HW-forced-indirect: 1.30 tok/s (O(n²) recompute + host islands, invariant to the numeric knobs); vLLM cannot serve Kimi-Linear-48B at bf16 on one GB10 (oracle capture needed util 0.82 for a single-seq eager run) so a direct `vllm bench throughput` arm is infeasible. Row STAYS ACTIVE.
78 changes: 78 additions & 0 deletions .agents/specs/kimi-linear.md
Original file line number Diff line number Diff line change
Expand Up @@ -752,6 +752,84 @@ is not token-exact). Row STAYS `ACTIVE`.

---

## 14. W7-speed STRICT lever MEASURED — bf16 regime recovers 106→120/128, plateaus; device islands remain the residual (2026-08-07, `row/KIMI-LINEAR-STRICT-SPEED`)

The recorded path to STRICT ("device islands + bf16 residual stream = the same work as
speed") was implemented as three env-gated numeric knobs in `kimi_linear_device.cpp`
(default OFF → the f32 vehicle is byte-identical, CPU gate `test_kimi_linear_forward`
**13/13·656** in the CUDA binary) and MEASURED on GB10 (the full 48.9B model, the §12
128-token gate vs the STRICT deterministic golden). Clean-from-`origin/main` CUDA build
(`-DVLLM_CPP_CUDA=ON -DVLLM_CPP_TRITON=ON -DVLLM_CPP_CUDA_ARCHITECTURES=121a -DVLLM_CPP_
CUTLASS_DIR=…cutlass-4.5.0`, nvcc 13.0.88, Release, 14 GDN AOT symbols linked). Memory
safe throughout every reload (min-avail ≥ 115 GiB, both flock locks, reclaim-wait, no
reboot).

### Knobs (`kimi_linear_device.cpp`)
- `VT_KIMI_BF16_RESIDUAL` — carry the residual stream in bf16 like vLLM's
`fused_add_rms_norm` (residual stored bf16, block outputs bf16, RMSNorm variance over
the f32 pre-store sum). Implemented as in-place `CastBf16`→`CastF32` rounds at exactly
vLLM's rounding points (embed out, each block out, residual after each add) keeping f32
STORAGE so the islands still read f32 — byte-matching vLLM's fused-add order.
- `VT_KIMI_BF16_ISLANDS` — round the host-fallback island INPUTS (KDA q/k/v/g1/beta,
NoPE-MLA q/kv/kpe) to bf16 (RNE) before the recurrence/softmax.
- `VT_KIMI_ISLAND_F32ACC` — f32 (not f64) accumulation in the islands. **MEASURED NEGATIVE**,
kept as a documented-negative A/B knob.

### Measurement (token match /128 vs the deterministic golden; STRICT required)
| Config | env | /128 | verdict |
|---|---|---|---|
| control | (none) | **106** | reproduces §13 baseline exactly |
| residual only | `BF16_RESIDUAL` | 106 | net-zero — SHUFFLES flips (fixes p2, BREAKS p3 into a `163586×` repeat loop) |
| islands only | `BF16_ISLANDS` | 106 | fixes p2 but destabilizes p3 (repeat loop) — net-zero |
| **residual + islands** | `BF16_RESIDUAL BF16_ISLANDS` | **120** | **BEST** — p0–p6 all 16/16 exact; only p7 flips |
| + island-output bf16 | `…BF16_ISLANDS(out)` | 90 | **REGRESSION** (reverted) |
| + f32 accumulation | `…ISLAND_F32ACC` | 91–106 | **NEGATIVE** (reverted from the ISLANDS path) |

### Flip ledger / razor verdict (the deterministic golden arbitrates every flip)
- The two levers INTERACT: island bf16-input rounding fixes p2 but destabilizes p3 into a
degenerate repeat (`163586×`); the bf16 residual stream then RE-stabilizes p3 (kills the
repeat). Together they make **p0–p6 all 16/16 token-exact** (was 6/8 → 7/8 fully exact).
- The SOLE remaining divergence at 120/128 is **p7 position 8**: the golden deterministically
emits `18705`, our island emits `58084` (a genuine near-tie), then the greedy path cascades
(8/16 on p7). A single near-tie flip across the whole 8-prompt battery.
- Further precision-"matching" (rounding the island OUTPUT to bf16 → 90; f32 accumulation →
91–106) is a COIN-FLIP that regresses, because it is not vLLM's ACTUAL GDN-Triton / FA2
kernel arithmetic — it just perturbs which near-ties flip. Host-precision-matching PLATEAUS
at 120/128.

### Verdict + default
**NO arm reaches STRICT** (best 120/128 is a DIVERGENCE; the golden is K=3 deterministic so
STRICT — not the distributional gate — is required). Per parity-enablers, `VT_KIMI_DEVICE_
COMPUTE` and all three knobs STAY **OFF** (a near-tie is not token-exact). Row STAYS `ACTIVE`.

### The named residual — the device islands (why host-matching cannot close p7)
The one principled path to STRICT is routing the two islands through vLLM's ACTUAL device
kernels, but it is NOT a drop-in (the mission's own assessment, now proven by measurement):
- **KDA:** vLLM's decay is **per-k-channel** `g[T,H,D]` (`kimi_gdn_linear_attn.py` +
`third_party/flash_linear_attention/ops/kda.py`), but our `vt::GdnDecode`/`GdnPrefill`
carry only a **per-HEAD scalar** decay `g/beta[T,Hv]` (`include/vt/ops.h:1797,1846`). So the
vendored GDN Triton-AOT cubins CANNOT express KDA — a NEW per-channel-decay GDN kernel
(`g[T,H,D]` + the `-exp(A_log)*softplus(f_b(f_a(x))+dt_bias)` gate) is required.
- **NoPE-MLA:** needs the paged `mla::ForwardMlaAttentionBlock` (FA2) over the runner's het-KV
(the born-on-runner residual), not a host softmax.
These are ALSO the speed levers — the correctness vehicle re-computes the WHOLE sequence every
decode step (O(n²)) with a host Download/upload per KDA/MLA layer per step. That is why the
measured tok/s is invariant to the numeric knobs.

### Speed (HW-forced-indirect — vLLM cannot serve this model on one GB10 at bf16)
Steady tok/s over 127 steps (single-load, medians, cold first-leg discarded): control **1.31**,
islands **1.30**, resIsl **1.30** (first-step ~0.62 s, steady ~0.77 s/step). The extra bf16
rounding casts cost ~0.15 s/step vs the §13 1.59 baseline; ALL configs are the same O(n²)
full-recompute + host-island rate. **Honest bar:** the §12 oracle golden capture itself needed
`gpu_memory_utilization=0.82` (~97.6 GiB) with only 15 GiB min-avail for a SINGLE-seq eager
run — vLLM cannot SERVE Kimi-Linear-48B at bf16 on ONE GB10 with any KV headroom, so a direct
`vllm bench throughput` arm is HW-infeasible; the comparison is recorded as HW-forced-indirect
(our absolute 1.30 tok/s + the per-step GPU-active anchor). No isolated host-tail lever (grouped
MoE seam, on-GPU sampling) moves the needle without the device-island + paged-incremental-decode
rewrite, which is the SAME W7-speed residual. Scoped as the named follow-up, not forced.

---

## Structured contract (machine-readable — mirrors deepseek-v4-flash.md)

## Scope
Expand Down
23 changes: 23 additions & 0 deletions .agents/state.md
Original file line number Diff line number Diff line change
Expand Up @@ -40773,3 +40773,26 @@ sensitivity. Clean-reference disambiguation blocked (bf16 132GiB = OOM on 1 GB10
= 23G-disk-blocked); path forward = an official modelopt-NVFP4 ckpt. fp4-resident Marlin arm's separate
grid untouched (wiring-gated-only). Diagnostic hooks landed byte-inert. Records: spec §8.12 + §8.2 row,
STATUS/BENCHMARKS/FEATURES H3 rows, benchmark-record, NOW. Box left clean.

## 2026-08-07T09:45 — Kimi-Linear-48B W7-speed STRICT lever MEASURED: bf16 regime 106→120/128, plateaus; device islands remain the residual
<!-- state: 2026-08-07T09:45 -->
KIMI-LINEAR-STRICT-SPEED (`row/KIMI-LINEAR-STRICT-SPEED`, helper) — implemented + MEASURED the recorded
STRICT path ("device islands + bf16 residual stream") as three env-gated numeric knobs in
`kimi_linear_device.cpp` (default OFF → f32 vehicle byte-identical, CPU gate `test_kimi_linear_forward`
13/13·656 in the CUDA binary). Clean CUDA build on GB10 (sm_121a, cutlass-4.5.0, 14 GDN AOT symbols),
full 48.9B model, §12 128-token gate vs the STRICT deterministic golden, both flock locks + reclaim-waits,
min-avail ≥115 GiB, no reboot. RESULT (token match /128): control 106; `VT_KIMI_BF16_RESIDUAL` alone 106
(net-zero — shuffles flips, BREAKS p3 into a `163586×` repeat); `VT_KIMI_BF16_ISLANDS` alone 106 (fixes p2
but destabilizes p3); **residual + islands = 120/128 BEST** (p0–p6 all 16/16; only p7 pos-8 flips: golden
`18705` vs ours `58084`, a single near-tie that cascades). The two levers INTERACT — island input-rounding
fixes p2 but repeats p3; the bf16 residual re-stabilizes p3. Further precision-matching MEASURED NEGATIVE
(island-output bf16 → 90, reverted; f32 accumulation → 91–106, kept as documented-negative A/B knob):
host-precision-matching PLATEAUS at 120/128 because it is not vLLM's ACTUAL GDN-Triton/FA2 arithmetic.
VERDICT: NO arm STRICT (golden is K=3-deterministic → STRICT, not distributional); default STAYS OFF
(parity-enablers). NAMED residual (also the speed lever): the device islands — a NEW per-channel-decay GDN
kernel (`g[T,H,D]` per `kda.py`; `vt::GdnDecode`/`GdnPrefill` carry only per-head `g[T,Hv]`, ops.h:1797,1846
— NOT a drop-in) + the paged `mla::ForwardMlaAttentionBlock` (FA2). Speed: 1.30 tok/s (O(n²) full-recompute
+ host islands; invariant to the numeric knobs, confirming the cost is the recompute structure not the
arithmetic); vLLM can't SERVE Kimi-Linear-48B at bf16 on ONE GB10 (oracle capture needed util 0.82 for a
single-seq eager run) → HW-forced-indirect. SHAs e048f4ee/3d6f81d1/(revert) on the row branch; PR pending.
Records: spec §14, STATUS/BENCHMARKS/FEATURES Kimi rows, benchmark-record, NOW. Box left clean.
2 changes: 1 addition & 1 deletion docs/BENCHMARKS.md
Original file line number Diff line number Diff line change
Expand Up @@ -309,7 +309,7 @@ built on it rather than keeping the flattering one.
| Multimodal image, audio, video | Correctness gated, speed unmeasured | Per-modality speed grids |
| `/v1/videos` OpenAI (Sora) shape | **No number owed**: a CPU serving-surface change (request aliases, the MP4 content route, and reference conditioning wiring), unit-gated only, no kernel or generation path touched | Video generation speed stays the MiniMax-H3 FP4 row below |
| Qwen3-dense decode CUDA-graph | Token-exact pass, ~4.3% e2e directional | Steady-state per-step tok/s |
| Kimi-Linear-48B-A3B (KDA+MLA+MoE) | Full-model GB10 e2e RUNS (bf16-resident §13), NEAR-TIE 106/128, pool math CLOSES; default OFF | Full model RUNS on GB10 (bf16-resident, RSS peak 1.7 GiB, min-avail 21 GiB, no OOM). Token NEAR-TIE 106/128 (6/8 prompts exact, numerics vs deterministic oracle). 1.59 tok/s. Detail: spec §13 |
| Kimi-Linear-48B-A3B (KDA+MLA+MoE) | e2e RUNS (bf16-resident §13); bf16-regime knobs 106→120/128 (7/8 exact), NOT STRICT; default OFF | bf16 residual+island-inputs → 120/128 best (control/each-alone 106; output-bf16 & f32-accum NEGATIVE); 1 near-tie left. 1.30 tok/s (O(n²)); vLLM HW-can't-serve bf16 on 1 GB10. Residual = device islands. §14 |
| vLLM 0.26 re-benchmark | Pending | Re-run the binding grids on the advanced pin |
| MiniMax-H3 FP4 speed (W-FP4a) | **Measured GB10 (`row/H3-FP4-GPU-E2E`).** Marlin W4A16 byte-exact vs bf16; fp4 a memory win, 0.8x bf16/forward. Real-ckpt fp4-resident e2e RUNS (mp4/wav) | fp4 speed CLOSED. Detail: benchmark-record + spec §8 |
| MiniMax-H3 render coherence (`row/H3-RENDER-CLOSE` #77) | **CLOSED: a COHERENT scene on GB10.** #70/#74 white was wrong-PARTITION usage (t2va on the ref2va ckpt); t2va on the FL2VA GGUF renders a prompt-matched orange cat (adj-cos 0.95 vs 0.06, no patch-grid) | Verified first: t2va inputs byte-exact vs upstream; CUDA device==host at seq 1920. Follow-up `H3-TASK-PARTITION-GUARD`: the task/partition mismatch now RAISES 1:1 with `_resolve_task` (spec §8.6-8.7) |
Expand Down
3 changes: 3 additions & 0 deletions docs/ENVIRONMENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -64,6 +64,9 @@ portable/reference path. In normal operation leave them unset.
| `VT_MARLIN_DENSE` | off (opt-in) | `=1` routes the E=1 dense NVFP4/MXFP4 projections (`dense_nvfp4_gemm.h` `MatmulNvfp4MarlinD`/`GateUpFusedMarlinD`) through vLLM's OWN dense marlin GEMM (`vt::MarlinDenseGemm`) instead of the single-expert MoE-marlin route. The dense kernel is direct-A + tile-per-CTA with vLLM's own dense fp32-C_tmp reduce, so at M≤8 it runs the sms-wide (48-CTA) grid `VT_MARLIN_E1_PAR1` targets but WITHOUT that flag's par-regroup ULP — the byte-preserving fix for the strict-32B token flip (row `KERNEL-MARLIN-DENSE-PORT`). Reuses the same marlin resident + workspace (shared `marlin_permute` repack). **DEFAULT OFF** until the strict token battery proves oracle byte-match and the binding beats the MoE route; then flipped ON per the parity-enabler policy. CUDA-only (needs `VT_MARLIN_NVFP4`) |
| `VT_MM_DECODE_EAGER` | off (graph on) | Set to `1` to force the eager per-step multimodal (Qwen3.6-27B image/video) decode instead of routing it through the captured dense decode graph. Rollback / A-B knob; the graphed path is token-exact with the eager path |
| `VT_KIMI_DEVICE_COMPUTE` | off (opt-in) | `=1` routes the Kimi-Linear-48B-A3B runner path (`KimiLinearModel::ForwardDevice`) through the W7 DBuf-resident device COMPUTE (`ForwardDeviceCompute`, the whole KDA/NoPE-MLA + MoE hybrid over pooled DBufs via the shared vt:: ops) instead of the default W6 host-reference compose. Default OFF keeps the CPU-verified host-ref-compose seam as production until the device compute is GPU-verified against the SACRED oracle; the device compute is CPU-gated (`test_kimi_linear_forward`, device==W2 reference within f32-accumulation tolerance, greedy-token-identical) but its GPU numerics are a NAMED pending. The flag exists so the device path CAN be exercised as the runner path for that verification |
| `VT_KIMI_BF16_RESIDUAL` | off (opt-in) | `=1` carries the Kimi-Linear device-compute residual stream in bf16 like vLLM's `fused_add_rms_norm` (residual/block-outputs bf16, RMSNorm variance over the f32 pre-store sum), via in-place f32→bf16→f32 rounds. W7-speed STRICT-lever A/B (spec §14). Default OFF → byte-identical. MEASURED: alone net-zero; WITH `VT_KIMI_BF16_ISLANDS` → 120/128 (best, still a near-tie, NOT STRICT) |
| `VT_KIMI_BF16_ISLANDS` | off (opt-in) | `=1` rounds the Kimi-Linear host-fallback island INPUTS (KDA q/k/v/g1/beta, NoPE-MLA q/kv/kpe) to bf16 (RNE) before the recurrence/softmax, toward vLLM's GDN-Triton/FA2 kernel precision. W7-speed STRICT-lever A/B (spec §14). Default OFF → byte-identical. MEASURED best config paired with `VT_KIMI_BF16_RESIDUAL` (106→120/128) |
| `VT_KIMI_ISLAND_F32ACC` | off (opt-in) | `=1` computes the Kimi-Linear island recurrence/softmax in f32 accumulation (not f64). W7-speed A/B knob, **MEASURED NEGATIVE** (91–106/128; kept as a documented-negative A/B, spec §14). Default OFF → byte-identical |
| `VT_WHISPER_ENC_EAGER` | off (flash-tiled attention on) | Set to `1` to force the naive per-key block-reduction attention in the Voxtral/Whisper audio encoder instead of the default flash-tiled kernel. Rollback / A-B knob; token-identical to the default path |
| `VT_WHISPER_ENC_WARP` | off (flash-tiled attention on) | Set to `1` to force the warp-scoped online-softmax attention (`vt::AttentionDenseFast`, the pre-flash default) in the Voxtral/Whisper audio encoder instead of the default flash-tiled kernel (`vt::AttentionDenseFlash`). Rollback / A-B knob; the flash-tiled path is bit-identical to the warp path (encoder self-attention ~1.82x faster) |
| `VT_WHISPER_ENC_REMARSHAL` | off (encoder weights resident) | Set to `1` to disable device-resident encoder weights and re-marshal (host f32->bf16 convert + H2D upload) all Whisper/Voxtral encoder weights on EVERY forward, restoring the pre-residency behavior. Rollback / A-B knob; byte-identical output (moves data only). Default residency uploads each encoder weight once and reuses it, removing ~648 ms of per-call host marshalling from the encoder forward |
Expand Down
Loading
Loading