Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
34 commits
Select commit Hold shift + click to select a range
0fce293
spec(protocol): issue-native tracking — control plane to GitHub issue…
mudler Aug 6, 2026
d41e7bc
plan(protocol): P0 live-state audit — reconcile 160 live rows against…
mudler Aug 6, 2026
37c1892
tools(audit): load live matrix rows and gather Git evidence (P0 step 1)
mudler Aug 6, 2026
3f6390c
plan(protocol): correct the P0 census — 188 live rows, and 2 matrices…
mudler Aug 6, 2026
fb9c1c9
spec(protocol): correct the census — 714 rows, 188 live, 2 matrices w…
mudler Aug 6, 2026
8601e9b
tools(audit): cover all seven matrices, not the five MATRIX_PATHS gates
mudler Aug 6, 2026
e26535e
plan(protocol): repair 4 review findings in the P0 tool design
mudler Aug 6, 2026
805da16
tools(audit): anchor ID matching, surface parse errors, refuse a sile…
mudler Aug 6, 2026
fee5403
plan(protocol): put the anchoring tests in Task 1, where they belong
mudler Aug 6, 2026
030e118
tools(audit): pure evidence classifier for ACTIVE rows (P0 step 2)
mudler Aug 6, 2026
9bdf10b
plan(protocol): classifier must fail loudly on an ungathered branch
mudler Aug 6, 2026
732fdd6
tools(audit): index unmerged_by_branch, pin the multi-branch reason p…
mudler Aug 6, 2026
5db478f
plan(protocol): pin sorted() with two branches on the SAME side of th…
mudler Aug 6, 2026
3d5f314
tools(audit): advisory PARTIAL missing-modes flag (P0 step 3)
mudler Aug 6, 2026
83156ce
plan(protocol): pin CHECK_FAILS_ON exactly, and pin both halves of th…
mudler Aug 6, 2026
90792fe
plan(protocol): name the marker that fired, and escape the marker list
mudler Aug 6, 2026
698fa8d
tools(audit): name the marker that fired; escape markers before compi…
mudler Aug 6, 2026
547516e
plan(protocol): take the markers as an argument so the escaping is te…
mudler Aug 6, 2026
4cb357c
plan(protocol): restore Task 4, which an index-based edit had deleted
mudler Aug 6, 2026
2ef64fe
tools(audit): report, JSON and check modes (P0 step 4)
mudler Aug 6, 2026
4aaa9a8
plan(protocol): Task 4 ships 38 tests — the briefed vague-flag test p…
mudler Aug 6, 2026
4174bde
record(audit): live-state audit findings, no corrections applied yet …
mudler Aug 6, 2026
795da24
plan(protocol): Task 6 must retire 11 coordination claims atomically
mudler Aug 6, 2026
7211dba
plan(protocol): name the claim-retirement mechanic, and re-fetch befo…
mudler Aug 6, 2026
0bb5ac7
plan(protocol): scope Task 6 to the 10 abandoned ACTIVE rows only
mudler Aug 6, 2026
5de5f40
record(engine): live-state audit corrections — 3 rows off stale ACTIV…
mudler Aug 6, 2026
0222a04
record(model): live-state audit corrections — 3 rows off stale ACTIVE…
mudler Aug 6, 2026
00c267f
record(kernel): live-state audit corrections — 2 rows off stale ACTIV…
mudler Aug 6, 2026
8ad293f
record(quant): live-state audit correction — 1 row off stale ACTIVE (…
mudler Aug 6, 2026
7d3fad9
record(backend): live-state audit correction — 1 row off stale ACTIVE…
mudler Aug 6, 2026
95d1b27
record(state): live-state audit checkpoint — the ACTIVE claim set is …
mudler Aug 6, 2026
10fa813
plan(protocol): squashing does NOT clear the doc gate — the owed surf…
mudler Aug 6, 2026
5d8ed26
gate(audit): ACTIVE rows stay reconciled with Git reality (P0 step 7)
mudler Aug 6, 2026
cb41041
record(audit): name the gate's self-blinding hazard, ship the ACTIVE …
mudler Aug 6, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 9 additions & 9 deletions .agents/NOW.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,8 +38,7 @@ both gate models, reproduced 2–3x on an idle box. See [gates.md](gates.md) and
(0.26.0.dev0).

Method rules hardened (AGENTS.md): the STRUCTURAL lens (same kernel, different
throughput ⇒ audit the context; scan the REFERENCE's own rationale; per-shape
MEASUREMENT arbitrates; distrust aggregate bytes/time and CROSS-TOOL comparisons).
throughput ⇒ audit the context; per-shape MEASUREMENT arbitrates).

## Next actions

Expand All @@ -52,19 +51,20 @@ MEASUREMENT arbitrates; distrust aggregate bytes/time and CROSS-TOOL comparisons
f32-out caller) once the Laguna fix proves the mechanism.
4. **Restore `local-ai-worker`** on dgx when the GPU campaign ends
(`docker update --restart=always` + `docker start`).
5. **Protocol substrate — partly done.** Claim triage DONE; `docs/STATUS.md`
under a shrink-only ratchet; roadmap compacted; `AGENTS.md` tiered. REMAINING:
anchor backfill (98 rows `SPIKE`/`ACTIVE`, need code/test anchors; 6 model rows
need a DECISION, architecture unregistered); record-era rollover BLOCKED on
`check-agent-record.py` binding `DONE` rows to `parity-ledger.md` LINE anchors
(re-anchor by ROW ID first; `state.md`/`benchmark-record.md` can roll now).
5. **Protocol substrate — partly done.** Claim triage + live-state audit DONE
(10 unevidenced rows → `READY`, 11 claims retired, 9 amended); `STATUS.md`
ratcheted; roadmap compacted; `AGENTS.md` tiered. REMAINING: anchor backfill
(6 model rows need a DECISION); record-era rollover BLOCKED on `DONE` rows
bound to `parity-ledger.md` LINE anchors (re-anchor by ROW ID).
★ The gate SELF-BLINDS on those same 10 (audit §➁a); its fix owes an 8-row
adjudication. workflow.md now states the `ACTIVE` precondition.

**Operator/helper protocol**
([spec](specs/operator-helper-protocol.md)): roles DECLARED then MATERIALIZED
into a lock or worktree+PR; operator merges PRs first and does features only via
sub-agents; helpers use worktrees on `row/<ROW-ID>` and open a DRAFT PR at the
START, which IS the claim. **W0-W5 LANDED**; role discipline ENFORCING,
`--require-role` still opt-in. Queue: 4 rows. Backfill: 79 rows, 30 anchored; blocker is claim FAMILIES.
`--require-role` still opt-in. Queue: 10 rows — 6 are audit-vacated, with LANDED gate anchors; READ before picking. Backfill: 79 rows, 30 anchored; blocker is claim FAMILIES.
**Upstream inventory** ([spec](specs/upstream-derived-inventory-2026-08-05.md),
drift-gated, arch parity BOTH ways): SM060/061/070 below vLLM's floor =
OUT-OF-SCOPE; COMP-*/DISTRIBUTED-* are REAL unported work; **all 362 archs now have rows**; llama.cpp's 11 extra devices are IN SCOPE, spike-gated
Expand Down
2 changes: 1 addition & 1 deletion .agents/backend-matrix.md
Original file line number Diff line number Diff line change
Expand Up @@ -270,7 +270,7 @@ the rest are `SPIKE`.
| ID | Item | Upstream | Our code | Tests/evidence | Spike/spec | State | Owner |
|---|---|---|---|---|---|---|---|
| `BACKEND-DISTRIBUTED-COMM` | The unifying `vt::Communicator` / process-group abstraction — rank/world_size + AllReduce(sum/max/min/prod)/AllGather/Send/Recv, stream-ordered (each takes a `Queue&`). **W1 LANDED**: abstraction (`include/vt/communicator.h`) + a CPU in-process multi-rank transport (`src/vt/communicator.cpp`, N ranks = N host threads over one barrier+staging+mailbox) proven by `tests/vt/test_communicator.cpp` (2/4-rank AllReduce-sum + AllGather exact on every rank, Send/Recv rendezvous, RED-verified; 8 cases/50 assertions). `world_size==1` ⇒ every collective a byte-identical no-op (asserted). **W2 LANDED**: collectives now ROUTE through `OpProvider`/`OpId` (`kAllReduce`/`kAllGather`/`kSend`/`kRecv`, keyed on the queue's DeviceType) — the CPU in-process reduce registered on kCPU (`test_communicator` still 50/50 through the OpId path), the NCCL provider on kCUDA. W2+ residuals: RDMA/TCP (Spark), MLX-ring (kMETAL) transports | vLLM `device_communicators/base_device_communicator.py:147` (DeviceCommunicatorBase interface, the port template) + `distributed/parallel_state.py:358` (GroupCoordinator dispatch; world_size==1 bypass :638) | LANDED: `include/vt/communicator.h` + `src/vt/communicator.cpp` (sibling of `vt::Queue` `include/vt/device.h:50`); OpId routing via `include/vt/op_provider.h:108` (`OpId::kAllReduce/…`); stream-order hooks reused `include/vt/backend.h:87-104` | CPU exact-gate (`test_communicator`, 50/50, via OpId path) | [scale-out spike](specs/scale-out-distributed.md) | `ACTIVE` | `CLAIM-SCALE-OUT-W2` |
| `BACKEND-DISTRIBUTED-TP` | Tensor parallel (intra-node multi-GPU) — sharded Column/Row/QKV linears, vocab-parallel embed + LM head, attention-head split, MoE expert-parallel; all-reduce after o_proj/MLP-down and the EP combine. **W2 LANDED (CPU-gated)**: the `TensorParallel`/`TpShard`/`TpAllReduceSum` wiring (`include/vllm/model_executor/models/tensor_parallel.h`) threaded into the Qwen3-dense forward (o_proj all-reduce `dense_attn_block.h`, MLP-down `qwen3.cpp`) + the MergedColumn shard at the loader chokepoint (`dense_weight_loaders.h`), proven by `tests/vt/test_tp_forward.cpp` — sharded-matmul + RowParallel all-reduce **== the unsharded tp=1 forward** over the W1 CPU communicator (RED-verified: dropping the all-reduce fails 24 assertions). `tp_size==1`/nullptr ⇒ every helper a byte-identical no-op (asserted). RESIDUAL (HW-gated, no ≥2-GPU box): QKV head-aware KV replication, vocab/LM-head + MoE-EP sharding, and a real TP-2 GPU forward | vLLM `layers/linear.py:418` (Column, out-dim shard) / `:1612` (Row, all-reduce :1766) / `:1021` (QKV heads :1074) + `vocab_parallel_embedding.py:198` + `fused_moe/expert_map_manager.py:22` | LANDED `include/vllm/model_executor/models/tensor_parallel.h`; seams `dense_attn_block.h` (o_proj all-reduce) + `qwen3.cpp` (MLP-down); weight chokepoint `dense_weight_loaders.h:131` (column shard) | CPU multi-rank TP gate (`test_tp_forward`, 60/60, RED-verified) | [scale-out spike](specs/scale-out-distributed.md) | `ACTIVE` | `CLAIM-SCALE-OUT-W2` |
| `BACKEND-DISTRIBUTED-TP` | Tensor parallel (intra-node multi-GPU) — sharded Column/Row/QKV linears, vocab-parallel embed + LM head, attention-head split, MoE expert-parallel; all-reduce after o_proj/MLP-down and the EP combine. **W2 LANDED (CPU-gated)**: the `TensorParallel`/`TpShard`/`TpAllReduceSum` wiring (`include/vllm/model_executor/models/tensor_parallel.h`) threaded into the Qwen3-dense forward (o_proj all-reduce `dense_attn_block.h`, MLP-down `qwen3.cpp`) + the MergedColumn shard at the loader chokepoint (`dense_weight_loaders.h`), proven by `tests/vt/test_tp_forward.cpp` — sharded-matmul + RowParallel all-reduce **== the unsharded tp=1 forward** over the W1 CPU communicator (RED-verified: dropping the all-reduce fails 24 assertions). `tp_size==1`/nullptr ⇒ every helper a byte-identical no-op (asserted). RESIDUAL (HW-gated, no ≥2-GPU box): QKV head-aware KV replication, vocab/LM-head + MoE-EP sharding, and a real TP-2 GPU forward | vLLM `layers/linear.py:418` (Column, out-dim shard) / `:1612` (Row, all-reduce :1766) / `:1021` (QKV heads :1074) + `vocab_parallel_embedding.py:198` + `fused_moe/expert_map_manager.py:22` | LANDED `include/vllm/model_executor/models/tensor_parallel.h`; seams `dense_attn_block.h` (o_proj all-reduce) + `qwen3.cpp` (MLP-down); weight chokepoint `dense_weight_loaders.h:131` (column shard) | CPU multi-rank TP gate (`test_tp_forward`, 60/60, RED-verified) | [scale-out spike](specs/scale-out-distributed.md) | `READY` | - |
| `BACKEND-DISTRIBUTED-PP` | Pipeline parallel — PP stage split (`PPMissingLayer` analogue) + inter-stage `IntermediateTensors` send/recv over the comm layer + multi-worker executor fan-out | vLLM `models/utils.py:785` (PPMissingLayer) / `:798` (make_layers) + `distributed/utils.py:127` (get_pp_indices) + `parallel_state.py:957` (send_tensor_dict) | fan-out seam `src/vllm/v1/executor/executor.cpp:7-34` (direct single-worker call today) | - | [scale-out spike](specs/scale-out-distributed.md) | `SPIKE` | `CLAIM-SCALE-OUT-SPIKE` |
| `BACKEND-DISTRIBUTED-DP` | Data parallel — N independent engine replicas over the SAME weights + a DP coordinator (global "request wave" so all DP ranks step together) + a per-step token-count all-reduce; DP×EP is the large-scale DeepSeek serving topology (DP-replicated attention + EP-sharded experts). NOT part of `world_size` (DP is outside: `world_size_across_dp = world_size × DP`) | vLLM `v1/engine/coordinator.py:23` (`DPCoordinator`, wave :33-56) + `v1/worker/dp_utils.py:164` (`coordinate_batch_across_dp`; per-step `num_tokens_across_dp` all-reduce :53) + group `distributed/parallel_state.py:1866` + flags `config/parallel.py:129-145` | reuses W1 `Communicator::AllReduce` for the token-count sync; NEW engine-replica executor + coordinator; **depends on the multi-worker executor (`executor.cpp:7-34`)** | - | [parallelism-modes spike](specs/parallelism-modes.md) | `SPIKE` | `CLAIM-PARALLELISM-MODES-SPIKE` |
| `BACKEND-DISTRIBUTED-EP` | Expert parallel — each rank owns a DISJOINT subset of WHOLE experts (vs TP's every-rank-all-experts intermediate-dim shard); all-to-all dispatch (tokens→expert-owner) + all-to-all combine. EP group = DP×PCP×TP ranks (MoE-only), a re-grouping NOT a new `world_size` knob; EPLB is a separate same-ranks group. DeepEP HT/LL/V2 all-to-all backends present upstream | vLLM `fused_moe/expert_map_manager.py:22` (`determine_expert_map`, local_num_experts :69) + `use_ep` `fused_moe/config.py:1204` + EP group `distributed/parallel_state.py:1892` + all-to-all backends `fused_moe/all2all_utils.py:44` (`prepare_finalize/deepep_{ht,ll,v2}.py`); flags `--enable-expert-parallel` `config/parallel.py:165` + `--all2all-backend` `:188` | NEW `OpId::kAllToAll` (dispatch/combine) on the same `vt::Communicator` + the expert-subset loader (`qwen3_moe_weights.cpp:35-38` per-expert loop); `vt::MoeCombine` (`include/vt/ops.h:104`) becomes the combine | - | [parallelism-modes spike](specs/parallelism-modes.md) | `SPIKE` | `CLAIM-PARALLELISM-MODES-SPIKE` |
Expand Down
Loading