Skip to content

release: ship complete W1-W13 binary matrix - #196

Merged
localai-bot merged 18 commits into
mainfrom
row/ENG-RELEASE-BINARIES
Aug 9, 2026
Merged

release: ship complete W1-W13 binary matrix#196
localai-bot merged 18 commits into
mainfrom
row/ENG-RELEASE-BINARIES

Conversation

@localai-bot

@localai-bot localai-bot commented Aug 9, 2026

Copy link
Copy Markdown
Collaborator

Scope

Tracks #117.

This is the single delivery PR for the complete primary downloadable vllm-server release matrix. Required W1-W11/W13 are implemented here; W5 arrived through #141. W12 remains the accepted optional, non-primary single-SM diagnostic lane and is not used to bypass either fat-CUDA bundle.

The pipeline now builds eight host/backend tuples, validates freshly extracted bytes, creates an immutable verified handoff, generates JSON/Markdown indexes from embedded manifests, attests the archives, and publishes every explicitly authenticated archive/checksum/provenance triplet. This fixes the test release that attached only a handoff and no binary.

Matrix

  • W1 ten-SM per-source CUDA gencode
  • W2 six exact multi-SM Triton AOT trees and four portable fallbacks
  • W3 x86_64 adaptive CPU inventory/completion
  • W4 aarch64 adaptive CPU inventory/completion
  • W5 release manifest/schema (merged via release: add versioned binary manifest (W5) #141)
  • W6 canonical deterministic server package
  • W7 staged archive validator and supply-chain metadata
  • W8 least-privilege dry-run/tag, provenance, protected publish
  • W9 adaptive CPU bundles
  • W10 x86_64/aarch64 CUDA fat bundles
  • W11 Metal/MLX/Vulkan bundles
  • W12 explicitly optional/non-primary; not emitted by the primary matrix
  • W13 generated release index/docs/retention

Verification

  • local clean CPU archive and all adaptive x86 tiers — PASS
  • local Vulkan archive: backend 35/35 and cross-device 11/11 — PASS
  • release archive/manifest/metadata/index/workflow and mutation suites — PASS
  • staged and post-commit scripts/agent-preflight.sh — PASS
  • PR-size waiver WAIVER-PR-SIZE-003 — PASS
  • current hosted CI, including corrected ten-SM CUDA proof — RUNNING
  • full eight-tuple release dry run — PENDING after merge because GitHub only dispatches workflows present on the default branch

ROCm remains blocked by the accepted contract; musl stays experimental preview. Build-only lanes do not acquire runtime/correctness/performance claims.

@mudler
mudler force-pushed the row/ENG-RELEASE-BINARIES branch from 0394402 to 80deeca Compare August 9, 2026 10:38
@localai-bot
localai-bot marked this pull request as ready for review August 9, 2026 10:50
@localai-bot
localai-bot requested a review from mudler August 9, 2026 10:50
@localai-bot
localai-bot marked this pull request as draft August 9, 2026 14:19
@mudler
mudler force-pushed the row/ENG-RELEASE-BINARIES branch from 80deeca to 654163c Compare August 9, 2026 14:20
@localai-bot localai-bot changed the title release: add canonical server package (W6) release: ship complete W1-W13 binary matrix Aug 9, 2026
@mudler
mudler force-pushed the row/ENG-RELEASE-BINARIES branch from c4e4adb to d693523 Compare August 9, 2026 15:07
mudler added 9 commits August 9, 2026 15:32
Record the W6 install and archive checkpoint after diagnosing the empty tagged release.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-5 [Codex]
Install a static-core vllm-server component, add deterministic stage and archive targets, and gate the extracted binary in CPU CI while preserving the library install.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-5 [Codex]
Record the developer-directed single-PR delivery topology while preserving independently verified work-unit checkpoints.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-5 [Codex]
Pin PR 196 as the single delivery surface and reconcile the live claim without advancing release evidence beyond W6.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-5 [Codex]
Split architecture-specific sm12x NVFP4 bodies from portable CUDA sources, assign exact per-source gencode intersections, and add a hosted compile/cubin audit with mutation coverage. The real NVIDIA toolkit build remains pending on PR CI.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-5 [Codex]
Use single-token nvcc -gencode options for per-source properties. This repairs the quoting failure observed in hosted CUDA run 31320289475 while retaining the exact ten-SM source matrix.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-5 [Codex]
Embed and namespace all six vendored Triton cubin trees, dispatch launchers and loaders by exact runtime SM, and preserve portable fallback on release SMs without an AOT tree. Add toolkit-free selector, namespace, ELF identity, and archive mutation gates.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-5 [Codex]
Register disabled creation mutations for the new W1 and W2 release auditors so the PR-size gate proves red-before/green-after behavior. Re-pin the byte-tight status ratchet after rebasing the release checkpoint onto current main.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-5 [Codex]
Bind the sanctioned size waiver to PR #196 only. The accepted release topology and explicit maintainer direction require W1-W13 to remain in one review and delivery PR.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-5 [Codex]
@mudler
mudler force-pushed the row/ENG-RELEASE-BINARIES branch from d693523 to 50f8de3 Compare August 9, 2026 15:40
mudler added 9 commits August 9, 2026 16:15
FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-5 [Codex]
FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-5 [Codex]
FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-5 [Codex]
FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-5 [Codex]
FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-5 [Codex]
FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-5 [Codex]
FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-5 [Codex]
FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-5 [Codex]
FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:GPT-5 [Codex]
@localai-bot
localai-bot marked this pull request as ready for review August 9, 2026 18:22
@localai-bot
localai-bot merged commit 4d2fdae into main Aug 9, 2026
15 checks passed
mudler added a commit that referenced this pull request Aug 9, 2026
…message

Main CI has been red on EVERY merged PR. Two independent defects, both landed
by #178's push (run 31332846716); `cuda-fat-build` in that run was cancelled by
concurrency, not failing.

THE GATE REDDENED MAIN FOR OBEYING IT. `check-role-discipline.py` judges every
commit in the `before..after` range on its own message. A PR landed with a REAL
merge commit pushes the merge AND the branch commits under it: the merge names
the PR, the branch commits were never required to, so each one read as a direct
push. `6603356a` (#178), `e73cbbae` (#204) and `1a02ab4f` (#196) all failed this
way, in both `documentation-checkpoint` and `agent-record`. Arrival is now judged
ONCE, on the commit that lands the change: `merged_pr_content` exempts what a
row/* PR merge brings in. Squash-merges are untouched -- their one commit carries
"(#N)" and passes on its own message. NOT a weakening, and gated as such: only
the SIDE parents count, so `--not parents[0]` keeps a commit pushed straight to
main from being laundered by merging a PR on top, and a merge naming no row and
no PR exempts nothing. Four unit checks build real git history for those cases,
plus the exact `3bbee96e..0cf3dbb` range CI ran, pinned with `has_reached_main`
forced TRUE -- from a `row/*` worktree everything reports as pending PR
disposition, and the test would have passed against the defect it exists to
catch. Suite 47/47, and 5/5 red without the fix.

A STALE MESSAGE IN A TEST. `6603356a` taught `LoadMergedBf16RawNK` to accept
F8_E4M3 shards and rewrote its rejection to name the supported dtypes;
`test_qwen27_dense_forward.cpp:229` still asserted the old "expected BF16", so
`build-test-cpu` and both sanitizer legs failed on it. The expectation now reads
the message the loader raises, and the FP8 merge path that arrived WITHOUT a test
in this file gets one: a mixed BF16+FP8 merged parameter, expectations hand-
computed from E4M3 bytes and the scale (never re-derived through the same
dequant helper the loader calls), plus the per-channel-scale rejection.
7/7, 333 assertions.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Claude:claude-opus-5 [ClaudeCode]
mudler added a commit that referenced this pull request Aug 9, 2026
…message (#210)

Main CI has been red on EVERY merged PR. Two independent defects, both landed
by #178's push (run 31332846716); `cuda-fat-build` in that run was cancelled by
concurrency, not failing.

THE GATE REDDENED MAIN FOR OBEYING IT. `check-role-discipline.py` judges every
commit in the `before..after` range on its own message. A PR landed with a REAL
merge commit pushes the merge AND the branch commits under it: the merge names
the PR, the branch commits were never required to, so each one read as a direct
push. `6603356a` (#178), `e73cbbae` (#204) and `1a02ab4f` (#196) all failed this
way, in both `documentation-checkpoint` and `agent-record`. Arrival is now judged
ONCE, on the commit that lands the change: `merged_pr_content` exempts what a
row/* PR merge brings in. Squash-merges are untouched -- their one commit carries
"(#N)" and passes on its own message. NOT a weakening, and gated as such: only
the SIDE parents count, so `--not parents[0]` keeps a commit pushed straight to
main from being laundered by merging a PR on top, and a merge naming no row and
no PR exempts nothing. Four unit checks build real git history for those cases,
plus the exact `3bbee96e..0cf3dbb` range CI ran, pinned with `has_reached_main`
forced TRUE -- from a `row/*` worktree everything reports as pending PR
disposition, and the test would have passed against the defect it exists to
catch. Suite 47/47, and 5/5 red without the fix.

A STALE MESSAGE IN A TEST. `6603356a` taught `LoadMergedBf16RawNK` to accept
F8_E4M3 shards and rewrote its rejection to name the supported dtypes;
`test_qwen27_dense_forward.cpp:229` still asserted the old "expected BF16", so
`build-test-cpu` and both sanitizer legs failed on it. The expectation now reads
the message the loader raises, and the FP8 merge path that arrived WITHOUT a test
in this file gets one: a mixed BF16+FP8 merged parameter, expectations hand-
computed from E4M3 bytes and the scale (never re-derived through the same
dequant helper the loader calls), plus the per-channel-scale rejection.
7/7, 333 assertions.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Claude:claude-opus-5 [ClaudeCode]

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants