Skip to content

update aarch64 libraries to release/0.5.0 branch#7

Merged
juney-nvidia merged 1 commit into
release/0.5.0from
xiaoweis/update-release-0.5.0
Oct 18, 2023
Merged

update aarch64 libraries to release/0.5.0 branch#7
juney-nvidia merged 1 commit into
release/0.5.0from
xiaoweis/update-release-0.5.0

Conversation

@Shixiaowei02

Copy link
Copy Markdown
Collaborator

No description provided.

@juney-nvidia
juney-nvidia merged this pull request into release/0.5.0 Oct 18, 2023
@sjtu-cz sjtu-cz mentioned this pull request Nov 7, 2023
@poweiw poweiw added invalid and removed invalid labels Jun 3, 2025
greg-kwasniewski1 pushed a commit to greg-kwasniewski1/TensorRT-LLM that referenced this pull request Jun 10, 2025
…DIA#7)

* example of inductor pattern matcher for RoPE with explicit cos/sin matcher

Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com>

* move to utils

Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com>

* add usage of scalar_workaround, support op_ignore_type

Signed-off-by: Ubuntu <201670829+Fridah-nv@users.noreply.github.com>

* minor

Signed-off-by: Ubuntu <201670829+Fridah-nv@users.noreply.github.com>

* update all 3 types of RoPE matcher to use inductor pattern matcher

Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com>

* address feedback and refine code/doc

Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com>

* minor

Signed-off-by: Ubuntu <201670829+Fridah-nv@users.noreply.github.com>

* fix 2e2 for llama4 and ds rope, remove legalize_graph in canonicalize_graph, update ds rope impl to match with the exported graph

Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com>

* deprecate previous rope matcher

Signed-off-by: Ubuntu <201670829+Fridah-nv@users.noreply.github.com>

---------

Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com>
Signed-off-by: Ubuntu <201670829+Fridah-nv@users.noreply.github.com>
danielafrimi added a commit to danielafrimi/TensorRT-LLM that referenced this pull request Jun 30, 2025
# This is the 1st commit message:

kernel

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

wip

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

remove prints

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

test pass

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

test refactor with more use cases

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

refacor

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

refacor_2

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

add tuner wip

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

autotuner works

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

bfloat16 works. moer changes to the thop file

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

is tune for autotuner is True --> gets real tactics configs

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

wip

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

wip

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

zeros + quant mode is works

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

act int8

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

removed fp8 for now

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

wip

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

w4a16 linear module

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

wip

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

changed cutalss for sm==89

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

wip

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

test linear work

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

add license

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

works!

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

refactor + linear test pass

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

preprocess in load weights

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

wip

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

wip

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

wip

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

wip

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

refactor + rebase

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

wip

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

wip

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

Blackwell not supported

Signed-off-by: Daniel Afrimi <dafrimi@nvidia.com>

wip

Signed-off-by: Daniel Afrimi <dafrimi@nvidia.com>

skip blackwell

Signed-off-by: Daniel Afrimi <dafrimi@nvidia.com>

wip

Signed-off-by: Daniel Afrimi <dafrimi@nvidia.com>

works

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

# This is the commit message NVIDIA#2:

rebased

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

# This is the commit message NVIDIA#3:

align with my pld worked version of linear

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

# This is the commit message NVIDIA#4:

wip

Signed-off-by: Ubuntu <dafrimi@nvidia.com>

# This is the commit message NVIDIA#5:

refactor

Signed-off-by: Daniel Afrimi <danielafrimi8@gmail.com>

# This is the commit message NVIDIA#6:

refactor

Signed-off-by: Daniel Afrimi <danielafrimi8@gmail.com>

# This is the commit message NVIDIA#7:

refactor

Signed-off-by: Daniel Afrimi <danielafrimi8@gmail.com>

# This is the commit message NVIDIA#8:

refactor

Signed-off-by: Daniel Afrimi <danielafrimi8@gmail.com>

# This is the commit message NVIDIA#9:

sys path

Signed-off-by: Daniel Afrimi <danielafrimi8@gmail.com>

# This is the commit message NVIDIA#10:

sys path

Signed-off-by: Daniel Afrimi <danielafrimi8@gmail.com>
litaotju pushed a commit to litaotju/TensorRT-LLM that referenced this pull request Jul 19, 2025
Signed-off-by: Yuxian Qiu <142763828+yuxianq@users.noreply.github.com>
litaotju pushed a commit to litaotju/TensorRT-LLM that referenced this pull request Jul 24, 2025
Signed-off-by: Yuxian Qiu <142763828+yuxianq@users.noreply.github.com>
Signed-off-by: Fanrong Li <23290157+lfr-0531@users.noreply.github.com>
yuxianq added a commit to yuxianq/TensorRT-LLM that referenced this pull request Jul 28, 2025
Signed-off-by: Yuxian Qiu <142763828+yuxianq@users.noreply.github.com>
Signed-off-by: Fanrong Li <23290157+lfr-0531@users.noreply.github.com>
zongfeijing pushed a commit to zongfeijing/TensorRT-LLM that referenced this pull request Jul 31, 2025
Signed-off-by: Yuxian Qiu <142763828+yuxianq@users.noreply.github.com>
Signed-off-by: Fanrong Li <23290157+lfr-0531@users.noreply.github.com>
tongyuantongyu pushed a commit to tongyuantongyu/TensorRT-LLM that referenced this pull request Dec 5, 2025
[None][chore] Add accuracy test for allreduce strategy
karljang added a commit to karljang/TensorRT-LLM that referenced this pull request Feb 19, 2026
- Revert utils.py docstring changes (NVIDIA#1)
- Remove disable_deep_gemm flag, use module-level quant_config=None (NVIDIA#2)
- Replace Flux2SwiGLU with shared swiglu from _torch/modules (NVIDIA#3)
- Unify Flux2PosEmbed into FluxPosEmbed with parameterized axes_dim (NVIDIA#4)
- Replace Flux2FeedForward with shared GatedMLP (NVIDIA#6)
- Revert pipeline.py post_load_weights; add create_weights() in FLUX
  transformer __init__ to match WAN's __post_init__ pattern (NVIDIA#7)
- Refactor _get_quant_algo_for_layer to encapsulate hasattr logic (NVIDIA#9)
- Remove partial docstrings that just repeat class names (NVIDIA#10)
- Simplify bias flags to direct booleans (NVIDIA#11)
- Unify apply_rotary_emb: delete FLUX copy, FluxPosEmbed outputs
  [1,S,1,D] matching WAN format so shared function unchanged (NVIDIA#12)
- Remove dynamic=True from maybe_compile (NVIDIA#13)
- Add @torch.compiler.disable on FluxPosEmbed.forward
- Tighten FLUX.2 PSNR test threshold from 12 to 20 dB

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Signed-off-by: Kanghwan Jang <861393+karljang@users.noreply.github.com>
chienchunhung added a commit to chienchunhung/TensorRT-LLM that referenced this pull request Apr 30, 2026
…DIA#7; reposition deadline as fallback

Restructure the report around a cleaner, more honest framing of the
final terminal wedge driver:

- Promote the NIXL UCX-internal pthread_mutex_lock deadlock (previously
  documented loosely as #7a / #7b in the run8 post-mortem) into a
  proper Signature NVIDIA#7 in the canonical signature list, with its own
  detailed section in Failure Signatures including the gdb stack
  evidence, mechanism, reproducer, and the explicit "this is NOT a
  TRT-LLM bug" framing.
- Retire #7a as a real bug (misnamed trace marker; counts match the
  signature #1 fix path firing correctly) and fold #7b's symptom into
  NVIDIA#7's mechanism description.
- Reposition the kv_transfer_timeout_ms deadline work (Next Steps
  item 7 + the Effort Estimate subsection) explicitly as a TRT-LLM-side
  fallback / mitigation for signature NVIDIA#7, NOT the ultimate fix. The
  ultimate fix is the NIXL/UCX root-cause bug (Next Steps item 8).
- Update the Executive Summary to say "six TRT-LLM bugs plus a seventh
  that lives one architectural layer below in NIXL/UCX" instead of
  "five distinct bugs plus a caveat".
- Add a row for NVIDIA#7 to the Signature ↔ PR Map: status "identified,
  classified, documented, not a TRT-LLM bug"; ultimate fix tracked
  via the NIXL/UCX bug; TRT-LLM-side fallback tracked via the
  deadline work.
- Update Phase 10 narrative and the Signature NVIDIA#6 status block to use
  the cleaner NVIDIA#7 framing.

Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Made-with: Cursor
chienchunhung added a commit to chienchunhung/TensorRT-LLM that referenced this pull request Apr 30, 2026
… PRs landed

Apply nine consistency fixes against the post-NVIDIA#13671/NVIDIA#13672/NVIDIA#13673/NVIDIA#13674
state of the investigation:

1. Front-matter Status block: replace the "sig NVIDIA#6 root-caused, validation in
   flight" wording with the post-run8 picture (all 6 TRT-LLM PRs in review;
   NVIDIA#7 is an out-of-scope NIXL bug; deadline work is the TRT-LLM-side
   fallback for NVIDIA#7).
2. Front-matter Branches in this worktree: add the four new sig #4 / #5 / NVIDIA#6
   branches.
3. Front-matter Related PRs: add NVIDIA#13674 / NVIDIA#13671 / NVIDIA#13672 / NVIDIA#13673 with
   chained-on-NVIDIA#13640 callout for NVIDIA#13673.
4. "Configurations that did not reproduce": #5 and NVIDIA#6 now do reproduce in
   single-process unit tests via the new tests added by NVIDIA#13672 and NVIDIA#13673;
   only #3 and NVIDIA#7 remain field-only.
5. Phase 6 close: the sig #4 regression test is no longer isolated in
   local/rc11-disagg-repro - it is now in the chained NVIDIA#13674 / NVIDIA#13671 pair.
6. Signature NVIDIA#6 section: drop the "(suspected)" qualifier and the "(most
   likely)" hedging on Where-it-lives - both are confirmed by run7 and run8.
   Rename the section header to describe the actual failure shape (recv
   buffer index leak via !isReady early return wedging
   assignBufferIndexForRecv) rather than the early control-path-stall
   hypothesis. Mirror the rename in the Signature - PR Map row.
7. File / Branch Index "New unit tests": add the new sig #5
   (test_cancel_queued_gen_request_fulfills_receiver_future) and sig NVIDIA#6
   (test_cancelled_after_ready_does_not_leak_recv_buffer_index, NIXL
   backend) tests.
8. Signature #3 status hypothesis: add a one-paragraph note that NIXL
   (signature NVIDIA#7) is now also a candidate cause of the half-initialized
   state, so a future field hit is not misattributed to a fresh TRT-LLM
   bug.
9. Phase 5 narrative: add a forward link explaining that the underlying
   terminal driver of the Phase-5 wedge was already NVIDIA#7 (NIXL), but #4 was
   the visible TRT-LLM-side symptom because the gen event loop was
   self-blocking before any of the later layers could surface.

Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Made-with: Cursor
chienchunhung added a commit to chienchunhung/TensorRT-LLM that referenced this pull request Apr 30, 2026
… section

Fold two related pieces of analysis into the report as a new section
between the Investigation Timeline and "Why the Existing Tests Did
Not Catch This":

(1) Signature taxonomy refining the naive "burst -> timeout ->
cancellation -> bug" framing. Four-of-seven signatures (#1, #3, #5, NVIDIA#6)
are direct cancellation-handling bugs; #4 is a structural latent
blocking bug that cancellations expose; #2 is an eviction-driven bug
that burst traffic exposes via memory pressure; NVIDIA#7 is a NIXL-internal
contention bug that the same load shape happens to trigger but which
is not strictly a cancellation bug. Includes a refined trigger chain
diagram and two precise corrections (burst alone is not the trigger;
"cancellation" is one of several entry points to cleanup paths).

(2) Cascade map distinguishing two kinds of inter-signature tangling:
- Type 1 (a fix produces a new signature): only one case, the #1 fix
  produces NVIDIA#6 by making the receiver-side !isReady early-return path
  reachable in production where a latent recv-buffer leak existed.
  This is why NVIDIA#6 PR (NVIDIA#13673) is explicitly chained on #1 fix PR
  (NVIDIA#13640).
- Type 2 (a fix exposes a pre-existing signature): three cases where
  #4 fix exposes #5 / NVIDIA#6 and NVIDIA#6 fix exposes NVIDIA#7 because the upstream
  fix removes the masking effect on the downstream bug. These are not
  regressions of the fixes; they were latent pre-existing issues.
- Subtler third relationship: #4 is structurally a defensive catcher
  for any upstream bug that produces a never-resolving receiver
  future. The #4 fix is independently valuable as defence in depth,
  not just a symptomatic patch.

Also includes a fix-to-file mapping showing that the fixes do not
overlap in code; the only structural dependency is the NVIDIA#6 -> #1 chain
enforced by the PR base.

Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Made-with: Cursor
chienchunhung added a commit to chienchunhung/TensorRT-LLM that referenced this pull request May 1, 2026
NVIDIA#7 confirmed independent of TRT-LLM-side fix strategy

Add Phase 11 to the Investigation Timeline documenting the
pr13056_run1 experiment: ran the same 1P1D long-prompt burst harness
against an independent fix stack (a comprehensive end-to-end
shared_ptr<LlmRequest> + BufferIndexHolder RAII + deadline-enforcement
refactor) on the same rc11 base. Same outcome as run8: NO RECOVERY
after 180s idle, with a gdb-confirmed pthread_mutex_lock frame in
CacheSender::Impl::response on the ctx worker, alongside the same
NIXL plugin threads in the same process.

Coverage map shows the two stacks converge on the same TRT-LLM-side
bugs (#1, #4, #5, NVIDIA#6) via different mechanisms (surgical patches vs
comprehensive refactor), eliminating two alternative hypotheses:
- our chained PRs introduced a regression that masquerades as NVIDIA#7
- comprehensive deadline enforcement alone clears the field reproducer

Both refuted by the experiment. The independent stack's defensive
diagnostics ([buf] CANCEL: 0, [buf] STILL_WAITING: 0, kNETWORK_ERROR:
0, broken-promise: 0, deadline-driven failures: 0) confirm the wedge
is below where any TRT-LLM-side deadline can reach.

Update Next Steps item 8 to reference the pr13056_run1 stack dump as
a second independent reproducer for the NIXL/UCX bug filing — much
stronger evidence than run8 alone because it shows the deadlock is
independent of any TRT-LLM fix strategy.

Update Phase 10 title to remove the (current) qualifier since
Phase 11 is now the latest phase.

Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Made-with: Cursor
chienchunhung added a commit to chienchunhung/TensorRT-LLM that referenced this pull request May 1, 2026
…as a TRT-LLM mCondMutex bug, not a NIXL plugin bug

Rerun the same 1P1D long-prompt burst harness on TRT-LLM's direct UCX
backend (the path that bypasses the NIXL plugin entirely) and capture
the gdb evidence. The wedge fires identically with libnixl.so not
loaded, in a process where the only transport library is
libtensorrt_llm_ucx_wrapper.so. Same dataTransResp thread, same
CacheSender::Impl::response() frame, same pthread_mutex_lock at the
same point in the loop.

This falsifies the Phase 10/11 framing of sig NVIDIA#7 as a NIXL UCX-plugin
internal mutex deadlock. The mutex is in TRT-LLM-owned code (most
plausibly mCondMutex at the top of CacheSender::Impl::response()'s
loop, line 684 of dataTransceiver.cpp), exposed by both NIXL and
direct UCX backends. The bug is fixable in TRT-LLM.

Cover three sub-experiments in Phase 12:
- pr13056 + UCX direct: aborts at first request with HTTP 500
  (unrelated UCX-direct path regression in the comprehensive refactor)
- pr13056 + UCX direct + UCX_TLS=all: same fast HTTP 500
  (rules out TLS configuration)
- rc11+chained-PRs + UCX direct: same silent NO RECOVERY wedge as
  run8 and pr13056_run1 — gdb shows pthread_mutex_lock inside
  CacheSender::Impl::response() with no NIXL frames anywhere

Update the report to reflect the broader sig NVIDIA#7 framing throughout:
- Rewrite the Failure Signatures section for NVIDIA#7 (cross-variant
  analysis, mutex source-code inference, ctx-mpi4py-exit variant)
- Add Phase 12 to the timeline with all three sub-experiments
- Update the Executive Summary signature table row and caveat
- Reorder Next Steps: surgical mutex fix is now item 7
  (top priority), deadline enforcement is item 8 (now both directly
  addresses NVIDIA#7 AND defence-in-depth), NIXL/UCX bug filing is item 9
  (secondary), rename trace marker is item 10
- Update the Effort Estimate to reflect the dual role of deadline
  enforcement (direct fix for NVIDIA#7 + defence-in-depth)
- Update cross-references (Next Steps item N renumbering) throughout

Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Made-with: Cursor
chienchunhung added a commit to chienchunhung/TensorRT-LLM that referenced this pull request May 1, 2026
…ifestations broaden sig NVIDIA#7 to a CacheSender::Impl::* bug class

run9 (rc11 + our fixes + UCX) and run10 (PR NVIDIA#13056 + UCX) both used a
gdb capture loop on the dataTransResp thread to pin down the exact mutex
behind sig NVIDIA#7. The wedge changed character in both runs:

- run9 ctx mpi worker SIGSEGVs at iter 92 of the burst inside
  _PyObject_GenericGetAttrWithDict (Python C-API), downstream of the
  sig #1 fix path firing cleanly.
- run10 ctx mpi worker SIGSEGVs synchronously inside
  CacheSender::Impl::handleAsyncSend(AsyncSendResource&) on the very
  first sanity-probe request - no concurrency, no burst, no
  cancellation. Zero UCX/NIXL frames in the stack; root cause is most
  plausibly a null shared_ptr<LlmRequest> deref at line 594 since
  PR NVIDIA#13056's ownership model is necessary but does not enforce
  Response::mRequest non-null at producer or consumer.

Both findings are cleanly TRT-LLM-internal and broaden sig NVIDIA#7 from a
single pthread_mutex_lock deadlock to a class of CacheSender::Impl::*
bugs with at least four observed manifestations across two transports
and three fix bundles.

Touched sections: Status block, caveat block, signature table (sig NVIDIA#7
row), sig NVIDIA#7 section variants/fix/status, Phase 12 cross-reference,
new Phase 13 section, Next Steps items 7 / 7a / 7b.

Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Made-with: Cursor
chienchunhung added a commit to chienchunhung/TensorRT-LLM that referenced this pull request May 3, 2026
…le structure

Split the 3193-line single-file report into 13 reader-friendly files
under docs/investigations/nvbug-6104831-disagg-permanent-wedge/:

- README.md: top-level orientation, status, navigation, suggested
  reading paths.
- 01-background.md: architecture diagrams, request lifecycle
  walkthrough, state machine, cancellation flow.
- 02-failure-signatures.md: the seven failure signatures (#1-NVIDIA#7).
- 03-defect-class-stack.md: NEW - the L1-L8 defect-class layering
  that frames the four-approach comparison.
- 04-reproduction.md: how to reproduce locally, load shape, run
  archive index.
- 05-investigation-timeline.md: chronological story (Phases 0-14).
- 06-fix-approaches/: side-by-side comparison plus one file per
  approach (A chained, B PR NVIDIA#13056, C PR NVIDIA#13495, D combo).
- 07-architectural-reflections.md: seven invariants + retrospective.
- 08-next-steps-and-pr-map.md: PR map, outstanding work, deadline
  effort estimate, run archive index.

The L1-L8 defect class stack is the new framework introduced here.
It re-frames the seven signatures as the visible faces of eight
underlying invariant gaps and is used to explain why the combo
approach (D) is the only stack that recovers cleanly across the
test matrix. The five Mermaid diagrams from the request walkthrough
move into 01-background.md as the prerequisite reading for the rest
of the investigation.

The single-file form was hard to navigate at this size; the split
follows the user's reading-path needs (cold reader, single-PR
reviewer, fix-path picker, retrospective reader) called out in the
README.

Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
chienchunhung added a commit to chienchunhung/TensorRT-LLM that referenced this pull request May 3, 2026
… colored fix components

Two improvements to 00-tldr.md based on review feedback:

1. Add a 7-row table after the architecture section that gives a
   one-line "where it lives" + "symptom" for each signature #1-NVIDIA#7.
   The TL;DR previously mentioned the seven signatures by number
   without describing them, forcing the reader to jump to
   02-failure-signatures.md to make sense of references like
   "sig #4" or "sig NVIDIA#7".

2. Color-code the four fix components in the combo diagram:
   - PR NVIDIA#13056: blue (lifetime + cancel-flag + RAII)
   - PR NVIDIA#13495: orange (NIXL TransferStatus::release)
   - eval-order fix: green
   - Python idempotency guards: purple

   Each L1-L8 layer node is tinted with the lighter shade of
   whichever fix closes it, so the "closes" arrows are reinforced
   by colour-matching. Makes the fix-to-layer mapping visually
   scannable at a glance.

Note added underneath the diagram explaining the colour scheme.

Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
chienchunhung added a commit to chienchunhung/TensorRT-LLM that referenced this pull request May 20, 2026
…view

Apply 5 review-driven edits to §16 to address residual concerns:

Walk ordering (concern NVIDIA#8): change GMS RO and MX-receiver alias walks
from per-module to top-level model.setup_aliases(). Matches the §7
mitigation contract from ai-dynamo/dynamo PR NVIDIA#7053 ("Call
model.post_load_weights() (top-level only) before
materialize_module_from_gms()"). transform_weights() and
cache_derived_state() walks remain per-module since those bodies live
on submodules. New "Why setup_aliases() is top-level-only" callout
documents the asymmetry.

Lifecycle of _weights_transformed (concern #4): new subsection
specifying explicit set/reset/orthogonality rules. Includes a 2x2
truth table showing _weights_removed and _weights_transformed track
different lifecycles and can take any combination. Reset is the
orchestrator's responsibility (e.g., ModelLoader.reload() resets the
flag before re-binding tensors); subclasses do not manage reset.

Hard preconditions (concern #5): promote MX source-identity
completeness from "open question" to "hard precondition P1." Lists
transform-affecting parameters that MX identity must cover
(attn_backend, quant backend list, FP8/NVFP4 fusion strategy, TP/EP
layout, model revision). Specifies an in-tree backend-fingerprint
fail-safe as the fallback if upstream MX cannot guarantee
completeness. P2 documents that orchestrator owns _weights_transformed
reset. Removes redundant open question NVIDIA#6 from the table.

Cosmetic fixes (concerns #2, NVIDIA#6): "four stages" -> "three per-module
stages plus orchestrator-managed per-process finalization."
cache_derived_state description softened to "reserved for
data-dependent state where it exists; many existing modules will have
empty bodies."

Scope clarifications (concerns #1, #3, NVIDIA#7):
- Tiny PR scope reframed as "duck-typed helpers, not inheritance"
  with citations to existing model_loader.py walker pattern. Lists
  4 walker helpers (_setup_aliases, _walk_transform, _walk_cache_state,
  _walk_full_post_load).
- Migration callout: when migrating a subclass, the old
  post_load_weights() override MUST be removed; otherwise the new
  staged calls silently no-op.
- Family PR #2 (Linear/Attention) gains a "quant-method callback
  decision" note with default = keep quant_method.post_load_weights
  callback name (no rename).

No code changes. Drives Tiny prep PR scope and family-PR migration
sequence. References:
- TRT-LLM PR NVIDIA#13926 (GMS-only)
- TRT-LLM PR NVIDIA#14151 (MX shim refactor)
- ai-dynamo/dynamo PR NVIDIA#7053 (upstream GMS prototype)

Signed-off-by: Chien-Chun Hung <chienchunh@nvidia.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
SpencerGarnets added a commit to ai-blaise/TensorRT-LLM that referenced this pull request Jun 2, 2026
Full {2,4,8 GPU} x {16,32 conc} x {7 context} matrix, both pure NVFP4, both
production-optimized, matched per-rank work. Honest peak 1.65x (G=8 conc32,
a2a exposed); range 1.05-1.65x by a2a overlap. Below 1.84x against the strong
fused TRTLLMGen native baseline.

Decomposition (req NVIDIA#7 baseline-optimality): at G=8 the all-to-all is 30% of EP
total (35-37us of 112-122us); local compute 76-82us. WarpDecode eliminates the
a2a + fuses the rest. Matched-shape local-only fusion win is 1.03-1.26x
(native runner already heavily fused).

Documents the legitimate path to 1.84-2.5x: the output-owned sub-128 kernel
(current grouped kernels locked at tile 128/256, padding 4-64x at decode
row counts; a no-pad output-owned kernel cuts local 74us->~35us = ~3.5x).
liyuhannnnn added a commit to liyuhannnnn/TensorRT-LLM that referenced this pull request Jun 8, 2026
…ement device tensor

The dense blockscaled GEMM kernel expects alpha as a device Tensor (it reads
alpha[0]); the run script passed a Python float 1.0, which fails compilation
with "argument NVIDIA#7 (alpha) expects Tensor, got float". Build a single-element
float32 device tensor via from_dlpack(torch.ones((1,), ...).cuda())
.mark_layout_dynamic() and use it at all three call sites (compile, reference
check, and the JitArguments benchmark closure).

Signed-off-by: Yuhan Li <51736452+liyuhannnnn@users.noreply.github.com>
liyuhannnnn added a commit to liyuhannnnn/TensorRT-LLM that referenced this pull request Jun 8, 2026
…ement device tensor

The dense blockscaled GEMM kernel expects alpha as a device Tensor (it reads
alpha[0]); the run script passed a Python float 1.0, which fails compilation
with "argument NVIDIA#7 (alpha) expects Tensor, got float". Build a single-element
float32 device tensor via from_dlpack(torch.ones((1,), ...).cuda())
.mark_layout_dynamic() and use it at all three call sites (compile, reference
check, and the JitArguments benchmark closure).

Signed-off-by: Yuhan Li <51736452+liyuhannnnn@users.noreply.github.com>
SpencerGarnets added a commit to ai-blaise/TensorRT-LLM that referenced this pull request Jun 10, 2026
… DeepEP-LL + fp4out shipped; FC2 N-tile killed; eager c16 profile replaces the Indexer-dominant premise

Update the optimization-candidates plan + per-component docs for everything
shipped/learned since the doc was last touched (authoritative record: the
commit messages of 29f492b..fd705a6):

- SHIPPED: B1 cuBLASLt NVFP4 backends (4220bf4, ~1.62 ms/tok ~6% TPOT,
  bit-identical, + the kv_a_proj_with_mqa re-creation fix and the
  everywhere-extension in 3e03d66, +1.62 ms/tok incremental); I6 fp16
  indexer logits (config-matched, top-k -15..-22% at kv>=33k, recall at the
  fp8 noise floor, downstream cos 1.0); G1 gated-norm/glue fusions
  (fused_lowrank_gate -91.6% on the ~7 ms/step fp32 SGEMM+cast soup found
  hiding in the dense-proj bucket; HISA per-step invariant memo; fused
  sigmoid-mul) + 33e801f CuTe DSL lowrank-gate (-36% vs Triton, default
  impl); I7 prod decode top-k -> vanilla C++ (841f987, ~1.7x at prod live
  kv, width-override premise measured false); 51918fb DeepEP LL enablement
  (overlay guard + park/restore; 45 vs 115 us/layer at token_limit=16,
  INVERTS at 64); fd705a6 _FP4OUT_MIN_M lift (K3, exact vs TRUE-f32, OOB
  demo clean, ~100 us/step).

- KILLED with evidence: the FC2 N-tile lever (N=160 numerically broken --
  SFB miscompute, cos 0.790 vs a TRUE f32 reference at decode AND prefill;
  the original gate was a fused-vs-sequential self-comparison, blind by
  construction; 256 already optimal; megakernel FC2_N default fixed 160->256
  in e105fd7 and the validator de-blinded); MoE tuner mispick at bs=16
  (none); attention-tactic retune at TP bs=16 (optimal; TP-rank attention
  faster than ADP at equal load); NVFP4-quantizing the dense MLA proj (cos
  0.63-0.83); I4 score->topk fusion (net-zero under graphs). bf16 AB-swap
  stays killed.

- NEW/REVISED: M3 DeepEP LL production flip (env delta + the c16-vs-max_batch
  sizing decision); C2 KVarN delta-restore (impl opt-in; bs=16 pre-replay
  scan ~4 ms/step amortized at TP -- the biggest TP-regime cost; 5-scenario
  equivalence verification pending); G2 gated-norm -> PRE_MOE_FUSION (gate
  measured 0.500 +/- 0.0025, ~1-1.5 ms/step); I5 indexer wk+wp fused GEMM
  (build() gating vs production loader under investigation); S2 moot under
  the WarpDecode+TP production plan (held); SM1/SM3 schedule hoists already
  overlapped at c16 (correct but flat).

- Methodology: the eager c16 profile (MoE ~60% incl EP comm ~40%, dense proj
  19.8%, glue 13.5%, Indexer 4%, HISA 0.7%) replaces the stale 'Indexer
  50-74%' premise; DP4-vs-TP16 regime distinction recorded; production
  targets WarpDecode+TP. Production-target table refreshed (HISA gate scales
  via max_gen_kv_len, cuBLASLt backends, fp16 logits, CuTe lowrank gate,
  fp4out fusion on at decode).

Per-component: indexer.md (premise, NVIDIA#6 live-kv-only dispatch, new NVIDIA#7 fp16
logits, invariant memo), warpdecode.md (megakernel FC2 256 + validator
de-blind, DeepEP LL under WarpDecode, N=160 do-not-redo), kvarn.md (C1/C2
restore-cost section), nvfp4_fusions.md (#13b swiglu+FP4-out guard lift),
README.md (index rows + premise).

Signed-off-by: Spencer Garnets <210263209+SpencerGarnets@users.noreply.github.com>
KleinBlueC pushed a commit to KleinBlueC/TensorRT-LLM that referenced this pull request Jul 7, 2026
…E. All 8 acceptance criteria hold at runtime against...

QA verdict: APPROVE (weighted_score 8.6)
Change vs main: 7 files changed, 1376 insertions(+), 8 deletions(-)
Build iterations run: 13

Code-quality verification: clean after 2 round(s)
  - round 1 (REJECT):
      ROUND MIS-SCOPED — patch not actually reviewed. This round's /simplify, checklist, and /code-review executed with cwd = the agent-flow framework repo (branch kleinc/code_quality_improvement_v1) and reviewed the code-quality-stage feature itself, NOT this task's PR patch. The ChatGLM3 code is on branch agent-team/chatglm3-6b-bringup-v1 in the TensorRT-LLM snapshot repo (a different directory), so `/simplify`/`/code-review` (which diff the cwd's HEAD) never touched it. Root cause: the Quality agent runs in the framework repo's cwd, not the snapshot repo; another round will misfire identically unless the Quality agent runs with the TensorRT-LLM task branch as its working directory.

      I re-reviewed the ACTUAL patch (git diff main in TensorRT-LLM). Real, unaddressed findings:
      - TESTS (duplication): `_patch_chatglm3_hf_tied_compat()` duplicated in all 4 new test files; `_find_cuda_graph_runner()` and `_assert_cuda_graph_hard_path()` duplicated across test_chatglm3_gsm8k.py and test_chatglm3_replay.py; CHATGLM3_CKPT const, skip decorators, and RuntimeCfg class repeated. Consolidate to a shared util/conftest (keep all coverage).
      - modeling_chatglm.py (simplification): redundant post-normalization reads of attention_bias/mlp_bias/head_dim via `getattr(...) or ...` after normalize_chatglm_config already set them; normalize_chatglm_config returns config but both callers discard it; a few verbose comments.
      - ALTITUDE (larger, behavior-sensitive; flagged not applied): custom imperative load_weights vs framework BaseWeightMapper; ChatGLM-specific normalization special-cased in shared config_utils.py; partial_rotary_factor/rope_fusion hardcoded.

      No fixes applied: the round did not validly review this patch, and edits to modeling/test code cannot be validated against the GSM8K accuracy criterion in this environment (risking a criterion regression). Recommend fixing the Quality agent's working directory (must be the snapshot repo/task branch) before the next round, then applying the test-dedup and post-normalization simplifications above.
  - round 2 (REJECT):
      Round 2 finally assessed the ACTUAL task patch (git diff main...agent-team/chatglm3-6b-bringup-v1 in the TensorRT-LLM repo), unlike round 1 whose skill steps mis-scoped to the agent-flow framework repo (cwd). Note: the automated stage still mis-scopes — /simplify and /code-review diff the cwd (agent-flow), not this task's snapshot repo — so I targeted the correct repo manually.

      /code-review (correctness): CLEAN — returned []. No correctness bugs survive verification. The high-risk items are validated: partial-RoPE, QKV/GQA split sizes (32 heads / 2 KV / head_dim 128), and transformer.* weight-name mapping are exercised by the GSM8K-parity acceptance gate QA approved; the fused-residual path is completed by the final self.norm(hidden_states, residual); embedding matches the Qwen sibling and load_weights handles attention-DP (tp_size = 1 if enable_attention_dp else mapping.tp_size). Speculative "other ChatGLM variant" concerns (_set_missing fill-if-None, head_dim None if kv_channels absent, missing dense/output biases) are not reachable for chatglm3-6b and are guarded by strict _take/consumed accounting.

      /simplify (round 2): applied one behavior-preserving cleanup (dropped the unused `return config` from normalize_chatglm_config, annotation -> None); it was subsequently reverted in the working tree, so the patch is unchanged.

      OUTSTANDING (why REJECT): the checklist's test item is not satisfied — heavy test-helper duplication across the 4 new test files (`_patch_chatglm3_hf_tied_compat` in all 4; `_find_cuda_graph_runner`/`_assert_cuda_graph_hard_path`/`RuntimeCfg` across gsm8k+replay; repeated CHATGLM3_CKPT/skip decorators), ~100+ duplicated lines to consolidate into a shared util; plus minor post-normalization config-read redundancy in modeling_chatglm.py. NOT applied here because tensorrt_llm is not importable in this environment and there is no GSM8K/test rig, so a test-file move cannot be validated for collection/pass (risking the acceptance criterion), and the working-tree edit I did make was reverted.

      RECOMMENDATION: this cannot be resolved by another automated round as currently wired — fix the Quality agent's working directory so /simplify and /code-review target the snapshot repo/task branch, then apply the test-helper de-duplication with the test rig available (or the author applies it directly). Correctness and framework-conformance are clean; the residual is a non-behavioral maintainability cleanup.
  - round 1 (REJECT):
      Code-quality round 1 for the chatglm3-6b bringup patch (modeling_chatglm.py, config_utils.py, _torch/models/__init__.py + 4 new test files).

      /simplify applied (safe, behavior-preserving, verified against repo convention):
      - modeling_chatglm.py: consume the normalized config fields instead of re-deriving booleans inline — `bias=config.attention_bias`, `dense_bias=config.mlp_bias`, GatedMLP `bias=config.mlp_bias`, and `head_dim=config.head_dim` at both sites (attention + load_weights). Confirmed no base infra reads these fields and that the model already hard-depends on normalize_chatglm_config having run; matches Llama/Phi3/Nemotron/Parakeet.
      - Both integration tests: removed the unreachable BFS-over-object-graph fallback in `_find_cuda_graph_runner`, keeping the deterministic `_executor.engine.model_engine.cuda_graph_runner` chain (verified in source; the runner is at a fixed location in single-process mode and in a subprocess otherwise, where BFS also fails).

      Codex checklist pass: the working tree carried minor docstring/comment trims in the ChatGLM3 test files (removed helper docstrings on `_patch_chatglm3_hf_tied_compat`, `_pick_metric_key`, `greedy_generate`, and an inline comment); no behavioral change.

      /code-review (high) found 10 items. Fixed this turn: (NVIDIA#4) `_pick_metric_key` in test_chatglm3_gsm8k.py now raises loudly instead of silently gating on an arbitrary non-exact_match metric — strict hardening, no change to any real gsm8k run.

      OUTSTANDING (why not auto-fixed):
      - CORRECTNESS, latent/out-of-target: (NVIDIA#1) load_weights only handles chatglm3-6b's bias layout — it unconditionally fetches the QKV bias and never loads dense/MLP biases, so ChatGLM2/3 variants with add_qkv_bias=False or add_bias_linear=True would KeyError or trip the strict unconsumed-weight ValueError. (NVIDIA#2) normalize_chatglm_config never maps ChatGLM `rope_ratio` to `rope_theta` (get_hf_rope_theta reads config.rope_theta), so chatglm3-6b-32k/128k silently get theta=10000 and degrade at long context. Both are no-ops for base chatglm3-6b (add_qkv_bias=True/add_bias_linear=False/rope_ratio=1) so QA passed; fixing them adds untestable variant paths (no GPU/checkpoint here) and is scope creep on a base bringup — hand to the coder to fix+test or narrow the module's "ChatGLM2/3" claim.
      - TEST ROBUSTNESS: (NVIDIA#3) the CUDA-graph hard-path tests require TLLM_WORKER_USE_SINGLE_PROCESS=1 but only name it in an assert string and never set it; without it, ENABLED-cfg tests hard-fail on `assert runner is not None` and BASELINE silently skips its checks. Not auto-fixed because gating/skipping would weaken the author's intended hard-path assertion — needs a CI-env or fixture decision.
      - ALTITUDE/REUSE (larger refactors, unsafe without test execution): (NVIDIA#5) hand-rolled load_weights + consumed-set + ignorable-suffix lists reimplement BaseWeightMapper/register_mapper (cf. Glm4WeightLoader in modeling_glm.py); (NVIDIA#9) normalize special-cased in shared load_pretrained_config + duplicated in __init__ vs a ChatGLMConfigLoader; (NVIDIA#6) `_patch_chatglm3_hf_tied_compat` (4 copies) + RuntimeCfg/_build_llm/_find_cuda_graph_runner/_assert_cuda_graph_hard_path (2 copies) duplicated across two separate, GPU-gated test trees; (NVIDIA#10) _GreedyHFLM._model_generate hand-rolls HF greedy decode.
      - EFFICIENCY: (NVIDIA#7) test_chatglm3_source_activation_replay reloads the 6B HF model + rebuilds TRT model for both parametrized scenarios though only the trailing cuda_graph block differs (use the _HFRef cache pattern).
      - SCOPE CREEP: (NVIDIA#8) the diff reflows unrelated bart/minimaxm3/qwen imports (+ a mistral import in config_utils.py); left as-is because formatter ownership is ambiguous (line 33 is 89 chars vs isort/yapf's 80) and reverting risks fighting the pre-commit hook.

      Rejecting: a substantive edit was made this turn and several actionable findings remain (test env-var robustness is the most important; the two latent correctness gaps and the altitude/reuse refactors should be triaged by the coder).
  - round 2 (APPROVE):
      Code-quality round 2 for the chatglm3-6b bringup patch.

      /simplify (round 2) applied 3 trivial, behavior-neutral cosmetic cleanups in modeling_chatglm.py: inlined the single-use `head_dim` local in ChatGLMAttention (`head_dim=config.head_dim` passed directly to super()), inlined the single-use `head_dim` temp in normalize_chatglm_config, and dropped the noise `: set` annotation on `consumed`. Reuse/efficiency/altitude agents found nothing new safely-actionable and confirmed the two "small fold" hypotheses don't hold (base `skip_modules` matches model module names, not the checkpoint keys `_IGNORABLE_*` filters; the second normalize_chatglm_config call is load-bearing for the direct-ModelConfig test path) — the only remaining altitude/reuse items are large refactors.

      Codex checklist pass (round 2): no net changes — independent verification confirmed the only diffs since round 1 were the 4 quality edits above/below, all behavior-neutral.

      /code-review (high) re-surfaced the same stable finding set (the substantive code was unchanged; no coder commits landed between rounds). I acted on the most actionable one this turn:
      - FIXED (NVIDIA#3, test robustness): the CUDA-graph hard-path tests required TLLM_WORKER_USE_SINGLE_PROCESS=1 (their ENABLED config hard-asserts the in-process runner) but never set it, so they would spuriously fail/silently-skip when run with a checkpoint outside a harness that exports it. Added an autouse `monkeypatch.setenv` fixture to both integration files (test_chatglm3_gsm8k.py, test_chatglm3_replay.py). Verified the env is read at runtime (utils.py:380, os.environ.get inside a function) so the fixture takes effect before _build_llm; the fix aligns each test with its own documented requirement and cannot regress the already-validated path.
      - Previously fixed (round 1): _pick_metric_key now fails loudly instead of gating on an arbitrary metric; consumed bias/head_dim inline derivation replaced with normalized fields; BFS runner-finder fallback removed.

      RESIDUAL (documented follow-ups — not safely resolvable by the automated quality loop):
      - Out-of-target latent correctness: (NVIDIA#1) load_weights only handles the chatglm3-6b/chatglm2-6b bias layout (add_qkv_bias=True/add_bias_linear=False) — it unconditionally fetches the QKV bias and never loads dense/MLP biases, so ChatGLM variants with add_qkv_bias=False (KeyError) or add_bias_linear=True (unloaded bias + unconsumed-weight ValueError) fail; (NVIDIA#2) rope_ratio is never mapped to rope_theta, so long-context chatglm3-6b-32k/128k (rope_ratio>1) silently use theta=10000. Both are no-ops for the bringup target (base chatglm3-6b/chatglm2-6b, rope_ratio=1); fixing needs variant checkpoints to test and expands scope beyond the bringup — for the coder/a follow-up.
      - Larger altitude/reuse/efficiency refactors needing test-execution access: (NVIDIA#4/NVIDIA#7) adopt BaseWeightMapper/register_mapper for load_weights + a ChatGLMConfigLoader/config-subclass to remove the shared-runtime normalize special-case and its second call; (NVIDIA#5) consolidate the 4x/2x-duplicated test helpers across the two GPU-gated test trees; (NVIDIA#6) cache the 6B HF/TRT model in the parametrized attention test; (NVIDIA#8) the _GreedyHFLM hand-rolled reference decode.

      APPROVE rationale: base chatglm3-6b is fully correct (QA-approved, acceptance criteria met); two thorough rounds applied all safe in-scope polish culminating in the NVIDIA#3 test-robustness fix; the quality loop has converged (round 2 reproduced round 1's findings with no coder changes); and every remaining finding is either out-of-target latent correctness for untested ChatGLM variants or a larger refactor requiring test-execution access — none of which another automated /simplify→/code-review round can safely resolve. Residuals are surfaced above for human/coder triage.

QA report:
DECISION: APPROVE. All 8 acceptance criteria hold at runtime against the current committed code (HEAD 2323424, iter11). Reference is the checkpoint's own trust_remote_code modeling_chatglm.py — an independent HF reference sharing zero code with the TRT-LLM path.

PER-CRITERION (runtime evidence):
1. Bootstrap chain — PASS. slurm/core_gpu_tests.4147161.log: "Built target build_wheel_targets", "Successfully installed tensorrt_llm-1.3.0rc21 transformers-5.5.4" (pip install -e .[devel]), "Claude Code successfully installed!", "claude-agent-sdk 0.2.93"; all inside srun pyxis container with --container-mounts=<repo>:<repo> --container-workdir=<repo> under set -eo pipefail (tests ran only because bootstrap rc=0).
2. config_registration_and_weight_accounting — PASS (TIER1 rc=0). Real ckpt loads unedited, arch resolves to ChatGLMForCausalLM, TRTLLM backend + KVCacheManagerV2, strict state-dict accounting (raises on missing/unexpected/shape-mismatch; only rotary_pos_emb.inv_freq tolerated).
3. source_activation_replay — PASS both graph modes. layer0 cosine=1.000000 mean_abs=0.00000, layer27 cosine=0.999999; decode graph-vs-eager max_abs=0.0 cosine=1.0. Companion partial_rope_boundary_and_theta PASS (dims [0:64] rotate, [64:128] pass-through, theta=10000^(-i/32)).
4. source_logit_replay — PASS both configs. trt_tok==hf_tok (5231/30910/4802), cosine~0.99999; enabled num_captured_graphs=8, baseline 0.
5. generation_parity — PASS both configs. 5 prompts x 32 steps, token_mismatches=0, min per-step cosine~0.99998.
6. llm_api_smoke — PASS both configs (nonempty deterministic; V2 + TRTLLM asserted; enabled enabled=True graphs=8, baseline graphs=0).
7. gsm8k accuracy_canary — PASS (slurm/gsm8k_gate.4145862.log): HF=40.00, baseline=40.00 gap 0.00, enabled=40.00 gap 0.00; enabled num_captured_graphs=32.
8. gsm8k full_trtllm_eval — PASS: HF=52.77, baseline=52.99 gap 0.23, enabled=53.07 gap 0.30 (both < 2.0 tol), enabled captured 32 CUDA graphs; RESULT: PASS. Absolute ~53% matches published ChatGLM3-6B GSM8K, ruling out a shared-bug artifact.

PROVENANCE: Core log 4147161 started 05:55, after the 05:33 HEAD commit -> definitively iter11 (criteria 1-6). git diff iter10->iter11 (2d58cc7..2323424) touched ONLY the two test files (additive CUDA-graph hard-path assertions); tensorrt_llm/ is byte-identical. GSM8K log 4145862 prints the iter11-only [cuda_graph_hard_path] output -> it ran the current test file against the unchanged runtime code (live bind-mount imported iter11 mid-bootstrap). Criteria 7-8 are therefore verified against current code.

INDEPENDENT RERUN THIS TURN: I resubmitted both jobs. Core (4149172) matched the prior post-commit run and was cancelled as redundant. My GSM8K resubmit (4149173) FAILED in ~1.5 min inside build_wheel.py (FileNotFoundError: inherit_graph_6.md5 in copy_resolving_symlink) — a concurrent-build race from my own simultaneous submission of two bootstrap-heavy jobs on one live repo mount, NOT a ChatGLM defect. Accepted evidence is the sequential post-commit runs + full code/provenance review.

RED-TEAM (all addressed): independent HF reference (no shared helpers); partial GPT-J RoPE covered analytically + behaviorally (layer0 exact); CUDA-graph silent fallback ruled out by in-process CUDAGraphRunner introspection (runner.enabled + >=1 torch.cuda.CUDAGraph; baseline 0); KV-cache/mask/scale drift ruled out by 32-step parity (0 mismatches) + GSM8K absolute-score match; transformers 5.x bridges (all_tied_weights_keys={}, max_length, num_hidden_layers) are loader-compat only and do not alter ChatGLM forward semantics.

EVALUATION SCORES:
- Functionality (2.0): 9 — every criterion passes with exact argmax parity, cosine >0.9999, GSM8K within 0.3 pts of HF on both configs.
- Code Quality (1.0): 8 — clean, well-documented 373-line model reusing existing modules; narrow config hook; strict weight accounting; minor incidental import reformatting in __init__.py; a pre-existing (non-diff) ruff E501 at __init__.py:130.
- Performance (1.5): 8 — parity-first bring-up (no perf gate in task.yaml); production TRTLLM backend + CUDA graph + overlap scheduler + fused QKV/gate-up + KVCacheManagerV2 all working.
- Completeness (1.0): 9 — complete, buildable, no placeholders, full state-dict accounting, both runtime configs.
- Technical Sophistication (1.0): 9 — correct unfused partial GPT-J RoPE with TRTLLM backend + CUDA-graph capture, compact MQA KV, config/transformers bridges.
Weighted = (9x2.0 + 8x1.0 + 8x1.5 + 9x1.0 + 9x1.0)/6.5 = 56/6.5 = 8.6.

STRENGTHS: genuine independent HF reference; real CUDA-graph hard-path proof (not just config flags); strict weight accounting; large GSM8K margin; minimal composable design (no new schemas). WEAKNESSES: multi-GPU/perf/chunked-prefill deferred (out of scope per task.yaml); harness must run the two jobs sequentially to avoid the concurrent-build race; pre-existing E501 worth a one-line fix. RECOMMENDATION: APPROVE — no code defect identified; task.yaml completion_criteria (GSM8K within 2 pts with and without CUDA graph + overlap scheduler) is met on both configs.</summary>
</invoke>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants