Mirror of linear changes - #2
Open
kwyss-nvidia wants to merge 56 commits into
Open
Conversation
13 tasks
kwyss-nvidia
force-pushed
the
kwyss/subchannel_recipe_linear
branch
2 times, most recently
from
March 12, 2025 21:39
07d55ea to
8bb7d63
Compare
kwyss-nvidia
force-pushed
the
kwyss/cublas_gemm_github_mr
branch
from
March 12, 2025 21:42
c3eebe7 to
b848509
Compare
kwyss-nvidia
force-pushed
the
kwyss/subchannel_recipe_linear
branch
from
March 13, 2025 00:06
8bb7d63 to
365a4d9
Compare
kwyss-nvidia
force-pushed
the
kwyss/cublas_gemm_github_mr
branch
from
March 13, 2025 00:07
b848509 to
1058efc
Compare
kwyss-nvidia
force-pushed
the
kwyss/subchannel_recipe_linear
branch
3 times, most recently
from
March 15, 2025 00:22
6c70366 to
08aa4de
Compare
kwyss-nvidia
force-pushed
the
kwyss/cublas_gemm_github_mr
branch
2 times, most recently
from
March 17, 2025 17:24
eee37bf to
ce4ca80
Compare
kwyss-nvidia
force-pushed
the
kwyss/subchannel_recipe_linear
branch
2 times, most recently
from
March 17, 2025 17:33
51fbe41 to
78c194d
Compare
kwyss-nvidia
force-pushed
the
kwyss/cublas_gemm_github_mr
branch
from
March 19, 2025 22:42
ce4ca80 to
5ebc93a
Compare
kwyss-nvidia
force-pushed
the
kwyss/subchannel_recipe_linear
branch
from
March 19, 2025 22:43
78c194d to
8f4f0f0
Compare
kwyss-nvidia
force-pushed
the
kwyss/subchannel_recipe_linear
branch
2 times, most recently
from
April 1, 2025 19:43
1d112ac to
48648a9
Compare
kwyss-nvidia
force-pushed
the
kwyss/cublas_gemm_github_mr
branch
from
April 1, 2025 19:45
5aa279e to
8466c36
Compare
kwyss-nvidia
force-pushed
the
kwyss/subchannel_recipe_linear
branch
from
April 1, 2025 21:46
ca005ab to
e35f2b6
Compare
kwyss-nvidia
force-pushed
the
kwyss/cublas_gemm_github_mr
branch
from
April 1, 2025 21:48
8466c36 to
e788ca2
Compare
kwyss-nvidia
force-pushed
the
kwyss/subchannel_recipe_linear
branch
2 times, most recently
from
April 1, 2025 23:23
22828fe to
413331d
Compare
kwyss-nvidia
force-pushed
the
kwyss/cublas_gemm_github_mr
branch
from
April 1, 2025 23:23
e788ca2 to
9ac89ea
Compare
kwyss-nvidia
force-pushed
the
kwyss/subchannel_recipe_linear
branch
4 times, most recently
from
April 2, 2025 18:52
db5b49e to
8d59b0a
Compare
kwyss-nvidia
force-pushed
the
kwyss/cublas_gemm_github_mr
branch
from
April 2, 2025 18:53
9ac89ea to
fa019d5
Compare
kwyss-nvidia
force-pushed
the
kwyss/subchannel_recipe_linear
branch
from
April 2, 2025 19:19
8d59b0a to
3424dc7
Compare
kwyss-nvidia
force-pushed
the
kwyss/cublas_gemm_github_mr
branch
from
April 2, 2025 19:20
fa019d5 to
cd3e414
Compare
Signed-off-by: Keith Wyss <kwyss@nvidia.com>
Signed-off-by: Keith Wyss <kwyss@nvidia.com>
* Use dummy wgrads for lower memory consumption Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com> Signed-off-by: Vasudevan Rengasamy <vrengasamy@nvidia.com> * Bug fix to avoid sharing gradients. Signed-off-by: Vasudevan Rengasamy <vrengasamy@nvidia.com> * Disable automatic use of batch_p2p_comm for CP2 Signed-off-by: Vasudevan Rengasamy <vrengasamy@nvidia.com> * Change weight to origin_weight for LN_LINEAR Signed-off-by: Vasudevan Rengasamy <vrengasamy@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci Signed-off-by: Vasudevan Rengasamy <vrengasamy@nvidia.com> --------- Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com> Signed-off-by: Vasudevan Rengasamy <vrengasamy@nvidia.com> Co-authored-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
Signed-off-by: zhongboz <zhongboz@nvidia.com>
Signed-off-by: Keith Wyss <kwyss@nvidia.com>
Signed-off-by: Keith Wyss <kwyss@nvidia.com>
Signed-off-by: Keith Wyss <kwyss@nvidia.com>
Signed-off-by: Keith Wyss <kwyss@nvidia.com>
* Minor stylistic tweaks and typo fixes Review suggestions from @ptrendx Signed-off-by: Tim Moon <tmoon@nvidia.com> * Fix incorrect col strides for MXFP8 matrices Signed-off-by: Tim Moon <tmoon@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Signed-off-by: Tim Moon <tmoon@nvidia.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Signed-off-by: Keith Wyss <kwyss@nvidia.com>
Signed-off-by: Keith Wyss <kwyss@nvidia.com>
Signed-off-by: Keith Wyss <kwyss@nvidia.com>
Signed-off-by: Keith Wyss <kwyss@nvidia.com>
Signed-off-by: Keith Wyss <kwyss@nvidia.com>
kwyss-nvidia
force-pushed
the
kwyss/subchannel_recipe_linear
branch
from
April 8, 2025 23:35
d7775fc to
b62d555
Compare
Apply MR comment change. Co-authored-by: Tim Moon <4406448+timmoon10@users.noreply.github.com> Signed-off-by: kwyss-nvidia <kwyss@nvidia.com>
kwyss-nvidia
force-pushed
the
kwyss/subchannel_recipe_linear
branch
from
April 9, 2025 00:05
8fc753d to
67e790b
Compare
Signed-off-by: Keith Wyss <kwyss@nvidia.com>
kwyss-nvidia
force-pushed
the
kwyss/subchannel_recipe_linear
branch
from
April 9, 2025 01:32
6948759 to
ea9e46b
Compare
Signed-off-by: Keith Wyss <kwyss@nvidia.com>
Signed-off-by: Keith Wyss <kwyss@nvidia.com>
* scaling enum abstract * rm NVTE_ from ScalingMode names * rework scaling mode enum in grouped gemm * fix norm sharding --------- Signed-off-by: Phuong Nguyen <phuonguyen@nvidia.com>
…r op backward (NVIDIA#1646) Explicitly specify quantized tensor usages needed for linear op backward Signed-off-by: Tim Moon <tmoon@nvidia.com>
* Debug checkpointing with te.Sequential Signed-off-by: Tim Moon <tmoon@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Signed-off-by: Tim Moon <tmoon@nvidia.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Signed-off-by: Tim Moon <tmoon@nvidia.com>
Signed-off-by: Keith Wyss <kwyss@nvidia.com>
Signed-off-by: Xin Yao <yaox12@outlook.com>
Signed-off-by: Xin Yao <yaox12@outlook.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A more reviewable mirror of the changes from NVIDIA#1559