Skip to content

fix: support ITensor amex in quantize and add export_torch_mode() wrapper - #4431

Merged
zewenli98 merged 4 commits into
mainfrom
evanli/support_ITensor_amex_in_quantize
Aug 5, 2026
Merged

fix: support ITensor amex in quantize and add export_torch_mode() wrapper#4431
zewenli98 merged 4 commits into
mainfrom
evanli/support_ITensor_amex_in_quantize

Conversation

@zewenli98

@zewenli98 zewenli98 commented Jul 24, 2026

Copy link
Copy Markdown
Collaborator

Description

  1. support ITensor amex in quantize
  2. add export_torch_mode() wrapper in torch_tensorrt.compile to modelopt quant models

Fixes #4429 #2976

Type of change

Please delete options that are not relevant and/or add your own.

  • Bug fix (non-breaking change which fixes an issue)

Checklist:

  • My code follows the style guidelines of this project (You can use the linters)
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas and hacks
  • I have made corresponding changes to the documentation
  • I have added tests to verify my fix or my feature
  • New and existing unit tests pass locally with my changes
  • I have added the relevant labels to my PR in so that relevant reviewers are notified

@zewenli98
zewenli98 requested a review from narendasan July 24, 2026 01:44
@zewenli98 zewenli98 self-assigned this Jul 24, 2026
@meta-cla meta-cla Bot added the cla signed label Jul 24, 2026
@github-actions github-actions Bot added component: conversion Issues re: Conversion stage component: core Issues re: The core compiler component: converters Issues re: Specific op converters component: api [Python] Issues re: Python API component: dynamo Issues relating to the `torch.compile` or `torch._dynamo.export` paths labels Jul 24, 2026
@github-actions
github-actions Bot requested a review from apbose July 24, 2026 01:44

@narendasan narendasan left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you explain what this context is and maybe add some tests?

@zewenli98

Copy link
Copy Markdown
Collaborator Author

Can you explain what this context is and maybe add some tests?

  1. In quantize(), amex could be trt.ITensor which doesn't have amax.numel()
  2. ModelOpt requires with export_torch_mode() wrapper to make inserted quantizers traceable while exporting pytorch models. See: torch_tensorrt.compile() fails with "fake tensor in exported program constant's list" for ModelOpt quantized models unless wrapped in export_torch_mode() #4429

@narendasan

Copy link
Copy Markdown
Collaborator

Can you explain what this context is and maybe add some tests?

1. In quantize(), `amex` could be trt.ITensor which doesn't have `amax.numel()`

2. ModelOpt requires `with export_torch_mode()` wrapper to make inserted quantizers traceable while exporting pytorch models. See: [torch_tensorrt.compile() fails with "fake tensor in exported program constant's list" for ModelOpt quantized models unless wrapped in export_torch_mode() #4429](https://github.com/pytorch/TensorRT/issues/4429)

I see, is there any harm in always having this context? should we catch cases where the modelopt is missing but the context is needed and throw an error/warning?

@zewenli98

Copy link
Copy Markdown
Collaborator Author

I see, is there any harm in always having this context? should we catch cases where the modelopt is missing but the context is needed and throw an error/warning?

I remember ModelOpt is not required in torch-trt. If we need to check all models with ModelOpt API, the package should be required rather than opt in.

Comment thread py/torch_tensorrt/dynamo/_tracer.py Outdated
if importlib.util.find_spec("modelopt") is not None:
from modelopt.torch.quantization.utils import export_torch_mode, is_quantized

if is_quantized(mod):

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we throw a warning if its a quantized model but model opt is not available? Does this matter for torchao or other sources?

@zewenli98 zewenli98 Aug 3, 2026

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

well since the function is_quantized is imported from modelopt, it must be available at this point. If we want to throw errors/warnings, I think we might need to write a helper function to detect quantizers to override is_quantized. For other sources like torchao, if they insert quantizers or similar indicators, we can use the same mechanism.

@zewenli98

Copy link
Copy Markdown
Collaborator Author

@narendasan I optimized codes so that if modelopt is not installed, it will fallback to check module names. And it will error out for quantized model if modelopt is requirement. Also, it's easy to generalize for other sources like torchao.

@narendasan narendasan left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lgtm

@github-actions github-actions Bot added the component: lowering Issues re: The lowering / preprocessing passes label Aug 4, 2026
@zewenli98
zewenli98 merged commit f528ed6 into main Aug 5, 2026
119 of 141 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cla signed component: api [Python] Issues re: Python API component: conversion Issues re: Conversion stage component: converters Issues re: Specific op converters component: core Issues re: The core compiler component: dynamo Issues relating to the `torch.compile` or `torch._dynamo.export` paths component: lowering Issues re: The lowering / preprocessing passes

Projects

None yet

2 participants