# Release Manager @cp5555 # Endgame - [x] Code freeze: Aug. 30, 2024 - [x] Bug Bash date: Sept. 02, 2024 - [x] Release date: Oct. 09, 2024 # Main Features ## SuperBench Improvement 1. - [x] Add CUDA 12.4 dockerfile (#619) 2. - [x] Improve document (#628 and #632) 3. - [x] Update omegaconf version to 2.3.0 (#631) 4. - [x] Fix MSCCL build error in CUDA12.4 docker build pipeline (#633) 5. - [x] Update Docker Exec Command for Persistent HPCX Environment (#635) 6. - [x] Use types-setuptools to replace types-pkg_resources (#637) 7. - [x] Update Docker Exec Command for Persistent HPCX Environment (#635) 8. - [x] Fix bug of failure test and warning of pandas in data diagnosis (#638) 9. - [x] Limit protobuf version to be 3.20.x (#645) 10. - [x] Update hpcx link in cuda11.1 dockerfile to fix CI (#648) 11. - [x] Upgrade nccl version and install ucx to fix bug in cuda 12.4 docker file (#646) 12. - [x] Add ROCm6.2 dockerfile (#647) 13. - [x] Use identical metric names described in result-summary.md and micro-benchmarks.md (#651) 14. - [x] Support Azure H100 NDv5 configuration and AMD MI300 configuration (#652) 15. - [ ] Remove pytest and protobuf version constraint (<=7.4.4) ## Micro-benchmark Improvement 1. - [x] Add hipblasLt tuning to dist-inference cpp implementation (#616) 2. - [x] Add support for NVIDIA L4/L40/L40s GPUs in gemm-flops (#634) 3. - [x] Upgrade mlc to v3.11 (#620) 4. - [ ] Support cuDNN Backend API in cudnn-function. ## Model Benchmark Improvement 1. Support VGG, LSTM, and GPT-2 small in TensorRT Inference Backend 16. Support VGG, LSTM, and GPT-2 small in ORT Inference Backend 17. Support more TensorRT parameters (Related to #366) ## Result Analysis
Release Manager
@cp5555
Endgame
Main Features
SuperBench Improvement
types-setuptoolsastypes-pkg_resourcesis Yanked #637)Micro-benchmark Improvement
Model Benchmark Improvement
Result Analysis