mlas/arm64: add NEON conv asm kernels and tune NCHWC kernel selection - #27099
Conversation
|
Interesting contribution - thank you! A few questions -
|
|
Hi @aviralagrawal, thank you vey much for your prompt feedback.
Compared to direct GEMM implementation of pointwise convolution asm kernel computes 1x1 conv directly:
As usual there are trade-offs so direct GEMM would be faster when output count is small because then asm kernel drops to single-output path which has less ILP and won't be able to reuse filter loads, non-unit stride and non-contigious output regions hence why we have heuristics checking for stride width and height and very large K/M when GEMM blocking can make better use of caches then a fixed 4-output tile. This is best illustrated if we extract pointwise convolutions from mobilnet that we ran and we can see that on average asm implementation is 1.07x faster, and significant speed ups are when number of channels is high and K/M are small (in the image those are H and W dimensions). In convolution heavy networks the convolutions that are dominant are ones with large number of channels and low height and width so we see visible performance improvements as optimisations from this PR are weighted in that direction.
For benchmarking we used the model from: https://github.com/onnx/models/blob/main/validated/vision/classification/mobilenet/model/mobilenetv2-7.onnx
Running |
|
Thanks @milpuz01 for the detailed description & comment! A couple questions from my side:
|
|
Hi @Rohanjames1997, thank you very much for your comments.
No, particular reason. Mostly because the focus for this PR was on MobileNet model and lack of bandwidth. Thank you for sharing the model where
Yes, I think that is great idea and would be interesting to hear from @hariharans29 too what other testing we should make to try to make these kernels default. As you can see above this change is not going to accelerate all possible pointwise convolutions for example but on average it will show the improvements so if we could agree on a set of performance targets we can use that to drive the decision. Also thank you for your code review I will address them in a separate commit. |
Unfortunately, I don't have a comprehensive list of performance targets to be met to make the feature default. Since, the performance testing may not include all possible Conv shapes, I would like to err on the side of caution and atleast provide one release timeline heads-up to the users before considering making the feature default. I would also encourage you to open a discussion to solicit feedback from other ORT users on ARM if they see speed-up for their models with this feature. It would provide greater confidence and a strong data point to turn it on by default. Thanks for this contribution, we will review it shortly ! |
Thanks @hariharans29. I agree with erring on the side of caution. If this PR goes through and it is in the main release is it possible to add a note that we would like to make |
Thanks @milpuz01. The PR should go through in main eventually but I don't think it will go in 1.24.0 unfortunately as the release branch is cut and the bar to take in new code at this point is critical bug fixes and urgent customer asks only. I will try to take this in for 1.24.1 when it happens and sure I will add a note about considering making it default in one of the future releases, but ultimately, as discussed in the comment #27099 (comment), I expect the NchwcFloatKernel needs optimizations before considering that. |
There was a problem hiding this comment.
Pull request overview
Adds new AArch64 NEON assembly micro-kernels for NCHW, depthwise NCHWc, and pointwise NCHWc convolution, integrates them into the MLAS build, and updates NCHWc kernel-selection heuristics to prefer the asm kernels in selected shapes.
Changes:
- Add new AArch64
.Sconvolution micro-kernels (NCHW, depthwise NCHWc, pointwise NCHWc) and wire them into the MLAS build. - Update ARM64 platform init and NCHWc execution heuristics to select asm kernels for pointwise (stride-1, larger tiles) and depthwise (wider outputs).
- Remove the old intrinsics wrapper for the NCHW float kernel in the NCHWc NEON source file.
Reviewed changes
Copilot reviewed 8 out of 8 changed files in this pull request and generated 4 comments.
Show a summary per file
| File | Description |
|---|---|
| cmake/onnxruntime_mlas.cmake | Adds new AArch64 asm sources to the ARM NEON NCHWc MLAS build setup. |
| onnxruntime/core/mlas/lib/snchwc.cpp | Adds ARM64 heuristics to prefer asm depthwise/pointwise kernels in “safe” cases. |
| onnxruntime/core/mlas/lib/sconv_nchwc_kernel_neon.cpp | Removes the old NCHW float kernel wrapper implementation from the NCHWc NEON source file. |
| onnxruntime/core/mlas/lib/platform.cpp | Switches ARM64 NCHW conv kernel default to asm; updates commentary around kernel choices. |
| onnxruntime/core/mlas/lib/mlasi.h | Declares new asm kernel entry points for ARM64 NEON NCHWc. |
| onnxruntime/core/mlas/lib/aarch64/SconvKernelNeon.S | Adds new NCHW convolution asm micro-kernel. |
| onnxruntime/core/mlas/lib/aarch64/SconvDepthwiseKernelNeon.S | Adds new depthwise NCHWc asm micro-kernel (fast/slow path for padding). |
| onnxruntime/core/mlas/lib/aarch64/SconvPointwiseKernelNeon.S | Adds new pointwise NCHWc asm micro-kernel (multi-output reuse). |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
|
/azp run Linux QNN CI Pipeline,Win_TRT_Minimal_CUDA_Test_CI,Windows ARM64 QNN CI Pipeline,Windows GPU Doc Gen CI Pipeline |
|
Azure Pipelines successfully started running 4 pipeline(s). |
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 8 out of 8 changed files in this pull request and generated 3 comments.
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
|
This change looks good to me. Thanks. Can you please address remaining comments (Copilot + mine) so that it can be merged ? FYI - I have called out the NCHWc layout support on ARM in the 1.24 release notes, so that the community can give a try and share feedback/issues if any - https://github.com/microsoft/onnxruntime/releases. CC: @Rohanjames1997 |
|
Nice 🚀 |
Thanks for bringing this to my attention. I am not sure how the contributors list is generated myself. I ll pass along the information for folks to take a look. Meanwhile, I have added Rohan manually to the list and apologies. EDIT: Filed as issue for tracking: #27274 |
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 8 out of 8 changed files in this pull request and generated no new comments.
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
### Description Initially the NCHWc code was built only on Mac CIs to keep the build path regression-free. There were some Linux-specific paths introduced in #26838 and there is more community interest in contributing to these code paths. See #27099. Hence, it makes sense to keep these code paths built and tested on Linux and Windows too. ### Motivation and Context Improve CI quality with regards to ARM64 NCHWc builds CC: @Rohanjames1997
|
/azp run Linux QNN CI Pipeline,Win_TRT_Minimal_CUDA_Test_CI,Windows ARM64 QNN CI Pipeline,Windows GPU Doc Gen CI Pipeline |
|
Can you please rebase with main ? |
|
Azure Pipelines successfully started running 4 pipeline(s). |
Signed-off-by: Milos Puzovic <milos.puzovic@arm.com>
Signed-off-by: Milos Puzovic <milos.puzovic@arm.com>
Description Enables the file mapping of weights as well as the overall context bin. This feature is currently only enabled for ARM64 WIN devices Motivation and Context Currently, when reading the context bin, ORT allocates a large buffer on the heap. Assuming the same model is used, each ORT session will allocate a buffer for the context bin. This is incredibly wasteful when large models are used. Instead, WIN file mapping can be leveraged to map the context bin, then every time a context needs to be created with the context bin, the pointer to the context bin can be retrieved and used instead of some pre-allocated buffer, thus making QNN EP more memory-efficient. In the case of multiple ORT sessions, the context bin will only be loaded once for all sessions, increasing memory efficiency and overall initialization performance. This is very useful regarding the use of LLMs going forward. --------- Co-authored-by: quic_calvnguy <quic_calvnguy@quic_inc.com>
…ft#27151) Previously in `MatMulReadFnSource()` we use duplicated code to read data from two inputs `a` and `b`. This patch implements another overload of `MatMulReadFnSource()` to only read data from one input to reduce duplicated code and get ready for further use.
Signed-off-by: Milos Puzovic <milos.puzovic@arm.com>
…ck spill Signed-off-by: Milos Puzovic <milos.puzovic@arm.com>
Fix bad merge
Signed-off-by: Milos Puzovic <milos.puzovic@arm.com>
2d05853 to
bd38b0e
Compare
Signed-off-by: Milos Puzovic <milos.puzovic@arm.com>
Just rebased. |
|
/azp run Linux QNN CI Pipeline,Win_TRT_Minimal_CUDA_Test_CI,Windows ARM64 QNN CI Pipeline,Windows GPU Doc Gen CI Pipeline |
|
Azure Pipelines successfully started running 4 pipeline(s). |
…microsoft#27099) ## Overview This PR adds ARM64 NEON assembly micro‑kernels for NCHW, depthwise, and pointwise convolution, wires them into the MLAS build, and adds shape‑based selection heuristics for NCHWC depthwise/pointwise to favor the asm kernels in safe cases (stride‑1 pointwise; wider depthwise outputs). The BF16 path is unchanged. ## Key changes - cmake/onnxruntime_mlas.cmake - Add new AArch64 assembly sources for NCHW, depthwise, and pointwise conv to the MLAS build. - onnxruntime/core/mlas/lib/aarch64/SconvKernelNeon.S - New vectorised NCHW convolution micro‑kernel. - onnxruntime/core/mlas/lib/aarch64/SconvDepthwiseKernelNeon.S - New vectorised depthwise micro‑kernel (fast path for in‑bounds loads, slow path for padding). - onnxruntime/core/mlas/lib/aarch64/SconvPointwiseKernelNeon.S - New vectorised pointwise micro‑kernel (multi‑output reuse). - onnxruntime/core/mlas/lib/mlasi.h, onnxruntime/core/mlas/lib/platform.cpp - Declare/register new asm kernels and prefer them on ARM64. - onnxruntime/core/mlas/lib/snchwc.cpp - Heuristics: use pointwise asm when StrideHeight == 1 && StrideWidth == 1 and OutputThisIteration >= 4; use depthwise asm when OutputWidth >= 4. - onnxruntime/core/mlas/lib/sbconv_kernel_neon.cpp - Include fix for the conv kernel flags header. ## Performance Numbers below are expressed as multipliers vs the non‑NCHWC baseline (same model and perf_test settings): Baseline (no `--enable_arm_neon_nchwc`) - 8 cores: 1.00× - 16 cores: 1.00× With `--enable_arm_neon_nchwc` (no asm additions/heuristics) - 8 cores: 1.18× - 16 cores: 1.24× With this PR (asm kernels + heuristics) - 8 cores: 1.77× - 16 cores: 2.54× ## Testing - `./build.sh --config Release --build_shared_lib --parallel --compile_no_warning_as_error --skip_submodule_sync --skip_tests --enable_pybind --build_wheel --enable_arm_neon_nchwc` - `OMP_NUM_THREADS=8 ./build/Linux/Release/onnxruntime_perf_test -I -m times -r 1000 --x 8 ~/mobilenetv2-7.onnx` --------- Signed-off-by: Milos Puzovic <milos.puzovic@arm.com> Co-authored-by: xhcao <xinghua.cao@intel.com> Co-authored-by: quic-calvnguy <quic_calvnguy@quicinc.com> Co-authored-by: quic_calvnguy <quic_calvnguy@quic_inc.com> Co-authored-by: Jiawei Shao <jiawei.shao@intel.com>
… into build (#27788) ### Description This change introduces a new AArch64 assembly owner for the NCHWc float convolution path and keeps the existing C++ entrypoint stable, with build wiring so the same kernel entry is used transparently by existing call sites. The optimization strategy is centered on three ideas: 1. Route the stable NCHWc C++ entrypoint to an AArch64 assembly implementation. 2. Split execution into left/middle/right regions, with a 4-output center tile and 3/2/1 remainder kernels in the middle region. 3. Duplicate hot loops for flags == 0 so the no-post-op path avoids repeated accumulate/bias/ReLU flag checks inside the hottest loops. ### Motivation and Context This improves throughput while preserving existing behavior and integration points and address comment #27099 (comment) from that PR (cc: @Rohanjames1997 and @hariharans29). Per-convolution improvements from convolutions in the model from: #25580 (comment) per core count when running on AWS Graviton 4 (all of these convolution have common N = 1, KH = 3, KW = 3, DH = 1, DW = 1 and G = 1) <table> <thead> <tr> <th>Cores</th> <th>IC</th><th>OC</th><th>IH</th><th>IW</th><th>OH</th><th>OW</th> <th>SH</th><th>SW</th><th>PT</th><th>PL</th><th>PB</th><th>PR</th> <th>P90 Improvement</th><th>P99 Improvement</th> </tr> </thead> <tbody> <tr><td rowspan="5">1</td><td>32</td><td>32</td><td>192</td><td>192</td><td>192</td><td>192</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>+47.57%</td><td>+43.64%</td></tr> <tr><td>32</td><td>96</td><td>192</td><td>192</td><td>96</td><td>96</td><td>2</td><td>2</td><td>0</td><td>0</td><td>1</td><td>1</td><td>+47.56%</td><td>+44.69%</td></tr> <tr><td>48</td><td>192</td><td>96</td><td>96</td><td>48</td><td>48</td><td>2</td><td>2</td><td>0</td><td>0</td><td>1</td><td>1</td><td>+47.25%</td><td>+45.12%</td></tr> <tr><td>48</td><td>192</td><td>96</td><td>96</td><td>96</td><td>96</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>+48.73%</td><td>+47.07%</td></tr> <tr><td>64</td><td>256</td><td>48</td><td>48</td><td>48</td><td>48</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>+47.73%</td><td>+47.76%</td></tr> <tr><td rowspan="5">2</td><td>32</td><td>32</td><td>192</td><td>192</td><td>192</td><td>192</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>+47.72%</td><td>+42.78%</td></tr> <tr><td>32</td><td>96</td><td>192</td><td>192</td><td>96</td><td>96</td><td>2</td><td>2</td><td>0</td><td>0</td><td>1</td><td>1</td><td>+48.41%</td><td>+47.76%</td></tr> <tr><td>48</td><td>192</td><td>96</td><td>96</td><td>48</td><td>48</td><td>2</td><td>2</td><td>0</td><td>0</td><td>1</td><td>1</td><td>+47.13%</td><td>+44.62%</td></tr> <tr><td>48</td><td>192</td><td>96</td><td>96</td><td>96</td><td>96</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>+48.58%</td><td>+46.65%</td></tr> <tr><td>64</td><td>256</td><td>48</td><td>48</td><td>48</td><td>48</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>+47.59%</td><td>+47.32%</td></tr> <tr><td rowspan="5">4</td><td>32</td><td>32</td><td>192</td><td>192</td><td>192</td><td>192</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>+45.18%</td><td>+38.51%</td></tr> <tr><td>32</td><td>96</td><td>192</td><td>192</td><td>96</td><td>96</td><td>2</td><td>2</td><td>0</td><td>0</td><td>1</td><td>1</td><td>+46.34%</td><td>+45.25%</td></tr> <tr><td>48</td><td>192</td><td>96</td><td>96</td><td>48</td><td>48</td><td>2</td><td>2</td><td>0</td><td>0</td><td>1</td><td>1</td><td>+46.17%</td><td>+42.45%</td></tr> <tr><td>48</td><td>192</td><td>96</td><td>96</td><td>96</td><td>96</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>+48.20%</td><td>+45.25%</td></tr> <tr><td>64</td><td>256</td><td>48</td><td>48</td><td>48</td><td>48</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>+47.32%</td><td>+46.67%</td></tr> <tr><td rowspan="5">8</td><td>32</td><td>32</td><td>192</td><td>192</td><td>192</td><td>192</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>+40.07%</td><td>+25.57%</td></tr> <tr><td>32</td><td>96</td><td>192</td><td>192</td><td>96</td><td>96</td><td>2</td><td>2</td><td>0</td><td>0</td><td>1</td><td>1</td><td>+42.93%</td><td>+45.06%</td></tr> <tr><td>48</td><td>192</td><td>96</td><td>96</td><td>48</td><td>48</td><td>2</td><td>2</td><td>0</td><td>0</td><td>1</td><td>1</td><td>+44.50%</td><td>+38.86%</td></tr> <tr><td>48</td><td>192</td><td>96</td><td>96</td><td>96</td><td>96</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>+47.31%</td><td>+45.22%</td></tr> <tr><td>64</td><td>256</td><td>48</td><td>48</td><td>48</td><td>48</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>+47.02%</td><td>+46.42%</td></tr> <tr><td rowspan="5">16</td><td>32</td><td>32</td><td>192</td><td>192</td><td>192</td><td>192</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>+34.06%</td><td>+28.48%</td></tr> <tr><td>32</td><td>96</td><td>192</td><td>192</td><td>96</td><td>96</td><td>2</td><td>2</td><td>0</td><td>0</td><td>1</td><td>1</td><td>+38.12%</td><td>+39.83%</td></tr> <tr><td>48</td><td>192</td><td>96</td><td>96</td><td>48</td><td>48</td><td>2</td><td>2</td><td>0</td><td>0</td><td>1</td><td>1</td><td>+42.17%</td><td>+35.48%</td></tr> <tr><td>48</td><td>192</td><td>96</td><td>96</td><td>96</td><td>96</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>+45.25%</td><td>+42.73%</td></tr> <tr><td>64</td><td>256</td><td>48</td><td>48</td><td>48</td><td>48</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>1</td><td>+46.44%</td><td>+45.84%</td></tr> </tbody> </table> And end-to-end performance improvement when running the model on different core count: | Model | Cores | P90 Improvement | P99 Improvement | |---|---:|---:|---:| | shareable_model.onnx | 1 | +24.35% | +24.17% | | shareable_model.onnx | 2 | +24.14% | +23.82% | | shareable_model.onnx | 4 | +23.40% | +23.22% | | shareable_model.onnx | 8 | +22.72% | +22.55% | | shareable_model.onnx | 16 | +20.19% | +19.89% | Running `bench_sconv_nchwc.cpp` benchmark on AWS Graviton 4 `c8g.16xlarge` the following is obtained: ``` $ ./build/Linux/Release/onnxruntime_mlas_benchmark --benchmark_filter='SCONV_NCHWC_DIRECT/DirectNchwcCases' The number of inputs is very large. QNBITGEMM<float, 4> will be repeated at least 128 times. The number of inputs is very large. QNBITGEMM<float, 8> will be repeated at least 128 times. 2026-03-23T17:35:58+00:00 Running ./build/Linux/Release/onnxruntime_mlas_benchmark Run on (64 X 2000 MHz CPU s) CPU Caches: L1 Data 64 KiB (x64) L1 Instruction 64 KiB (x64) L2 Unified 2048 KiB (x64) L3 Unified 36864 KiB (x1) Load Average: 0.01, 0.13, 0.36 -------------------------------------------------------------------------------------------------------------------------------------------------------- Benchmark Time CPU Iterations -------------------------------------------------------------------------------------------------------------------------------------------------------- SCONV_NCHWC_DIRECT/DirectNchwcCases/IC:32/OC:32/IH:192/IW:192/KH:3/KW:3/PT:1/PL:1/PB:1/PR:1/S:1/D:1/real_time 9169970 ns 9170100 ns 76 SCONV_NCHWC_DIRECT/DirectNchwcCases/IC:32/OC:96/IH:192/IW:192/KH:3/KW:3/PT:0/PL:0/PB:1/PR:1/S:2/D:1/real_time 7048829 ns 7047903 ns 99 SCONV_NCHWC_DIRECT/DirectNchwcCases/IC:48/OC:192/IH:96/IW:96/KH:3/KW:3/PT:0/PL:0/PB:1/PR:1/S:2/D:1/real_time 5422228 ns 5421912 ns 129 SCONV_NCHWC_DIRECT/DirectNchwcCases/IC:48/OC:192/IH:96/IW:96/KH:3/KW:3/PT:1/PL:1/PB:1/PR:1/S:1/D:1/real_time 22050813 ns 22047140 ns 31 SCONV_NCHWC_DIRECT/DirectNchwcCases/IC:64/OC:256/IH:48/IW:48/KH:3/KW:3/PT:1/PL:1/PB:1/PR:1/S:1/D:1/real_time 9856375 ns 9855928 ns 71 ``` Running the same benchmark on the same system without this PR we get: ``` Running ./build/Linux/Release/onnxruntime_mlas_benchmark Run on (64 X 2000 MHz CPU s) CPU Caches: L1 Data 64 KiB (x64) L1 Instruction 64 KiB (x64) L2 Unified 2048 KiB (x64) L3 Unified 36864 KiB (x1) Load Average: 24.34, 13.45, 5.35 -------------------------------------------------------------------------------------------------------------------------------------------------------- Benchmark Time CPU Iterations -------------------------------------------------------------------------------------------------------------------------------------------------------- SCONV_NCHWC_DIRECT/DirectNchwcCases/IC:32/OC:32/IH:192/IW:192/KH:3/KW:3/PT:1/PL:1/PB:1/PR:1/S:1/D:1/real_time 17576371 ns 17576571 ns 40 SCONV_NCHWC_DIRECT/DirectNchwcCases/IC:32/OC:96/IH:192/IW:192/KH:3/KW:3/PT:0/PL:0/PB:1/PR:1/S:2/D:1/real_time 14319665 ns 14316035 ns 49 SCONV_NCHWC_DIRECT/DirectNchwcCases/IC:48/OC:192/IH:96/IW:96/KH:3/KW:3/PT:0/PL:0/PB:1/PR:1/S:2/D:1/real_time 11710396 ns 11707113 ns 60 SCONV_NCHWC_DIRECT/DirectNchwcCases/IC:48/OC:192/IH:96/IW:96/KH:3/KW:3/PT:1/PL:1/PB:1/PR:1/S:1/D:1/real_time 47237327 ns 47217289 ns 15 SCONV_NCHWC_DIRECT/DirectNchwcCases/IC:64/OC:256/IH:48/IW:48/KH:3/KW:3/PT:1/PL:1/PB:1/PR:1/S:1/D:1/real_time 21110794 ns 21103617 ns 33 ``` which shows the following speed up: | Benchmark row | First run (ns) | Second run (ns) | Speedup (first vs second) | |---|---:|---:|---:| | IC32 OC32 IH192 IW192 KH3 KW3 PT1 PL1 PB1 PR1 S1 D1 | 9,169,970 | 17,576,371 | 1.92x | | IC32 OC96 IH192 IW192 KH3 KW3 PT0 PL0 PB1 PR1 S2 D1 | 7,048,829 | 14,319,665 | 2.03x | | IC48 OC192 IH96 IW96 KH3 KW3 PT0 PL0 PB1 PR1 S2 D1 | 5,422,228 | 11,710,396 | 2.16x | | IC48 OC192 IH96 IW96 KH3 KW3 PT1 PL1 PB1 PR1 S1 D1 | 22,050,813 | 47,237,327 | 2.14x | | IC64 OC256 IH48 IW48 KH3 KW3 PT1 PL1 PB1 PR1 S1 D1 | 9,856,375 | 21,110,794 | 2.14x | --------- Signed-off-by: Milos Puzovic <milos.puzovic@arm.com>

Overview
This PR adds ARM64 NEON assembly micro‑kernels for NCHW, depthwise, and pointwise convolution, wires them into the MLAS build, and adds shape‑based selection heuristics for NCHWC depthwise/pointwise to favor the asm kernels in safe cases (stride‑1 pointwise; wider depthwise outputs). The BF16 path is unchanged.
Key changes
Performance
Numbers below are expressed as multipliers vs the non‑NCHWC baseline (same model and perf_test settings):
Baseline (no
--enable_arm_neon_nchwc)With
--enable_arm_neon_nchwc(no asm additions/heuristics)With this PR (asm kernels + heuristics)
Testing
./build.sh --config Release --build_shared_lib --parallel --compile_no_warning_as_error --skip_submodule_sync --skip_tests --enable_pybind --build_wheel --enable_arm_neon_nchwcOMP_NUM_THREADS=8 ./build/Linux/Release/onnxruntime_perf_test -I -m times -r 1000 --x 8 ~/mobilenetv2-7.onnx