Skip to content

Harden intrinsic compile-time validation - #3213

Open
maleadt wants to merge 1 commit into
mainfrom
tb/static_assert_intrinsics
Open

Harden intrinsic compile-time validation#3213
maleadt wants to merge 1 commit into
mainfrom
tb/static_assert_intrinsics

Conversation

@maleadt

@maleadt maleadt commented Jul 22, 2026

Copy link
Copy Markdown
Member

No description provided.

@github-actions

github-actions Bot commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

CUDA.jl Benchmarks

Details
Benchmark suite Current: 7ff74bf Previous: 069cdff Ratio
array/accumulate/Float32/1d 97377 ns 98074 ns 0.99
array/accumulate/Float32/dims=1 70765 ns 74919 ns 0.94
array/accumulate/Float32/dims=1L 1598543 ns 1600002 ns 1.00
array/accumulate/Float32/dims=2 136147 ns 140367 ns 0.97
array/accumulate/Float32/dims=2L 657859 ns 659745 ns 1.00
array/accumulate/Int64/1d 116864 ns 117818 ns 0.99
array/accumulate/Int64/dims=1 74675 ns 79251 ns 0.94
array/accumulate/Int64/dims=1L 1715389 ns 1717606 ns 1.00
array/accumulate/Int64/dims=2 146884 ns 153507 ns 0.96
array/accumulate/Int64/dims=2L 986104 ns 986800 ns 1.00
array/broadcast 17420 ns 18060 ns 0.96
array/broadcast launch 8094.666666666667 ns
array/construct 888.7916666666666 ns 869.1052631578947 ns 1.02
array/copy 16188 ns 15967 ns 1.01
array/copyto!/cpu_to_gpu 206376 ns 208599 ns 0.99
array/copyto!/gpu_to_cpu 240029 ns 241280 ns 0.99
array/copyto!/gpu_to_gpu 8681.333333333334 ns 9209.333333333334 ns 0.94
array/iteration/findall/bool 130551 ns 132021 ns 0.99
array/iteration/findall/int 144064 ns 145286 ns 0.99
array/iteration/findfirst/bool 67599 ns 67274 ns 1.00
array/iteration/findfirst/int 68729 ns 68529 ns 1.00
array/iteration/findmin/1d 58915 ns 64309 ns 0.92
array/iteration/findmin/2d 98691 ns 99940 ns 0.99
array/iteration/logical 181155 ns 185815 ns 0.97
array/iteration/scalar 60535 ns 63136 ns 0.96
array/permutedims/2d 47344 ns 48431 ns 0.98
array/permutedims/3d 48641 ns 50026 ns 0.97
array/permutedims/4d 48921 ns 49849 ns 0.98
array/random/rand/Float32 11774 ns 11669 ns 1.01
array/random/rand/Int64 22116 ns 22883 ns 0.97
array/random/rand!/Float32 7589 ns 7731.75 ns 0.98
array/random/rand!/Int64 19663 ns 20210 ns 0.97
array/random/randn/Float32 32341 ns 32574 ns 0.99
array/random/randn!/Float32 24107 ns 23189 ns 1.04
array/reductions/mapreduce/Float32/1d 30985 ns 31371 ns 0.99
array/reductions/mapreduce/Float32/dims=1 36964 ns 37491 ns 0.99
array/reductions/mapreduce/Float32/dims=1L 49830 ns 49977 ns 1.00
array/reductions/mapreduce/Float32/dims=2 54080 ns 54508 ns 0.99
array/reductions/mapreduce/Float32/dims=2L 65782 ns 66329 ns 0.99
array/reductions/mapreduce/Int64/1d 37774 ns 37643 ns 1.00
array/reductions/mapreduce/Int64/dims=1 39752 ns 39961 ns 0.99
array/reductions/mapreduce/Int64/dims=1L 87904 ns 87917 ns 1.00
array/reductions/mapreduce/Int64/dims=2 56658 ns 57047 ns 0.99
array/reductions/mapreduce/Int64/dims=2L 82234 ns 82751 ns 0.99
array/reductions/reduce/Float32/1d 31216 ns 31482 ns 0.99
array/reductions/reduce/Float32/dims=1 37048 ns 37630 ns 0.98
array/reductions/reduce/Float32/dims=1L 49811 ns 50073 ns 0.99
array/reductions/reduce/Float32/dims=2 54199 ns 54518 ns 0.99
array/reductions/reduce/Float32/dims=2L 67403 ns 67963 ns 0.99
array/reductions/reduce/Int64/1d 37958 ns 38054 ns 1.00
array/reductions/reduce/Int64/dims=1 39663 ns 39983 ns 0.99
array/reductions/reduce/Int64/dims=1L 87774 ns 87988 ns 1.00
array/reductions/reduce/Int64/dims=2 56605 ns 56860 ns 1.00
array/reductions/reduce/Int64/dims=2L 82201 ns 82191 ns 1.00
array/reverse/1d 16240 ns 16098 ns 1.01
array/reverse/1dL 69003 ns 68844 ns 1.00
array/reverse/1dL_inplace 67229 ns 66918 ns 1.00
array/reverse/1d_inplace 10096 ns 9757 ns 1.03
array/reverse/2d 18984 ns 19108 ns 0.99
array/reverse/2dL 72288 ns 72561 ns 1.00
array/reverse/2dL_inplace 66818 ns 66785 ns 1.00
array/reverse/2d_inplace 9659 ns 9397 ns 1.03
array/sorting/1d 2639299 ns 2656939 ns 0.99
array/sorting/2d 1027118 ns 1038211 ns 0.99
array/sorting/by 3182201 ns 3192363 ns 1.00
cuda/synchronization/context/auto 1005.6 ns 1015.3 ns 0.99
cuda/synchronization/context/blocking 777.1979166666666 ns 784.6960784313726 ns 0.99
cuda/synchronization/context/nonblocking 5615.666666666667 ns 5655.166666666667 ns 0.99
cuda/synchronization/stream/auto 852.8032786885246 ns 857.3174603174604 ns 0.99
cuda/synchronization/stream/blocking 663.6792452830189 ns 654.7730061349694 ns 1.01
cuda/synchronization/stream/nonblocking 5493.5 ns 5491 ns 1.00
integration/byval/reference 147365 ns 147245 ns 1.00
integration/byval/slices=1 151324 ns 149267 ns 1.01
integration/byval/slices=2 294603 ns 292191 ns 1.01
integration/byval/slices=3 437511 ns 434868 ns 1.01
integration/cudadevrt 104557 ns 104357 ns 1.00
integration/volumerhs 9137644 ns 9315717 ns 0.98
kernel/indexing 12436 ns 12183 ns 1.02
kernel/indexing_checked 13280 ns 13078 ns 1.02
kernel/launch 1939.3 ns 2032.3333333333333 ns 0.95
kernel/occupancy 668.304347826087 ns 667.8466257668712 ns 1.00
kernel/rand 13531 ns 14903 ns 0.91
latency/import 4091768416 ns 4116528672 ns 0.99
latency/precompile 4794097643 ns 4814427609 ns 1.00
latency/ttfp 5137544706 ns 5198634459 ns 0.99

This comment was automatically generated by workflow using github-action-benchmark.

@kshyatt

kshyatt commented Jul 23, 2026

Copy link
Copy Markdown
Member

(Rebased to run this on top of the Crayons fix)

@kshyatt
kshyatt force-pushed the tb/static_assert_intrinsics branch from 74a5d1a to 7ff74bf Compare July 23, 2026 08:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants