Translate prototype_source/semi_structured_sparse.rst - #889
Conversation
| Semi-structured sparsity derives its name from its unique sparsity pattern, where n out of every 2n elements are pruned. We most often see n=2, hence 2:4 sparsity | ||
| Semi-structured sparsity is particularly interesting because it can be efficiently accelerated on GPUs and doesn't degrade model accuracy as much as other sparsity patterns. | ||
| ๋ฐ๊ตฌ์กฐ์ ํฌ์์ฑ์ 2n๊ฐ์ ์์ ์ค n๊ฐ์ ์์๊ฐ ์ ๊ฑฐ๋๋ ๋ ํนํ ํฌ์์ฑ ํจํด์์ ๊ทธ ์ด๋ฆ์ ๋ฐ์์ต๋๋ค. ๊ฐ์ฅ ์ผ๋ฐ์ ์ผ๋ก n=2๊ฐ ์ ์ฉ๋๋ฏ๋ก 2:4 ํฌ์์ฑ์ด๋ผ๊ณ ๋ถ๋ฆฝ๋๋ค. | ||
| ๋ฐ๊ตฌ์กฐ์ ํฌ์์ฑ์ ํนํ GPU์์ ํจ์จ์ ์ผ๋ก ๊ฐ์ํํ ์ ์๊ณ , ๋ค๋ฅธ ํฌ์์ฑ ํจํด๋ณด๋ค ๋ชจ๋ธ์ ์ ํ๋๋ฅผ ๋ ์ ํ์ํค๊ธฐ ๋๋ฌธ์ ํฅ๋ฏธ๋กญ์ต๋๋ค. |
There was a problem hiding this comment.
"๊ฐ์ํ๋ ์ ์๊ณ "๊ฐ ์ข๋ ์์ฐ์ค๋ฌ์ด ๊ฒ ๊ฐ์ต๋๋ค!
line 84์์ ์๋ํ๋ก ์ฌ์ฉํ์ ๊ฒ๊ณผ ํต์ผ์ฑ๋ ์๊ธธ ๊ฒ ๊ฐ์์ ๐
| answers = [] | ||
|
|
||
| # Loop through all features associated with that example | ||
| # ํด๋น ์์ ์ ์ฐ๊ด๋ ๋ชจ๋ ํผ์ฒ ๋ฐ๋ณตํ๊ธฐ |
There was a problem hiding this comment.
TRANSLATION_GUIDE.md ์ ๋ฐ๋ผ feature๋ ํน์ง์ด๋ผ๊ณ ๋ฒ์ญํ๋ ๊ฒ์ ์ด๋จ๊น์?
hyoyoung
left a comment
There was a problem hiding this comment.
์ ๋ฐ์ ์ผ๋ก ์๋์ด์์ผ๋
๋ช๊ฐ์ง ํ์ธํด๋ด์ผํ ๋ถ๋ถ์ด ์์ต๋๋ค.
ํ์ธํ์ ์์ ๋ถํ๋๋ฆฝ๋๋ค.
| (ํ๋กํ ํ์ ) ๋ฐ๊ตฌ์กฐ์ (2:4) ํฌ์์ฑ(semi-structured (2:4) sparsity)์ ์ด์ฉํ BERT ๊ฐ์ํํ๊ธฐ | ||
| ================================================================= | ||
| **Author**: `Jesse Cai <https://github.com/jcaip>`_ | ||
| **์ ์**: `Jesse Cai <https://github.com/jcaip>`_ **๋ฒ์ญ**: `Dabin Kang <https://github.com/dabinishere>`_ |
There was a problem hiding this comment.
์ ์์ ๋ฒ์ญ์ ์ค๋ด๋ฆผ์ ํด์ 2์ค๋ก ๋ง๋ค์ด์ฃผ์ธ์
|
|
||
| The natural handoff point between these two problems are zeroed-out dense tensors. Our inference solution is designed to compress and accelerate tensors in this format. | ||
| We anticipate many users coming up with custom masking solution, as this is an active area of research. | ||
| ์ด ๋ ๋ฌธ์ ์ฌ์ด์ ์์ฐ์ค๋ฌ์ด ํธ๋์คํ(handoff) ํฌ์ธํธ๋ 0์ผ๋ก ๋ ๋ฐ์ง ํ ์์ ๋๋ค. ์ด ํ์์ ํ ์๋ฅผ ์์ถํ๊ณ ๊ฐ์ํํ๋๋ก ์ถ๋ก ์ ์ค๊ณํ์ต๋๋ค. |
There was a problem hiding this comment.
์ฉ์ด์ง์์ tensor๋ ์ผ๋ฐ์ ์ผ๋ก ๋ฒ์ญํ์ง ์์ต๋๋ค
| # ํ๊ฐ๋ฅผ ์ํด ๋น๊ตํ ๋ฐฐ์น ํฌ๊ธฐ | ||
| batch_sizes = [4, 16, 64, 256] | ||
| # 2:4 sparsity require fp16, so we cast here for a fair comparison | ||
| # 2:4 ํฌ์์ฑ์ fp16์ด ํ์ํ๋ฏ๋ก ๊ณต์ ํ ๋น๊ต๋ฅผ ์ํด ์ฌ๊ธฐ์ ์บ์คํ |
There was a problem hiding this comment.
์ฌ๊ธฐ์ ์บ์คํ ์ ๋ณํ์ด๋ ํ ๋ณํ์ด๋ผ๊ณ ์์ญํ๋๊ฒ ์ข์ ๊ฑฐ ๊ฐ์ต๋๋ค
| ๋ฐ๊ตฌ์กฐ์ ํฌ์์ฑ์ ํนํ GPU์์ ํจ์จ์ ์ผ๋ก ๊ฐ์ํํ ์ ์๊ณ , ๋ค๋ฅธ ํฌ์์ฑ ํจํด๋ณด๋ค ๋ชจ๋ธ์ ์ ํ๋๋ฅผ ๋ ์ ํ์ํค๊ธฐ ๋๋ฌธ์ ํฅ๋ฏธ๋กญ์ต๋๋ค. | ||
|
|
||
| With the introduction of `semi-structured sparsity support <https://pytorch.org/docs/2.1/sparse.html#sparse-semi-structured-tensors>`_, it is possible to prune and accelerate a semi-structured sparse model without leaving PyTorch. | ||
| We will explain this process in this tutorial. |
There was a problem hiding this comment.
์ 3๋ฌธ์ฅ์ ๋ฒ์ญ๋์๋๋ฐ ๋จ์์๋ ๊ฒ์ผ๊น์?
|
|
||
| By the end of this tutorial, we will have sparsified a BERT question-answering model to be 2:4 sparse, fine-tuning it to recover nearly all F1 loss (86.92 dense vs 86.48 sparse). | ||
| Finally, we will accelerate this 2:4 sparse model for inference, yielding a 1.3x speedup. | ||
| ์ด ํํ ๋ฆฌ์ผ์ ๋๋ด๋ฉด, BERT ์ง๋ฌธ-๋ต๋ณ ๋ชจ๋ธ์ 2:4 ํฌ์์ฑ์ผ๋ก ๊ฐ์ง์น๊ธฐํ๊ณ , ๊ฑฐ์ ๋ชจ๋ F1 ์์ค(๋ฐ์ง 86.92 vs ํฌ์ 86.48)์ ๋ณต๊ตฌํ๋๋ก ๋ฏธ์ธ ์กฐ์ ํ ์ ์์ต๋๋ค. |
There was a problem hiding this comment.
์ด ํํ ๋ฆฌ์ผ์ ๋๋ด๋ฉด, ~~ ํ๋ ๋ชจ๋ธ์ด ์์ฑ๋ฉ๋๋ค. ์ ๋๋ ์ด๋จ๊น์?
| @@ -1,31 +1,31 @@ | |||
| (prototype) Accelerating BERT with semi-structured (2:4) sparsity | |||
| (ํ๋กํ ํ์ ) ๋ฐ๊ตฌ์กฐ์ (2:4) ํฌ์์ฑ(semi-structured (2:4) sparsity)์ ์ด์ฉํ BERT ๊ฐ์ํํ๊ธฐ | |||
There was a problem hiding this comment.
semi-structure๊ฐ ๋ฐ๊ตฌ์กฐ์ธ์ง ๋ฐ์ ํ์ธ์ง, ์กฐ๊ธ ๊ฒฐ์ ํ๊ธฐ ์ด๋ ค์ด ๋ฌธ์ ์ธ๋ฐ
DB์ฉ์ด๋ผ๋ฉด ๋ฐ์ ํ์ ์ด์ธ๋ฆด๊ฑฐ ๊ฐ์ต๋๋ค
https://ko.wikipedia.org/wiki/%EB%B0%98%EC%A0%95%ED%98%95_%EB%8D%B0%EC%9D%B4%ED%84%B0
์ด๋ค๊ฒ ๋ ๋์๊น์?
jaeseong98
left a comment
There was a problem hiding this comment.
๋ฒ์ญํ์๋๋ผ ์๊ณ ๋ง์ผ์ จ์ต๋๋ค!
| 2. At the same time, semi-structured sparsity tends to have a milder impact on model accuracy compared to other sparse formats, especially when accounting for more advanced pruning / fine-tuning methods. | ||
| NVIDIA has shown in their `white paper <https://arxiv.org/abs/2104.08378>`_ that a simple paradigm of magnitude pruning once to be 2:4 sparse and then retraining the model yields nearly identical model accuracies. | ||
| ๊ฐ๊ธฐ ๋ค๋ฅธ ์ฅ๋จ์ ์ ๊ฐ์ง ์ฌ๋ฌ ๊ฐ์ง ํฌ์์ฑ ๊ตฌ์กฐ๋ค์ด ์์ต๋๋ค. ํนํ 2:4 ๋ฐ๊ตฌ์กฐ์ ํฌ์ ๋ ์ด์์์ ๋ ๊ฐ์ง ์ด์ ๋ก ํฅ๋ฏธ๋กญ์ต๋๋ค: | ||
| 1. ์ด์ ์ ํฌ์ ํ์๊ณผ ๋ฌ๋ฆฌ ๋ฐ๊ตฌ์กฐ์ ํฌ์์ฑ์ GPU์์ ํจ์จ์ ์ผ๋ก ๊ฐ์ํ๋๋๋ก ์ค๊ณ๋์์ต๋๋ค. |
There was a problem hiding this comment.
๊ฐ์ํํ๋ค์ ๊ฐ์ํ๋๋ค๊ฐ ์์ฌ์๋๊ฒ ๊ฐ์์ ๋ ์ค ํ๋๋ก ํต์ผ์ฑ์ ์ฃผ๋ฉด ์ข์ ๊ฒ ๊ฐ์ต๋๋ค!
๋ผ์ด์ ์ค ๋์
๋ณ๊ฒฝํด์ฃผ์๋ ๋ด์ฉ์ BSD 3ํญ ๋ผ์ด์ ์ค๊ฐ ์ ์ฉ๋จ์ ๋์ํด์ฃผ์ ์ผ ํฉ๋๋ค.
๋ ์์ธํ ๋ด์ฉ์ ๊ธฐ์ฌํ๊ธฐ ๋ฌธ์๋ฅผ ์ฐธ๊ณ ํด์ฃผ์ธ์.
๋์ํ์๋ฉด ์๋
[ ]๋ฅผ[x]๋ก ๋ง๋ค์ด์ฃผ์ธ์.๊ด๋ จ ์ด์ ๋ฒํธ
์ด Pull Request์ ๊ด๋ จ์๋ ์ด์ ๋ฒํธ๋ฅผ ์ ์ด์ฃผ์ธ์.
์ด์ ๋๋ PR ๋ฒํธ ์์ #์ ๋ถ์ด์๋ฉด ์ ๋ชฉ์ ๋ฐ๋ก ํ์ธํ์ค ์ ์์ต๋๋ค. (์. #999 )
PR ์ข ๋ฅ
์ด PR์ ํด๋น๋๋ ์ข ๋ฅ ์์
[ ]์[x]๋ก ๋ณ๊ฒฝํด์ฃผ์ธ์.PR ์ค๋ช
์ด PR๋ก ๋ฌด์์ด ๋ฌ๋ผ์ง๋์ง ๋๋ต์ ์ผ๋ก ์๋ ค์ฃผ์ธ์.
prototype_source/semi_structured_sparse.rst ๋ฌธ์๋ฅผ ๋ฒ์ญํ์์ต๋๋ค.