Deep Learning
Why did the output dimension of all the embedding and sublayers of original Transformer (Attention is All You Need 2017) need to be the same?
Because of the residual connections.
Proposed by @GokuMohandas
Because of the residual connections.
Proposed by @GokuMohandas
Card 1 of 197. Answer: Because of the residual connections.