Self-Attention in Transformers
Medium
neural-networks
Google
OpenAI
Anthropic
In the transformer architecture, self-attention computes query, key, and value vectors. What does the dot product of queries and keys represent, and why is it scaled?
Multiple Choice
Correct!
Incorrect — try again next time!