Self-Attention in Transformers

Medium
neural-networks Google OpenAI Anthropic

In the transformer architecture, self-attention computes query, key, and value vectors. What does the dot product of queries and keys represent, and why is it scaled?

Multiple Choice

Correct!
Incorrect — try again next time!