Vanishing Gradient Problem

Easy
neural-networks Google DeepMind Meta

In a deep network using sigmoid activations, gradients can become extremely small in early layers during backpropagation. Which of the following best explains why this happens?

Multiple Choice

Correct!
Incorrect — try again next time!