Vanishing Gradient Problem
Easy
neural-networks
Google
DeepMind
Meta
In a deep network using sigmoid activations, gradients can become extremely small in early layers during backpropagation. Which of the following best explains why this happens?
Multiple Choice
Correct!
Incorrect — try again next time!