AdaGrad Limitations

Hard
neural-networks Google DeepMind NVIDIA

AdaGrad adapts the learning rate for each parameter by dividing by the square root of the sum of all past squared gradients. Why can this become problematic in deep learning, and how does RMSProp fix it?

Multiple Choice

Correct!
Incorrect — try again next time!