AdaGrad Limitations
Hard
neural-networks
Google
DeepMind
NVIDIA
AdaGrad adapts the learning rate for each parameter by dividing by the square root of the sum of all past squared gradients. Why can this become problematic in deep learning, and how does RMSProp fix it?
Multiple Choice
Correct!
Incorrect — try again next time!