Gradient
Descent
Roll downhill by following −∇f
Loss surface
Optimizer
Vanilla gradient descent
Momentum
Nesterov momentum
RMSProp
Adam
Learning rate α
0.100
Step size along the gradient.
Momentum β
0.85
Speed
6 /s
Play
Step
Reset
Race all optimizers
Gradient arrow
Contour floor
Surface grid
Click anywhere on the surface to drop the ball there.
θ
t+1
= θ
t
−
α
∇f(θ
t
)
iteration
0
position (x, y)
—
loss f(θ)
—
gradient ∇f
—
‖∇f‖
—
step size
—
loss vs. iteration
—
Drag
to orbit ·
Scroll
to zoom ·
Click surface
to reposition ·
Space
play/pause ·
R
reset