Skip to content
← Calculus & Optimisation

/knowledge/notes/gradient-descent-variants

Concept note · Calculus & Optimisation

Gradient Descent Variants

Optimisation Algorithms

Studied
Statistical Machine LearningCOMP90051
When
2023 S1
Applied in
Human or Machine?
Read / Refreshed
~6 min read2026-10-15

Stochastic gradient descent has many variants. The choice between them affects convergence speed, memory use, and generalization. This note covers the main optimizers: SGD, Momentum, RMSprop, and Adam, and shows when to use each.

01

The idea

All gradient-based optimizers update parameters in the direction of the negative gradient. They differ in how they use past gradients. Vanilla SGD takes one step per gradient. Momentum accumulates gradients, smoothing updates. RMSprop adapts step size per parameter based on historical gradient magnitude. Adam combines both ideas: momentum and per-parameter adaptation.

The practical choice depends on your problem. SGD is simple and often generalises well. Adam converges fast and requires little tuning. RMSprop is between them. Momentum is rarely used alone in modern practice.

02

The maths

Let ∇L\nabla L be the gradient of loss LL. The update rules:

  • SGD: θ←θ−α∇L\theta \leftarrow \theta - \alpha \nabla L
  • Momentum: v←βv+∇Lv \leftarrow \beta v + \nabla L; θ←θ−αv\theta \leftarrow \theta - \alpha v
  • RMSprop: s←βs+(1−β)(∇L)2s \leftarrow \beta s + (1-\beta)(\nabla L)^2; θ←θ−α∇Ls+ϵ\theta \leftarrow \theta - \alpha \frac{\nabla L}{\sqrt{s + \epsilon}}
  • Adam: Combines momentum and RMSprop with bias correction. Uses first moment (mean of gradients) and second moment (mean of squared gradients).

The hyperparameters α\alpha (learning rate), β\beta (momentum decay), and batch size all affect convergence. Larger batches reduce variance but require more memory. Higher learning rates converge faster but may diverge.

03

Try it

The widget shows a synthetic loss surface and four optimizers racing to the minimum. Adjust learning rate and batch size, then click each optimizer to see its trajectory. Adam typically converges fastest.

Optimiser paths on loss surface
Learning rate 0.100, batch size 32
Optimizer comparison on a 2D loss surface. Pick learning rate and batch size, then click to race.
  • High learning rate: fast convergence but risk of divergence.
  • Large batch: smoother gradients, slower overall (fewer updates).
  • Adam usually wins on complicated surfaces; SGD may generalise better.

04

Where I used it

05

Easy to get wrong

06

Sources

First studied in COMP90051 (2024), expanded in 2026.