/knowledge/notes/gradient-descent-variants
Concept note · Calculus & Optimisation
Gradient Descent Variants
Optimisation Algorithms
- Studied
- Statistical Machine LearningCOMP90051
- When
- 2023 S1
- Applied in
- Human or Machine?
- Read / Refreshed
- ~6 min read2026-10-15
Stochastic gradient descent has many variants. The choice between them affects convergence speed, memory use, and generalization. This note covers the main optimizers: SGD, Momentum, RMSprop, and Adam, and shows when to use each.
01
The idea
All gradient-based optimizers update parameters in the direction of the negative gradient. They differ in how they use past gradients. Vanilla SGD takes one step per gradient. Momentum accumulates gradients, smoothing updates. RMSprop adapts step size per parameter based on historical gradient magnitude. Adam combines both ideas: momentum and per-parameter adaptation.
The practical choice depends on your problem. SGD is simple and often generalises well. Adam converges fast and requires little tuning. RMSprop is between them. Momentum is rarely used alone in modern practice.
02
The maths
Let be the gradient of loss . The update rules:
- SGD:
- Momentum: ;
- RMSprop: ;
- Adam: Combines momentum and RMSprop with bias correction. Uses first moment (mean of gradients) and second moment (mean of squared gradients).
The hyperparameters (learning rate), (momentum decay), and batch size all affect convergence. Larger batches reduce variance but require more memory. Higher learning rates converge faster but may diverge.
03
Try it
The widget shows a synthetic loss surface and four optimizers racing to the minimum. Adjust learning rate and batch size, then click each optimizer to see its trajectory. Adam typically converges fastest.
- High learning rate: fast convergence but risk of divergence.
- Large batch: smoother gradients, slower overall (fewer updates).
- Adam usually wins on complicated surfaces; SGD may generalise better.
04
Where I used it
05
Easy to get wrong
06
Sources
- Adam: A Method for Stochastic OptimizationarXiv 1412.6980
First studied in COMP90051 (2024), expanded in 2026.