AI Research Scientist · Optimization Algorithms & Schedules
Cosine annealing and other LR decay shapes
Cosine annealing and other LR decay shapes
Cosine annealing
Cosine annealing decays the learning rate smoothly along a cosine curve from an initial value toward a minimum : No abrupt jumps — just a smooth taper.- Rate of change varies — barely moves near the start and end, steepest at the midpoint of training.
- Plots as a smooth S-curve with no kinks or corners.
- Rate of change is constant — the LR drops by the same fixed amount every step, from to a final value.
- Plots as a straight line the whole way, with a single slope.
- 1LR stays flat at 0.1, then cuts to 0.05 exactly at epoch 10.WhyStep decay holds the rate constant within an interval, then applies a fixed multiplicative drop only at milestones.
- 2The loss curve, which had flattened, drops sharply right after epoch 10 before settling onto a new plateau.WhyThe abrupt LR cut shrinks the step size and gradient-noise floor all at once, so loss falls quickly - a gradual taper spreads that same improvement out smoothly.
Shape shapes dynamics: cosine's smooth taper lets loss improve smoothly throughout training. Step decay's abrupt multiplicative cuts shrink the step size all at once, producing a visible sharp drop in loss right at each milestone before it plateaus again. Linear sits between — no discontinuity, but its constant slope doesn't linger at low LR near the end the way cosine does. Spotting the shape — S-curve, straight line, or staircase — tells you what the loss curve will do.
Recall check from the same lesson
If a training loss curve plateaus and then shows sudden sharp drops at exactly epochs 30, 60, and 90, with smooth loss in between, the schedule most likely used step decay with milestones at those epochs rather than cosine annealing.
Review the explanation
Answer: True. Step decay's abrupt multiplicative LR cuts at fixed milestones shrink the step size and gradient-noise floor right at those points, producing the classic sharp loss drops onto a new plateau; cosine's continuous taper has no discontinuities, so its loss improves smoothly with no milestone-aligned events.
Sources
One sitting · 20–30 minutes
A focused session on your AI Research Scientist interview
LearnBench starts from what you already know — skip what you have, master what you’re missing.
Start now