LearnBenchStart learning →

AI Research Scientist · Generalization Theory & Phenomena

Grokking and delayed generalization

Real lesson card · Page 1 of 4

Grokking and delayed generalization

Setup
A transformer trained on modular addition (mod 97) hits 99% train accuracy by step 1,000, while test accuracy stays stuck below 10% (chance is about 1%). It jumps from 8% to 96% around step 9,000.
  1. 1
    Compare the train and test curves.
    WhyThey split early: memorization finishes long before test accuracy moves.
  2. 2
    Notice the test-accuracy rise is a sharp jump, not a slope.
    WhyThat abruptness after a long flat stretch is grokking’s signature.
Takeaway
Grokking is near-perfect train accuracy held for many steps while test accuracy stays flat, then jumps abruptly rather than gradually.

Recall check from the same lesson

Because weight decay and simplified internal representations reliably appear around the time a model groks, researchers have proven that weight decay is the mechanistic cause of grokking's abrupt generalization jump.

Sources

· Editorial policy

One sitting · 20–30 minutes

A focused session on your AI Research Scientist interview

LearnBench starts from what you already know — skip what you have, master what you’re missing.

Start now

More Generalization Theory & Phenomena questions