LearnBenchStart learning →

AI Research Scientist · Scaling Laws & Stability

Depth vs width tradeoffs in transformer scaling

Real lesson card · Page 1 of 4

Depth vs width tradeoffs in transformer scaling

Depth vs. width allocation

A fixed parameter budget can be spent on more layers (depth) or on wider layers (width) — same total size, different shape.
Example
A 1B-parameter budget could become 48 thin layers or 12 wide ones; both fit the budget but train and perform differently.

Recall check from the same lesson

If a very wide, shallow transformer underperforms a balanced depth/width model at the same parameter budget, the most likely cause is that the wide model has too few layers to build up sequential composition, not that it has too few total parameters.

Sources

· Editorial policy

One sitting · 20–30 minutes

A focused session on your AI Research Scientist interview

LearnBench starts from what you already know — skip what you have, master what you’re missing.

Start now

More Scaling Laws & Stability questions