AI Research Scientist · Generalization Theory & Phenomena
Information bottleneck and generalization bounds
Information bottleneck and generalization bounds
mutual-information bound
It bounds generalization error using , the mutual information between input and representation . Lower means more compression, less capacity to overfit.- Weight-normbigger signals more capacity to overfit.
- Marginwider decision margins imply lower complexity.
- Mutual informationhigher signals more retained input detail.
- Compressionless compression implies higher overfit risk.
Both families chase the same target: a proxy for how much a model 'remembers' versus generalizes. Norms track weight magnitude; mutual information tracks retained input detail. The goal is one complexity score meant to track the generalization gap — recognize new bounds by what they compress, not their formula.
Recall check from the same lesson
Because information-bottleneck and norm-based bounds are grounded in solid theory, a tighter bound value for a deep network reliably predicts it will achieve lower test error in practice.
Review the explanation
Answer: False. A known limitation of both bound families is that they can be vacuous or fail to correlate with actual generalization behavior in deep networks, so a tighter-looking bound value does not reliably predict better test performance.
Sources
One sitting · 20–30 minutes
A focused session on your AI Research Scientist interview
LearnBench starts from what you already know — skip what you have, master what you’re missing.
Start now