LearnBenchStart learning →

AI Research Scientist · Attention Mechanisms & Kernels

Attention head specialization and interpretability signals

Real lesson card · Page 1 of 4

Attention head specialization and interpretability signals

Attention head specialization

The empirical finding that individual heads in a trained transformer consistently attend in distinctive, function-like ways rather than uniformly.
Example
One head in many transformers reliably attends from a verb to its subject, e.g. from ran to dog in ‘The dog ran fast.’

Recall check from the same lesson

If a head reliably attends from adjectives to the nouns they modify across many inputs, this alone proves the head is the mechanism the model uses to combine adjective meaning into the noun's representation.

Sources

· Editorial policy

One sitting · 20–30 minutes

A focused session on your AI Research Scientist interview

LearnBench starts from what you already know — skip what you have, master what you’re missing.

Start now

More Attention Mechanisms & Kernels questions