Interpretability · main findings · updated 3 October 2026

Embodied Cognition in Transformers

Image schemas and metaphorical mappings in language models trained on text alone.

Language models write about rising spirits and heavy hearts, arguments that collapse and plans that move forward, as if the bodily scaffolding under those phrases were available to them. How does a system trained on nothing but text come by that?

This page was rewritten on 3 October 2026 after an audit found two faults in the instrument behind several earlier findings. What survived re-testing is here, with new results; the earlier version is kept, with corrections marking what was withdrawn. Pythia 410M–1.4B · GPT-2 medium · Llama-3.2-1B base and instruct.

In 80 secondsHAPPY IS UP, explained

An animated explainer of the first result below. The tilt of the scale and the positions of the grey dots are measured values (Pythia 1.4B, layer 12). Open it full-size.

1 · The questionLakoff's schemas: the body as the source of abstract thought

George Lakoff and Mark Johnson argued that abstract thought is structured by image schemas: UP-DOWN, IN-OUT, BALANCE, FORCE, SOURCE-PATH-GOAL. Recurring patterns of bodily experience, projected metaphorically into abstract domains. HAPPY IS UP. MORE IS UP. PURPOSES ARE DESTINATIONS.

From there a familiar argument runs: LLMs have no body. No body, no image schemas; no image schemas, no embodied cognitive structure; so LLM meaning is structurally defective.

My hypothesis: as well as explicit descriptions of the physical world, human language encodes a lot of implicit information about having a body, and a model that compresses human language hard enough reconstructs embodiment as a projection from it.

So: take the metaphors Lakoff describes as foundational and test whether they transfer to the transformer, or whether the concepts come apart.

2 · First experimentUP is HAPPY, for transformers too

Build an UP direction from purely spatial word pairs (up/down, rise/fall, above/below, climb/descend), each word read inside eight neutral sentences, single-token words only, word frequency projected out. Add it to the residual stream while the model completes prompts. The completions get happier.

  • Judged against directions built the same way from random words, which turn out to move things far more than "zero" suggests. On happiness, UP beats every one of them: all 30 in Pythia, and all 20 at each GPT-2 layer.
  • In Pythia 1.4B it is happiness only: status and quantity don't clear the random-word bar.
  • In GPT-2 medium happiness clears it at all three layers tried, status at two, and quantity (MORE IS UP) at one.
  • In Pythia 410M the UP direction leans toward the valence direction more than random directions do (cos ≈ +0.2, 97th–99th percentile across layers), but is mostly something else: valence itself moves happiness more than twice as far.
Effect of steering a clean UP direction on happiness, status and quantity, as a multiple of the 95th percentile of random-word directions, in Pythia 1.4B and three layers of GPT-2 medium
Steering effect (strength +8 minus −8) divided by the 95th percentile of 20–30 random-word directions. Above the dashed line = beyond random. Pale bars don't clear it. Happiness is scored without the two vertical candidate words ("uplifted", "low"); removing them changed nothing.

3 · The objectionOf course UP is HAPPY in the text

The corpus was written by embodied humans who think in these metaphors; a model that compresses their language inherits their shadow. Steering shows the direction is causally live inside the model, not that it is anything more than the corpus's fingerprint.

That objection stands; nothing here refutes it. What the next results add is that the inheritance is structured: the schemas relate to each other the way the theory says, and the shape of a literal journey is reused for change and for tasks.

4 · Is it a system?The schemas couple to each other as the theory predicts

Lakoff's claim was never about isolated correspondences: the schemas form a coherent system. So the predictions were written down first: six couplings the embodied logic implies (UP↔LIGHT-DARK, LIGHT-DARK↔BALANCE, FORCE↔DIFFICULTY, FORWARD-BACK↔PATH, UP↔BALANCE, UP↔FORCE), then the full 8×8 matrix at every layer of Pythia 410M.

+0.17
mean coupling, the 6 predicted pairs
+0.01
mean coupling, the 22 unpredicted pairs

The bar is not zero. The same words, scrambled into eight fake schemas of the same sizes, give a difference of at most +0.08 (95th percentile); the real schemas give +0.16. Four predicted pairs are positive at all 24 layers; two (UP↔BALANCE, UP↔FORCE) are near zero. Most connected: LIGHT-DARK. Least: BALANCE.

Predicted minus unpredicted schema couplings across 24 layers of Pythia 410M, against a scrambled-schema null band
Clean axes, Pythia 410M. Above the scrambled range at every layer except the last.

5 · NewSource-path-goal: the shape of a walk, reused for change and for tasks

Train a simple probe on sentences about literal journeys, with no "from" or "to" in them ("The walk began at the barn and ended at the river"), to tell the start from the end. Then show it things that aren't walks.

ShownReads start / endChance tops out at
"His poverty slowly turned into wealth" · Pythia67–73%61–64%
"From poverty to wealth" (words it never saw) · Pythia98%
"Turn this draft into a summary" · Llama base75%66%
same, Llama instruct72%67%

Swap the two nouns and the probe follows the roles, not the words. The same nouns with no change in the sentence sit at chance. GPT-2 shows it a few layers later than the layer named in advance; Llama instruct falls one point short of the pre-registered bar at the named layer and clears it at an earlier one.

And the role is relational. "Write a haiku" names a thing asked for with nothing to start from, and the haiku is not marked as a goal. A goal is the far end of something.

Cartoon: a walker going from a barn to a river, and below it an empty bowl becoming a full one (poverty to wealth), with one magnifying glass reading both
One probe, trained on walks. PURPOSES ARE DESTINATIONS (Lakoff & Johnson, 1980).

6 · In progressWhat happens to a goal while the model carries it out?

In Llama-3.2-1B-Instruct, which form was asked for (haiku, limerick, list, email, joke, story) can be read perfectly from the last token before the answer begins, and swapping that state in from another prompt nudges the answer toward the other form. Then, subtracting what the text itself shows:

7 · Where this leaves thingsEmbodied-style structure, smaller than I claimed and still there

None of this shows the structure is more than the corpus's fingerprint. It shows the fingerprint is organised the way Lakoff said human thought is, and that a model reuses the shape of a walk for things that are not walks. Open: whether a goal is spent as it's reached; a layer sweep for the steering result; larger models.

Niamh McCombe, 2026. The research was conducted as a collaboration between the author and Claude (Anthropic) across many sessions. Every experiment since the audit was pre-registered, with odds, before it ran.

Earlier version of this page (September 2026), with corrections · code and pre-registrations