Someone begins a familiar sentence:

“I’m not angry, but…”

Your brain has already written several possible endings.

None involves a calm discussion of local weather.

People do not wait passively for words or events to arrive. We anticipate the next sentence, the mood of a room, the outcome of a meeting and whether the footsteps behind us are getting closer.

Then reality differs.

The supposedly calm person laughs.

The difficult meeting ends early.

The footsteps turn into a doorway.

What does the mind do with the error?

Prediction is part of perception

One influential family of theories describes the brain as a predictive system.

Instead of constructing experience only from incoming sensation, the brain may generate expectations about the causes of sensation and compare them with what arrives. Differences—prediction errors—help adjust perception, action or the model.

This family includes predictive coding, predictive processing and the free-energy principle. The ideas overlap but are not identical, and proposals that they explain almost everything the brain does remain ambitious and debated. (Friston, 2010)

The modest claim is enough for me:

Prior context helps the mind anticipate what comes next.

This is visible in language.

“Peanut butter and…” makes one word far more likely than most of the dictionary.

It is visible socially.

The colleague who writes “Quick question” at 4:57 p.m. activates a model built from previous suffering.

And it is visible in perception. Expectations help resolve noisy or incomplete input.

The world does not enter a blank system.

It enters an argument already in progress.

The autoregressive analogy

Large language models generate text by predicting a next token from prior context, then using the generated token as context for the next prediction.

Humans also anticipate language.

That makes the analogy useful for about six minutes.

Humans do not merely predict the next word.

We predict:

  • what the speaker intends;
  • whether they are joking;
  • what our response will cost;
  • how the relationship may change;
  • and what kind of situation we are in.

We can stop the sentence, change the subject, refuse the premise or decide silence is morally preferable.

An autoregressive model is a flashlight on one feature of human cognition, not a blueprint of a person.

Consequences teach too

Prediction is not the whole story.

Behaviour changes through consequences.

A joke receives laughter.

A child cries and receives attention.

An employee disagrees and is excluded from the next meeting.

A person begins learning not only what happens, but which actions are worth repeating.

Reinforcement learning provides a useful formal language: an agent takes actions, receives outcomes and updates expected value. A key signal is reward prediction error—the difference between an expected and received outcome.

Classic neuroscience research found patterns in dopamine-neuron activity in monkeys consistent with errors in predicted reward: unexpected reward increased firing; once a cue predicted the reward, activity shifted toward the cue; omission produced a decrease. (Schultz, Dayan & Montague, 1997)

This does not mean dopamine is simply “pleasure,” or that human life is one giant points programme.

It shows that prediction and consequence meet inside biological learning.

The two systems overlap

Suppose a new employee offers a critical opinion.

The room becomes quiet.

Several things can be learned.

A predictive model updates:

This group treats disagreement as danger.

An action value updates:

Speaking openly has a cost here.

Next time, the person predicts the silence and withholds the comment.

The absence of disagreement then allows the manager to predict:

The team agrees with me.

One person’s learning environment becomes another person’s evidence.

Working model—human interpretation runs through every step

The same outcome can teach different lessons when people assign it different meaning.

This is how culture can be reinforced without an explicit rule. Nobody writes “Do not challenge leadership.” The consequences teach it more efficiently.

Rewards are interpreted, not merely received

Simple reinforcement-learning analogies become fragile here.

Humans argue with rewards.

A child can decide praise is manipulative.

An employee can accept punishment as evidence of integrity.

A person can endure short-term pain because of a moral promise, imagined future or identity.

We also learn socially. Someone else’s punishment can update our behaviour. A story about an ancestor can shape a choice. Culture can make sacrifice rewarding and applause suspicious.

The outcome is not a scalar delivered from the environment.

It is interpreted by a person with language, embodiment, history and commitments.

Two people can receive the same promotion.

One learns:

My work matters.

The other:

They are buying my silence.

Same event. Different model. Different reward.

A provisional synthesis

I currently find this description useful:

Humans are predictive world-model builders operating inside reinforcement loops.

The prediction system estimates what is happening and what may follow.

Consequences influence which actions, interpretations and goals become more likely.

But the two cannot be cleanly separated.

What counts as a reward depends on the model.

What the model predicts depends on the history of reward, punishment and social meaning.

Attention connects them. We cannot update from every event, so some outcomes receive weight and others disappear into noise.

That leads to the harder question.

What I still cannot figure out

When the world violates a prediction, surprise is not guaranteed to produce learning.

We can notice an error and call it luck, manipulation, an exception or somebody else’s fault.

Are we learning from reality—or mostly learning to predict the rewards and punishments created by the people around us?

Until the next strange question,

Osagie

My current hunch: Prediction and reinforcement are intertwined: we forecast the world, act inside it and let interpreted consequences reshape both behaviour and expectation.

Most likely reason I am wrong: The computational analogy may compress distinct neural, social and reflective processes until the description becomes elegant but difficult to test.