Late one August, during my first week in Canada, I was trying to find my first Master’s class in Computer Science.

I had arrived in Saskatoon from Nigeria and was still learning the campus.

My timetable knew exactly when the lecture started. It had shown considerably less interest in helping me find the building.

I was lost.

There were plenty of people on the path, and I could have asked the nearest one.

I did not.

Instead, I ran a tiny, completely unapproved classification model over the crowd:

Which face has the highest probability of helping a confused new international student with an accent?

I looked for the warmest expression.

Someone smiled. I selected them.

I approached openly, they responded warmly, and their directions got me to class on time.

By that evening, my model was one-for-one.

An impeccable sample size if the goal is to become confident much faster than the evidence permits.

But on the walk back to my apartment, I began wondering what exactly the interaction had proved.

Perhaps I had identified a helpful person. Facial expressions are information, and the simplest explanation may be that someone who looked friendly was friendly. But perhaps my expectation had also entered the interaction.

I had approached this person with relief and openness because I already expected patience. They encountered someone treating them as safe and helpful, then responded in kind.

Did I observe their friendliness?

Or did we partly create the evidence together?

This was not yet a stereotype in the strict sense. I had made a quick prediction from one visible cue.

But it revealed a harmless version of the loop:

expectation → selection → approach → response → stronger expectation

Now attach the same mechanism to a heavier label.

A teacher thinks one child is difficult.

Not dangerous. Not incapable. Just difficult.

So she watches him a little more closely.

When he whispers, she notices. When another child whispers, she is helping a friend. When he questions an instruction, he is challenging authority. When another child does it, she is curious.

The boy notices the surveillance.

He becomes defensive. The teacher intervenes earlier. He stops asking for help because asking now feels like entering a courtroom in which the verdict has already been typed.

By June, the teacher has an impressive collection of evidence.

The child really did argue more. He really did withdraw. He really did make the classroom harder to manage.

What neither of them can easily see is that the belief may not have remained only a description.

It became part of the environment.

A belief can enter the thing it predicts

Some predictions are passive.

I can predict rain without making the cloud nervous.

Predictions about people are different. People detect expectations, respond to treatment and adapt to the roles available to them.

This creates the possibility of a loop:

expectation → treatment → response → interpretation → stronger expectation

The original belief need not be completely false. The child may be impulsive, the colleague awkward or the department chronically late.

The question is whether the expectation merely observed the pattern—or helped deepen it.

Expectations alter treatment

Teachers, managers and parents distribute small resources all day:

  • patience;
  • eye contact;
  • difficult assignments;
  • freedom;
  • follow-up questions;
  • benefit of the doubt;
  • warmth after a mistake.

Those differences can affect performance.

Research on teacher expectations finds self-fulfilling effects, though they are more conditional than the popular “believe in a child and their IQ will soar” version. A North Carolina study associated higher, plausibly exogenous teacher expectations with higher later test scores; reviews suggest effects vary by teacher, student and context. (Hill & Jones, 2021; Rubie-Davies & Hattie, 2024)

The careful conclusion is not that expectations determine destiny.

It is that expectations sometimes change the conditions under which ability is expressed and developed.

The same behaviour can arrive with two biographies

Imagine a confident statement in a meeting.

From the person already considered leadership material, it is decisive.

From the person considered difficult, it is aggressive.

Or consider a quiet employee.

If we think they are thoughtful, the silence contains depth.

If we think they are weak, the same silence contains nothing.

The behaviour is ambiguous. The stereotype supplies a story.

Once supplied, the story directs memory. Stereotype-consistent events become examples. Inconsistent events become exceptions:

He handled that client well, but that was an easy account.

She wrote excellent code, but she had help.

The teenager arrived on time, but only because we reminded him.

The stereotype owns the general rule. Reality is permitted occasional guest appearances.

Then the target adapts

People do not always respond to expectations by becoming them.

They may:

  • withdraw from the setting;
  • conceal parts of themselves;
  • work twice as hard;
  • resist the authority imposing the label;
  • perform exaggerated competence;
  • or seek a group in which the same trait is interpreted differently.

Each response can be absorbed by the belief. Withdrawal proves they lack commitment. Resistance proves they are difficult. Overperformance proves the standard was useful. Leaving proves they were never a fit.

A strong stereotype is not merely a prediction. It is an explanation generator. Whatever happens, it finds a way to remain employed.

What about stereotype threat?

One possible mechanism is that making a negative group stereotype salient changes performance in the moment. Anxiety, self-monitoring or working-memory demands may consume resources that would otherwise go to the task.

This idea became famous as stereotype threat.

It also became more certain in public conversation than the evidence permits.

The literature contains influential experiments and serious concerns about publication bias and boundary conditions. A meta-analysis of girls’ mathematics performance found a small average effect but warned that publication bias may have inflated it; a large preregistered UK replication found no evidence for the tested effect. (Flore & Wicherts, 2015; Inglis & O’Hagan, 2022)

So I would not use stereotype threat as a universal key.

Effects may depend on whether a person knows the stereotype, identifies with the domain, expects evaluation, finds the task difficult and interprets the cue as relevant.

Context is not a footnote here.

Context is the proposed mechanism.

A small experiment I want to run

I keep wondering what happens if someone who already lacks confidence in mathematics must say:

“I am not good at mathematics.”

immediately before solving a problem.

Perhaps the statement activates a negative identity and redirects attention toward anticipated failure.

Perhaps nothing happens.

Perhaps saying it actually reduces pressure because the person no longer has to hide.

That is why this belongs in an experiment, not in a motivational poster.

A reasonable design would measure mathematics confidence, then randomly assign low-confidence participants to a negative self-description, neutral verbalization, positive reframe or no statement before the same task. The self-label question deserves its own essay because it is narrower than the social loop here.

For now, the important point is that a belief can affect a person through several routes at once:

  1. it changes how others treat them;
  2. it changes how ambiguous behaviour is interpreted;
  3. it changes what gets remembered;
  4. and it may become psychologically active for the person being judged.

But sometimes the stereotype is detecting a pattern

This is the uncomfortable counterargument, and avoiding it would make the essay useless.

Groups can differ on measured outcomes. Individuals can develop stable reputations for reasons. Teachers often know their students better than an outside observer does.

Not every expectation is prejudice.

Not every confirmed prediction is self-fulfilling.

The harder problem is causal untangling. If an expectation was partly accurate and partly productive, how much of the later evidence belongs to the person, how much to the treatment and how much to the situation they built together?

Removing the label may not remove the original difference.

Keeping it may prevent us from discovering how much could change.

What I still cannot figure out

The cleanest test would compare the same person under different expectations.

Life rarely gives us that control group.

By the time a stereotype has operated for years, treatment and adaptation have become part of the evidence. The loop has written itself into the world.

When a stereotype appears to predict behaviour, how can we tell whether it observed the behaviour—or helped produce it?

Until the next strange question,

Osagie