While visiting family in Edmonton, I watched bedtime turn into a kitchen negotiation. It was 7:42 p.m. The child was supposed to be getting ready for bed and was instead requesting a cookie with the moral urgency of someone seeking the release of a political prisoner.

I watched the parent perform a calculation that probably happens in millions of homes every evening.

Give them the cookie and happiness rises immediately. Bedtime may even happen without a forty-minute diplomatic crisis.

But tomorrow, the child may remember:

Delaying bedtime sometimes produces cookies.

Refuse and the boundary remains intact. But perhaps the child is tired, hungry and operating without the emotional infrastructure required for a principled lecture on consistency.

Offer half a cookie and everybody gets a compromise—and learns that persistent negotiation produces half a cookie.

At first, this looked like a parent trying to choose the option with the lowest cost.

Then I realized there was no single cost.

The parent was trying to preserve sleep tonight, the boundary tomorrow, the child’s trust, the household’s peace and whatever patience remained after 7 p.m.

The child had a clearer objective:

Cookie now.

That did not make the child irrational.

It meant the two people were optimizing across different time horizons.

And every decision changed the next decision. Giving in might alter tomorrow’s negotiation. Holding the boundary might teach predictability—or simply overwhelm an exhausted child who was no longer capable of learning the intended lesson.

The “loss function” was not sitting still long enough to be solved.

That was when I began wondering:

Is parenting an optimization problem?

The suspicious similarity

Optimization sounds like something done by engineers, economists and people who describe buying groceries as a “resource-allocation problem.”

But the basic idea is ordinary:

You have an outcome you want, several actions you could take, limited resources and uncertainty about what each action will produce.

Parenting contains all of those ingredients.

You want your child to be healthy, curious, kind, resilient, independent, confident, disciplined and preferably asleep before you lose the will to continue.

You also have limited time, money, patience, information and energy.

Every day brings small decisions that may influence what happens next:

  • Do I help or let them struggle?
  • Do I reward this?
  • Do I punish that?
  • Do I insist, negotiate or explain?
  • Is this a teachable moment, or am I too tired for a teachable moment?

In the simplest possible model, parenting looks like this:

Choose the actions most likely to produce a good adult.

This seems reasonable.

It is also where the trouble begins.

Problem one: What exactly are we optimizing?

Suppose your child receives a difficult homework assignment.

You could help extensively and optimize for tonight: correct answer, completed work, everyone in bed. But perhaps you remove the struggle through which persistence develops.

You could refuse and optimize for independence. Except the task may genuinely be beyond the child’s ability; now they are learning that difficult work feels lonely and impossible. Optimize for confidence and approval may become the goal. Optimize for discipline and you may produce obedience in the room but poor judgment when you leave it.

Parenting is not a single-objective problem.

It is a multi-objective optimization problem in which the objectives regularly conflict.

You are trying to increase:

  • immediate safety and long-term independence;
  • emotional security and tolerance for frustration;
  • curiosity, competence and moral judgment;
  • and the probability that this person will still voluntarily call you when they are thirty-four.

There is no obvious formula for combining those things.

How many units of short-term happiness equal one unit of future resilience?

How much independence should be sacrificed for safety?

What is the exchange rate between achievement and peace of mind?

Anyone who gives you a precise answer is either selling a parenting course or has never met two different children.

Problem two: Children learn the reward function

Parents quickly discover that consequences shape behaviour.

Praise, attention, privileges and boundaries can all affect what children repeat. A meta-analysis of 77 parent-training evaluations for families with children aged zero to seven found larger effects associated with components including positive parent–child interaction, emotional communication, consistent discipline and practising new skills during the program. It did not establish one universal set of techniques for every child or outcome. (Kaminski et al., 2008)

This sounds like reinforcement learning:

  1. The child does something.
  2. The environment responds.
  3. The child updates their expectations.
  4. The behaviour becomes more or less likely.

Put away your toys, receive praise. Hit your brother, lose screen time. Scream in the supermarket, receive a snack because your parent has a meeting in twelve minutes.

The system learns.

Unfortunately, the system may learn something different from what the parent intended.

The parent thinks:

I gave her a snack to end this unusual emergency.

The child may learn:

Supermarket screaming has a surprisingly attractive return on investment.

This is sometimes called reward misspecification in AI: the designer rewards a measurable behaviour that is only an imperfect substitute for the result they actually wanted. It is the same problem explored in why every metric eventually learns to lie: once the proxy becomes the target, the system learns to produce the proxy.

Parents do this constantly. We reward grades when we mean learning, quietness when we mean self-control, winning when we mean effort and obedience when we mean judgment.

The child optimizes for the signal they actually receive, not the value the parent believed the signal represented.

This is the parenting version of an AI system finding a loophole in its objective function.

Except the AI system rarely asks for another cookie while maintaining direct eye contact.

Problem three: Short-term success can produce long-term failure

Imagine that a child dislikes reading.

A parent offers five dollars for every completed book.

Reading increases. Optimization achieved. But what increased—a love of stories, or the discovery that thin books have excellent financial yield?

Research on extrinsic rewards and intrinsic motivation suggests that the effects of rewards depend on their form and context. A meta-analysis of 128 experiments found that expected tangible rewards tied to engagement, completion or performance reduced later free-choice engagement on average, while positive feedback increased it. (Deci, Koestner & Ryan, 1999)

That does not mean rewards are always harmful. It means the observable behaviour and the internal reason for doing it are not the same variable.

A parent may successfully optimize:

Number of books completed this month

while accidentally weakening:

Desire to read when nobody is paying.

The same problem appears everywhere. A child can be optimized to say “sorry” without becoming sorry, to share only when watched, or to condemn lying while becoming sophisticated at evidence management.

The easiest outcomes to measure are often not the outcomes we actually care about.

This is true in companies, schools, governments—and apparently kitchens.

Problem four: The system moves—and argues back

Even if you discover the perfect strategy on Monday, your child may become a different person by Thursday. What works at three may fail at seven. What motivates one sibling may embarrass another. A calm-afternoon rule may collapse during hunger, jealousy or the mysterious psychological event caused by a banana breaking in half.

The child, parent, school and family are all changing. A strategy is never applied to the same child by the same parent in the same environment. This is why rigid formulas are so attractive and so unreliable: they offer the fantasy that a stable rule can control a moving system.

The system also changes the optimizer. A difficult week reduces patience; reduced patience increases conflict; conflict changes the child’s behaviour, which further reduces patience. Research increasingly treats parent–child development as transactional rather than one-directional. A longitudinal study found bidirectional associations between parenting stress and children’s problems across several ages, though group-level relationships cannot diagnose one difficult week. (Chiang & Bai, 2023)

The model is not:

Parent acts → child develops.

It is closer to:

Parent acts → child responds → parent adapts → child adapts to the adaptation → both become tired → somebody loses a shoe.

The child is another learning system, constantly testing how serious the warning is, whether “five more minutes” means five and which parent is more likely to approve the request.

Parents study children. Children study parents. Both run experiments. Only one side usually admits it.

The danger of optimizing the wrong time horizon

One way to judge a parenting decision is:

Did it work?

But “work” needs a time horizon.

Did it work for five minutes, a week or ten years?

Taking over a difficult task may end tonight’s frustration but reduce tomorrow’s competence.

Strict control may produce immediate compliance but leave the child dependent on external authority.

Allowing every natural consequence may produce independence, or it may simply make childhood unnecessarily dangerous.

Parents are continually choosing between immediate reward and delayed development.

And delayed outcomes are painfully difficult to observe.

If your child becomes thoughtful at thirty, which conversation caused it? If they become resilient, was it because you allowed failure—or because they knew you would catch them afterward?

Parenting has one of the longest and noisiest feedback loops imaginable.

By the time some results appear, the experiment has been running for decades and there is no control group.

Siblings do not count. They will object strongly to the assumption that they received identical treatment.

Perhaps the real objective is not behaviour

There is a deeper problem with the optimization framing.

It encourages us to ask:

How do I make my child behave correctly?

But perhaps the better question is:

How does my child gradually learn to choose well without me?

That changes the objective.

The goal is no longer to produce the desired action every time. It is to help build the internal machinery that eventually produces good choices:

  • judgment and emotional regulation;
  • curiosity, empathy and confidence;
  • the ability to recover after getting things wrong.

Parental autonomy support—giving children meaningful choices, acknowledging their perspective and reducing unnecessary control—has been associated with children’s motivation and well-being. A meta-analysis of the pathways to student motivation found positive links between parental autonomy support and autonomous motivation, while also emphasizing that much of the evidence is correlational and that associations vary by age and form of motivation. (Bureau et al., 2022)

This suggests that some parenting decisions should not optimize the child’s immediate action at all.

They should optimize the child’s capacity to act.

That may require tolerating temporary inefficiency.

A child choosing clothes may produce an objectively questionable combination. Pouring milk may create a secondary ecosystem beneath the table. Solving a disagreement may take longer than an adult issuing a ruling.

From the perspective of immediate performance, adult intervention is often superior.

From the perspective of learning, adult efficiency may be the problem.

So, is parenting an optimization problem?

Yes—but only in the way that running a country, sustaining a marriage or becoming a good person is an optimization problem.

The framing reveals something useful.

Parents really do make repeated decisions under uncertainty.

Rewards really do shape behaviour.

Short-term and long-term outcomes really do conflict.

Measured targets can replace the values they were meant to represent.

And strategies must adapt as the child and environment change.

But the framing also breaks down.

A child is not a product to be optimized.

There is no agreed objective function.

Many of the most important outcomes cannot be measured.

The system changes while we interact with it.

And the supposed target of optimization is a person with their own values, temperament and emerging goals.

Perhaps parenting is not the optimization of a child.

Perhaps it is the gradual optimization of an environment:

An environment safe enough for mistakes.

Structured enough for learning.

Warm enough for honesty.

Demanding enough for growth.

Flexible enough for individuality.

And stable enough that the child does not need to spend all their energy predicting whether love is still available.

The parent does not produce the finished person.

The parent helps shape the conditions under which that person learns to shape themselves.

Which is less satisfying than a formula.

But probably more true.

What I still cannot figure out

Optimization requires some definition of success.

Parents usually say they want their children to be “happy.”

But happiness today can conflict with capability tomorrow, and capability alone does not guarantee a meaningful life.

So what is the outcome parenting should ultimately optimize?

Happiness?

Independence?

Character?

Adaptability?

A secure relationship?

Or is choosing one final objective the first mistake?

Until the next strange question,

Osagie