It is 7:42 p.m.
Your child is supposed to be getting ready for bed.
Instead, they are standing in the kitchen requesting a cookie with the moral urgency of someone seeking the release of a political prisoner.
You have several options.
You could give them the cookie.
This would immediately increase happiness in the house. Unfortunately, it may also teach them that delaying bedtime produces cookies.
You could refuse.
This protects the bedtime rule, but may initiate a forty-minute diplomatic crisis involving tears, floor-based protest and accusations that you have never allowed them to eat anything.
You could offer half a cookie.
This feels like compromise until you realize that you may have just taught them that sufficiently persistent negotiation produces half a cookie.
Congratulations.
You may think you are parenting.
But you appear to be running an optimization algorithm.
The suspicious similarity
Optimization sounds like something done by engineers, economists and people who describe buying groceries as a “resource-allocation problem.”
But the basic idea is ordinary:
You have an outcome you want, several actions you could take, limited resources and uncertainty about what each action will produce.
Parenting contains all of those ingredients.
You want your child to be healthy, curious, kind, resilient, independent, confident, disciplined and preferably asleep before you lose the will to continue.
You also have limited time, money, patience, information and energy.
Every day, you make hundreds of small decisions that may influence what happens next:
- Do I help or let them struggle?
- Do I reward this?
- Do I punish that?
- Do I insist?
- Do I negotiate?
- Do I explain?
- Do I pretend I did not hear it?
- Is this a teachable moment, or am I simply too tired for a teachable moment?
In the simplest possible model, parenting looks like this:
Choose the actions most likely to produce a good adult.
This seems reasonable.
It is also where the trouble begins.
Problem one: What exactly are we optimizing?
Suppose your child receives a difficult homework assignment.
You could help extensively and optimize for tonight’s result.
The work gets completed. The answer is correct. Everyone gets to bed.
But perhaps you have reduced the amount of struggle required for the child to learn persistence.
So you could refuse to help and optimize for independence.
Except the task may genuinely be beyond their current ability. Now they are not learning persistence. They are learning that difficult work feels lonely and impossible.
Maybe you should optimize for confidence.
But too much praise could make approval itself the goal.
Perhaps you should optimize for discipline.
But too much control may produce obedience in the room and poor judgment when you are no longer in it.
Parenting is not a single-objective problem.
It is a multi-objective optimization problem in which the objectives regularly conflict.
You are trying to increase:
- immediate safety,
- long-term independence,
- emotional security,
- competence,
- curiosity,
- social belonging,
- tolerance for frustration,
- moral judgment,
- physical health,
- and the probability that this person will still voluntarily call you when they are thirty-four.
There is no obvious formula for combining those things.
How many units of short-term happiness equal one unit of future resilience?
How much independence should be sacrificed for safety?
What is the exchange rate between achievement and peace of mind?
Anyone who gives you a precise answer is either selling a parenting course or has never met two different children.
Problem two: Children learn the reward function
Parents quickly discover that consequences shape behaviour.
Praise, attention, privileges, boundaries and natural consequences can all affect what children repeat. Reviews of parenting interventions have found that components such as positive reinforcement, praise and natural or logical consequences are associated with stronger effects in programs addressing disruptive behaviour.
This sounds like reinforcement learning:
- The child does something.
- The environment responds.
- The child updates their expectations.
- The behaviour becomes more or less likely.
Put away your toys, receive praise.
Hit your brother, lose screen time.
Scream in the supermarket, receive a snack because your parent has a meeting in twelve minutes.
The system learns.
Unfortunately, the system may learn something different from what the parent intended.
The parent thinks:
I gave her a snack to end this unusual emergency.
The child may learn:
Supermarket screaming has a surprisingly attractive return on investment.
This is sometimes called reward misspecification in AI: the designer rewards a measurable behaviour that is only an imperfect substitute for the result they actually wanted.
Parents do this constantly.
We reward grades when we mean to reward learning.
We reward quietness when we mean to encourage self-control.
We reward winning when we mean to encourage effort.
We reward obedience when we mean to teach judgment.
The child optimizes for the signal they actually receive, not the value the parent believed the signal represented.
This is the parenting version of an AI system finding a loophole in its objective function.
Except the AI system rarely asks for another cookie while maintaining direct eye contact.
Problem three: Short-term success can produce long-term failure
Imagine that a child dislikes reading.
A parent offers five dollars for every completed book.
Reading increases.
Optimization achieved.
But what exactly increased?
Perhaps the child discovered a love of stories.
Or perhaps the child discovered that thin books have excellent financial yield.
Research on extrinsic rewards and intrinsic motivation suggests that the effects of rewards depend on their form and context; expected, tangible rewards tied to task completion can sometimes reduce intrinsic motivation, particularly for activities a person already finds interesting.
That does not mean rewards are always harmful. It means the observable behaviour and the internal reason for doing it are not the same variable.
A parent may successfully optimize:
Number of books completed this month
while accidentally weakening:
Desire to read when nobody is paying.
The same problem appears everywhere.
A child can be optimized to say “sorry” without becoming sorry.
They can be optimized to share when an adult is watching.
They can learn the sentence “lying is wrong” while becoming extremely sophisticated at evidence management.
The easiest outcomes to measure are often not the outcomes we actually care about.
This is true in companies, schools, governments—and apparently kitchens.
Problem four: The environment will not stay still
Even if you discovered the perfect parenting strategy on Monday, your child may become a different person by Thursday.
What works at age three may fail at age seven.
What motivates one sibling may embarrass another.
What helps during a calm afternoon may collapse during hunger, illness, jealousy or the mysterious psychological event caused by a banana breaking in half.
In machine learning terms, the data distribution keeps shifting.
The child is developing.
The parent is changing.
The school changes.
Friends appear.
New fears emerge.
The family becomes busier.
A strategy is not applied to the same child repeatedly. It is applied to a moving person in a moving environment by another moving person.
This may be one reason rigid parenting formulas are so attractive and so unreliable. They offer the fantasy that a stable rule can control a changing system.
Problem five: The optimizer is inside the system
Most optimization problems do not argue back.
Parenting does.
The child’s behaviour changes the parent, just as the parent’s behaviour changes the child.
A difficult week may reduce a parent’s patience. Reduced patience may increase conflict. Increased conflict may worsen the child’s behaviour, which further reduces the parent’s patience.
Research on parent–child development increasingly treats the relationship as transactional rather than one-directional. Longitudinal studies have found bidirectional relationships between parenting stress and child behaviour problems: children affect parental responses, and parental states and behaviour affect children in return.
So the model is not:
Parent acts → child develops.
It is closer to:
Parent acts → child responds → parent adapts → child adapts to the adaptation → both become tired → somebody loses a shoe.
The child is not merely the object being optimized.
The child is another learning system, constantly forming theories about the parent:
- How serious is that warning?
- Does “five more minutes” really mean five?
- Which parent is more likely to approve this request?
- What happens if I ask while they are on the phone?
- Is the consequence still active if everyone forgets why it started?
Parents study children.
Children study parents.
Both run experiments.
Only one side usually admits it.
The danger of optimizing the wrong time horizon
One way to judge a parenting decision is:
Did it work?
But “work” needs a time horizon.
Did it work for the next five minutes?
The next week?
The next ten years?
Taking over a difficult task may end tonight’s frustration but reduce tomorrow’s competence.
Strict control may produce immediate compliance but leave the child dependent on external authority.
Allowing every natural consequence may produce independence, or it may simply make childhood unnecessarily dangerous.
Parents are continually choosing between immediate reward and delayed development.
And delayed outcomes are painfully difficult to observe.
If your child becomes thoughtful at thirty, which conversation caused it?
If they become anxious, which protection was excessive?
If they become resilient, was it because you allowed failure—or because they knew you would catch them afterward?
Parenting has one of the longest and noisiest feedback loops imaginable.
By the time some results appear, the experiment has been running for decades and there is no control group.
Siblings do not count. They will object strongly to the assumption that they received identical treatment.
Perhaps the real objective is not behaviour
There is a deeper problem with the optimization framing.
It encourages us to ask:
How do I make my child behave correctly?
But perhaps the better question is:
How does my child gradually learn to choose well without me?
That changes the objective.
The goal is no longer to produce the desired action every time. It is to help build the internal machinery that eventually produces good choices:
- judgment,
- emotional regulation,
- curiosity,
- empathy,
- confidence,
- and the ability to recover after getting things wrong.
Parental autonomy support—giving children meaningful choices, acknowledging their perspective and reducing unnecessary control—has been associated with children’s motivation and well-being.
This suggests that some parenting decisions should not optimize the child’s immediate action at all.
They should optimize the child’s capacity to act.
That may require tolerating temporary inefficiency.
A child choosing their own clothes may produce an objectively questionable combination.
A child pouring their own milk may create a secondary milk ecosystem beneath the table.
A child solving their own disagreement may take considerably longer than an adult simply issuing a ruling.
From the perspective of immediate performance, adult intervention is often superior.
From the perspective of learning, adult efficiency may be the problem.
So, is parenting an optimization problem?
Yes—but only in the way that running a country, sustaining a marriage or becoming a good person is an optimization problem.
The framing reveals something useful.
Parents really do make repeated decisions under uncertainty.
Rewards really do shape behaviour.
Short-term and long-term outcomes really do conflict.
Measured targets can replace the values they were meant to represent.
And strategies must adapt as the child and environment change.
But the framing also breaks down.
A child is not a product to be optimized.
There is no agreed objective function.
Many of the most important outcomes cannot be measured.
The system changes while we interact with it.
And the supposed target of optimization is a person with their own values, temperament and emerging goals.
Perhaps parenting is not the optimization of a child.
Perhaps it is the gradual optimization of an environment:
An environment safe enough for mistakes.
Structured enough for learning.
Warm enough for honesty.
Demanding enough for growth.
Flexible enough for individuality.
And stable enough that the child does not need to spend all their energy predicting whether love is still available.
The parent does not produce the finished person.
The parent helps shape the conditions under which that person learns to shape themselves.
Which is less satisfying than a formula.
But probably more true.
What I still cannot figure out
Optimization requires some definition of success.
Parents usually say they want their children to be “happy.”
But happiness today can conflict with capability tomorrow, and capability alone does not guarantee a meaningful life.
So what is the outcome parenting should ultimately optimize?
Happiness?
Independence?
Character?
Adaptability?
A secure relationship?
Or is choosing one final objective the first mistake?
Until the next strange question,
Osagie