Suppose I offer you two banknotes.
Both have 95 printed in the corner.
One was issued by School A.
The other came from School B.
You are told they represent the same thing: excellent performance in Grade 12 mathematics.
Then you learn that students carrying the first note have historically lost eight points in first-year engineering, while students carrying the second have lost eighteen.
Do you continue to treat the notes as equal?
Or do you admit that high-school grades are not one currency at all?
Perhaps every school is issuing local currency with the same symbol.
Universities are left to discover the exchange rate.
The conversion is already happening
In the previous question, I wondered whether first-year university has become the expensive calibration we postponed in high school.
Waterloo Engineering has made part of that calibration explicit.
Its admissions FAQ says the adjustment factor is based on historical performance from a given school: the difference between students’ incoming high-school averages and their averages after first year at Waterloo Engineering, using data from the past six or more years.
The factor is only a small part of the decision. Waterloo says academic performance, its Admission Information Form and the online interview have more impact. (Waterloo Engineering admissions FAQ)
A simplified illustration:
The grade exchange
This can sound punitive.
It can also sound like basic measurement.
If two thermometers consistently produce different readings in the same room, pretending they are identical is not fairness.
It is an unwillingness to calibrate.
What exactly is being adjusted?
The careless interpretation is:
Waterloo knows which schools are easy.
The data supports a narrower statement:
Students from different schools have historically shown different average gaps between admission marks and first-year Waterloo marks.
That gap may reflect grading, curriculum alignment, access to advanced courses, assessment difficulty, study habits, applicant selection, small cohorts, socioeconomic differences, language, migration or the university environment itself.
The factor predicts a transition.
It does not reveal one clean cause.
Calling it an “inflation detector” would give the model more certainty than it has.
Why raw grades are not neutral
The usual objection is:
Just assess each student on their own merit.
I agree.
But a raw school grade is not the student separated from context. It is already a joint product:
What travels inside a grade
Ignoring systematic differences does not remove them.
It rewards whichever local system converts knowledge into the largest number.
That creates the same metric problem I keep finding elsewhere: when a proxy controls access to a scarce reward, the environment learns how to improve the proxy. (Why every metric eventually learns to lie)
An exchange rate is an attempt to recover the underlying value.
The danger is that the recovery mechanism becomes another proxy worth gaming.
A fairer exchange rate would admit uncertainty
Imagine a school sent three students to a program over six years.
Their first-year outcomes should not determine the fate of the fourth with the confidence of a school that sent three hundred.
A sensible model would shrink small-sample estimates toward a wider provincial or national baseline. The less evidence we have about a school, the less dramatic its special adjustment should be.
In plain language:
Do not pretend three former students have revealed the permanent character of a school.
The public Waterloo page says it uses historical data from six or more years, but it does not describe current minimum cohort sizes, statistical smoothing or fallback rules. That is not evidence that such safeguards are absent. It is a reason not to invent details.
If I were designing a transparent version, I would want a published purpose, visible uncertainty, regular updating, a neutral fallback for small schools, an appeal route, an equity audit and a modest role in the final decision.
This is less satisfying than one precise number.
It is more honest.
Why not adjust by subject and program?
A school may prepare students exceptionally well for writing and poorly for calculus.
Another may produce strong physics students but graduates who struggle with independent project work.
One school-wide exchange rate assumes the currency loses the same value in every market.
That seems unlikely.
A more relevant model would connect:
- the high-school subject;
- the university program or gateway course;
- the outcome we actually care about;
- enough historical data to support the comparison.
For engineering, calculus, physics and chemistry readiness may matter. For history, a timed analytical writing task may reveal more. Design could use a constrained portfolio exercise; nursing might use a situational judgment task.
The point is not to turn every application into a week-long audition.
It is to ask for direct evidence when the grade is an indirect proxy for the capability.
The non-STEM problem
Engineering offers relatively clear prerequisites and large first-year courses. The exchange-rate idea becomes harder where outcomes are interpretive.
What is the common unit for:
- making a historical argument;
- reading a difficult text;
- evaluating a source;
- developing an original design;
- listening to a patient;
- revising a weak draft?
Standardized multiple-choice testing would often destroy the very capability we wanted to observe.
But “difficult to standardize” does not mean “impossible to compare.”
Universities could use short, common performance tasks scored with clear rubrics:
- annotate a source and explain what would change your interpretation;
- revise a paragraph and describe the choices;
- compare two plausible arguments;
- solve an unfamiliar problem while showing the reasoning;
- respond to a realistic scenario with incomplete information.
These tasks would not replace school grades.
They would add a small amount of shared evidence where comparability matters most.
Transparency creates its own problem
If universities publish the exact conversion mechanism, schools may teach toward it.
If they keep it secret, students cannot evaluate whether it is fair.
This is not a reason to choose secrecy automatically.
We publish tax rules even though people adapt to them. We publish safety standards even though manufacturers design around them. The answer to gaming is not necessarily an unknowable system; it can be a measure that is broad enough, updated enough and direct enough that gaming it resembles doing the thing we value.
If the common task rewards reasoning on unfamiliar problems, preparing students for it may improve reasoning on unfamiliar problems.
That is a much healthier target than “make the local average larger.”
What universities are optimizing
Universities do not only want to identify potential. They also want likely success, manageable failure rates, efficient review, diverse cohorts, strong enrolment and public legitimacy.
Those goals conflict.
A statistically predictive factor can still be socially indefensible.
A perfectly individualized review can be too expensive to operate at scale.
A transparent process can expose uncomfortable differences universities would rather handle quietly.
This is why admissions systems often become a pile of compromises wearing a decimal point.
The student did not choose the issuer
The strongest objection remains.
A seventeen-year-old usually did not design their school’s grading policy.
Why should they lose admission points because previous students from that school struggled?
They should not be treated as the average of earlier cohorts.
But the alternative is not context-free individual judgment. If the only evidence is a locally produced grade, the school has already influenced the decision.
The humane response is to let the student produce additional evidence:
- a common readiness task;
- a strong performance in a relevant external course;
- a contextual statement;
- evidence of improvement;
- a diagnostic result that leads to a supported pathway rather than rejection.
Adjustment should change the university’s uncertainty.
It should not erase the student.
Perhaps the exchange should happen after admission
There is another possibility.
Use calibration less to decide who gets in and more to decide what happens next.
A student could be admitted conditionally to a program and complete a low-stakes diagnostic before classes. The result might lead to:
- a summer bridge course;
- a slower first-term sequence;
- targeted tutoring;
- a co-requisite workshop attached to calculus or writing;
- an early opportunity to demonstrate readiness again.
Now the exchange rate is not merely a discount applied to a teenager’s achievement.
It becomes a map of where support is likely to be needed.
That seems closer to the purpose of education.
It also leads to the most uncomfortable question in this series.
If a student reaches university with a visible gap, why did we wait until admission to notice it?
Did we eliminate failure—or just move it forward?
Until the next strange question,
Osagie