Last week, I found my daughter hunched over a stack of SAT prep books, frantically memorizing vocabulary words and solving practice math problems. The dining room table had become a war zone of flash cards, practice tests, and coffee-stained notebooks. As I watched her stress over a practice test that might shape her college future, I couldn't help but question whether this was really the best way to evaluate her potential.
"Do you think this actually measures how well I'll do in college?" she asked me one evening, frustrated after a practice test.
Her question stayed with me, and I recalled the news from around the time COVID hit, when more universities were making SAT scores optional for admissions. Since then, many colleges have reinstated the SAT as a requirement, but the question lingered. Was standardized tests truly measuring aptitude, or had it become something else entirely?
The SAT was originally designed to identify academic potential regardless of background. In theory, it would provide colleges with a standardized way to compare students from different high schools with varying grading standards. A noble goal.
But something happened along the way. As the SAT became increasingly important for college admissions, an entire industry grew around test preparation. Students began studying specifically for the test rather than focusing on genuine learning. Families with means hired tutors and purchased prep materials. Schools in wealthier districts offered SAT prep courses.
The result? The test began measuring something different: not just academic potential, but access to test preparation resources and strategies. This phenomenon is a perfect example of something economists and statisticians call "Goodhart's Law."
British economist Charles Goodhart observed in 1975: "When a measure becomes a target, it ceases to be a good measure."
"When a metric is used to evaluate performance, it turns into a target and loses its reliability as an indicator of the underlying reality it was meant to represent."
This simple insight explains why metrics often fail when we place too much emphasis on them. When we make a measurement important enough, people naturally focus on improving that specific number, often at the expense of the underlying goal the metric was supposed to track.
The SAT illustrates this perfectly. When colleges made SAT scores crucial for admissions, students and schools redirected their energy toward improving those specific numbers. The test no longer purely measured academic potential; it began measuring test-taking skill and preparation.
This pattern repeats across many domains. In each case, the metric was initially useful but became distorted when it became a target.
- Schools focus on standardized test scores, leading teachers to "teach to the test" rather than develop deeper understanding.
- Companies track call center metrics like call duration, so representatives rush customers off the phone.
- Hospitals evaluated on readmission rates might delay admitting patients who need care.
This problem affects organizations across industries. Let me share some real examples of how different organizations have encountered and addressed Goodhart's Law:
Google's OKR Framework
Google recognized early that single measurements could distort behavior. Their solution was to implement Objectives and Key Results (OKRs), setting multiple objectives with several key results for each. This prevents gaming any single metric while still providing clear direction. Instead of saying "increase user numbers" (which might lead to pursuing low-quality users), they might set multiple related metrics: user growth, engagement rates, and retention. This creates a more balanced approach to improvement.
Amazon's Customer Focus
Amazon focuses on customer-centric outcome metrics rather than internal process metrics. Instead of primarily measuring how efficiently their representatives handle calls (which might encourage rushing customers), they prioritize metrics like Customer Satisfaction Score and Order Defect Rate that directly reflect customer experiences. This keeps the focus on real customer outcomes rather than internal efficiency alone.
So how can we avoid these pitfalls when setting goals and measurements? Based on real-world examples, here are three approaches that address the issue.
Use Multiple Metrics Instead of Single Measures
Using several complementary metrics prevents the gaming of any single number. When metrics might conflict with each other, optimizing for one at the expense of others becomes difficult.
Focus on Outcome Measures Rather Than Process Measures
Measuring the actual results you want to achieve rather than the process to get there gives people flexibility to find the best approach while keeping the end goal clear.
Create a Culture That Emphasizes Underlying Goals
When people understand and believe in the broader purpose, they're less likely to game metrics at the expense of actual progress.
As my daughter continues preparing for her SATs, I've gained a new perspective. While the test still matters for many colleges, it's just one data point. What really matters is her genuine learning, critical thinking skills, and readiness for higher education.
The more thoughtful colleges seem to understand this too. Their move toward holistic admissions processes that consider course rigor, extracurricular commitment, essays, and recommendations alongside test scores, reflects an understanding that no single metric can capture a student's potential.
Perhaps the most important lesson is that we shouldn't confuse the measure with the goal. The goal isn't a high SAT score. It's developing the knowledge, skills, and character needed for success in college and beyond.