The answers you need are the questions that keep reappearing—because they govern your incentives, agency, and stakes.


Previous · Part 6
Trust That Holds: Consistency, Competence, and Intent

Next · Part 8
The Slow Art of Preservation — What We Keep, What We Let Go, What We Evolve
This builds on Part 6: Trust That Holds: Consistency, Competence, and Intent
Continue with Part 8: The Slow Art of Preservation — What We Keep, What We Let Go, What We Evolve
What measurement really does
Measurement is never neutral. It doesn’t merely describe behaviour—it recruits it. The moment you count something, you signal what matters. People then do what gets counted, not what gets you closer to the outcome you actually care about.
So the question isn’t only what should we measure?
It’s: what are we willing to train our system to do?
The hidden bargain: definition over intention
Most measurement failures aren’t caused by bad maths. They’re caused by unspoken bargains.
A metric carries three embedded assumptions:
- What you can measure well is what you’ll value.
- What improves is what matters.
- What you can’t track doesn’t meaningfully exist.
When those assumptions go unexamined, measurement becomes a kind of moral outsourcing. You stop asking whether the result is good—and start asking whether the number moved.
Measures vs. motivations
Some metrics align naturally with good behaviour. Others tug against it.
Consider two worlds:
-
A metric that reflects the underlying goal
Example: measuring defects to improve product quality. -
A metric that reflects a convenient proxy
Example: measuring “time spent” to improve learning.
Proxies can be useful. But they always change incentives. When the proxy becomes the prize, the goal gets rearranged to fit the measurement.
A practical framework: measure the control points
Instead of asking “what can we measure?”, ask “where does behaviour get controlled?”
Measurement should be placed where it creates learning loops:
- You can observe change.
- You can attribute it to a decision or action.
- You can adjust based on it.
- You can verify improvement without relying on the same metric forever.
The two-level metric stack (so you don’t get fooled)
A robust measurement approach uses two layers:
- Leading indicators: early signals that the outcome will improve.
- Lagging indicators: what you ultimately want, measured later.
Lagging metrics tell you whether you succeeded. Leading metrics tell you whether you’re on the path. Using only one layer is how you get surprises—or slow self-deception.
Where leading indicators usually go wrong
Leading indicators can become “mini-goals.” Teams then optimise the lead without fixing the root.
So for every leading metric, ask:
- What would have to be true in the real world for this to improve?
- What behaviours would inflate the metric without creating the truth?
- How would we detect that inflation early?
Avoid the “optimise-to-numb” trap
Some things shouldn’t be optimised because optimisation kills the value you were trying to preserve. (And yes, this is different from “don’t measure.” It’s “measure carefully, or measure around it.”)
Examples of values that resist simple optimisation:
- Trust
- Safety
- Creativity
- Fairness
- Learning
In these domains, measurement can still help—if it supports judgement rather than replacing it.
Measurement as judgement: keep humans in the loop
Not every important variable needs a number. Many do better with structured judgement: audits, peer review, post-mortems, qualitative checks, and red-team perspectives.
A mature measurement culture uses numbers to inform judgement—not to replace it.
A decision rule: measure outcomes you can protect
Here’s a clean filter you can use when deciding whether to measure something:
Will measuring this increase your ability to protect the outcome?
Or will it merely increase your ability to perform the outcome?
If measurement makes it easier to protect quality, safety, integrity, and long-term value—great.
If it mostly makes it easier to manage optics, throughput, and short-term compliance—beware.
The anti-fragile scoreboard
Good measurement systems are resilient to the team’s creativity. They don’t just ask “did we hit the target?” They ask “did we hit the target for the right reasons?”
That implies:
- Multiple metrics (not one headline number)
- Counterfactual thinking (“what else changed?”)
- Periodic metric audits
- Explicit definitions and failure modes
Diagram: Outcome leads to Control Points; Control Points leads to Leading Indicators; Control Points leads to Mechanisms to Protect; Leading Indicators leads to Actions & Experiments; Mechanisms to Protect leads to Qualitative Checks / Audits; Actions & Experiments leads to Lagging Indicators; Lagging Indicators leads to Review & Adjustment; Qualitative Checks / Audits leads to Review & Adjustment; Review & Adjustment leads to Control Points.
Diagram: Outcome leads to Control Points; Control Points leads to Leading Indicators; Control Points leads to Mechanisms to Protect; Leading Indicators leads to Actions & Experiments; Mechanisms to Protect leads to Qualitative Checks / Audits; Actions & Experiments leads to Lagging Indicators; Lagging Indicators leads to Review & Adjustment; Qualitative Checks / Audits leads to Review & Adjustment; Review & Adjustment leads to Control Points.
Measurement governance: the part people skip
The most overlooked piece is who decides what gets measured, and how often the system is reconsidered.
A measurement strategy decays unless it is owned.
Minimum governance habits:
- Version your metrics (definitions change; behaviour follows)
- Set review intervals (metrics don’t get to live forever)
- Separate roles (those who measure aren’t always the ones who benefit)
- Publish the “why” behind each metric
What to do this week
Choose one metric you currently rely on—especially one that shapes behaviour daily. Then run the diagnostic.
Metric:
Outcome it protects:
Control points it influences:
What would success actually require in the real world?
Most likely gaming/failure modes:
Leading indicator(s):
Lagging indicator(s):
Judgement check(s) we’ll run:
Review cadence:
If this resonates, see how to apply it to your own work with the interactive Dispatch agent.
Be first to like this dispatch



