What Should Be Measured, and What Should Not?
Philosophy17 April 2026Published by Pen & Muse

What Should Be Measured, and What Should Not?

Back to Dispatches
6 min read · 1,002 words
Dispatch SeriesPart 7 of 15
The Great Questions Series I — The Questions That Don’t Go Away

The answers you need are the questions that keep reappearing—because they govern your incentives, agency, and stakes.

Series PositionPart 7 of 15
What Should Be Measured, and What Should Not?
Trust That Holds: Consistency, Competence, and Intent

Previous · Part 6

Trust That Holds: Consistency, Competence, and Intent

The Slow Art of Preservation — What We Keep, What We Let Go, What We Evolve

Next · Part 8

The Slow Art of Preservation — What We Keep, What We Let Go, What We Evolve

This builds on Part 6: Trust That Holds: Consistency, Competence, and Intent

Continue with Part 8: The Slow Art of Preservation — What We Keep, What We Let Go, What We Evolve

What measurement really does

Measurement is never neutral. It doesn’t merely describe behaviour—it recruits it. The moment you count something, you signal what matters. People then do what gets counted, not what gets you closer to the outcome you actually care about.

So the question isn’t only what should we measure?
It’s: what are we willing to train our system to do?

The hidden bargain: definition over intention

Most measurement failures aren’t caused by bad maths. They’re caused by unspoken bargains.

A metric carries three embedded assumptions:

  • What you can measure well is what you’ll value.
  • What improves is what matters.
  • What you can’t track doesn’t meaningfully exist.

When those assumptions go unexamined, measurement becomes a kind of moral outsourcing. You stop asking whether the result is good—and start asking whether the number moved.

Measures vs. motivations

Some metrics align naturally with good behaviour. Others tug against it.

Consider two worlds:

  1. A metric that reflects the underlying goal
    Example: measuring defects to improve product quality.

  2. A metric that reflects a convenient proxy
    Example: measuring “time spent” to improve learning.

Proxies can be useful. But they always change incentives. When the proxy becomes the prize, the goal gets rearranged to fit the measurement.

A practical framework: measure the control points

Instead of asking “what can we measure?”, ask “where does behaviour get controlled?”

Measurement should be placed where it creates learning loops:

  • You can observe change.
  • You can attribute it to a decision or action.
  • You can adjust based on it.
  • You can verify improvement without relying on the same metric forever.
1
Define the outcome you truly want (the thing you’d fund even if no one could measure it).
2
Identify the smallest set of control points that influence that outcome.
3
Choose metrics that track movement through those control points, not just the final surface result.
4
Build a feedback loop: act → measure → learn → adjust.
5
Run a stress test: What behaviours would this metric reward unintentionally?

The two-level metric stack (so you don’t get fooled)

A robust measurement approach uses two layers:

  • Leading indicators: early signals that the outcome will improve.
  • Lagging indicators: what you ultimately want, measured later.

Lagging metrics tell you whether you succeeded. Leading metrics tell you whether you’re on the path. Using only one layer is how you get surprises—or slow self-deception.

Where leading indicators usually go wrong

Leading indicators can become “mini-goals.” Teams then optimise the lead without fixing the root.

So for every leading metric, ask:

  • What would have to be true in the real world for this to improve?
  • What behaviours would inflate the metric without creating the truth?
  • How would we detect that inflation early?

Avoid the “optimise-to-numb” trap

Some things shouldn’t be optimised because optimisation kills the value you were trying to preserve. (And yes, this is different from “don’t measure.” It’s “measure carefully, or measure around it.”)

Examples of values that resist simple optimisation:

  • Trust
  • Safety
  • Creativity
  • Fairness
  • Learning

In these domains, measurement can still help—if it supports judgement rather than replacing it.

Measurement as judgement: keep humans in the loop

Not every important variable needs a number. Many do better with structured judgement: audits, peer review, post-mortems, qualitative checks, and red-team perspectives.

A mature measurement culture uses numbers to inform judgement—not to replace it.

A decision rule: measure outcomes you can protect

Here’s a clean filter you can use when deciding whether to measure something:

Will measuring this increase your ability to protect the outcome?
Or will it merely increase your ability to perform the outcome?

If measurement makes it easier to protect quality, safety, integrity, and long-term value—great.
If it mostly makes it easier to manage optics, throughput, and short-term compliance—beware.

The anti-fragile scoreboard

Good measurement systems are resilient to the team’s creativity. They don’t just ask “did we hit the target?” They ask “did we hit the target for the right reasons?”

That implies:

  • Multiple metrics (not one headline number)
  • Counterfactual thinking (“what else changed?”)
  • Periodic metric audits
  • Explicit definitions and failure modes

Diagram: Outcome leads to Control Points; Control Points leads to Leading Indicators; Control Points leads to Mechanisms to Protect; Leading Indicators leads to Actions & Experiments; Mechanisms to Protect leads to Qualitative Checks / Audits; Actions & Experiments leads to Lagging Indicators; Lagging Indicators leads to Review & Adjustment; Qualitative Checks / Audits leads to Review & Adjustment; Review & Adjustment leads to Control Points.

Diagram: Outcome leads to Control Points; Control Points leads to Leading Indicators; Control Points leads to Mechanisms to Protect; Leading Indicators leads to Actions & Experiments; Mechanisms to Protect leads to Qualitative Checks / Audits; Actions & Experiments leads to Lagging Indicators; Lagging Indicators leads to Review & Adjustment; Qualitative Checks / Audits leads to Review & Adjustment; Review & Adjustment leads to Control Points.

Measurement governance: the part people skip

The most overlooked piece is who decides what gets measured, and how often the system is reconsidered.

A measurement strategy decays unless it is owned.

Minimum governance habits:

  • Version your metrics (definitions change; behaviour follows)
  • Set review intervals (metrics don’t get to live forever)
  • Separate roles (those who measure aren’t always the ones who benefit)
  • Publish the “why” behind each metric

What to do this week

Choose one metric you currently rely on—especially one that shapes behaviour daily. Then run the diagnostic.

Checklist0/6
Metric Diagnostic (copy/paste)

Metric:
Outcome it protects:
Control points it influences:
What would success actually require in the real world?
Most likely gaming/failure modes:
Leading indicator(s):
Lagging indicator(s):
Judgement check(s) we’ll run:
Review cadence:

1
Pick one existing metric.
2
Complete the diagnostic.
3
Decide: keep, replace, or supplement.
4
If you keep it, add a protection: a judgement check or a paired metric.
5
Implement the change and observe behaviour for a full cycle.

If this resonates, see how to apply it to your own work with the interactive Dispatch agent.

Be first to like this dispatch

More in PhilosophyView all →
Keep Reading, Then Step Inside
Cover of Dead Reckoning
Book + Immersive Experience£3.99

Dead Reckoning

Pen & Muse

## Dead Reckoning — Season One You come to Threshwick to do what you always do: read the accounts, write the report, recommend a path out of a failing marina. It should be routine. Professional. Contained. But the books don’t balance. Offshore companies pay for work that never happens. Boats slip in without logs. Money moves cleanly around something no one will name. And the family who runs this place aren’t villains — just people who made one decision in a bad year and have been living inside its math ever since. In *Dead Reckoning*, you don’t watch the story; you become the person whose signature makes it real. Every conversation shifts your standing. Every discovery changes what you owe. Every choice pulls you closer to the truth — or deeper into the arrangement. You are not a detective. You are not a criminal. You are a professional paid to make the numbers make sense. Until you understand the real brief: You weren’t hired to find the truth. You were hired to make sure the truth balances. And now you have to decide what your signature is worth. --- ### What This Experience Is *Dead Reckoning* is a live, choice-driven coastal thriller where you work inside a closed town, a family with something to protect, and a network that already knows your name. There is no fixed path and no safe neutral. What matters is what you learn, what you can prove, and how long you can keep doing your job before someone decides for you.

Platform Access

Interested in building narratives using our proprietary architecture? Join the creator waitlist.