Public agencies cannot rely on profit to tell them whether they are succeeding. They need measures that reflect their mission: what they produce, how efficiently and correctly they produce it, and whether it changes anything for the people they serve.
This guide introduces the logic model, shows how to choose and define measures, explains how to pair them so they cannot be gamed, and walks through a scorecard for a permit office. Legal reporting requirements vary by jurisdiction, so check the rules that apply to your agency.
Before You Start
Why Public-Sector Measurement Is Different
No Profit Signal
A business can tell whether it is succeeding by its margin. A public agency must define success through mission results, service, and cost.
Many Stakeholders
Residents, legislators, oversight bodies, and staff may each want different measures and different definitions of success.
Outcomes Are Slow and Shared
Public safety or health results depend on many actors and take years to appear, so agencies measure earlier outputs and processes as well.
Measures Change Behavior
A number that is reported and rewarded gets attention, including gaming. Design measures with that in mind.
The Logic Model: Inputs to Impact
A logic model links what an agency puts in to what it produces and what changes as a result. Measures can be placed at each stage.
| Type of measure | Question it answers | Example (permit office) |
|---|---|---|
| Input | What resources are used? | Staff hours, budget |
| Activity or process | How is the work done? | Applications reviewed per week, review time |
| Output | What was produced? | Permits issued, inspections completed |
| Efficiency | How much output per resource? | Cost per permit, permits per reviewer |
| Quality | Was it done correctly? | First-time-complete rate, errors found at inspection |
| Outcome | What changed for people? | Code compliance, project completion time, complaints |
What Makes a Good Measure
- Tied to purpose. It reflects what the service is for, not just what is easy to count.
- Clearly defined. Write down the numerator, denominator, data source, owner, and frequency. See Metric Governance.
- Balanced. Pair a speed measure with a quality measure and, where relevant, an equity or satisfaction measure, so no one improves one at the expense of the others.
- Actionable. Someone can do something about it, and it is reviewed often enough to act.
- Hard to game. Consider how a measure could be met without serving the resident, and add a check.
Worked Example: A Scorecard for a Permit Office
A permit office builds a small scorecard with six measures. It records the baseline, the target, and the latest actual, and calculates progress toward the target. The values are illustrative.
| Measure | Type | Better is | Baseline | Target | Actual | Progress |
|---|---|---|---|---|---|---|
| Days from application to permit | Efficiency | Lower | 26.5 | 16 | 21 | 52% |
| First-time-complete rate | Quality | Higher | 65% | 85% | 74% | 45% |
| Cost per permit | Efficiency | Lower | $310 | $250 | $290 | 33% |
| Applicant satisfaction | Outcome | Higher | 72% | 85% | 84% | 92% |
| Cases waiting | Output | Lower | 1,200 | 300 | 900 | 33% |
| Inspections on time | Quality | Higher | 80% | 95% | 96% | 107% |
Progress is (actual − baseline) / (target − baseline). For cycle time: (21 − 26.5) / (16 − 26.5) = 5.5 / 10.5 = 52%. The formula works for both "higher is better" and "lower is better" measures because the signs cancel.
Reading the scorecard: on-time inspections has met its target, and satisfaction is close. The measures furthest from target are cost per permit and the backlog, both at 33%, and first-time-complete at 45%. That points the team at the rework loop, since improving first-time-complete reduces both cost and backlog, and it is the measure the team can most directly influence.
Build your own with the Public Service Scorecard Builder, or use the Service Standards Scorecard Template for monthly tracking.
Guarding Against Gaming
Goodhart's law is often summarized as: when a measure becomes a target, it ceases to be a good measure. Common patterns include cherry-picking easy cases to keep average times low, closing cases early and reopening them, and redefining terms. Protective habits:
- Report the tail (90th percentile) and the age of the oldest case, not only the average.
- Use paired measures, such as cycle time with error rate.
- Audit definitions and a sample of records.
- Use measures to improve processes, not to rank individuals.
From Definition to Decision: The Life of a Measure
A measure is more than a number on a dashboard. It has a definition, a source, an owner, a method of checking, and a place in a decision. When any of these is missing, the measure will be disputed or ignored.
Write a definition sheet for each measure. Include the name, purpose, formula, what is included and excluded, data source, frequency, owner, and targets with their basis. Two people reading the sheet should calculate the same value.
| Field | Example for “permit decisions within 30 days” |
|---|---|
| Formula | Permits decided within 30 calendar days of a complete application / permits decided in the period |
| Clock start and stop | Starts on date a complete application is received; stops on the decision date |
| Exclusions | Applications withdrawn by the applicant; applications on legally required hold |
| Source and owner | Permit system report; owned by the permit office manager |
| Frequency and target | Monthly; target to be set with reference to past performance and capacity |
Validate. Check a sample of records against source documents, look for values that are implausible, and compare with other data. Fix definitions when the checks show ambiguity.
Using Measures for Learning and Decisions
Performance data are most useful when they are used in regular management conversations, not just published. A common model is a recurring performance review in which leaders and managers look at a small set of measures, ask why, and agree on actions.
- Ask about causes, not just whether the number is red or green. What changed? What do the staff say? What has been tried?
- Look at trends and variation, not only the latest value. A control or run chart helps to separate real change from noise, as in the SPC Control Charts Guide.
- Break results down by area, time, or group where it helps to reveal differences, taking care with privacy and small numbers.
- Follow up. Record actions, owners, and dates, and begin the next meeting with them.
- Use targets carefully. A target that is set without regard to capacity invites gaming. A target should be a prompt for learning and support.
Benchmark with care. Comparing with peers can show where improvement is possible, but differences in definitions, populations, and resources can make comparisons misleading. Understand the differences before drawing conclusions.
Publish with context. Public dashboards build trust when they explain what the measure means, how it is calculated, and what is being done. Numbers without context can be misread. Also review the set of measures periodically: retire those that no longer inform decisions, and add measures for new priorities. Figures in the examples are illustrative.
Self-Assessment Questions
- Can we describe our service's purpose in one sentence, and does each measure connect to it?
- Is every measure defined in writing, with an owner and a data source?
- Do we pair speed with quality and, where relevant, equity?
- Do we look at trends over time, not only a single period?
- Have we thought about how each measure could be gamed?
Common Mistakes
Measuring Only Outputs
Counting permits issued says nothing about whether construction is safer. Add outcome and quality measures.
Too Many Measures
A long list dilutes attention. Choose a handful that reflect purpose and review them regularly.
Targets Without Baselines
Without a baseline, a target is a guess. Measure the current state first.
Using Measures to Blame
If numbers are used to punish, people hide problems. Use them to find and fix process causes.
Performance Measurement in the Public Sector: Frequently Asked Questions
What is the difference between outputs and outcomes?
Outputs are what a program produces, such as permits issued or inspections completed. Outcomes are the changes that result for people or communities, such as safer buildings or shorter time to open a business. Outputs are easier to count, but outcomes show whether the work mattered, so agencies should track both.
What is a logic model?
A logic model is a simple diagram that connects inputs, activities, outputs, outcomes, and impact for a program. It shows the theory of how resources lead to results and helps decide where to place measures, from efficiency between inputs and outputs to effectiveness between outputs and outcomes.
How do we stop people from gaming a measure?
Pair measures so that improving one at the expense of another shows up, report the tail and the age of the oldest case instead of only averages, audit definitions and a sample of records, and use measures to improve processes rather than to punish individuals, which encourages honest reporting.
Sources and Further Reading
- W. K. Kellogg Foundation, Logic Model Development Guide.
- Robert S. Kaplan and David P. Norton, The Balanced Scorecard, and Kaplan's work on adapting it for public and nonprofit organizations.
- Harry Hatry, Performance Measurement: Getting Results.
- Discussions of Goodhart's law in the performance measurement literature, and the reporting requirements that apply to your jurisdiction.