Optimal Deployment Science

Measurement & the Score

The measurement discipline: what is measured, how a measurement becomes a judgment, who answers for each result, and how the whole picture rolls up into the optimal deployment score.

Measurement is retrospective. With a deployment plan running, the system produces a stream of events and measurement turns that stream into understanding. It is distinct from the forward-looking planning in Financial Sustainability, which decides what the system produces.

Measurement has two parts: activity and performance

Measurement answers two different questions and keeping them straight is essential:

  • Activity: "What happened?" The descriptive record: units activated, serviceable, and utilized; calls; task times and intervals. Activity is neutral: it counts and does not judge.
  • Performance: "Was it good or bad? Successful or not?" The evaluative layer: that same activity assessed against a standard or benchmark, with statistical rigor. Performance is where a number becomes a judgment.

You need activity before you can speak to performance, because you must know what happened before you can say whether it was good and activity without performance is just data. Both are measurement.

A third question, who is answerable for it?, is not measurement at all. That is accountability, governed separately by scope of control. Measuring something, and even judging it good or bad, does not by itself make it any one party's success or failure.

Readiness and response

Operations decomposes into readiness and response, measured separately (System Structure). Every subsystem is measured on both, not just patient transport. Taking patient transport as the example:

  • Readiness measurement. Unit hours delivered against unit hours committed, serviceability, and crew freshness.
  • Response measurement. What happened when calls came? Chute and turnout intervals, on-scene intervals, transfer-of-care intervals, and the operational response times.

Reporting response without readiness leaves readiness unmeasured. A count of calls responded to is pure response volume: it says nothing about readiness and rewards running more calls. The same readiness and response pair applies to each of the other subsystems, with its own metrics.

Unit-hour categories

Readiness is counted in unit hours and every unit hour falls into these categories. ODS uses only these terms.

  • Activated. The unit hour is scheduled and staffed.
  • Serviceable. Activated and fit for assignment: crew, equipment, and supplies all conforming.
  • Utilized. Serviceable and committed to a call.
  • Excessive task time. Utilized well beyond the interval the task should take. The threshold is set by the system. A common starting point is 150 percent of the allotted interval, not a fixed rule. Examples: waiting with the patient at an ED, extended on-scene waits.
  • Out of service. Activated but not serviceable.
  • Lost unit hours. The sum of out of service and excessive task time.

A fire company continuing care on scene until a slow transport unit arrives, or a crew waiting with a patient in a hospital hallway, is utilized with excessive task time, contributing lost unit hours (distinct from out of service). The readiness deficit does not belong to the committed unit. It belongs to the party that controls the delay: the ED at the hospital and, on scene, whoever is responsible for the deployment that determines when the transport unit arrives.

"Properly deployed," in the definition of readiness, includes position. Checking it asks two things: did the communications center assign the unit per the deployment plan and did the unit comply with the assignment? The first belongs to communications and the second to the unit's producer.

Measure with statistical rigor

Rigor is how activity becomes performance: how you judge what happened against a standard. On that judgment, ODS takes no side in the average-versus-fractile debate, because both a lone average and a lone fractile hide failure. An average conceals the spread: a mean of seven minutes says nothing about whether some neighborhoods wait twenty-five. A fractile conceals magnitude and the tail: a 90th-percentile compliant/non-compliant line scores a call that missed by one second and one that missed by an hour as the same. Each statistic is fine with its companion. An average needs a standard deviation; a fractile needs the full histogram.

When a system reports a single number, either of these scenarios could be behind it and the number cannot say which.

Drag the handle on any chart to show what the single number leaves out.

An average conceals the spread

The system reports 7:00 average operational response time

What the average leaves out

Scenario 1 It could be this

0725

Scenario 1. Every response is near seven minutes.

Scenario 2 Or it could be this

0725

Scenario 2. Some neighborhoods wait twenty-five minutes.

A fractile conceals magnitude and the tail

The system reports 90% compliant

What the fractile leaves out

Scenario 1 It could be this

01060Standard

Scenario 1. Every miss is within two minutes of the standard.

Scenario 2 Or it could be this

01060Standard

Scenario 2. Some misses run to most of an hour.

Illustration Both a lone average and a lone fractile hide failure. The distributions are drawn to show the idea and are not data from any system.

So ODS prescribes no single statistic. It uses a standardized battery of statistically rigorous tools for each metric (minimum, maximum, mean, standard deviation, fractiles, the full distribution), selected to fit the values and needs of the particular community, and reported against well-calibrated benchmarks. Statistical rigor and benchmark calibration are each specialist fields. A metric reported without rigor and calibration informs no one. A metric reduced to the wrong single summary conceals what it should show.

Response time: an input, not a verdict

Clinical response time
The time from the onset of symptoms or injury to the first meaningful clinical intervention. No single party controls it: the public, communications, and the responders each hold part of it.
Operational response time
The time from the request being initialized to a unit arriving on scene. The clock starts there because that is the first moment the standard that applies to the call can be known. There are two, first response time and transport response time. Operational response time is not clinical.

Operational response time cannot carry the weight of judging a system. The specific threshold has no demonstrated clinical meaning. The evidence base for the 8:59 tier is thin and contested (Hansen et al., 2025; see Citations). Even measured well, operational response time summarizes only the response component of one chain of events. Readiness in every subsystem and the clinical work are invisible to it.

The figure most often cited, 8:59, comes from cardiac arrest research. Seattle studies in the early 1970s tied survival to response in under eight minutes (Fitch, JEMS, 2005). Eisenberg, Bergner, and Hallstrom (JAMA, 1979) then measured two intervals, both starting at the patient's collapse: one ending at the initiation of CPR, the other at the provision of definitive care, which in that study meant defibrillation. When CPR began within four minutes and definitive care arrived within eight, 43% of patients survived. If either time was exceeded, survival fell sharply. The 59 seconds were added later as an accommodation for punch clocks. The finding concerned nontraumatic cardiac arrest alone and its clock was a clinical response time, from collapse to defibrillation. The standard built on it is an operational response time, from the request being initialized to arrival on scene, for every call type, which is a different clock.

Operational response time is an input into determining how much readiness a system needs. Its job is the axis the marginal utility curve is read against and setting the standard is where the goal is set from it. The judgment of the system is carried by the optimal deployment score, which covers readiness and response in every subsystem.

Operational response time does double duty, depending on who uses it:

  • The procuring entity or regulator setting the standard reads the marginal utility curve to set a defensible benchmark. It knows how much readiness it is demanding and can therefore reduce cost and redirect funds elsewhere.
  • The producer handed a benchmark it did not choose uses operational response time with the marginal utility curve to work out how much readiness that standard requires on top of anticipated utilization. That party carries the burden of meeting a standard selected for it.

These are two distinct uses: a process for benchmarking an appropriate response-time reliability target (setting the standard) and a use of operational response time to determine required readiness. This assigns operational response time to the right party for the right purpose and does not diminish it.

Reading compliance across tiers

A response-time standard is a matrix. Moving from dense areas to less dense ones, more time is allotted and fewer ambulances cover the ground. If the tiers were calibrated coherently, compliance would be roughly level across urban, suburban, and wilderness zones. If instead the urban life-threat tier is the hardest to hit while the outlying tiers come easily and hitting the urban tier sweeps up the rest by happenstance, the tiers are miscalibrated, with the urban tier too tight or the outlying tiers too loose. It does not signal that the producer is failing and it requires no accusation of bad faith: the numbers show the standard was not built coherently.

The functional difference to a patient between 90.02%, 90.01%, and 89.99% compliance is nil. An exemption should be granted on its merits and not because compliance is close to the line.

Scope of control: who is accountable for what

Measurement covers the whole system, regardless of who is responsible for any part of it. Accountability is the separate question of who answers for a given result. It follows control and three principles carry it:

  • A party is held to account only for what it controls.
  • A result one party controls belongs to that party, whatever its title. A body that holds several subsystems answers for what spans them.
  • A result no single party controls is measured and reported so the system can learn from it. No one is graded on it. Patient outcome is the clearest case.

Operational response time belongs to the procuring entity

Unit-hour procurement (quantity), the deployment plan (location), response-time accountability, and billing travel together.

No individual producer controls that clock; it is driven by caller interrogation, triage, dispatch, the posting plan, call classification, and zone definitions. Most of those are deployment and system-design decisions. For that reason a producer should not be held to operational response time. Operational response time against the standard belongs to the procuring entity, which holds the deployment plan (the three parties are defined in System Structure).

This is also why a "compliant" response can hide a broken one. In a system that allows the transport unit 30 to 40 minutes on distant or low-acuity calls, a fire company can arrive in 8 minutes and then continue care on scene for half an hour until the ambulance arrives. The fire company is utilized with excessive task time, contributing lost unit hours, while the contract still scores the response compliant. The number is met while the system is failing. Reading operational response time without the readiness picture of each subsystem (see System Structure) is how that failure stays invisible.

In-scope metrics live in every subsystem

Everything a party does control is fair to hold it to and each of the five subsystems along the patient's path has its own in-scope set, on both the readiness and response sides. These are the metrics that flow into the optimal deployment score. These are examples:

  • Public. Readiness: citizen recognition, public awareness and education. Response: bystander intervention (CPR/AED), 911 activation.
  • Communications. Readiness: communication capacity, deployment plan. Response: EMD and triage quality, the call-processing intervals (public safety answering point, transfer caller, initialize request, interrogate caller, dispatch first response, dispatch transport response), pre-arrival instructions. A request is initialized once the call taker has recorded when and where help is needed and the nature of the incident.
  • First response. Readiness: serviceable unit hours, mutual-aid agreements. Response: mobilizing (turnout in the fire service, out-of-chute in private EMS), en route, on-scene assessment, transfer of care, return to service.
  • Patient transport. Readiness: serviceable unit hours, mutual-aid agreements. Response: mobilizing, en route, on scene, preparing transport, transport, return to service.
  • Definitive care. Readiness: bed capacity. Response: preparing for the patient, transfer of care, providing care.
  • Operational response time. Two metrics, first response time and transport response time, each measured against the standard. Each spans communications and the subsystem that responds and belongs to no single subsystem. Its intervals are scored in the subsystems that own them and the whole clock is reported beside the score.

The common thread: a party answers for its own subsystem's readiness and response, not for another subsystem's and not for the operational response times. Wherever a party touches the patient (first response, patient transport, definitive care), clinical quality is measured as QA/QI: did the crew recognize the condition, select the correct protocol, and follow it? Each of these is real, measurable, and inside the party's control.

Split the clock where responsibility transfers

Where one subsystem depends on another to take the patient, a delay on one side impacts the other (System Structure). The measurement rule: cut an interval at the event where responsibility for the patient changes hands and put each resulting clock on the party that controls its delay. Neither side can then shave its clock without the delay landing on its own scorecard.

Patient transport to definitive care. The time a crew spends at the hospital is commonly reported as one number. It is two clocks, split at the transfer of care; the nurse signature may be used as a surrogate:

  1. Transfer of care. At destination until care is transferred. Under federal law (EMTALA), the hospital's obligation begins when the patient arrives and care is requested, whether or not the patient has been moved off the ambulance stretcher (CMS, 2006). In practice the crew stays with the patient until the hospital accepts care, so this clock measures how long definitive care takes to accept the patient. It is definitive care's metric, reported to the hospital.
  2. Return to service. Transfer of care to crew available. This is a different job: breaking down and rebuilding the rig. It is patient transport's metric, with its own standard.

The two clocks are reported separately and are not added together.

A practical refinement: set a total time-at-hospital budget split per side, for example 30 minutes as 15 + 15 (20 + 20 is also seen). The rationale is labor. One crew member stays with the patient and owns the transfer-of-care wait; the other breaks down and rebuilds the rig. If the hospital moves fast, the crew still gets its share to ready the unit.

First response to patient transport. The same split applies. There is value in capturing the transfer of patient care from first response to patient transport, which could be by way of a signature. Capturing it establishes a well-defined scope of control for both first response and patient transport.

A second metric sits beside it: first-responder on-scene time before transport arrival. This is the time a first-response unit spends on scene before the transport unit arrives. That time is not in first response's scope of control, so it is not first response's metric. It is part of transport response time: it is set by when the transport unit arrives and it belongs to whoever is responsible for the deployment. It answers a different question than the transfer of care does, so the two are reported separately.

Patient outcomes

Each subsystem's producer owns that subsystem's readiness and response metrics and, where it delivers care, recognition, protocol selection, and adherence. Medical direction owns protocol quality, the procuring entity owns operational response time, and no single party owns patient outcome.

Patient outcomes matter and ODS measures them. They are feedback to medical direction and the raw material for research into protocol efficacy: the loop by which a system learns whether its clinical protocols work and improves them.

Outcomes are not a metric to grade responders by. Outcome is dominated by the patient's underlying disease and physiology and by the whole system's time to care, none of which a responder controls. Ranking or penalizing crews on outcome drives risk-averse care and defensive documentation and lets a poor protocol pass as a poor crew.

What a responder is held to is the clinical work inside their control:

  • Recognition. Did the crew correctly identify the patient's condition?
  • Protocol selection. Did they choose the correct protocol for it?
  • Adherence. Did they follow that protocol correctly?

If recognition, selection, and adherence are all high but outcomes are still poor, the problem is the protocol rather than the crew and revising it is medical direction's responsibility, informed by the outcome data.

The optimal deployment score

A system that reports only operational response times is reporting response metrics and nothing else. The rest of what it does is left out of the picture.

Each of the five subsystems along the patient's path (public, communications, first response, patient transport, definitive care) has its own readiness metrics and its own response metrics. All of them flow into the score. The operational response times each span more than one subsystem. Their intervals are scored in the subsystems that own them and the whole clocks are reported beside the score. Every activity in the system affects the score: the public recognizing an emergency and calling, call answering and triage, first response, patient transport, definitive care. The score summarizes the whole system in one number.

Optimal DeploymentScienceOptimal deployment scoreSampleAgency nameReporting periodPublic91%Readiness90%Response92%89%Citizenrecognition91%BystanderinterventionCommunications91%Readiness90%Response92%89%Communicationcapacity91%DeploymentplanFirst response92%Readiness91%Response93%90%Mutual-aidagreements92%Serviceableunit hoursPatient transport78%Readiness91%Response65%90%Mutual-aidagreements92%Serviceableunit hoursDefinitive care58%Readiness96%Response20%96%BedcapacityOnset of symptomsActivation of 911 system92%CommunicationsPublic safety answering pointTransfer callerInitialize request90%Interrogate caller92%Dispatch first response94%Pre-arrival instructionsDispatch transport response92%First responseResponse94%Mobilizing94%En routeOn scene92%Patient assessment92%Transferring careReturning to servicePatient transportResponse93%Mobilizing93%En routeOn scene92%Patient assessmentPreparing transportTransportingNotify definitive careReturning to service10%Definitive carePreparing for patientAccepting patient20%Providing careThe patient's path through the system82%Optimal deployment scoreSample values. Each roll-up shown is a simple average. The scoring rubric, including how results are weighted, will be published separately.optimaldeployment.orgPage 1 of 1
Optimal deployment score: sample
Subsystem Readiness Response Score
Public 90% 92% 91%
Communications 90% 92% 91%
First response 91% 93% 92%
Patient transport 91% 65% 78%
Definitive care 96% 20% 58%
Optimal deployment score 82%
A sample optimal deployment score report. Sample values. Each roll-up shown is a simple average. The scoring rubric, including how results are weighted, will be published separately. Open the sample report on its own page.

Benchmark anchors come from the system's own plan. Each metric's target and fail point anchor to the standard the community chose and funded through the method's own chain (rubric, financial sustainability, deployment plan) and not to an external number the system has not adopted. NFPA and NEMSQA are respected sources for metric definitions and not for anchor values.

The loop

Measurement feeds the next production analysis and the next deployment plan. Where a producer's in-scope metrics chronically fail, it feeds the accountability mechanics of the readiness market.