Hospital AMR Pilot KPIs and Acceptance Criteria
A balanced hospital AMR pilot scorecard covering service, time, human work, quality, safety, reliability and evidence sufficiency.

A hospital AMR pilot should end with a decision, not a highlight reel. That requires agreed key performance indicators (KPIs), acceptance criteria, test conditions and stop rules before live operation begins. Otherwise, a team can collect impressive activity counts while missing delays, staff work, safety events or unreliable handoffs.
This guide gives hospital operations and PMO teams a practical acceptance framework. Thresholds must be set from the hospital's baseline, workflow risk and qualified review; the examples below are categories, not universal benchmarks.
The strongest scorecards stay small enough to manage, yet broad enough to expose hidden staff effort, rare failures and unsafe workarounds.
Write the decision before the KPI
State what the pilot must establish. For example: whether a defined logistics workflow can be completed within an acceptable service window, under representative traffic and staffing, without creating unacceptable safety, infection-control, privacy or workload consequences.
Then define the permitted scope: departments, routes, payloads, operating hours, interfaces, staff roles and excluded conditions. A metric is interpretable only inside that envelope. A high completion rate on quiet training routes does not establish readiness for normal hospital operations.
Use a balanced KPI set
NIST's robotic-system performance work separates useful dimensions such as task efficiency, completion quality, assurance and adaptability. A hospital pilot can apply the same measurement discipline without treating NIST as a hospital certification or setting. Measure the task, the environment, the human work and the quality of evidence together.
Scroll horizontally to compare all columns.
| Dimension | Example measure | Required definition | Evidence |
|---|---|---|---|
| Service | Eligible missions completed | Denominator and exclusion rules | Dispatch and handoff records |
| Time | Request-to-handoff duration | Start, stop and percentile | Timestamped event log |
| Human work | Interventions and staff minutes | Intervention categories | Observation plus incident log |
| Quality | Correct payload and destination | Verification and severity | Handoff record |
| Safety | Hazards, stops and near misses | Reporting and stop rules | Reviewed event file |
| Reliability | Availability in required window | Scope, exclusions and clock | System and service logs |

Make every metric reproducible
Create a metric dictionary. For each KPI record the question it answers, formula, numerator, denominator, unit, source system, owner, exclusions, reporting interval and acceptance threshold. Define whether durations use averages, medians or percentiles. Medians can hide severe delays; averages can be distorted by a few extreme events. Showing a distribution or percentile alongside the median is often more informative.
Separate eligible missions from dispatched missions and attempted missions. If staff cancel a request because the robot is unavailable, excluding it may make performance look better while service deteriorates. Record the reason for every exclusion and report excluded volume.
Measure interventions as work, not embarrassment
Interventions reveal the operating model. Classify them: planned training, traffic assistance, door or elevator help, localization recovery, payload handling, cleaning, technical reset and manual delivery. Record who intervened, why, duration, result and whether the same issue recurred.
A pilot may complete every mission only because staff quietly rescue it. That is not autonomous service. Conversely, a small number of planned interventions may be acceptable if they are explicitly designed, trained and included in the business case. The acceptance decision should reflect the approved future-state workflow.
Set thresholds and stop rules before the first mission
Acceptance criteria should include minimum performance, evidence sufficiency and mandatory conditions. Examples of mandatory conditions may include completion of required safety tests, approved cleaning procedure, correct access controls, trained operators and successful recovery from defined faults. A high overall score should not offset failure of a critical control.
Define stop-and-review triggers such as an injury, near miss, egress obstruction, loss of payload integrity, unapproved-area entry, privacy event or repeated uncontrolled behaviour. Specify who may stop the pilot, who investigates, what evidence is preserved and who authorizes restart.
Design representative test coverage
Map expected variation before deciding the sample: shifts, weekdays, traffic periods, routes, elevators, payload types, destinations and staff groups. Include normal operations plus approved challenge tests. Do not create unsafe conditions merely to test them; use a qualified, controlled method.
Compare with a credible baseline. Measure the current process using compatible definitions and a similar observation window. If the baseline uses request-to-delivery time while the pilot uses dispatch-to-arrival time, the comparison is invalid. Keep assumptions visible when complete baseline data are unavailable.
Run an evidence-led acceptance review
The final review should include the signed scope, configuration, metric dictionary, raw-data lineage, exception log, change record, training status, open risks and threshold result. Discuss data quality and missing periods. Separate confirmed findings from hypotheses and proposed future features.
Use three decisions: accept for the defined scope, revise and retest, or stop. Conditional acceptance should name the condition, owner and deadline. Rollout needs its own readiness gate because new floors, interfaces or departments can change the risk and performance envelope.
Use Warpify's hospital AMR workflow approach with the deployment timeline framework. When your baseline, owners and decision are clear, assess your hospital workflow.
Iven Wang
Iven Wang is the Co-Founder of Warpify Robotics, specializing in the commercialization and deployment of robotic solutions. With a background in electrical engineering and product management, he works with manufacturers, integrators, and enterprise clients across industrial inspection, security, logistics, and Robotics-as-a-Service.
Subscribe to our newsletter today
Get practical robotics deployment insights, case studies, and planning guidance from Warpify Robotics.




