Skip to content
Back to the lab

Measure healthcare call resolution beyond handling rates

A single automation or containment rate can hide materially different outcomes. Separate access, administrative task completion, callbacks, and human escalation.

Super Genius Labs Editorial · 4 min readUpdated

NHS England’s June 2026 telephony statistics show why an aggregate call-handling metric can obscure materially different outcomes. Across 5,292 GP practices and 33,128,213 inbound calls, 88.3% were classified as “dealt with.” The same release reports that 56.3% were answered, 25.0% ended during IVR before joining a queue, and 11.7% ended while callers waited for staff. NHS England Digital

Those figures do not establish whether an IVR exit represented successful self-service, a caller’s independent decision to leave, failed navigation, or an unresolved request. They do establish that “dealt with” and “answered” describe different measures in this dataset.

For buyers evaluating a healthcare AI receptionist, an automation or containment rate alone can leave the result ambiguous. A stronger pilot scorecard separates how a call ended from whether the caller’s administrative task was completed or a requested human connection succeeded.

One access model, several observable outcomes

Freeman Health System launched a centralized scheduling and call-center service on July 1, 2026, and subsequently acknowledged longer-than-expected waits and routing challenges reported by patients and provider offices. Freeman said its teams were reviewing call volumes, routing patterns, staffing, and system performance. Freeman Health System

This announcement documents one organization’s early rollout issues. It does not demonstrate that centralization or automation generally causes those issues. The narrower procurement lesson is to assess a changed access model across routing, waiting, staffing, and task outcomes rather than through one aggregate handling rate.

A sponsored industry article in Becker’s Hospital Review similarly argues for measuring whether patient needs are resolved instead of stopping at interaction handling, while retaining human intervention for work involving empathy or judgment. That is an industry argument, not evidence that a particular measurement framework or implementation produces better outcomes. Becker’s Hospital Review

Read the denominator before the rate

A 70% result can mean different things depending on whether the denominator is all inbound calls, calls that reached the automation, eligible administrative requests, or attempted tasks. Put the numerator, denominator, eligibility rules, observation period, outcome definition, data source, and known missing events beside every percentage.

That discipline keeps “calls contained,” “tasks completed,” and “human escalations connected” from collapsing into one measure.

Follow one call through the measurement model

The scorecard below is a Super Genius Labs proposal, not a framework established or validated by the cited sources. It follows a call from entry to the evidence available after the interaction.

Demand and entry. Record total inbound calls and the route presented to each caller: IVR, staff queue, or an automated administrative workflow. Keep IVR exits as routing outcomes until separate evidence identifies what the caller accomplished.

Human access. Report staff answers, abandonment while waiting, queue-time distribution when available, and detectable wrong-destination reroutes. Place those counts beside automation results so a containment rate does not conceal a change in human access.

Callbacks. Keep “offered,” “requested,” “attempted,” “caller reached,” and “administrative request completed or transferred” as separate events. A callback request is intake; it is not the later outcome.

Administrative work. For appointment booking, cancellation, rescheduling, practice information, or another supported administrative task, distinguish requested, understood, attempted, completed and confirmed, deferred, and unresolved states. These are administrative examples, not clinical judgment.

Human escalation. A transfer option and a completed transfer are different events. Track the request, attempt, connection, disconnection, and any callback or staff-work request created when a transfer is unavailable. A connected transfer also does not prove that the underlying request was resolved.

Resolution evidence. Where the organization can define the measure consistently, record confirmation that the administrative action was written, completion of deferred staff work, repeat contact about the same request within a defined period, and manual review of a disclosed sample. Repeat contact is a review signal, not automatic proof of failure; the caller may have a new question or changed circumstances.

Make the pilot reject ambiguity

Before accepting a headline rate, ask the vendor and operating team for a written definition of every terminal call outcome and the denominator behind every percentage. The evidence package should separate IVR exits, queue abandonment, staff answers, callbacks, automated task completion, deferred work, and transfer outcomes wherever those flows exist. It should also explain how the system distinguishes an attempted administrative action from a completed one, identifies failed or disconnected escalations, samples event labels for review, handles unclassified calls, and relates routing or staffing changes to the reporting period.

The NHS release provides aggregate telephony statistics; it does not evaluate an AI receptionist or explain the intent and ultimate outcome of every IVR or queue exit. Freeman’s update concerns one centralized access launch and does not isolate the effect of any single technology. The Becker’s article is sponsored commentary advancing a measurement position. None validates this measurement model or establishes that adopting it improves patient access.

Pilot acceptance can therefore use several explicitly defined measures rather than one containment target. Use Build to share the intended administrative tasks, escalation paths, terminal outcomes, and evidence boundaries for a specific workflow.