Skip to content
Back to the lab

Our view: a non-exhaustive healthcare call-resolution scorecard beyond handling rates

Our view is that a single automation, containment, or “dealt with” rate may be insufficient for evaluating healthcare voice automation. We infer a non-exhaustive scorecard that separates access, administrative task completion, and human escalation.

Super Genius Labs Editorial · 5 min read

Researched with AI assistance; prepared for review by Super Genius Labs editorial and a healthcare operations specialist before publication.

NHS England’s June 2026 telephony statistics show why an aggregate call-handling metric can obscure materially different outcomes. Across 5,292 GP practices and 33,128,213 inbound calls, 88.3% were classified as “dealt with.” The same release reports that 56.3% were answered, 25.0% ended during IVR before joining a queue, and 11.7% ended while callers waited for staff. NHS England Digital

Those figures do not establish whether an IVR exit represented successful self-service, a caller’s independent decision to leave, failed navigation, or an unresolved request. They do establish that “dealt with” and “answered” describe different measures in this dataset.

For buyers evaluating a healthcare AI receptionist, our view is that an automation or containment rate alone may be insufficient. A stronger pilot scorecard can separate how a call ended from whether the caller’s administrative task was completed or an requested human connection succeeded.

Centralization does not settle the measurement question

Freeman Health System launched a centralized scheduling and call-center service on July 1, 2026, and subsequently acknowledged longer-than-expected waits and routing challenges reported by patients and provider offices. Freeman said its teams were reviewing call volumes, routing patterns, staffing, and system performance. Freeman Health System

This announcement documents one organization’s early rollout issues. It does not demonstrate that centralization or automation generally causes those issues. We infer a narrower procurement lesson: a changed access model can be assessed across routing, waiting, staffing, and task outcomes rather than through one aggregate handling rate.

A sponsored industry article in Becker’s Hospital Review similarly argues for measuring whether patient needs are resolved instead of stopping at interaction handling, while retaining human intervention for work involving empathy or judgment. That is an industry argument, not evidence that a particular measurement framework or implementation produces better outcomes. Becker’s Hospital Review

Our non-exhaustive scorecard

We infer six measurement layers for a pilot. These are proposed operating categories, not categories established or validated by the cited sources.

1. Demand and entry

Record the number of inbound calls and the route presented to each caller. Possible measures include:

  • Total inbound calls.
  • Calls entering IVR.
  • Calls leaving during IVR.
  • Calls joining a staff queue.
  • Calls routed directly to an automated administrative workflow.

An IVR exit can remain a routing outcome until separate evidence identifies what the caller accomplished.

2. Human access

Separate staff access from broader handling labels:

  • Calls answered by staff.
  • Calls abandoned while waiting.
  • Median and distribution of queue time, if measured.
  • Calls routed to the wrong destination and subsequently rerouted, if detectable.

A pilot can report these counts alongside automation results so that containment does not conceal changes in human access.

3. Callback outcomes

If callbacks are offered, distinguish each stage:

  • Callback offered.
  • Callback requested.
  • Callback attempted.
  • Caller reached.
  • Administrative request completed or transferred for further work.

Counting a callback request as resolution would combine an intake event with a later outcome. Our view is to report both separately.

4. Administrative task outcomes

For every supported task type, distinguish:

  • Task requested.
  • Task understood well enough to attempt.
  • Task attempted.
  • Task completed and confirmed.
  • Task deferred for staff action.
  • Task failed or remained unresolved.

Examples may include appointment booking, cancellation, rescheduling, or a request for practice information. These are non-exhaustive administrative examples, not clinical judgment.

5. Human escalation

A transfer option and a completed transfer are different events. A pilot can measure:

  • Human assistance requested.
  • Transfer attempted.
  • Transfer connected.
  • Transfer abandoned or disconnected.
  • Transfer unavailable, followed by a callback or staff-work request.

This layer can also record the point at which escalation occurred without assuming that every connected transfer resolved the underlying request.

6. Resolution checks and repeat demand

Where an organization can define and measure them consistently, possible checks include:

  • Confirmation that the requested administrative action was recorded.
  • Deferred work completed by staff.
  • Repeat contact about the same request within a defined period.
  • Manual review of a disclosed sample of outcomes.

Repeat contact is not automatically proof of failure; a caller may have a new question or changed circumstances. It can instead serve as a review signal whose interpretation remains bounded.

Denominators belong next to every rate

A 70% result can mean different things depending on whether the denominator is all inbound calls, calls that reached the automation, eligible administrative requests, or attempted tasks.

For each reported percentage, our proposed scorecard records:

  • Numerator.
  • Denominator.
  • Eligibility and exclusion rules.
  • Observation period.
  • Outcome definition.
  • Data source.
  • Known missing events.

This makes it possible to compare “calls contained,” “tasks completed,” and “human escalations connected” without treating them as interchangeable.

A procurement and pilot checklist

Before accepting a headline handling rate, we would ask a vendor and operating team to provide:

  1. A written definition of every terminal call outcome.
  2. The numerator and denominator behind each percentage.
  3. Separate counts for IVR exits, queue abandonment, staff answers, callbacks, automated task completion, deferred work, and human-transfer outcomes where those flows exist.
  4. A method for distinguishing a completed administrative action from an attempted action.
  5. A method for identifying unavailable, failed, or disconnected escalation paths.
  6. A disclosed sampling method for checking whether event labels match reviewed outcomes.
  7. An explanation of missing events and calls that cannot be classified.
  8. Time-bounded reporting so routing, staffing, and configuration changes can be interpreted alongside the results.

Our view is that pilot acceptance can be tied to several explicitly defined measures rather than one containment target. The selected measures will depend on the administrative workflows in scope and the evidence the implementation can actually produce.

What remains unknown

The NHS release provides aggregate telephony statistics; it does not evaluate an AI receptionist or explain the intent and ultimate outcome of every call that ended during IVR or in a queue. Freeman’s update concerns one centralized access launch and does not isolate the effect of any single technology. The Becker’s article is sponsored commentary advancing a measurement position.

Accordingly, these sources do not validate our six-layer framework or establish that adopting it improves patient access. The framework is our inference for structuring procurement questions and pilot evidence.

Teams translating these questions into an instrumented agent workflow can work with Super Genius Labs. Any implementation discussion can begin with the intended administrative tasks, human-escalation paths, measurable terminal outcomes, and evidence boundaries.