The prior-authorization callback test for AI receptionist buyers
A proposed buyer test follows prior-authorization calls through waiting, denial, callback, and escalation—not just initial intake.
CMS has defined decision windows for certain prior-authorization requests, yet a recent practice poll does not indicate broadly faster turnaround.
Under its final rule, CMS requires covered payers—excluding QHP issuers on federally facilitated exchanges—to decide expedited non-drug requests within 72 hours and standard requests within seven calendar days. Beginning in 2026, those payers also have to provide a specific denial reason (CMS). These are outer decision windows for the requests covered by the rule, not evidence that every request finishes within them.
A September MGMA Stat poll adds a different, limited signal. Among 178 applicable responses, 44% of medical group leaders said payer turnaround had become slower in 2026, 40% said it was unchanged, 7% said it was faster, and 9% were unsure (MGMA). The poll reports respondents’ assessments; it does not measure every practice, payer, request type, or elapsed decision time.
For buyers evaluating Genius Care or another AI answering service, the operating question is therefore larger than whether the receptionist can answer the first call. Our proposed test asks whether the product can maintain an administrative request across waiting, denial, callback, and escalation states. It is a buyer-evaluation method, not a sourced standard or a validated product outcome.
The test begins after intake
Use two synthetic cases: one standard non-drug request and one expedited non-drug request. Give each case fictional identifiers and a predetermined timeline. Do not use real patient information in a sales demonstration unless the practice has separately established an appropriate environment and process.
For each case, begin with a caller asking for status. The receptionist receives enough information to identify the synthetic request, then encounters a sequence of later events: no decision yet, a promised callback, a denial with a stated reason, and a request that needs human attention.
The exercise is designed to expose state continuity. A fluent opening exchange earns little if the next caller, staff member, or system cannot determine what happened before.
Make the request state visible
Use one shared record for the run. The following fields are our proposed evaluation set, derived from the workflow described above rather than prescribed by CMS or MGMA:
| Field | What to inspect |
|---|---|
| Request identity | Whether the same synthetic request remains distinguishable across contacts |
| Request class | Whether standard and expedited scenarios stay correctly labeled |
| Current status | Whether pending, denied, escalated, and closed states remain distinct |
| Last contact | Whether the record identifies when the latest interaction occurred |
| Promised callback | Whether a stated callback time is captured without being converted into a completed callback |
| Denial reason | Whether the reason received is preserved accurately rather than summarized into a different meaning |
| Next action | Whether the record shows the next administrative step |
| Escalation owner | Whether a named role or queue owns the unresolved item |
The CMS rule makes request class and denial reason especially relevant test inputs: it specifies different windows for standard and expedited non-drug requests and, beginning in 2026, a specific-reason requirement for denials by covered payers (CMS). The table itself remains an SGL proposal; the source does not define a receptionist data model.
Walk the same cases through changing conditions
Run each synthetic case through five moments.
First, present the initial status inquiry. Record what the receptionist says, which fields it writes, and what remains unknown.
Next, advance the clock while leaving the request pending. Ask again from a different caller or channel. Look for invented certainty, a lost request class, or a callback treated as already completed.
Then supply a promised callback time. Interrupt the path before that time, at that time, and after it. The useful observation is not whether the product speaks confidently; it is whether the stored status and next action match the staged facts.
After that, introduce a denial and a specific synthetic reason. Ask the system to repeat the reason, route the item, and preserve the original wording. This tests administrative handling only. It does not test clinical judgment, the validity of a denial, or the merits of an appeal.
Finally, make the designated staff owner unavailable. Observe whether the item remains assigned, moves to an approved fallback, or becomes ownerless. Any preferred behavior here comes from the practice’s operating design, not from the two cited sources.
Read failures by consequence
A failed greeting and a lost denial reason are different defects. We suggest grouping observations by what the defect could obscure:
- Identity loss: The next interaction cannot locate or distinguish the request.
- State loss: Pending, denied, escalated, and closed become interchangeable.
- Time loss: A promised callback or staged deadline disappears or changes.
- Meaning loss: The denial reason is omitted or materially rewritten.
- Ownership loss: No role or queue remains accountable for the next action.
- Unsupported completion: The system presents an intended callback, escalation, or follow-up as completed without evidence from the test.
This classification is an analytical framework for comparing runs. Neither CMS nor MGMA evaluates receptionist products or establishes these failure categories.
Compare evidence, not performance theater
Ask each vendor to return the record produced by every step, the caller-facing response, the staff-facing handoff, and the final disposition. Repeat the same script after changing one fact, such as request class, denial reason, or callback time. The comparison can then focus on whether the product preserved and updated the right state rather than on which demonstration sounded smoother.
The CMS deadlines provide a bounded timing frame for covered non-drug requests, while the MGMA poll suggests that many respondents did not perceive faster turnaround in 2026 (CMS; MGMA). Neither source establishes that a receptionist product improves turnaround, reduces workload, or completes prior authorization.
That leaves a practical purchasing distinction: answering the first call is one event; preserving an unresolved administrative request across several days is a workflow. Buyers can use this proposed test to examine the latter. Teams that need a workflow tailored to their own queues, roles, and systems can also build from those operating constraints, with any product claims evaluated through current deployment evidence rather than inferred from the test design.
