Skip to content
Back to the lab

What evidence should healthcare AI operators retain for non-device software?

FDA’s patient-safety inquiry gives healthcare AI operators a bounded prompt: preserve intended use, observed benefits, safety signals, implementation lessons, governance practices, and evidence ownership without treating an operational record as a legal classification or proof of safety.

Super Genius Labs Editorial · 4 min read

FDA is requesting patient-safety input for its 2026 report on software functions excluded from the medical-device definition. Its named categories include administrative support, electronic patient records, data display, and limited clinical decision support; comments under docket FDA-2018-N-1910 are due August 13, 2026. FDA: Reports on Non-Device Software Functions

That inquiry creates a practical evidence question for healthcare AI operators: what records would let a reviewer understand what a software function was intended to do, how it was implemented, and what happened during use?

The categories named by FDA do not, by themselves, establish that a particular AI function falls outside the medical-device definition. Nor does a well-maintained operational record prove clinical safety. Classification and safety conclusions remain separate questions.

Preserve the function as it changed

The useful unit of analysis is a bounded software function, followed over time. One product may contain several functions with different intended users, inputs, outputs, escalation paths, and consequences.

The initial record should identify the function’s intended use, intended user, operating setting, inputs, outputs, and tasks outside its intent. It should also state whether the output completes an administrative action, displays information, or informs a person’s decision.

That description needs a change history. For each material version or configuration, operators can retain the active model or rule set, data sources, connectors, downstream dependencies, and the owner authorized to change or suspend the function. Human involvement should be equally concrete: where a person can review, alter, reject, escalate, or halt an output.

This decomposition is a house recommendation, not an FDA classification method. It prevents a broad product description from obscuring materially different functions.

Keep expectation, observation, and response distinct

Hall Render reports that FDA invited stakeholders to contribute operational experience, safety data, implementation lessons, or governance practices for the 2026 risks-and-benefits report. Hall Render: FDA Requests Input on Non-Device Software Functions and Patient Safety

Those evidence types are easier to inspect when an expected result is not merged with what was observed or what the team did next.

An expected-benefit record can state the proposed change, affected population or workflow, comparison point, measurement method, observation period, and evidence owner. A separate observation can record the actual result, including neutral or adverse findings and missing data. If no controlled comparison exists, the record should say so rather than imply causation.

A safety record needs the same separation. It can identify what occurred, how it was detected, and which function, version, and configuration were involved. It can then record the affected workflow, potential consequence, whether the event reached a user, patient, or downstream system, and the evidence known to be missing. Containment, escalation, follow-up, the assessor, and the person who accepted any residual uncertainty belong in the response record.

Implementation lessons should connect these stages. When testing or use leads to a rollback, training update, configuration change, or revised workflow, the retained decision should state what changed and why.

This structure is also a house recommendation. The supplied sources establish the inquiry and the requested kinds of input, but they do not prescribe a record format or demonstrate that any particular format is sufficient for a submission.

Make every assertion traceable

A reviewable collection does not need to collapse into one worksheet. It needs a reliable path from each assertion to dated source material.

For a function boundary, that path might lead to an intended-use statement, workflow map, or interface specification. For a version claim, it might lead to a release record, model identifier, prompt or rule revision, or connector configuration. Benefit observations should lead to the analysis output, denominator, observation period, and missing-data note. Safety signals should lead to an incident record, complaint, monitoring alert, or adjudication note.

Governance evidence should identify who approves changes, reviews signals, and can suspend operation. Evidence-ownership records should identify the canonical copy, retention period, access owner, and person responsible for maintaining each item. Open questions should remain visible as explicit limitations or planned investigations.

The distinction between design and operation matters throughout. A policy, interface, or architecture diagram documents intended behavior. An event record or measured observation addresses what happened in a particular instance. Neither should be presented as the other.

Different readers may use the same underlying facts for different purposes. Product leaders may examine scope and change history. Safety leaders may focus on signals, exposure, and escalation. Counsel may assess how those facts relate to applicable definitions and obligations. Preserving the facts without embedding a legal conclusion keeps those uses separate.

“The function summarizes information for staff review” is an inspectable description. Calling the entire product “non-device software” reaches a classification conclusion that these operational records cannot establish.

The next useful step is to choose one function, name its current owner, connect every benefit and safety assertion to a dated record, and mark what remains unknown. Teams designing the supporting systems can carry those same boundaries into how they build evaluation, monitoring, escalation, and change-control paths.

The result is not proof of classification or safety. It is a record that lets the next reviewer ask a specific question without first reconstructing the system from product language.