Skip to content
Back to the lab

How to read frontier-model disclosures without inventing an incident rate

A six-field comparison card helps buyers interpret frontier-model disclosures without turning report counts into false incident rates.

Super Genius Labs Editorial · 4 min read

OpenAI has published six initial model-misalignment reports under a work-in-progress framework for tracking, investigating, and disclosing such examples. The framework assigns reports to Ready for Disclosure, Minor Investigation, or Larger Investigation. OpenAI also says reports may appear before investigation or mitigation is complete and that the initial set is not a comprehensive account of known cases (OpenAI).

The Associated Press reported the six disclosures separately and described the framework as covering cases involving unauthorized action, inter-model coordination, or evasion of oversight (Associated Press).

A buyer can establish from this evidence that six reports were published. The same evidence does not establish how often misalignment occurs, how many cases are known, how many affected an external party, or whether corrective work is complete.

Build the comparison from six fields

Disclosure programs can differ in what they include and when they publish. Comparing their headline totals without those rules would combine unlike records.

The following card is an SGL diligence structure, not a framework published by either source.

FieldQuestion for the vendorWhat the answer clarifies
Disclosure scopeWhich behaviors, systems, environments, and severity levels can enter the program?The inclusion boundary behind the count.
Lifecycle coverageCan publication occur at discovery, during investigation, after mitigation, or at several stages?Whether reports represent early observations, closed cases, or both.
Investigation stateWhat is established, disputed, or still under examination?How much confidence to place in the current explanation.
External impactDid the behavior remain in evaluation, occur in deployment, or affect an external party?The operating consequence, if any.
Unanswered questionsWhich questions about cause, scope, recurrence, or attribution remain open?What the report has not yet resolved.
Corrective-action statusIs an action proposed, implemented, tested, monitored, or closed?Whether response and verification have progressed beyond publication.

The card keeps three objects separate: the observed behavior, the provider’s knowledge when it published, and the action taken afterward. A larger report count could reflect more events, broader inclusion criteria, earlier publication, greater transparency, or several of those conditions. The supplied evidence does not distinguish among them.

Interpret a report at its current state

OpenAI’s three tracks identify different investigation states. Its stated limits also mean that the first six reports are neither comprehensive nor necessarily published after mitigation (OpenAI). Each item should therefore be read as a dated record rather than a final incident determination.

For every disclosure, record its publication date, current track, latest substantive update, and stated corrective-action status. Preserve earlier versions or a change log when the vendor supplies them. A later update could narrow the affected system, revise an explanation, add evidence about external impact, or describe mitigation. These are review events to watch for, not outcomes established by the sources here.

The same discipline applies when no report exists. An empty public record could reflect no known cases, a narrow inclusion boundary, delayed publication, or the absence of a public program. Without the program’s scope and publication rules, the evidence does not select among those explanations.

Ask for the missing denominator

The comparison becomes useful when a vendor answers questions tied to each field:

  • What qualifies for entry, and what is excluded?
  • Does the program cover evaluations, internal deployments, customer-facing deployments, or all three?
  • Can publication precede root-cause analysis or completed mitigation?
  • How are preliminary findings corrected or expanded?
  • Is external impact stated separately from observed model behavior?
  • Which statuses distinguish proposed, implemented, tested, monitored, and closed corrective actions?
  • Which known cases fall outside the public reporting scope?

A missing answer is a diligence result, but it is not proof of a concealed failure. Record “not disclosed” or “not established,” along with the request and its date.

Use reporting quality as the decision object

The supported procurement question is not which vendor has the fewest incidents. It is whether each reporting program defines its scope, preserves investigation state, separates behavior from impact, identifies unresolved questions, and updates corrective-action status.

OpenAI’s framework provides one example of explicit investigation tracks and explicit limits on comprehensiveness and closure (OpenAI). The available evidence does not show how complete the program will become, how consistently its categories will be applied, or whether any disclosed mitigation will prevent recurrence.

To turn the card into a repeatable vendor record, assign an owner and evidence date to every field, then connect material exceptions to the wider system decisions documented during build. Disclosure totals can identify where to investigate; only the program’s scope and the report’s current state determine what those totals mean.