Skip to content
Back to the lab

A model score is sometimes a system score: an attribution sheet for multi-model agents

A named model may sit inside a routed, multi-agent harness. Record the models, routing, tools, benchmark configuration, access tier, and evidence owner behind the score.

Super Genius Labs Editorial · 3 min read

Microsoft describes MAI-Cyber-1-Flash as a cybersecurity model embedded in MDASH, a multi-model harness that orchestrates more than 100 agents. Microsoft says the specialized model handles up to 90% of MDASH tasks while larger models handle exceptional cases (Microsoft AI).

Secondary reporting adds relevant boundaries. SecurityWeek attributes the reported CyberGym performance to MAI-Cyber-1-Flash used with MDASH and GPT-5.4, and describes MDASH as combining multiple frontier and distilled models (SecurityWeek). Axios reports a routing pattern in which routine vulnerability analysis goes to the smaller specialized model and difficult tasks go to a larger frontier model (Axios).

These sources document or report a composite configuration. They do not isolate how much of the reported result belongs to the specialized model, the larger model, the routing policy, the agent harness, or their interaction.

Read the reported result as a stack

A result from this kind of configuration can function as a system score, even when discussion centers on one named model. That distinction matters when an evaluation team uses the result to select a model, reproduce a benchmark, or approve a release.

The evaluation artifact should identify both the named model and the surrounding system. This does not diminish the model's contribution. It makes the attribution boundary explicit.

Write the stack beside the score

For each result, record:

FieldWhat to record
Evaluated claimThe precise capability or outcome being assessed
Specialized modelExact model name and version
Fallback modelsEvery model eligible to receive routed work
Routing policyConditions that select, escalate, retry, or fall back
HarnessOrchestrator name and version
Agent topologyAgent count, roles, and coordination pattern
ContextProprietary data, retrieval sources, prompts, and other supplied material
Action spaceTools, permissions, environments, and available operations
Access tierPublic, preview, private, or otherwise restricted access
Benchmark configurationDataset version, task subset, scoring method, limits, and run count
Result scopeComponent score, routed-system score, or end-to-end score
Evidence ownerPerson or team accountable for the artifact
Reproduction statusWhat an independent evaluator can and cannot reproduce

These fields form a Super Genius Labs operating template, not a framework established by the cited sources. Teams can adapt it to the decision at hand.

Let the decision choose the granularity

Model selection

Compare like with like. A standalone-model result and a routed-system result answer different questions. Putting the result scope beside the score can reduce accidental comparisons between those categories.

Release review

Version the model, harness, routing policy, tools, and benchmark configuration together. If one component changes, the prior result may no longer describe the candidate system. The materiality of that change remains an evaluation judgment.

Incident or regression analysis

Retain the attribution sheet with the result. It gives reviewers a bounded record of which components were eligible to influence the evaluated behavior, without claiming that the sheet proves causal attribution.

The supplied evidence does not provide a component-level ablation of MAI-Cyber-1-Flash, MDASH, GPT-5.4, routing, or agent count. It also does not establish which parts of the reported configuration an external team can reproduce. Those unknowns belong beside any use of the result in procurement, architecture, or release decisions.

The sheet is useful when reviewers can see what shaped the result and what remains unreproducible. A concrete evaluation configuration can be shared through Build.