A model score is sometimes a system score: an attribution sheet for multi-model agents
A named model may sit inside a routed, multi-agent harness. Record the models, routing, tools, benchmark configuration, access tier, and evidence owner behind the score.
Microsoft describes MAI-Cyber-1-Flash as a cybersecurity model embedded in MDASH, a multi-model harness that orchestrates more than 100 agents. Microsoft says the specialized model handles up to 90% of MDASH tasks while larger models handle exceptional cases (Microsoft AI).
Secondary reporting adds relevant boundaries. SecurityWeek attributes the reported CyberGym performance to MAI-Cyber-1-Flash used with MDASH and GPT-5.4, and describes MDASH as combining multiple frontier and distilled models (SecurityWeek). Axios reports a routing pattern in which routine vulnerability analysis goes to the smaller specialized model and difficult tasks go to a larger frontier model (Axios).
These sources document or report a composite configuration. They do not isolate how much of the reported result belongs to the specialized model, the larger model, the routing policy, the agent harness, or their interaction.
Read the reported result as a stack
A result from this kind of configuration can function as a system score, even when discussion centers on one named model. That distinction matters when an evaluation team uses the result to select a model, reproduce a benchmark, or approve a release.
The evaluation artifact should identify both the named model and the surrounding system. This does not diminish the model's contribution. It makes the attribution boundary explicit.
Write the stack beside the score
For each result, record:
| Field | What to record |
|---|---|
| Evaluated claim | The precise capability or outcome being assessed |
| Specialized model | Exact model name and version |
| Fallback models | Every model eligible to receive routed work |
| Routing policy | Conditions that select, escalate, retry, or fall back |
| Harness | Orchestrator name and version |
| Agent topology | Agent count, roles, and coordination pattern |
| Context | Proprietary data, retrieval sources, prompts, and other supplied material |
| Action space | Tools, permissions, environments, and available operations |
| Access tier | Public, preview, private, or otherwise restricted access |
| Benchmark configuration | Dataset version, task subset, scoring method, limits, and run count |
| Result scope | Component score, routed-system score, or end-to-end score |
| Evidence owner | Person or team accountable for the artifact |
| Reproduction status | What an independent evaluator can and cannot reproduce |
These fields form a Super Genius Labs operating template, not a framework established by the cited sources. Teams can adapt it to the decision at hand.
Let the decision choose the granularity
Model selection
Compare like with like. A standalone-model result and a routed-system result answer different questions. Putting the result scope beside the score can reduce accidental comparisons between those categories.
Release review
Version the model, harness, routing policy, tools, and benchmark configuration together. If one component changes, the prior result may no longer describe the candidate system. The materiality of that change remains an evaluation judgment.
Incident or regression analysis
Retain the attribution sheet with the result. It gives reviewers a bounded record of which components were eligible to influence the evaluated behavior, without claiming that the sheet proves causal attribution.
The supplied evidence does not provide a component-level ablation of MAI-Cyber-1-Flash, MDASH, GPT-5.4, routing, or agent count. It also does not establish which parts of the reported configuration an external team can reproduce. Those unknowns belong beside any use of the result in procurement, architecture, or release decisions.
The sheet is useful when reviewers can see what shaped the result and what remains unreproducible. A concrete evaluation configuration can be shared through Build.
