Our view: a non-exhaustive release matrix for Opus 5 inference configurations
We infer that an Opus 5 deployment can be evaluated more precisely when model ID, thinking state, effort, output budget, endpoint, and fallback policy are versioned together. Our non-exhaustive matrix evaluates those inputs across quality, tool completion, latency, turn count, and cost.
Our view: a non-exhaustive containment review for agent evaluation environments
We infer that a sandbox label alone may be insufficient for tool-enabled agent evaluations. Our non-exhaustive review covers egress, intermediary services, credential reach, blast radius, action-volume detection, kill authority, forensic retention, and notification.
Our view: a non-exhaustive release-control loop built from agent traces
We propose a non-exhaustive engineering loop that connects production traces, contextual scoring, human review, regression cases, and predeployment checks. The model treats automated evaluation as bounded evidence rather than proof of correctness.
Our view: a six-control checklist for AI agent authorization
Our view is that enterprises can apply a non-exhaustive six-control authorization checklist while NIST’s agent-specific work remains under development.
Running an agent team in production
The boring infrastructure problems you only meet after the demo is over.
