Skip to content
Blog category

Engineering

How agent systems are built, observed, bounded, and kept useful in production.

5 notesLatest All notes RSS feed
Engineering

Our view: a non-exhaustive release matrix for Opus 5 inference configurations

We infer that an Opus 5 deployment can be evaluated more precisely when model ID, thinking state, effort, output budget, endpoint, and fallback policy are versioned together. Our non-exhaustive matrix evaluates those inputs across quality, tool completion, latency, turn count, and cost.

Super Genius Labs Editorial · 6 min read
Engineering

Our view: a non-exhaustive containment review for agent evaluation environments

We infer that a sandbox label alone may be insufficient for tool-enabled agent evaluations. Our non-exhaustive review covers egress, intermediary services, credential reach, blast radius, action-volume detection, kill authority, forensic retention, and notification.

Super Genius Labs Editorial · 7 min read
Engineering

Our view: a non-exhaustive release-control loop built from agent traces

We propose a non-exhaustive engineering loop that connects production traces, contextual scoring, human review, regression cases, and predeployment checks. The model treats automated evaluation as bounded evidence rather than proof of correctness.

Super Genius Labs Editorial · 6 min readUpdated
Engineering

Our view: a six-control checklist for AI agent authorization

Our view is that enterprises can apply a non-exhaustive six-control authorization checklist while NIST’s agent-specific work remains under development.

Super Genius Labs Editorial · 7 min read
Engineering

Running an agent team in production

The boring infrastructure problems you only meet after the demo is over.

Super Genius Labs Editorial · 2 min readUpdated