Our view: AI disclosure as a non-exhaustive set of tested interaction states
We infer that a footer or opening message alone may be insufficient for operating AI disclosure across channels, handoffs, and generated-media delivery. Our non-exhaustive matrix turns disclosure into owned, counsel-reviewable product states and regression checks.
Our view: a non-exhaustive release matrix for Opus 5 inference configurations
We infer that an Opus 5 deployment can be evaluated more precisely when model ID, thinking state, effort, output budget, endpoint, and fallback policy are versioned together. Our non-exhaustive matrix evaluates those inputs across quality, tool completion, latency, turn count, and cost.
Our view: a non-exhaustive containment review for agent evaluation environments
We infer that a sandbox label alone may be insufficient for tool-enabled agent evaluations. Our non-exhaustive review covers egress, intermediary services, credential reach, blast radius, action-volume detection, kill authority, forensic retention, and notification.
Our view: a non-exhaustive pre-purchase workflow audit for healthcare reception
We propose a non-exhaustive audit for tracing representative calls through scheduling rules, financial barriers, queues, callbacks, and ownership before evaluating voice automation.
Our view: a non-exhaustive release-control loop built from agent traces
We propose a non-exhaustive engineering loop that connects production traces, contextual scoring, human review, regression cases, and predeployment checks. The model treats automated evaluation as bounded evidence rather than proof of correctness.
Our view: a non-exhaustive healthcare call-resolution scorecard beyond handling rates
Our view is that a single automation, containment, or “dealt with” rate may be insufficient for evaluating healthcare voice automation. We infer a non-exhaustive scorecard that separates access, administrative task completion, and human escalation.
Our view: a six-control checklist for AI agent authorization
Our view is that enterprises can apply a non-exhaustive six-control authorization checklist while NIST’s agent-specific work remains under development.
Welcome to Super Genius Labs
Notes on why we built a lab around agent teams, not chatbots.
Genius Care: building a reliable voice receptionist
What the current receptionist can do, where it hands off, and what remains bounded.
Running an agent team in production
The boring infrastructure problems you only meet after the demo is over.
