Test the effective policy of every coding-agent surface
Turn shared coding-agent settings into surface-specific acceptance tests for policy source, precedence, refresh behavior, denial outcomes, and retained evidence.
Super Genius Labs Editorial · Aug 10, 2026 · 5 min readEngineeringA model score is sometimes a system score: an attribution sheet for multi-model agents
A named model may sit inside a routed, multi-agent harness. Record the models, routing, tools, benchmark configuration, access tier, and evidence owner behind the score.
Super Genius Labs Editorial · Aug 8, 2026 · 3 min readThinkingAI disclosure should survive the interaction
A footer or opening message can leave channel transitions, handoffs, or generated-media paths untested. Model disclosure as owned product states with reviewable regression evidence.
Super Genius Labs Editorial · Aug 7, 2026 · 5 min readEngineeringOpus 5 effort belongs in the release configuration
Version model ID, thinking state, effort, output budget, endpoint, and fallback policy together, then evaluate that tuple across quality, completion, latency, turns, and cost.
Super Genius Labs Editorial · Aug 6, 2026 · 5 min readEngineeringTreat agent evaluation sandboxes like connected production systems
A sandbox label does not describe every reachable system. Review egress, intermediaries, credentials, blast radius, kill authority, forensics, and notification before the run.
Super Genius Labs Editorial · Aug 5, 2026 · 6 min readThinkingAudit the workflow before buying healthcare voice automation
Map the scheduling rules, financial barriers, queues, callbacks, and ownership behind patient-access work before comparing healthcare voice products.
Super Genius Labs Editorial · Aug 4, 2026 · 3 min readEngineeringFrom agent traces to release decisions
A production trace becomes release evidence only after scoring, review, regression, and an explicit decision. Automated evaluation contributes evidence, not proof of correctness.
Super Genius Labs Editorial · Aug 3, 2026 · 4 min readThinkingMeasure healthcare call resolution beyond handling rates
A single automation or containment rate can hide materially different outcomes. Separate access, administrative task completion, callbacks, and human escalation.
Super Genius Labs Editorial · Aug 2, 2026 · 4 min readEngineeringSix authorization controls for AI agents
NIST’s agent-specific work is still developing. This six-control review connects established NIST baselines with OWASP’s finalized agentic risk guidance.
Super Genius Labs Editorial · Aug 1, 2026 · 7 min readThinkingWelcome to Super Genius Labs
Notes on why we built a lab around agent teams, not chatbots.
Super Genius Labs Editorial · May 21, 2026 · 1 min read