Skip to content
Blog archive

Page 4

Every field note remains reachable as the lab grows.

Page 4 of 510 notesAll notes RSS feed
Engineering

Test the effective policy of every coding-agent surface

Turn shared coding-agent settings into surface-specific acceptance tests for policy source, precedence, refresh behavior, denial outcomes, and retained evidence.

Super Genius Labs Editorial · Aug 10, 2026 · 5 min read
Engineering

A model score is sometimes a system score: an attribution sheet for multi-model agents

A named model may sit inside a routed, multi-agent harness. Record the models, routing, tools, benchmark configuration, access tier, and evidence owner behind the score.

Super Genius Labs Editorial · Aug 8, 2026 · 3 min read
Thinking

AI disclosure should survive the interaction

A footer or opening message can leave channel transitions, handoffs, or generated-media paths untested. Model disclosure as owned product states with reviewable regression evidence.

Super Genius Labs Editorial · Aug 7, 2026 · 5 min read
Engineering

Opus 5 effort belongs in the release configuration

Version model ID, thinking state, effort, output budget, endpoint, and fallback policy together, then evaluate that tuple across quality, completion, latency, turns, and cost.

Super Genius Labs Editorial · Aug 6, 2026 · 5 min read
Engineering

Treat agent evaluation sandboxes like connected production systems

A sandbox label does not describe every reachable system. Review egress, intermediaries, credentials, blast radius, kill authority, forensics, and notification before the run.

Super Genius Labs Editorial · Aug 5, 2026 · 6 min read
Thinking

Audit the workflow before buying healthcare voice automation

Map the scheduling rules, financial barriers, queues, callbacks, and ownership behind patient-access work before comparing healthcare voice products.

Super Genius Labs Editorial · Aug 4, 2026 · 3 min read
Engineering

From agent traces to release decisions

A production trace becomes release evidence only after scoring, review, regression, and an explicit decision. Automated evaluation contributes evidence, not proof of correctness.

Super Genius Labs Editorial · Aug 3, 2026 · 4 min read
Thinking

Measure healthcare call resolution beyond handling rates

A single automation or containment rate can hide materially different outcomes. Separate access, administrative task completion, callbacks, and human escalation.

Super Genius Labs Editorial · Aug 2, 2026 · 4 min read
Engineering

Six authorization controls for AI agents

NIST’s agent-specific work is still developing. This six-control review connects established NIST baselines with OWASP’s finalized agentic risk guidance.

Super Genius Labs Editorial · Aug 1, 2026 · 7 min read
Thinking

Welcome to Super Genius Labs

Notes on why we built a lab around agent teams, not chatbots.

Super Genius Labs Editorial · May 21, 2026 · 1 min read