Skip to content
Editorial team

Super Genius Labs Editorial

Research and operating analysis from Super Genius Labs.

46 notesLatest All notes RSS feed
Published analysis
Product

The prior-authorization callback test for AI receptionist buyers

A proposed buyer test follows prior-authorization calls through waiting, denial, callback, and escalation—not just initial intake.

Super Genius Labs Editorial · Sep 15, 2026 · 5 min read
Thinking

Meta Muse’s dedicated per-user VM does not prevent provider access

Meta documents a dedicated per-user VM for Muse, while provider-access prevention remains a planned confidential-computing capability.

Super Genius Labs Editorial · Sep 14, 2026 · 3 min read
Engineering

Classify AWS DevOps Agent changes by infrastructure lifecycle

Review AWS DevOps Agent configuration changes by lifecycle effect, with replacement checks and post-apply inventory verification.

Super Genius Labs Editorial · Sep 13, 2026 · 3 min read
Engineering

Test session continuity across a hosted agent harness and a self-hosted executor

A failure-injection drill can test identity, reconnection, retained state, idle shutdown, and clean restart across an agent runtime boundary.

Super Genius Labs Editorial · Sep 11, 2026 · 4 min read
Engineering

Turn Microsoft Agent Framework’s shared-client concurrency contract into a host-level test grid

Test shared chat clients against event-loop, thread, session, and streaming boundaries before approving a Python agent-host topology.

Super Genius Labs Editorial · Sep 9, 2026 · 6 min read
Thinking

OpenAI’s “automated research intern” is an internal measurement claim, not a portable productivity benchmark

OpenAI’s research-intern milestone separates agent runtime, spending, activity, and research progress into distinct measurement layers.

Super Genius Labs Editorial · Sep 8, 2026 · 4 min read
Engineering

Stronger restriction-following can coincide with lower monitorability

Review behavioral control and operational observability separately when approving an agent deployment.

Super Genius Labs Editorial · Sep 7, 2026 · 4 min read
Engineering

Commission the machine boundary before an AI agent touches the controls

Document the physical control boundary, evidence, and safe recovery conditions before an AI agent receives write access.

Super Genius Labs Editorial · Sep 5, 2026 · 5 min read
Thinking

AWS and Google Cloud agent governance: a source-bounded capability matrix

Compare AWS and Google Cloud’s documented agent-governance scope without mistaking vendor descriptions for operating evidence.

Super Genius Labs Editorial · Sep 4, 2026 · 4 min read
Engineering

Version skew at the persistent agent-host boundary

A mixed-version drill separates protocol negotiation, effective capabilities, and session recovery before an agent-host rollout.

Super Genius Labs Editorial · Sep 3, 2026 · 4 min read
Thinking

Anthropic is turning frontier-model safeguards into a product tier

Fable 5.1 and Mythos 5.1 show why model procurement now has to capture access, safeguards, routing, retention, and cost.

Super Genius Labs Editorial · Sep 2, 2026 · 4 min read
Engineering

Benchmark improvement is not alignment transfer: a review gate for automated alignment research

A review gate separates benchmark gains in automated alignment research from evidence that a method will transfer to another model or setting.

Super Genius Labs Editorial · Sep 1, 2026 · 5 min read
Engineering

“Local-first” agents need an escalation ledger

A proposed escalation ledger links each cloud approval to its workflow step, disclosed payload, destination, and returned result.

Super Genius Labs Editorial · Aug 31, 2026 · 5 min read
Engineering

An interoperability test for federated agent registries

A boundary-based test sequence separates ARD contract conformance from observed interoperability between federated agent registries.

Super Genius Labs Editorial · Aug 30, 2026 · 4 min read
Engineering

Tracing authority through AC2’s user-controlled signing flow

Trace AC2 signing from reviewed request to execution evidence while keeping credential custody and protocol limits explicit.

Super Genius Labs Editorial · Aug 29, 2026 · 4 min read
Engineering

Framework-agnostic agent evaluation still has an instrumentation contract

A six-part compatibility record shows whether an agent deployment supplies the telemetry that framework-agnostic evaluation expects.

Super Genius Labs Editorial · Aug 28, 2026 · 4 min read
Thinking

Google’s agent pricing now turns usage forecasts into commitment risk

Google’s new agent billing options make workload variability, unused spend, and commitment duration part of the platform decision.

Super Genius Labs Editorial · Aug 27, 2026 · 4 min read
Thinking

What a publication-date filter can prove

Date filters can enforce retrieval eligibility, but historical claims need a separate check of what each result contained at the cutoff.

Super Genius Labs Editorial · Aug 25, 2026 · 3 min read
Product

Is an AI answering service HIPAA compliant? What practices need to check

No answering service is HIPAA compliant on its own. What HIPAA requires of a phone vendor, 7 questions to ask an AI receptionist, and the gaps that fail audits.

Super Genius Labs Editorial · Aug 25, 2026 · 7 min read
Thinking

Treat frontier-model safety pauses as vendor-continuity signals

OpenAI’s disclosed Astra slowdown offers a bounded procurement signal: provider-controlled capability thresholds and research-environment changes can alter roadmap assumptions even while release timing and capability assessments remain unresolved.

Super Genius Labs Editorial · Aug 24, 2026 · 3 min read
Product

Answering service cost in 2026: human vs virtual vs AI

Answering service cost in 2026 from published pricing: per-minute live services, virtual receptionist tiers, flat-rate AI, plus a worked sizing example.

Super Genius Labs Editorial · Aug 24, 2026 · 7 min read
Thinking

The FDA’s GenAI-device paper is a question set, not a compliance checklist

FDA is considering a competency-based evaluation model for generative-AI medical devices. Its discussion paper offers diligence questions about the finished device, intended use, clinical confirmation, and postmarket monitoring—not adopted requirements.

Super Genius Labs Editorial · Aug 23, 2026 · 4 min read
Product

The best AI receptionist for small business: an honest 2026 comparison

Five AI receptionists compared on published pricing, industry fit, and what happens when the AI cannot handle the call. We build one of them, and it is not the right pick for most readers.

Super Genius Labs Editorial · Aug 23, 2026 · 5 min read
Thinking

Can the receptionist actually change the schedule? A write-path test for voice-AI buyers

A scheduling conversation is not the same as a completed scheduling transaction. This write-path test traces discovery, booking, rescheduling, and cancellation from caller request to committed record, exception, and handoff.

Super Genius Labs Editorial · Aug 22, 2026 · 5 min read
Engineering

Model session cardinality before choosing persistent agent compute

AgentCore Runtime Instances extend managed agent sessions to multi-day, specialized compute. An SGL worksheet turns session count, startup tolerance, storage retention, and deletion ownership into a concrete capacity decision.

Super Genius Labs Editorial · Aug 20, 2026 · 4 min read
Engineering

A migration worksheet for MCP’s stateless 2026-07-28 wire model

The MCP 2026-07-28 revision documents a stateless transport model, per-request metadata, and standardized HTTP requirements. This worksheet turns those documented changes into bounded inventory and test questions for teams upgrading remote servers, with the official C# SDK as an implementation lens.

Super Genius Labs Editorial · Aug 19, 2026 · 3 min read
Engineering

A two-plane worksheet for agent operating systems—and an explicit infrastructure boundary

A two-plane worksheet helps platform teams separate governance decisions from runtime coordination, then make deployment context and infrastructure ownership reviewable at every boundary.

Super Genius Labs Editorial · Aug 18, 2026 · 3 min read
Engineering

Give every agent runtime contract a control owner, an enforcement point, and an evidence owner

An SGL responsibility matrix assigns each runtime-contract clause to a control-plane owner, an enforcement point, an evidence artifact, and a predetermined failure disposition.

Super Genius Labs Editorial · Aug 17, 2026 · 6 min read
Engineering

Where to place the trust boundaries for shared agent memory

Persistent memory can carry influence across agent runs. An SGL decision map separates storage, provenance, retrieval, authority, and recovery so one session boundary does not bear the entire security model.

Super Genius Labs Editorial · Aug 16, 2026 · 5 min read
Engineering

Self-healing agent infrastructure needs a permission ladder, not an autonomy switch

A proposed permission ladder separates observation, diagnosis, approval, remediation, verification, and reversal so incident automation can gain authority one bounded action at a time.

Super Genius Labs Editorial · Aug 15, 2026 · 4 min read
Engineering

An admission gate for third-party agent skills

Third-party agent skills can combine runtime instructions with executable scripts. A bounded admission process can verify provenance, content, permissions, isolation, updates, and revocation before granting access.

Super Genius Labs Editorial · Aug 14, 2026 · 5 min read
Thinking

What evidence should healthcare AI operators retain for non-device software?

FDA’s patient-safety inquiry gives healthcare AI operators a bounded prompt: preserve intended use, observed benefits, safety signals, implementation lessons, governance practices, and evidence ownership without treating an operational record as a legal classification or proof of safety.

Super Genius Labs Editorial · Aug 13, 2026 · 4 min read
Engineering

A workflow-level benchmark for agent infrastructure

A workflow-level benchmark can expose host work, boundary crossings, tool spikes, queueing, and idle-capacity overlap that token throughput alone does not describe.

Super Genius Labs Editorial · Aug 12, 2026 · 4 min read
Thinking

Read Cortex AI Gateway as an announced control plane—not a verified operating result

Snowflake has announced one control layer for agent access, activity records, interoperability, and AI spending. Preview status leaves its behavior, coverage, and enforcement open to operator verification.

Super Genius Labs Editorial · Aug 11, 2026 · 3 min read
Engineering

Test the effective policy of every coding-agent surface

Turn shared coding-agent settings into surface-specific acceptance tests for policy source, precedence, refresh behavior, denial outcomes, and retained evidence.

Super Genius Labs Editorial · Aug 10, 2026 · 5 min read
Engineering

A model score is sometimes a system score: an attribution sheet for multi-model agents

A named model may sit inside a routed, multi-agent harness. Record the models, routing, tools, benchmark configuration, access tier, and evidence owner behind the score.

Super Genius Labs Editorial · Aug 8, 2026 · 3 min read
Thinking

AI disclosure should survive the interaction

A footer or opening message can leave channel transitions, handoffs, or generated-media paths untested. Model disclosure as owned product states with reviewable regression evidence.

Super Genius Labs Editorial · Aug 7, 2026 · 5 min read
Engineering

Opus 5 effort belongs in the release configuration

Version model ID, thinking state, effort, output budget, endpoint, and fallback policy together, then evaluate that tuple across quality, completion, latency, turns, and cost.

Super Genius Labs Editorial · Aug 6, 2026 · 5 min read
Engineering

Treat agent evaluation sandboxes like connected production systems

A sandbox label does not describe every reachable system. Review egress, intermediaries, credentials, blast radius, kill authority, forensics, and notification before the run.

Super Genius Labs Editorial · Aug 5, 2026 · 6 min read
Thinking

Audit the workflow before buying healthcare voice automation

Map the scheduling rules, financial barriers, queues, callbacks, and ownership behind patient-access work before comparing healthcare voice products.

Super Genius Labs Editorial · Aug 4, 2026 · 3 min read
Engineering

From agent traces to release decisions

A production trace becomes release evidence only after scoring, review, regression, and an explicit decision. Automated evaluation contributes evidence, not proof of correctness.

Super Genius Labs Editorial · Aug 3, 2026 · 4 min read
Thinking

Measure healthcare call resolution beyond handling rates

A single automation or containment rate can hide materially different outcomes. Separate access, administrative task completion, callbacks, and human escalation.

Super Genius Labs Editorial · Aug 2, 2026 · 4 min read
Engineering

Six authorization controls for AI agents

NIST’s agent-specific work is still developing. This six-control review connects established NIST baselines with OWASP’s finalized agentic risk guidance.

Super Genius Labs Editorial · Aug 1, 2026 · 7 min read
Thinking

Welcome to Super Genius Labs

Notes on why we built a lab around agent teams, not chatbots.

Super Genius Labs Editorial · May 21, 2026 · 1 min read
Product

Genius Care: building a reliable voice receptionist

What the current receptionist can do, where it hands off, and what remains bounded.

Super Genius Labs Editorial · May 15, 2026 · 1 min read
Engineering

Running an agent team in production

The boring infrastructure problems you only meet after the demo is over.

Super Genius Labs Editorial · May 10, 2026 · 2 min read