Skip to content
Blog archive

Page 2

Every field note remains reachable as the lab grows.

Page 2 of 310 notesAll notes RSS feed
Thinking

What evidence should healthcare AI operators retain for non-device software?

FDA’s patient-safety inquiry gives healthcare AI operators a bounded prompt: preserve intended use, observed benefits, safety signals, implementation lessons, governance practices, and evidence ownership without treating an operational record as a legal classification or proof of safety.

Super Genius Labs Editorial · 4 min read
Engineering

A workflow-level benchmark for agent infrastructure

A workflow-level benchmark can expose host work, boundary crossings, tool spikes, queueing, and idle-capacity overlap that token throughput alone does not describe.

Super Genius Labs Editorial · 4 min read
Thinking

Read Cortex AI Gateway as an announced control plane—not a verified operating result

Snowflake has announced one control layer for agent access, activity records, interoperability, and AI spending. Preview status leaves its behavior, coverage, and enforcement open to operator verification.

Super Genius Labs Editorial · 3 min read
Engineering

Test the effective policy of every coding-agent surface

Turn shared coding-agent settings into surface-specific acceptance tests for policy source, precedence, refresh behavior, denial outcomes, and retained evidence.

Super Genius Labs Editorial · 5 min read
Engineering

A model score is sometimes a system score: an attribution sheet for multi-model agents

A named model may sit inside a routed, multi-agent harness. Record the models, routing, tools, benchmark configuration, access tier, and evidence owner behind the score.

Super Genius Labs Editorial · 3 min read
Thinking

AI disclosure should survive the interaction

A footer or opening message can leave channel transitions, handoffs, or generated-media paths untested. Model disclosure as owned product states with reviewable regression evidence.

Super Genius Labs Editorial · 5 min readUpdated
Engineering

Opus 5 effort belongs in the release configuration

Version model ID, thinking state, effort, output budget, endpoint, and fallback policy together, then evaluate that tuple across quality, completion, latency, turns, and cost.

Super Genius Labs Editorial · 5 min readUpdated
Engineering

Treat agent evaluation sandboxes like connected production systems

A sandbox label does not describe every reachable system. Review egress, intermediaries, credentials, blast radius, kill authority, forensics, and notification before the run.

Super Genius Labs Editorial · 6 min readUpdated
Thinking

Audit the workflow before buying healthcare voice automation

Map the scheduling rules, financial barriers, queues, callbacks, and ownership behind patient-access work before comparing healthcare voice products.

Super Genius Labs Editorial · 3 min readUpdated
Engineering

From agent traces to release decisions

A production trace becomes release evidence only after scoring, review, regression, and an explicit decision. Automated evaluation contributes evidence, not proof of correctness.

Super Genius Labs Editorial · 4 min readUpdated